An infrared target recognition multi-dimension complexity representation method under a complex ground scene
By dynamically evaluating the complexity of infrared target recognition and tracking using Zernike moment eigenvectors and the likelihood transformation function, the problem of objective quantification of recognition and tracking tasks in complex ground scenarios is solved, and the system's adaptability and algorithm evaluation fairness are improved.
Patent Information
- Application Number
- CN202511758108.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-11-27
AI Technical Summary
Existing infrared target recognition and tracking technologies lack objective quantitative standards for recognition and tracking tasks in complex ground scenarios. They cannot dynamically characterize target feature degradation, motion complexity, and background interference intensity, resulting in unfair algorithm evaluation and the inability of the system to adaptively adjust.
Zernike moment feature vectors are used to extract target features, and tangential and normal complexity are combined to represent target motion. Local background interference is quantified by a susceptibility transformation function, and a dynamic weight fusion mechanism is constructed to comprehensively evaluate the complexity of identification and tracking.
It realizes multi-dimensional complexity representation of infrared target recognition in complex ground scenarios, improves the interpretability and adaptability of recognition and tracking technology, and supports online switching of algorithm strategies and dynamic allocation of system resources.
Smart Images

Figure CN121213898B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of infrared target recognition and tracking technology, and in particular to a method for characterizing the multi-dimensional complexity of infrared target recognition in complex ground scenarios. Background Technology
[0002] Infrared target recognition and tracking technology has significant application value in fields such as reconnaissance and security monitoring. Its core challenge lies in the ability to continuously and stably identify and track targets in complex scenarios. Current research mainly focuses on the design and optimization of tracking algorithms. However, a fundamental problem that has long been neglected is the lack of objective quantitative standards for the inherent difficulty of the recognition and tracking task itself. Existing evaluation systems generally rely on posterior statistical indicators, such as success rate curves and center position errors. These methods can only reflect the algorithm's performance in specific scenarios but cannot predict or explain the essential reasons for tracking failures in different scenarios, leading to a lack of fairness in algorithm comparisons.
[0003] Specifically, interference sources in complex ground scenes, such as high-intensity regions in a low signal-to-clutter background, target-like debris, sharp edges, and high-brightness pixel noise, dynamically weaken the discernibility of targets. The risk of tracking failure is further amplified when targets undergo violent maneuvers or encounter partial occlusion. Although the academic community has recognized the impact of these factors on tracking performance, current technologies still cannot characterize the combined effect. For example, the infrared sequence complexity measurement method proposed by Wang Xiaotian et al. (“Infrared Image Sequence Complexity Measurement for Automatic Target Tracking” Wang Xiaotian, Ma Wanchao, Zhang Kai, Li Shaoyi, Yan Jie. Journal of Northwestern Polytechnical University, 2019, 37(4):664-672.) constructs target confusion and target occlusion indices, but its feature space optimization relies on the grey relational method and does not cover the key dimension of the unpredictability of target motion; while extended methods based on traditional image indices, such as information entropy, discrete coefficient, PSNR / SSIM inter-frame variation, can describe the background clutter of a single frame or the continuity between frames, but it is difficult to unify the real-time contribution of motion mutation, background interference and feature degradation to the difficulty of recognition and tracking.
[0004] Furthermore, a more fundamental limitation lies in the fact that current research largely treats complexity as a static attribute or a posteriori result, failing to establish a mechanism capable of dynamically generating quantitative scores during the recognition and tracking process. This deficiency not only hinders the scientific evaluation of algorithm performance but also limits the intelligent evolution of tracking systems. For example, it prevents the dynamic switching of algorithm strategies or the activation of sensor collaboration based on real-time complexity. Summary of the Invention
[0005] The purpose of this invention is to provide a multi-dimensional complexity characterization method for infrared target recognition in complex ground scenarios. This method can integrate the multi-dimensional complexity characterization method of target feature degradation degree, target motion pattern complexity, and background and suspected target interference intensity, thereby improving the interpretability and adaptability of infrared recognition and tracking technology.
[0006] To achieve the above objectives, this invention provides a method for characterizing the multi-dimensional complexity of infrared target recognition in complex ground scenarios, comprising the following steps:
[0007] S1. Compare the differences between the current features of the target and the historical baseline features of the target to quantify feature degradation; use the simplified Zernike moments of the target from 0 to 4th order to extract feature vectors;
[0008] S2. Extract the target motion feature vector, use two components, tangential complexity and normal complexity, to represent the two complexities in the target motion, and calculate the target motion complexity accordingly.
[0009] S3. Extract the historical baseline features of the target and the local background features of the target to quantify their differences, and use the set suspicion degree conversion function to calculate the similarity between the historical features of the target and the local background features to characterize the suspicion degree of the local background.
[0010] S4. Extract the feature vectors and motion vectors of real targets and false moving targets, quantify the feature differences and inter-frame motion differences between real targets and false moving targets, and use the doubt degree conversion to represent the degree of interference of false moving targets on target recognition and tracking;
[0011] S5. Construct a dynamic weight fusion mechanism to integrate target feature degradation degree, target motion complexity, local background suspicion degree, and false target suspicion degree into a comprehensive recognition and tracking complexity.
[0012] Preferably, S1 specifically includes the following steps:
[0013] S11. Using the segmented target image frames and mask data, construct historical baseline features and select the latest continuous historical features. K For each frame of the target image, frames with a degradation degree greater than α are removed. The target region is then cropped according to the mask. The 0th to 4th order moduli of the Zernike moments are calculated to generate a 9-dimensional target feature vector. The target shape features are then scaled and normalized.
[0014] ; ;
[0015] ; ; ; ; ; ; ; ;
[0016] ;
[0017] The above describes the Zernike moment eigenvector calculation method, where... M ( x , y ) represents the coordinates of the mask data ( x , y The value at () I ( x , y ) represents the coordinates of the image data ( x , y The value at () I target ( x , y ) represents the coordinates of the target region in the image. x , y The value at () P target Represents image data; I norm ( x , y ) represents the coordinates in the image after grayscale normalization. x , y The value at () A This indicates the number of pixels in the target region within the currently calculated image data; c x Represents the x-coordinate of the target energy center; c y Represents the ordinate of the target energy center; R max Indicates the normalized maximum radius; x ′ represents the center-normalized x-coordinate; y ′ represents the center-normalized ordinate; r Indicates the pixel normalized radius; n This indicates the order of the Zernike moments; i Indicates the dimension index of the feature vector; V i ( r ) represents the Zernike radial polynomial; f i Indicates the first i The value of the first-order eigenvector;
[0018] S12. Select the latest historical continuous K After extracting features from the frame and performing scale normalization, we can obtain:
[0019] ; ; ;
[0020] in, k f represents the feature vector dimension index; k The 9-dimensional feature vector representing the target, where f k,1 Indicates the first k Frame image target energy features f k,1 ,…, f k,9 Indicates the first k Eight shape feature components of the target in the frame image; K This represents the latest consecutive historical frame count; A k Indicates the first k Total number of pixels in the frame target mask; Indicates the first k The first frame after target scale normalization i Each characteristic component; f base Indicates the historical baseline characteristics of the target;
[0021] S13. Perform feature extraction on the target region of the current frame and perform scale normalization to obtain f. curr This indicates the current characteristics of the target;
[0022] S14. Perform feature difference calculation:
[0023] ; ; ;
[0024] in, f base,1 Represents the historical baseline feature f of the target base The first component; f curr,1 Represents the historical baseline feature f of the target curr The first component; f base,i Represents the historical baseline feature f of the target base The i One component; f curr,i Represents the historical baseline feature f of the target curr The i One component; d e Indicates the degree of degradation of the target energy characteristic; d s γ represents the degradation degree of the target shape features; γ represents the weight of the target energy features. Cdegrad This indicates the degree of degradation of the target feature.
[0025] Preferably, S2 specifically includes the following steps:
[0026] S21, Extract the latest K The degradation degree of each target feature is less than α Image frames, if this K If the image frames are consecutive, the number of pixels in the target region can be directly extracted based on the target mask data. A k and the target in continuous K Position sequence in the frame:
[0027] ; ;
[0028] ; ;
[0029] ; ;
[0030] in, K The total number of frames calculated; M k Indicates the first k Each image data corresponds to a target mask data; A k Indicates the first k The number of pixels in the target region in each image data; Indicates the short-term average area of the target; I k ( x , y ) indicates the first k In the image data ( x , y Gray value at ) c k,x Indicates the first k Target location in image data x coordinate; c k,y Indicates the first k Target location in image data y Coordinates; p k express k The target gray-scale weighted center coordinates at each time step; p0, p1, ..., p K-1 Representing 0 to K- At time 1, the target's gray-scale weighted center coordinates; conversely, if this... KIf the image frames are not continuous, then after extracting the area coordinates as described above, linear interpolation is used to fill in the missing target locations based on the existing coordinate data, and the latest coordinates are extracted again according to the filled-in coordinates. The target location coordinates are p0, p1, ..., p1. K-1 ;
[0031] S22. Using the displacement representation between adjacent frames, calculate the velocity vector between adjacent frames of the target: ;in, v k express k arrive k+1 The change in target displacement at any given moment;
[0032] S23. Use the rate change between adjacent frames to represent the tangential complexity of a single frame;
[0033] ; ; ;
[0034] in, S k Represented as velocity vector v k The European mold, S k+1 Represented as velocity vector v k+1 The Euclidean model reflects the target in [ k , k +1] The speed of movement within the time window; S k This represents the rate difference between two adjacent frames; T k express[ k , k +1] The square of the rate difference within the time window;
[0035] S24. Calculate the change in motion direction between two adjacent frames:
[0036] ; ;
[0037] ; ;
[0038] in, v k,y Represents the velocity vector v k of y Axial components; v k,x Represents the velocity vector vk of x Axial components; θ k express k The target's direction of motion angle at any given moment; θ k+1 express k+ The target's direction of motion angle at moment 1; θ k Indicates the rotation angle of the target between two adjacent frames; This indicates the average rate between two frames; N k express k Approximate target normal acceleration at any given time;
[0039] S25. Combine time-memory weighted calculation of the overall target motion complexity:
[0040] ;
[0041] ; ; ;
[0042] in, M = K -3 represents the total number of components to be calculated; w k This represents the time memory weight of the corresponding component. k =0 corresponds to the earliest time component. k = M -1 corresponds to the latest time component; T k express[ k , k +1] The square of the rate difference within the time window; C T Indicates the target's tangential acceleration; N k express k The target's normal acceleration at any given moment is approximate; C N This represents an approximate acceleration in the target direction; C motion This represents the complexity of the target motion.
[0043] Preferably, step S3 specifically includes the following steps:
[0044] S31. Following the method shown in S1, construct the target baseline features based on historical data, and select the latest continuous historical data. K The target image of the frame does not contain any frames with a calculated degradation degree greater than [value missing]. αFrom the image frames, the target region is cropped according to the mask, the target feature vector is extracted, and scale normalization is performed to obtain the historical baseline feature vector f. base ;
[0045] S32. Perform local background region scanning on the target, and extract local background features for each scanning window and perform scale normalization:
[0046] ; ;
[0047] ;
[0048] ;
[0049] in, W t , H t These represent the width and height of the target region in the current frame, respectively. W e , H e These represent the width and height of the current frame scanning window, respectively. x , y These represent the horizontal and vertical step sizes of the window scan, respectively; a p , b p ) indicates the first p The coordinates of the scan window indicate the horizontal scan position. a p Each step length, vertical b p Each step length; P Indicates the total number of scanned windows; p Indicates the scan window number; b p,1 , b p,2 ,…, b p,9 Indicates the first p The nine feature values of each window after scale normalization; b p No. p Feature vectors of each window;
[0050] S33. Calculate the local background region likelihood, first calculating the current historical baseline vector f. base The differences are compared with the feature vectors of each window, and the differences are transformed into a likelihood level, then a weighted average is calculated based on the window coordinates:
[0051] ; ; ;
[0052] ; ;
[0053] ; ; ;
[0054] in, d e,p Indicates the first p Differences in energy characteristics among individual windows; d s,p Indicates the first p Differences in the shape characteristics of individual windows; d p Indicates the first p Differences in features among individual windows; S ( d () represents the suspicion conversion function; This represents the function for finding the median. MP This represents the set of highly suspected background window numbers; w p Indicates the first p The positional weight of each window; Indicates the first p The normalized position weights of each window; C bg Indicates the degree of suspicion of local background.
[0055] Preferably, step S4 specifically includes the following steps:
[0056] S41. Based on historical data, construct target baseline features and select the latest continuous historical data. K The frames contain real and fake target images, excluding those with a calculated degradation degree greater than [value missing]. α The image frames are processed, and the target regions are cropped according to the corresponding masks of the two targets. The feature vectors of real and fake targets are extracted according to S3 and scaled and normalized to obtain the historical baseline feature vectors of real and fake targets. , Simultaneously, according to S2, the positions of the two targets are extracted based on their mask, and their trajectories are obtained. , , ;
[0057] S42. Calculate the relative distance between the two targets in each frame based on their position trajectories: ;
[0058] like Larger than the set target relative local range If so, it will not affect target tracking and will directly skip the process of calculating the likelihood of a false target; otherwise, if First, calculate the difference in motion between the real and false targets:
[0059] ; ; ;
[0060] ; ; ; ; ;
[0061] in, L This indicates the relative local range of the set target; Indicates the true target number k The motion velocity vector of the frame; Indicates the false motion target. k The motion velocity vector of the frame; v k Indicate the true or false target number k Frame speed difference; d k Indicate the true or false target number k Frame scale-normalized relative distance; L This indicates the relative local area range of the set target; w s,k Indicates distance weight; This represents the distance weights after normalization; w t,k Indicates the weight of time memory; Indicates the joint weight; d m Indicates the difference in motion between real and false targets;
[0062] S43. Calculate the difference in features between true and false targets according to S1:
[0063] ; ;
[0064] ;
[0065] in, d base,e This indicates the difference in energy characteristics between true and false targets; d base,s Indicates the difference in shape features between real and false targets; d base Indicates the difference in features between true and false targets;
[0066] S44. Calculate the likelihood of a target being suspected using the differences in features and motion between real and fake targets:
[0067] ; ;
[0068] in, d total Indicates the total difference between true and false objectives; C dyn Indicates the degree of suspicion of a fake movement target.
[0069] Preferably, step S5 specifically includes the following steps:
[0070] S51. Obtain the latest identified segmentation. Historical frame data and their corresponding complexity For each complexity metric, kernel density estimation is performed, with the bandwidth parameter employing the Silverman criterion. The maximum probability impression value is then calculated using the probability density function. Furthermore, an innovative time-window memory fusion method is used to analyze the latest historical data. Frames and The frame is calculated, and the memory-weighted fusion based on the window entropy is used to obtain the basic weights corresponding to this complexity:
[0071] ; ; ;
[0072] ; ;
[0073] ; ;
[0074] ; ;
[0075] ; ;
[0076] ; ;
[0077] in, K This indicates that half of the latest historical frames have been selected. k Indicates the sequence number of the historical frame; Denotes the normalized function, where v min This represents the minimum value in vector v. v max This represents the maximum value in vector v; Denotes the normalization function, where vi Represents the first vector v i One portion, D This indicates the dimension of vector v; v (i) Representing complexity i The latest 2 K Frame data vector, where These represent four different levels of complexity. Representing complexity i The latest 1 to Frame data; Represents the complexity after standardization and normalization. i The latest Frame data vector; Represents the complexity after standardization and normalization. i The latest 1 to 2 K Frame data; h The bandwidth used for kernel density estimation is determined by the Silverman criterion. Represents the Gaussian kernel function; Representing complexity i The probability density function obtained by near-window kernel density estimation; Representing complexity i The probability density function obtained by full-window kernel density estimation; Represents the inverse normalized function; Represents the inverse normalization function; For complexity i The x-coordinate value at the point where the probability density near the window is maximized represents the complexity. i Impression value near the window; For complexity i The x-coordinate value at the point where the full-window probability density is maximized represents the complexity. i Full-window impression value; Representing complexity i Near-window data entropy; Representing complexity i Full-window data entropy; Representing complexity i Near-window weighting; Representing complexity i Full window weight; I (i) Representing complexity i The complexity impression value; Representing complexity i Dynamic fusion base weights;
[0078] S52. Calculate the dynamic adjustment factor based on the kernel density fitted probability prediction and the actual probability, and perform collaborative complexity enhancement on the dynamically adjusted weights. The specific calculation method is as follows:
[0079] ; ; ;
[0080] ;
[0081] ;
[0082] ; ;
[0083] ; ; ;
[0084] in, Representing complexity i The maximum density ratio of the current value in the near-window distribution; Representing complexity i The maximum density ratio of the current value in the full window distribution; Complexity i Near-window difference symbols; Complexity i The full window difference symbol; This indicates that the gain coefficient is dynamically adjusted. Representing complexity i The dynamically adjusted weights; β ij Representing complexity i With complexity j Synergistic factors; Representing complexity i The final dynamic fusion weight value; c (i) Representing complexity i Gains are influenced by the synergistic effect of other complexities; C (i) These represent the initial calculated values of the four complexity metrics, where... i =1,2,3,4; These represent the final values of the four complexity metrics; C total This represents the final dynamic blending complexity of the current window.
[0085] Therefore, this invention employs the aforementioned multi-dimensional complexity characterization method for infrared target recognition in complex ground scenarios. By simultaneously analyzing the unpredictability of target motion patterns, the interference intensity of local background on target identification, and the changes in the identifiability of target features affected by occlusion or noise, a dynamically generated recognition and tracking complexity scoring mechanism is constructed. This score can reflect the tracking difficulty of the current frame or sequence in real time and objectively, thus providing key theoretical basis and decision support for online switching of algorithm strategies, dynamic allocation of system resources, and scientific evaluation of cross-scenario algorithm performance. This invention is conducive to promoting the paradigm shift of recognition and tracking systems from passive response to active perception and adaptive decision-making.
[0086] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0087] Figure 1 This is a flowchart of an embodiment of the present invention;
[0088] Figure 2 These are actual data interference effect diagrams for frames 20 / 22 / 24 of this embodiment of the invention;
[0089] Figure 3 These are frames 1-61 of the embodiments of the present invention. C motion Time series variation curve;
[0090] Figure 4 These are frames 1-61 of the embodiments of the present invention. C degrad Time series variation curve;
[0091] Figure 5 These are frames 1-61 of the embodiments of the present invention. C bg Time series variation curve;
[0092] Figure 6 These are frames 1-61 of the embodiments of the present invention. C dyn Time series variation curve;
[0093] Figure 7 These are frames 1-61 of the embodiments of the present invention. C total Time series variation curve;
[0094] Figure 8 This is a comparison curve of the temporal changes in the complexity of each item in frames 1-61 of the embodiment of the present invention. Detailed Implementation
[0095] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0096] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0097] Example 1
[0098] This invention provides a method for characterizing the multi-dimensional complexity of infrared target recognition in complex ground scenes, the process of which is as follows: Figure 1 As shown.
[0099] Explanation of the test dataset and basic parameters in this embodiment:
[0100] This implementation example uses a sequence from the VOT-TIR infrared target tracking open-source dataset for testing. Infrared target images and corresponding mask data from frames 1 to 61 (frame range 1...61) of the sequence are selected to evaluate the overall complexity of the target tracking process. The core parameters used in the test are as follows to ensure consistency between the calculation logic and the theoretical method:
[0101] Start frame: START_FRAME=1, End frame: END_FRAME=61.
[0102] Number of historical frames selected: K=10 (used to construct historical baseline features and motion trajectories).
[0103] Occlusion detection threshold: α=0.5.
[0104] Energy feature weight: γ=0.382 (the proportion of energy features in the calculation of target feature degradation degree).
[0105] Number of local background scan windows: P =72 (Total number of scan windows in local background suspicion calculation), window size W e =1.5 W t , H e =1.5 H t( W t / H t (Target width / height).
[0106] False target identification range: L= 3 (Skip the calculation of suspected false targets when the relative distance is greater than this value).
[0107] Includes the following steps:
[0108] S1. Compare the current features of the target with the historical baseline features of the target to quantify feature degradation, mainly caused by occlusion or noise; use the simplified modulus of the target's 0th-4th order Zernike moments for feature vector extraction. The specific methods for feature vector extraction and target feature degradation degree are as follows:
[0109] S11. Using the segmented target image frames and mask data, construct historical baseline features and select the latest continuous historical features. K For each frame of the target image, frames with a degradation degree greater than α are removed. The target region is then cropped according to the mask. The 0th to 4th order moduli of the Zernike moments are calculated to generate a 9-dimensional target feature vector. The target shape features are then scaled and normalized.
[0110] ; ;
[0111] ; ; ; ; ; ; ; ;
[0112] ;
[0113] The above describes the Zernike moment eigenvector calculation method, where... M ( x , y ) represents the coordinates of the mask data ( x , y The value at () I ( x , y ) represents the coordinates of the image data ( x , y The value at () I target ( x , y ) represents the coordinates of the target region in the image. x , y The value at ()P target Represents image data; I norm ( x , y ) represents the coordinates in the image after grayscale normalization. x , y The value at () A This indicates the number of pixels in the target region within the currently calculated image data; c x Represents the x-coordinate of the target energy center; c y Represents the ordinate of the target energy center; R max Indicates the normalized maximum radius; x ′ represents the center-normalized x-coordinate; y ′ represents the center-normalized ordinate; r Indicates the pixel normalized radius; n This indicates the order of the Zernike moments; i Indicates the dimension index of the feature vector; V i ( r ) represents the Zernike radial polynomial; f i Indicates the first i The value of the eigenvector of order 1.
[0114] In this embodiment, based on the segmented target image and mask data, the latest continuous historical data is selected. K =10 frames, excluding frames with a degradation level >0.5. Taking frame 3 as an example, the calculation results of this step are explained:
[0115] The first two frames were selected as historical baseline frames, and the number of pixels in the target area were respectively A 1=560、 A 2=521, calculate the coordinates of the target energy center: C x1 =257.739、 C y1 =253.443 (frame 1); C x2 =282.973、 C y2 =253.033 (frame 2).
[0116] Normalized maximum radius R max1 =18.5、 R max2 =17.8, obtained through Zernike radial polynomials V i (r Calculate the eigenvalues of each order to obtain the historical baseline eigenvector:
[0117] .
[0118] S12. Select the latest historical continuous K After extracting features from the frame and performing scale normalization, we can obtain:
[0119] ; ; ;
[0120] in, k f represents the feature vector dimension index; k The 9-dimensional feature vector representing the target, where f k,1 Indicates the first k Frame image target energy features f k,1 ,…, f k,9 Indicates the first k Eight shape feature components of the target in the frame image; K This represents the latest consecutive historical frame count; A k Indicates the first k Total number of pixels in the frame target mask; Indicates the first k The first frame after target scale normalization i Each characteristic component; f base This indicates the historical baseline characteristics of the target.
[0121] Taking frame 2 as an example, the 9-dimensional feature vector is:
[0122] The shape features (dimensions 2-9) after normalization are as follows:
[0123] ;
[0124] Final historical baseline characteristic mean:
[0125] .
[0126] S13. Perform feature extraction on the target region of the current frame and perform scale normalization to obtain f. curr This represents the current feature of the target. Taking frame 3 as an example, the same feature extraction process as in previous frames is performed to obtain the current feature vector:
[0127] .
[0128] S14. Perform feature difference calculation:
[0129] ; ; ;
[0130] in, f base,1 Represents the historical baseline feature f of the target base The first component; f curr,1 Represents the historical baseline feature f of the target curr The first component; f base,i Represents the historical baseline feature f of the target base The i One component; f curr,i Represents the historical baseline feature f of the target curr The i One component; d e Indicates the degree of degradation of the target energy characteristic; d s γ represents the degradation degree of the target shape features; γ represents the weight of the target energy features. C degrad This indicates the degree of degradation of the target feature.
[0131] In this embodiment, the actual calculation result is as follows:
[0132] ;
[0133] ;
[0134] ;
[0135] Results Explanation: Frame 3 C degrad =0.2228<0.5, indicating no occlusion and no significant degradation of target features.
[0136] S2. Extract the target motion feature vector, using tangential complexity and normal complexity as two components to represent the two complexities in the target motion, and calculate the target motion complexity accordingly. This includes the following steps:
[0137] S21, Extract the latest K The degradation degree of each target feature is less than α Image frames, if this K If the image frames are consecutive, the number of pixels in the target region can be directly extracted based on the target mask data. A k and the target in continuous K Position sequence in the frame:
[0138] ; ;
[0139] ; ;
[0140] ; ;
[0141] in, K The total number of frames calculated; M k Indicates the first k Each image data corresponds to a target mask data; A k Indicates the first k The number of pixels in the target region in each image data; Indicates the short-term average area of the target; I k ( x , y ) indicates the first k In the image data ( x , y Gray value at ) c k,x Indicates the first k Target location in image data x coordinate; c k,y Indicates the first k Target location in image data y Coordinates; p k express k The target gray-scale weighted center coordinates at each time step; p0, p1, ..., p K-1 Representing 0 to K- At time 1, the target's gray-scale weighted center coordinates; conversely, if this... K If the image frames are not continuous, then after extracting the area coordinates as described above, linear interpolation is used to fill in the missing target locations based on the existing coordinate data, and the latest coordinates are extracted again according to the filled-in coordinates. The target location coordinates are p0, p1, ..., p1. K-1 .
[0142] Taking frame 6 as an example: using consecutive frames 1-5 (all non-occluded), extract the number of pixels in the target area and the gray-scale weighted center coordinates:
[0143] ;
[0144] Short-term average area .
[0145] S22. Using the displacement representation between adjacent frames, calculate the velocity vector between adjacent frames of the target: ;in, v k express k arrive k+1 The target displacement change at any given time. The calculation result in this embodiment is:
[0146] ;
[0147] ;
[0148] ;
[0149] .
[0150] S23. Use the rate change between adjacent frames to represent the tangential complexity of a single frame;
[0151] ; ; ;
[0152] in, S k Represented as velocity vector v k The European mold, S k+1 Represented as velocity vector v k+1 The Euclidean model reflects the target in [ k , k +1] The speed of movement within the time window; S k This represents the rate difference between two adjacent frames; T k express[ k , k +1] The square of the rate difference within the time window. In this embodiment, the calculation result is:
[0153] ;
[0154] ;
[0155] ;
[0156] ;
[0157] ;
[0158] .
[0159] S24. Calculate the change in motion direction between two adjacent frames:
[0160] ; ;
[0161] ; ;
[0162] in, v k,y Represents the velocity vector v k of y Axial components; v k,x Represents the velocity vector v k of x Axial components; θ k express k The target's direction of motion angle at any given moment; θ k+1 express k+ The target's direction of motion angle at moment 1; θ k Indicates the rotation angle of the target between two adjacent frames; This indicates the average rate between two frames; N k express k An approximation of the target's normal acceleration at any given time. The calculation result in this embodiment is:
[0163] ;
[0164] ;
[0165] ;
[0166] (Converted to 0.029 radians);
[0167] (Converted to -0.050 radians);
[0168] ;
[0169] ;
[0170] .
[0171] S25. Combine time-memory weighted calculation of the overall target motion complexity:
[0172] ;
[0173] ; ; ;
[0174] in, M = K -3 represents the total number of components to be calculated; w k This represents the time memory weight of the corresponding component. k =0 corresponds to the earliest time component. k = M -1 corresponds to the latest time component; T k express[ k , k +1] The square of the rate difference within the time window; C T Indicates the target's tangential acceleration; N k express k The target's normal acceleration at any given moment is approximate; C N This represents an approximate acceleration in the target direction; C motion This represents the complexity of the target motion. In this embodiment, it incorporates time-memory weighting. The calculation yields:
[0175] , ;
[0176] ;
[0177] ;
[0178] ;
[0179] Results Explanation: After post-normalization, the 6th frame is... C motion =0.0920, the target motion exhibits slight turning, low maneuverability, and low motion complexity.
[0180] S3. Extract the historical baseline features of the target and the local background features of the target, quantify their differences, and use a set suspicion conversion function to calculate the similarity between the historical features of the target and the local background features to characterize the suspicion of the local background. Specifically, this includes the following steps:
[0181] S31. Following the method shown in S1, construct the target baseline features based on historical data, and select the latest continuous historical data. K The target image of the frame. It does not contain frames with a calculated degradation greater than [value missing]. αImage frames containing such targets can contaminate the normal target baseline features. The target region is cropped based on the mask, the target feature vector is extracted, and scale normalization is performed to obtain the historical baseline feature vector f. base In this embodiment, frames 1-5 (non-occluded) are selected to construct historical baseline features, resulting in:
[0182] .
[0183] S32. Perform local background region scanning on the target, and extract local background features for each scanning window and perform scale normalization:
[0184] ; ;
[0185] ;
[0186] ;
[0187] in, W t , H t These represent the width and height of the target region in the current frame, respectively. W e , H e These represent the width and height of the current frame scanning window, respectively. x , y These represent the horizontal and vertical step sizes of the window scan, respectively; a p , b p ) indicates the first p The coordinates of the scan window indicate the horizontal scan position. a p Each step length, vertical b p Each step length; P Indicates the total number of scanned windows; p Indicates the scan window number; b p,1 , b p,2 ,…, b p,9 Indicates the first p The nine feature values of each window after scale normalization; b p No. p The feature vector of each window.
[0188] Centered on the target in frame 6 pUsing 6 = (309.237, 252.936) as the baseline, set the scan window parameters:
[0189] Target width W t =42, High H t =23, therefore the scanning window width W e =1.5×42=63、High H e =1.5×23=34.5 (rounded down to 35);
[0190] Horizontal / Vertical Step ;
[0191] Scan window coordinates ,common P =72 windows, extract the 9-dimensional feature vector b of each window. p .
[0192] S33. Calculate the local background region likelihood, first calculating the current historical baseline vector f. base The differences are compared with the feature vectors of each window, and the differences are transformed into a likelihood level, then a weighted average is calculated based on the window coordinates:
[0193] ; ; ;
[0194] ; ;
[0195] ; ; ;
[0196] in, d e,p Indicates the first p Differences in energy characteristics among individual windows; d s,p Indicates the first p Differences in the shape characteristics of individual windows; d p Indicates the first p Differences in features among individual windows; S ( d () represents the suspicion conversion function; This represents the function for finding the median. MP This represents the set of highly suspected background window numbers; w p Indicates the first p The positional weight of each window; Indicates the first pThe normalized position weights of each window; C bg Indicates the degree of suspicion of local background.
[0197] This embodiment uses the first window as an example:
[0198] Eigenvector difference calculation:
[0199] Then we have:
[0200] ; ;
[0201] ;
[0202] Suspicion level conversion: ;
[0203] Window coordinate weighted average calculation:
[0204] Window coordinates are After normalization, it becomes After weighted fusion, it becomes The calculation is simplified here by performing spatial weighted normalization on all scan windows to obtain the 6th frame. C bg =0.3219, the similarity between background and target features is moderate, and the degree of interference with tracking is low.
[0205] S4. Extract the feature vectors and motion vectors of the real target and the false moving target, quantify the feature differences and inter-frame motion differences between the real target and the false moving target, and use the suspicion level conversion to represent the degree of interference of the false moving target on target recognition and tracking. Specifically, this includes the following steps:
[0206] S41. Based on historical data, construct target baseline features and select the latest continuous historical data. K The frames contain images of real and fake targets. The target regions are cropped based on the corresponding masks of the two targets. Feature vectors of real and fake targets are extracted according to S3 and scaled normalized to obtain the historical baseline feature vectors of real and fake targets. , This process does not include cases where the calculated degree of degradation is greater than [a certain value]. α Image frames containing such targets can contaminate the normal target baseline features. Simultaneously, according to S2, the positions of the two targets are extracted based on their masks, resulting in their trajectories. , , In this embodiment, no test sequence was found that satisfied the condition "relative distance ≤ L= The false target "3" was selected. K =10 frames contain only real targets, with no false target trajectories. .
[0207] S42. Calculate the relative distance between the two targets in each frame based on their position trajectories: ;
[0208] like Larger than the set target relative local range If so, it will not affect target tracking and will directly skip the process of calculating the likelihood of a false target; otherwise, if First, calculate the difference in motion between the real and false targets:
[0209] ; ; ;
[0210] ; ; ; ; ;
[0211] in, L This indicates the relative local range of the set target; Indicates the true target number k The motion velocity vector of the frame; Indicates the false motion target. k The motion velocity vector of the frame; v k Indicate the true or false target number k Frame speed difference; d k Indicate the true or false target number k Frame scale-normalized relative distance; L This indicates the relative local area range of the set target; w s,k Indicates distance weight; This represents the distance weights after normalization; w t,k Indicates the weight of time memory; Indicates the joint weight; d m This indicates the difference in motion between real and false targets.
[0212] In this embodiment, the real target is... k Minimum relative distance between the frame and surrounding candidate regions Skip the calculation of motion differences for false targets.
[0213] S43. Calculate the difference in features between true and false targets according to S1:
[0214] ; ;
[0215] ;
[0216] in, d base,e This indicates the difference in energy characteristics between true and false targets; d base,s Indicates the difference in shape features between real and false targets; d base This indicates the difference in characteristics between true and false targets.
[0217] S44. Calculate the likelihood of a target being suspected using the differences in features and motion between real and fake targets:
[0218] ; ;
[0219] in, d total Indicates the total difference between true and false objectives; C dyn This indicates the likelihood of a false moving target. In this embodiment, because the dataset lacks valid false target data, when there are no false targets... C dyn =0, meaning the interference level is 0, the actual calculation result. C dyn =0.0000.
[0220] Results show that the false target suspicion score is 0.0000, indicating that at the minimum relative distance, without the interference of false moving targets, it does not interfere with the tracking of real targets.
[0221] S5. Construct a dynamic weight fusion mechanism to integrate target feature degradation, target motion complexity, local background suspicion, and false target suspicion into a comprehensive recognition and tracking complexity. Specifically, this includes the following steps:
[0222] S51. Obtain the latest identified segmentation. Historical frame data and their corresponding complexity For each complexity metric, kernel density estimation is performed, with the bandwidth parameter employing the Silverman criterion. The maximum probability impression value is then calculated using the probability density function. Furthermore, an innovative time-window memory fusion method is used to analyze the latest historical data. Frames and The frame is calculated, and the memory-weighted fusion based on the window entropy is used to obtain the basic weights corresponding to this complexity:
[0223] ; ; ;
[0224] ; ;
[0225] ; ;
[0226] ; ;
[0227] ; ;
[0228] ; ;
[0229] in, K This indicates that half of the latest historical frames have been selected. k Indicates the sequence number of the historical frame; Denotes the normalized function, where v min This represents the minimum value in vector v. v max This represents the maximum value in vector v; Denotes the normalization function, where v i Represents the first vector v i One portion, D This indicates the dimension of vector v; v (i) Representing complexity i The latest 2 K Frame data vector, where These represent four different levels of complexity. Representing complexity i The latest 1 to Frame data; Represents the complexity after standardization and normalization. i The latest Frame data vector; Represents the complexity after standardization and normalization. i The latest 1 to 2 K Frame data; h The bandwidth used for kernel density estimation is determined by the Silverman criterion. Represents the Gaussian kernel function; Representing complexity i The probability density function obtained by near-window kernel density estimation; Representing complexity i The probability density function obtained by full-window kernel density estimation; Represents the inverse normalized function; Represents the inverse normalization function; For complexity iThe x-coordinate value at the point where the probability density near the window is maximized represents the complexity. i Impression value near the window; For complexity i The x-coordinate value at the point where the full-window probability density is maximized represents the complexity. i Full-window impression value; Representing complexity i Near-window data entropy; Representing complexity i Full-window data entropy; Representing complexity i Near-window weighting; Representing complexity i Full window weight; I (i) Representing complexity i The complexity impression value; Representing complexity i The dynamic fusion of basic weights.
[0230] The specific calculation process for this step in this embodiment is as follows:
[0231] Get the latest 2 K The complexity data for 20 frames across 4 dimensions is as follows, and the specific calculation process for frame 6 is as follows:
[0232] Standardization and normalization: with C degrad For example, in 20 frames C degrad,min =0.0000、 C degrad,max =0.6373, therefore the 6th frame C degrad After standardization G ( v =0.286, after normalization .
[0233] Kernel density estimation: bandwidth h Determined by Silverman's criterion h =0.08, the x-coordinate of the point where the KDE has the highest probability in the near window (last 10 frames). The x-coordinate of the point with the highest probability in the KDE for the entire window (20 frames). .
[0234] Fusion of window entropy and weights: Near window entropy Full window entropy Therefore , ;
[0235] Impression value ;
[0236] Basic weights (The basic weights for other dimensions are calculated similarly).
[0237] S52. Calculate the dynamic adjustment factor based on the kernel density fitted probability prediction and the actual probability, and perform collaborative complexity enhancement on the dynamically adjusted weights. The specific calculation method is as follows:
[0238] ; ; ;
[0239] ;
[0240] ;
[0241] ; ;
[0242] ; ; ;
[0243] in, Representing complexity i The maximum density ratio of the current value in the near-window distribution; Representing complexity i The maximum density ratio of the current value in the full window distribution; Complexity i Near-window difference symbols; Complexity i The full window difference symbol; This indicates that the gain coefficient is dynamically adjusted. Representing complexity i The dynamically adjusted weights; β ij Representing complexity i With complexity j Synergistic factors; Representing complexity i The final dynamic fusion weight value; c (i) Representing complexity i Gains are influenced by the synergistic effect of other complexities; C (i) These represent the initial calculated values of the four complexity metrics, where... i =1,2,3,4; These represent the final values of the four complexity metrics; C total This represents the final dynamic blending complexity of the current window.
[0244] In this embodiment, the calculation process for this step is as follows:
[0245] Dynamic adjustment factor: Frame 6 C degrad Maximum density ratio in near-window distribution ,and Therefore Adjustment factor:
[0246] ;
[0247] Final weights: After normalization (The same applies to other dimensions);
[0248] Synergistic gain and overall complexity: Synergistic factor β ij = 0.1, therefore:
[0249] ;
[0250] ;
[0251] Results Explanation: Overall Complexity of Frame 6 C total =0.3097, which is at a moderate level, mainly contributed by background suspicion and feature degradation.
[0252] The overall result in this embodiment is as follows: Figure 2-8 As shown, Figure 2 These are the actual data interference effect diagrams for frames 20 / 22 / 24; Figure 3 This refers to frames 1-61. C motion Time series variation curve; Figure 4 It is frames 1-61 C degrad Time series variation curve; Figure 5 It is frames 1-61 C bg Time series variation curve; Figure 6 It is frames 1-61 C dyn Time series variation curve; Figure 7 It is frames 1-61 C total Time series variation curve; Figure 8 This is a comparison curve of the temporal changes in complexity of various components from frame 1 to frame 61.
[0253] from Figure 2-8 The results show that the multi-dimensional complexity calculation results of this embodiment are highly consistent with the actual scene characteristics of infrared target tracking: Figure 3 middleC motion A clear peak appears in the target acceleration and turning frames (such as frames 15 and 55), accurately reflecting the changes in trajectory maneuverability; Figure 4 middle C degrad The feature degradation caused by occlusion is significantly higher than the threshold of α=0.5 in occluded segments (such as frames 20-24, 45, and 53), which accurately identifies the feature degradation caused by occlusion. Figure 5 middle C bg Overall, it remains in the range of 0.25-0.45, which is consistent with the actual situation of moderate background interference in ground scenes; Figure 6 middle C dyn The value is always 0, which aligns with the objective scenario that the sequence has no valid false targets; Figure 7-8 middle C total The changing trend and C degrad , C motion The key fluctuations are synchronized, and the values are high in occluded and highly maneuverable frames and low in stable tracking frames, proving that the method can truly characterize the complexity changes of infrared target tracking in complex ground scenes, providing a quantitative basis for tracking algorithm performance optimization and scene difficulty assessment.
[0254] Therefore, this invention integrates the degree of target feature degradation, the complexity of target motion patterns, and the intensity of interference from background and suspected targets to construct a dynamically generated recognition and tracking complexity scoring mechanism. This mechanism can reflect the tracking difficulty of the current frame or sequence in real time and objectively, thereby providing key theoretical basis and decision support for online switching of algorithm strategies, dynamic allocation of system resources, and scientific evaluation of cross-scene algorithm performance. This is conducive to promoting the paradigm shift of recognition and tracking systems from passive response to active perception and adaptive decision-making.
[0255] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for multi-dimensional complexity characterization of infrared target recognition in complex ground scenes, characterized in that, Includes the following steps: S1. Compare the current features of the target with the historical baseline features of the target to quantify feature degradation; Feature vector extraction is performed using the simplified modulus of the target's Zernike moments from order 0 to 4. S2. Extract the target motion feature vector, use two components, tangential complexity and normal complexity, to represent the two complexities in the target motion, and calculate the target motion complexity accordingly. S3. Extract the historical baseline features of the target and the local background features of the target to quantify their differences, and use the set suspicion degree conversion function to calculate the similarity between the historical features of the target and the local background features to characterize the suspicion degree of the local background. Specifically, the following steps are included: S31. Following the method shown in S1, construct the target baseline features based on historical data, and select the latest continuous historical data. K The target image of the frame does not contain any frames with a calculated degradation degree greater than [value missing]. α From the image frames, the target region is cropped according to the mask, the target feature vector is extracted, and scale normalization is performed to obtain the historical baseline feature vector f. base ; S32. Perform local background region scanning on the target, and extract local background features for each scanning window and perform scale normalization: ; ; ; ; in, W t , H t These represent the width and height of the target region in the current frame, respectively. W e , H e These represent the width and height of the current frame scanning window, respectively. x , y These represent the horizontal and vertical step sizes of the window scan, respectively; a p , b p ) indicates the first p The coordinates of the scan window indicate the horizontal scan position. a p Each step length, vertical b p Each step length; P Indicates the total number of scanned windows; p Indicates the scan window number; b p,1 , b p,2 ,…, b p,9 Indicates the first p The nine feature values of each window after scale normalization; b p No. p Feature vectors of each window; S33. Calculate the local background region likelihood, first calculating the current historical baseline vector f. base The differences are compared with the feature vectors of each window, and the differences are transformed into a likelihood level, then a weighted average is calculated based on the window coordinates: ; ; ; ; ; ; ; ; in, d e,p Indicates the first p Differences in energy characteristics between individual windows; d s,p Indicates the first p Differences in window shape characteristics; d p Indicates the first p Differences in features among individual windows; S ( d ) represents the suspicion conversion function; Me {•} represents the median function; MP This represents the set of highly suspected background window numbers; w p Indicates the first p The positional weight of each window; Indicates the first p The normalized position weights of each window; C bg Indicates the degree of suspicion of a local background; S4. Extract the feature vectors and motion vectors of real targets and false moving targets, quantify the feature differences and inter-frame motion differences between real targets and false moving targets, and use the doubt degree conversion to represent the degree of interference of false moving targets on target recognition and tracking; S5. Construct a dynamic weight fusion mechanism to integrate target feature degradation degree, target motion complexity, local background suspicion degree, and false target suspicion degree into a comprehensive recognition and tracking complexity.
2. The method for multi-dimensional complexity characterization of infrared target recognition in complex ground scenes according to claim 1, characterized in that, S1 specifically includes the following steps: S11. Using the segmented target image frames and mask data, construct historical baseline features and select the latest continuous historical features. K For each frame of the target image, frames with a degradation degree greater than α are removed. The target region is then cropped according to the mask. The 0th to 4th order moduli of the Zernike moments are calculated to generate a 9-dimensional target feature vector. The target shape features are then scaled and normalized. ; ; ; ; ; ; ; ; ; ; ; The above describes the Zernike moment eigenvector calculation method, where... M ( x , y ) represents the coordinates of the mask data ( x , y The value at () I ( x , y ) represents the coordinates of the image data ( x , y The value at () I target ( x , y ) represents the coordinates of the target region in the image. x , y The value at ); P target Represents image data; I norm ( x , y ) represents the coordinates in the image after grayscale normalization. x , y The value at () A This indicates the number of pixels in the target region within the currently calculated image data; c x Represents the x-coordinate of the target energy center; c y Represents the ordinate of the target energy center; R max Indicates the normalized maximum radius; x ′ represents the center-normalized x-coordinate; y ′ represents the center-normalized ordinate; r Indicates the pixel normalized radius; n This indicates the order of the Zernike moments; i Indicates the dimension index of the feature vector; V i ( r ) represents the Zernike radial polynomial; f i Indicates the first i The value of the eigenvector of order 1; S12, Select the latest historical continuous K After extracting features from the frame and performing scale normalization, we can obtain: ; ; ; in, k f represents the feature vector dimension index; k The 9-dimensional feature vector representing the target, where f k,1 Indicates the first k Frame image target energy features f k,1 ,…, f k,9 Indicates the first k The eight shape feature components of the target in the frame image; K This represents the latest consecutive historical frame count; A k Indicates the first k Total number of pixels in the frame target mask; Indicates the first k The first frame after target scale normalization i Each characteristic component; f base Indicates the historical baseline characteristics of the target; S13. Perform feature extraction on the target region of the current frame and perform scale normalization to obtain f. curr This indicates the current characteristics of the target; S14. Perform feature difference calculation: ; ; ; in, f base,1 Represents the historical baseline feature f of the target base The first component; f curr,1 Represents the historical baseline feature f of the target curr The first component; f base,i Represents the historical baseline feature f of the target base The i One component; f curr,i Represents the historical baseline feature f of the target curr The i One component; d e Indicates the degree of degradation of the target energy characteristic; d s γ represents the degradation degree of the target shape features; γ represents the weight of the target energy features. C degrad This indicates the degree of degradation of the target feature.
3. The method for multi-dimensional complexity characterization of infrared target recognition in complex ground scenes according to claim 2, characterized in that, S2 specifically includes the following steps: S21, Extract the latest K The degradation degree of each target feature is less than α Image frames, if this K If the image frames are consecutive, the number of pixels in the target region can be directly extracted based on the target mask data. A k and the target in continuous K Position sequence in the frame: ; ; ; ; ; ; in, K The total number of frames calculated; M k Indicates the first k Each image data corresponds to a target mask data; A k Indicates the first k The number of pixels in the target region in each image data; Indicates the target short-term average area; I k ( x , y ) indicates the first k In the image data ( x , y Gray value at ) c k,x Indicates the first k Target location in image data x coordinate; c k,y Indicates the first k Target location in image data y Coordinates; p k express k The target gray-scale weighted center coordinates at each time step; p0, p1, ..., p K-1 Representing 0 to K- At time 1, the target's gray-scale weighted center coordinates; conversely, if this... K If the image frames are not continuous, then after extracting the area coordinates as described above, linear interpolation is used to fill in the missing target locations based on the existing coordinate data, and the latest coordinates are extracted again according to the filled-in coordinates. The target location coordinates are p0, p1, ..., p1. K-1 ; S22. Using the displacement representation between adjacent frames, calculate the velocity vector between adjacent frames of the target: ; in, v k express k arrive k+1 The target displacement change at any given moment; S23. Use the rate change between adjacent frames to represent the tangential complexity of a single frame; ; ; ; in, S k Represented as velocity vector v k The European mold, S k+1 Represented as velocity vector v k+1 The Euclidean model reflects the target in [ k , k +1] The speed of movement within the time window; S k This represents the rate difference between two adjacent frames; T k express[ k , k +1] The square of the rate difference within the time window; S24. Calculate the change in motion direction between two adjacent frames: ; ; ; ; in, v k,y Represents the velocity vector v k of y Axial components; v k,x Represents the velocity vector v k of x Axial components; θ k express k The target's direction of motion angle at any given moment; θ k+1 express k+ The target's direction of motion angle at moment 1; θ k Indicates the rotation angle of the target between two adjacent frames; This indicates the average rate between two frames; N k express k Approximate target normal acceleration at any given time; S25. Combine time-memory weighted calculation of the overall target motion complexity: ; ; ; ; in, M = K -3 represents the total number of components to be calculated; w k This represents the time memory weight of the corresponding component. k =0 corresponds to the earliest time component. k = M -1 corresponds to the latest time component; T k express[ k , k +1] The square of the rate difference within the time window; C T Indicates the target's tangential acceleration; N k express k The target's normal acceleration at any given moment is approximate; C N This represents an approximate acceleration in the target direction; C motion This represents the complexity of the target motion.
4. The method for multi-dimensional complexity characterization of infrared target recognition in complex ground scenes according to claim 3, characterized in that, S4 specifically includes the following steps: S41. Based on historical data, construct target baseline features and select the latest continuous historical data. K The frames contain real and fake target images, excluding those with a calculated degradation degree greater than [value missing]. α The image frames are processed, and the target regions are cropped according to the corresponding masks of the two targets. The feature vectors of real and fake targets are extracted according to S3 and scaled and normalized to obtain the historical baseline feature vectors of real and fake targets. , Simultaneously, according to S2, the positions of the two targets are extracted based on their mask, and their trajectories are obtained. , , ; S42. Calculate the relative distance between the two targets in each frame based on their position trajectories: ; like Larger than the set target relative local range If so, it will not affect target tracking and will directly skip the process of calculating the likelihood of a false target; otherwise, if First, calculate the difference in motion between the real and false targets: ; ; ; ; ; ; ; ; in, L This indicates the relative local range of the set target; Indicates the true target number k The motion velocity vector of the frame; Indicates the false motion target. k The motion velocity vector of the frame; v k Indicate the true or false target number k Frame speed difference; d k Indicate the true or false target number k Frame scale-normalized relative distance; L This indicates the relative local area range of the set target; w s,k Indicates distance weight; This represents the distance weights after normalization; w t,k Indicates the weight of time memory; Indicates the joint weight; d m Indicates the difference in motion between real and false targets; S43. Calculate the difference in features between true and false targets according to S1: ; ; ; in, d base,e This indicates the difference in energy characteristics between true and false targets; d base,s Indicates the difference in shape features between real and false targets; d base Indicates the difference in features between true and false targets; S44. Calculate the likelihood of a target being suspected using the differences in features and motion between real and fake targets: ; ; in, d total Indicates the total difference between true and false objectives; C dyn Indicates the degree of suspicion of a fake movement target.
5. The method for multi-dimensional complexity characterization of infrared target recognition in complex ground scenes according to claim 4, characterized in that, S5 specifically includes the following steps: S51. Obtain the latest identified segmentation. Historical frame data and their corresponding complexity For each complexity metric, kernel density estimation is performed, with the bandwidth parameter employing the Silverman criterion. The maximum probability impression value is then calculated using the probability density function. Furthermore, an innovative time-window memory fusion method is used to analyze the latest historical data. Frames and The frame is calculated, and the memory-weighted fusion based on the window entropy is used to obtain the basic weights corresponding to this complexity: ; ; ; ; ; ; ; ; ; ; ; ; ; in, K This indicates that half of the latest historical frames have been selected. k Indicates the sequence number of the historical frame; Denotes the normalized function, where v min This represents the minimum value in vector v. v max This represents the maximum value in vector v; Denotes the normalization function, where v i Represents the first vector v i One portion, D This indicates the dimension of vector v; v (i) Representing complexity i The latest 2 K Frame data vector, where These represent four different levels of complexity. Representing complexity i The latest 1 to Frame data; Represents the complexity after standardization and normalization. i The latest Frame data vector; Represents the complexity after standardization and normalization. i The latest 1 to 2 K Frame data; h The bandwidth used for kernel density estimation is determined by the Silverman criterion. Represents the Gaussian kernel function; Representing complexity i The probability density function obtained by near-window kernel density estimation; Representing complexity i The probability density function obtained by full-window kernel density estimation; Represents the inverse normalized function; Represents the inverse normalization function; For complexity i The x-coordinate value at the point where the probability density near the window is maximized represents the complexity. i Impression value near the window; For complexity i The x-coordinate value at the point where the full-window probability density is maximized represents the complexity. i Full-window impression value; Representing complexity i Near-window data entropy; Representing complexity i Full-window data entropy; Representing complexity i Near-window weighting; Representing complexity i Full window weight; I (i) Representing complexity i The complexity impression value; Representing complexity i Dynamic fusion base weights; S52. Calculate the dynamic adjustment factor based on the kernel density fitted probability prediction and the actual probability, and perform collaborative complexity enhancement on the dynamically adjusted weights. The specific calculation method is as follows: ; ; ; ; ; ; ; ; ; ; in, Representing complexity i The maximum density ratio of the current value in the near-window distribution; Representing complexity i The maximum density ratio of the current value in the full window distribution; Complexity i Near-window difference symbols; Complexity i The full window difference symbol; This indicates that the gain coefficient is dynamically adjusted. Representing complexity i The dynamically adjusted weights; β ij Representing complexity i With complexity j Synergistic factors; Representing complexity i The final dynamic fusion weight value; c (i) Representing complexity i Gains are influenced by the synergistic effect of other complexities; C (i) These represent the initial calculated values of the four complexity metrics, where... i =1,2,3,4; These represent the final values of the four complexity metrics; C total This represents the final dynamic blending complexity of the current window.
Citation Information
Patent Citations
A visible light-infrared image registration method based on salient region features and edge degree
CN106447704A
Infrared moving small target detection method of a complex scene
CN109345472A