A method and system for processing anti-occlusion in tracking conditions
Through the triple filtering mechanism of multi-peak states, Kalman filtering and dynamic threshold judgment, the problems of multi-peak misjudgment and occlusion loss in target tracking in complex scenes are solved, and efficient tracking is achieved in scenes with similar target interference and occlusion.
Patent Information
- Application Number
- CN202511052796.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing target tracking technologies are prone to multi-peak states in complex scenes, and are lost or have low recapture success rates when the target is occluded, affecting the continuity and robustness of tracking.
It adopts a triple filtering mechanism of multi-peak state, collaborative calibration of Kalman filter and tracking model, dynamic threshold judgment, feature reinitialization after tracking loss, and search area expansion technology, and achieves accurate target identification and stable tracking through IOU calculation and confidence threshold adjustment.
In scenarios with similar target interference and occlusion, it can accurately identify and lock the tracking target, maintain the continuity and stability of tracking, and improve the success rate of recapture of the target after occlusion.
Smart Images

Figure CN120563567B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to an anti-occlusion processing method and system in a tracking situation. Background Art
[0002] In the field of target tracking, with the development of deep learning technology, deep learning-based tracking algorithms have been widely used in scenarios such as intelligent surveillance, autonomous driving, and robotic vision. Currently, mainstream tracking methods typically combine target detection models with motion prediction algorithms (such as Kalman filtering) to achieve continuous tracking by outputting the target position in real time. To improve the model's adaptability to dynamic scenes, researchers have introduced techniques such as attention mechanisms and temporal feature fusion, further optimizing the real-time performance and accuracy of tracking. This allows the tracking system to achieve high accuracy in ideal environments with no obstructions and minimal interference.
[0003] However, existing technologies still have obvious limitations in complex scenarios: when similar targets interfere with the tracking process, the confidence map output by the model is prone to form a multi-peak state, leading to misjudgment of the tracked target; when the target is partially or completely occluded, due to the lack of effective multi-peak screening and dynamic adjustment mechanisms, the tracking system is prone to lose the target or experience trajectory jumps; in addition, the recapture strategy after the target is lost often relies on fixed parameters and is difficult to adapt to the complex environment after the occlusion is lifted, resulting in a low recapture success rate, affecting the continuity and robustness of tracking. Summary of the Invention
[0004] The object of the present invention is to overcome one or more deficiencies of the prior art and to provide an anti-occlusion processing method and system in a tracking situation.
[0005] The object of the present invention is achieved through the following technical solutions:
[0006] A method for processing anti-occlusion in a tracking situation includes the following steps:
[0007] After the user selects and clicks the tracking target, the tracking model and Kalman filter are initialized with the location information of the tracking target;
[0008] In the tracking state of the first 5 frames, the inference results obtained by post-processing the model output are input into the Kalman filter and its state is updated; after 5 frames, the inference results of each frame and the Kalman filter prediction results of the previous frame are calculated to obtain the IOU result, and the high confidence threshold, low confidence threshold and IOU threshold are set at the same time; when the IOU result is higher than the IOU threshold, the tracking result is the inference result of each frame;
[0009] When the IOU result is lower than the IOU threshold, it is processed according to the relationship between the maximum value of the score map obtained by the model output and the high confidence threshold and the low confidence threshold;
[0010] If tracking is determined to be lost, the tracking state of the previous frame is obtained, and the model is re_initialized based on the target features of the previous frame. At the same time, the search area is double-upsampled for inference, and the above determination process is repeated to recapture the target. During the recapture process, the high confidence threshold and the low confidence threshold are temporarily increased by 10%. If it is still determined to be lost, the Kalman filter prediction result is used as the tracking result. If the state persists for 20 frames, the tracking state is exited.
[0011] Furthermore, when the IOU result is lower than the IOU threshold and the maximum value of the score map output by the model is higher than the high confidence threshold, the tracking result is the inference result of each frame.
[0012] Furthermore, when the IOU result is lower than the IOU threshold and the maximum value of the score map obtained by the model output is lower than the high confidence threshold and higher than the low confidence threshold, it is determined whether it is a multi-peak state. If it is not a multi-peak state, the minimum circumscribed rectangle of the inference result and the Kalman filter prediction result of the previous frame is used as the tracking result.
[0013] Furthermore, when it is determined to be a multi-peak state, the corresponding target frame coordinates of the main peak and the remaining secondary peaks are obtained respectively, and the IOU calculation is performed with the Kalman filter prediction results of the previous frame. At the same time, the mapping area of the Kalman filter prediction results on the score map is obtained, and the intersection area of each peak and the mapping area is obtained. The tracking score is obtained based on the IOU result of each peak and the intersection area. The calculation formula is as follows:
[0014] ;
[0015] Among them, m and n are the width and height of the intersection area respectively, A is the intersection area of the score map, and x and y are the upper left coordinates of the intersection area. is the iou result for each peak, and P is the peak confidence of the main peak.
[0016] Furthermore, a score threshold is set. When the highest tracking score is higher than the score threshold, the target frame coordinates with the highest tracking score are used as the tracking result. If there is no score higher than the score threshold, the tracking is judged to be lost.
[0017] Furthermore, when the IOU result is lower than the IOU threshold and the maximum value of the score map output by the model is lower than the low confidence threshold, tracking is judged to be lost.
[0018] Furthermore, the method for determining the multi-peak state includes the following steps:
[0019] On the score map, find peak points that are significantly higher than those in the neighborhood. When multiple peak points appear, obtain the peak confidence of the main peak and each secondary peak, and filter out secondary peaks that are less than 40% of the peak confidence of the main peak.
[0020] Calculate the peak distances between the remaining secondary peaks and the main peak, filter out the secondary peaks whose peak distances to the main peak are less than one-quarter of the score map size, and merge the secondary peaks whose peak distances to the main peak are less than one-eighth of the score map size;
[0021] Obtain the rectangular coordinates based on the main peak and each secondary peak as the diagonal line, and obtain the minimum value of the score map within the rectangular coordinates as the inter-peak valley depth. If the inter-peak valley depth is greater than 30% of the confidence level of the main peak, filter out this peak.
[0022] If secondary peaks still exist after three filtrations, it is determined to be a multi-peak state.
[0023] Furthermore, the formula for fusing sub-peaks is: ,in, is the main peak area, The main peak position, is the secondary peak area, It is the secondary peak position.
[0024] In some embodiments, a processing system for anti-occlusion in a tracking situation includes:
[0025] The initialization module is used to initialize the tracking model and Kalman filter with the position information of the tracking target after the user selects and clicks the tracking target;
[0026] The Kalman filter update module is used to input the inference results obtained by post-processing the model output into the Kalman filter under the tracking state of the first five frames and update its state;
[0027] The calculation and judgment module is used to calculate the IOU of each frame's inference result and the Kalman filter prediction result of the previous frame after 5 frames to obtain the IOU result, set the high confidence threshold, the low confidence threshold and the IOU threshold at the same time, and judge the tracking result based on the relationship between the IOU result and the maximum value of the score map and each threshold;
[0028] The tracking loss processing module is used to obtain the tracking status of the previous frame when it is determined that the tracking is lost, re_init the model based on the target features of the previous frame, and at the same time, perform two-fold upsampling on the search area for inference, and repeat the above determination process to recapture the target. During the recapture process, the high confidence threshold and the low confidence threshold are temporarily increased by 10%. If it is still determined to be lost, the Kalman filter prediction result is used as the tracking result. If the state persists for 20 frames, the tracking state is exited.
[0029] Furthermore, the calculation and judgment module further includes a multi-modal state judgment submodule for judging the multi-modal state.
[0030] The beneficial effects of the present invention are:
[0031] (1) Through the triple filtering mechanism of multi-peak states and the tracking score quantization selection technology, the effect of accurately identifying and locking the tracking target in the interference scene of similar targets is achieved;
[0032] (2) Through the coordinated calibration of the Kalman filter and the tracking model and the dynamic threshold judgment technology, the tracking continuity and stability are maintained when the target's motion state changes;
[0033] (3) Through feature re-initialization after tracking loss, search area expansion and threshold dynamic adjustment technology, the effect of efficiently recapture the target after occlusion and extend the effective tracking time is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A flowchart of the steps of a method for processing anti-occlusion in a tracking situation;
[0035] Figure 2 This is a schematic diagram of the tracking process of two white cars in an unobstructed scene;
[0036] Figure 3 Schematic diagram of the multi-modal tracking and recapture process of two white cars in an occluded scene. DETAILED DESCRIPTION
[0037] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.
[0038] Example 1:
[0039] See Figure 1 , provides an anti-occlusion processing method in tracking, comprising the following steps:
[0040] After the user selects and clicks the tracking target, the tracking model and Kalman filter are initialized with the location information of the tracking target;
[0041] In the tracking state of the first 5 frames, the inference results obtained by post-processing the model output are input into the Kalman filter and its state is updated; after 5 frames, the inference results of each frame and the Kalman filter prediction results of the previous frame are calculated to obtain the IOU result, and the high confidence threshold, low confidence threshold and IOU threshold are set at the same time; when the IOU result is higher than the IOU threshold, the tracking result is the inference result of each frame;
[0042] When the IOU result is lower than the IOU threshold, it is processed according to the relationship between the maximum value of the score map obtained by the model output and the high confidence threshold and the low confidence threshold;
[0043] If tracking is determined to be lost, the tracking state of the previous frame is obtained, and the model is re_initialized based on the target features of the previous frame. At the same time, the search area is double-upsampled for inference, and the above determination process is repeated to recapture the target. During the recapture process, the high confidence threshold and the low confidence threshold are temporarily increased by 10%. If it is still determined to be lost, the Kalman filter prediction result is used as the tracking result. If the state persists for 20 frames, the tracking state is exited.
[0044] When the IOU result is lower than the IOU threshold and the maximum value of the score map output by the model is higher than the high confidence threshold, the tracking result is the inference result of each frame.
[0045] When the IOU result is lower than the IOU threshold and the maximum value of the score map obtained by the model output is lower than the high confidence threshold and higher than the low confidence threshold, it is determined whether it is a multi-peak state. If it is not a multi-peak state, the minimum circumscribed rectangle of the inference result and the Kalman filter prediction result of the previous frame is used as the tracking result.
[0046] When it is determined to be a multi-peak state, the corresponding target frame coordinates of the main peak and the remaining secondary peaks are obtained respectively, and the IOU calculation is performed with the Kalman filter prediction results of the previous frame. At the same time, the mapping area of the Kalman filter prediction results on the score map is obtained, and the intersection area of each peak and the mapping area is obtained. The tracking score is obtained based on the IOU result of each peak and the intersection area. The calculation formula is as follows:
[0047] ;
[0048] Among them, m and n are the width and height of the intersection area respectively, A is the intersection area of the score map, and x and y are the upper left coordinates of the intersection area. is the iou result for each peak, and P is the peak confidence of the main peak.
[0049] Set a score threshold. When the highest tracking score is higher than the score threshold, the target frame coordinates with the highest tracking score are used as the tracking result. If there is no target frame coordinate higher than the score threshold, the tracking is considered lost.
[0050] When the IOU result is lower than the IOU threshold and the maximum value of the score map output by the model is lower than the low confidence threshold, tracking is judged to be lost.
[0051] The method for determining the multi-peak state includes the following steps:
[0052] On the score map, find peak points that are significantly higher than those in the neighborhood. When multiple peak points appear, obtain the peak confidence of the main peak and each secondary peak, and filter out secondary peaks that are less than 40% of the peak confidence of the main peak.
[0053] Calculate the peak distances between the remaining secondary peaks and the main peak, filter out the secondary peaks whose peak distances to the main peak are less than one-quarter of the score map size, and merge the secondary peaks whose peak distances to the main peak are less than one-eighth of the score map size;
[0054] Obtain the rectangular coordinates based on the main peak and each secondary peak as the diagonal line, and obtain the minimum value of the score map within the rectangular coordinates as the inter-peak valley depth. If the inter-peak valley depth is greater than 30% of the confidence level of the main peak, filter out this peak.
[0055] If secondary peaks still exist after three filtrations, it is determined to be a multi-peak state.
[0056] The formula for fusing sub-peaks is: ,in, is the main peak area, The main peak position, is the secondary peak area, It is the secondary peak position.
[0057] A system for processing anti-occlusion in a tracking situation is provided, which is used to execute a method for processing anti-occlusion in a tracking situation, including the following steps:
[0058] The initialization module is used to initialize the tracking model and Kalman filter with the position information of the tracking target after the user selects and clicks the tracking target;
[0059] The Kalman filter update module is used to input the inference results obtained by post-processing the model output into the Kalman filter under the tracking state of the first five frames and update its state;
[0060] The calculation and judgment module is used to calculate the IOU of each frame's inference result and the Kalman filter prediction result of the previous frame after 5 frames to obtain the IOU result, set the high confidence threshold, the low confidence threshold and the IOU threshold at the same time, and judge the tracking result based on the relationship between the IOU result and the maximum value of the score map and each threshold;
[0061] The tracking loss processing module is used to obtain the tracking status of the previous frame when it is determined that the tracking is lost, re_init the model based on the target features of the previous frame, and at the same time, perform two-fold upsampling on the search area for inference, and repeat the above determination process to recapture the target. During the recapture process, the high confidence threshold and the low confidence threshold are temporarily increased by 10%. If it is still determined to be lost, the Kalman filter prediction result is used as the tracking result. If the state persists for 20 frames, the tracking state is exited.
[0062] The calculation and judgment module further includes a multi-modal state judgment submodule, which is used to execute the step of judging the multi-modal state according to claim 7.
[0063] Example 2: Parallel tracking process of two white cars in the upper right area of the left image without any obstruction:
[0064] See Figure 2 The upper right image on the left shows a vehicle tracking scenario within the monitoring area: Two white cars (denoted as target A and interfering vehicle B) are traveling side by side in the same lane, with a stable spacing and no mutual obstruction. There are no other large vehicles interfering with the vehicle. Target A is the target for tracking. Interfering vehicle B is similar in model and color to target A, which may cause potential interference signals in the confidence map output by the model, but it does not form a multi-peak state. Detailed implementation steps:
[0065] Target initialization and system startup. After the monitoring system is started, the user clicks on target A (the white sedan on the left in the upper right area of the left image) through the interactive interface. The system immediately triggers the initialization module: extracting the real-time position information of target A (marked by a rectangular box in the image, containing the coordinates of the upper left and lower right corners), and synchronously inputting these coordinates into the tracking model and Kalman filter. Tracking model initialization: Loading pre-trained vehicle detection weights, setting the resolution of the input image and the region of interest (ROI) for target detection, limiting the detection range to the upper right area of the left image to reduce interference from irrelevant areas. Kalman filter initialization: Based on the initial position of target A, establish a motion state model (the state vector contains position coordinates, horizontal and vertical velocities), initialize the process noise covariance matrix and the measurement noise covariance matrix, and set the initial prediction step size.
[0066] Kalman filter calibration for the first five frames: Within frames 1 to 5, the system executes the Kalman filter update module. In each frame, the tracking model infers the upper right region of the left image and outputs a detection bounding box (inference result) for target A. After post-processing to remove duplicate frames, the result is input into the Kalman filter. The Kalman filter uses a measurement update step to fuse the inference results with the prediction results, dynamically adjusting the state vector and covariance matrix. For example, after the first frame is input, the filter corrects the initial velocity parameters; after the third frame is input, the predicted deviation of the motion direction is optimized. By the end of the fifth frame, the filter parameters, such as the velocity prediction error and position measurement error, are calibrated to form a stable motion state prediction model.
[0067] In the sixth frame, the IOU exceeds the threshold, and target A and jammer B maintain a stable distance, with no sudden changes in relative position. The calculation module calculates the IOU (intersection-over-union) between the inference result of the current frame (target A's detection bounding box) and the Kalman filter prediction result of the previous frame. The result exceeds the preset IOU threshold. The inference result of the current frame is directly used as the tracking output and fed back to the Kalman filter to update its state vector (correcting the position and velocity parameters). At this point, the detection bounding box of jammer B is automatically filtered out by the model because its confidence level is lower than that of target A and no valid peak is formed.
[0068] In the 10th frame, the IOU is below the threshold but the confidence is high. Due to a slight lane deviation, the inference result and the Kalman filter prediction from the previous frame have an IOU below the threshold. The model outputs the score map, and the peak confidence (maximum value) corresponding to Target A is above the high confidence threshold (indicating that the model recognizes Target A with high reliability). Although the IOU is below the threshold, the high confidence feature is clear, so the system directly uses the inference result of the current frame as the tracking result and updates the Kalman filter state. During this process, although the score map peak of the interfering vehicle B exists, it does not affect the tracking of Target A because it does not reach the interference threshold.
[0069] In frame 15, the IOU is below the threshold and the confidence level is between the high and low thresholds, indicating a non-multimodal scenario. Target A and interfering vehicle B simultaneously change direction slightly, causing the IOU between the inference result and the Kalman prediction of Target A to further decrease. The score map maximum value is between the high and low confidence thresholds (the detection confidence level decreases slightly due to the change in direction). The system activates the multimodal state judgment logic and analyzes the score map. The peak confidence level of interfering vehicle B is less than 40% of that of Target A. After the first filtering step, it is removed and determined to be non-multimodal. The system then uses the minimum bounding rectangle of the current frame's inference result (Target A's detection bounding box) and the previous frame's Kalman prediction result as the tracking result. This preserves the model's real-time position information while also using the motion continuity constraints of the Kalman prediction to avoid tracking box jumps caused by direction changes.
[0070] During the 20th to 50th frames of continuous tracking, target A and jammer vehicle B remained parallel and unobstructed, and the system entered the stable tracking phase. In most frames, the inference result and the Kalman filter's prediction of the inter-connection-of-union (IOU) exceeded the threshold, and the inference result was directly adopted and the filter state updated. In a few frames, target A experienced slight jitter due to road bumps, and the IOU briefly fell below the threshold. However, the maximum score map value remained above the high-confidence threshold, maintaining tracking continuity. During tracking, the Kalman filter continuously optimized the prediction model based on the target's motion trends (e.g., constant speed and small turns), ensuring a smooth output trajectory with no noticeable drift.
[0071] Final processing of the target leaving the area When target A approaches the edge of the upper right area in the left figure, the system still performs tracking according to the above logic until the detection box of target A completely exceeds the monitoring area. The tracking is automatically terminated, the Kalman filter state is cleared, and it waits for the next target initialization.
[0072] This example fully encompasses initialization, calibration, and tracking logic, validating the effectiveness of high-confidence priority determination, non-multimodal state processing, and system module collaboration. The initialization module and Kalman filter update module collaborate to ensure the accuracy of the tracking starting point. The calculation and judgment module, through dynamic thresholds for IOU and confidence, stably locks onto the target despite interference from similar vehicles, demonstrating the reliability of this method's tracking logic.
[0073] Example 3: Two white cars in the upper right area of the left image block the multi-peak tracking process: For this example, refer to Figure 3 The upper right corner of the left image shows a complex scene within the monitoring area: the target is a white car on the right side of the area (target A). A white car on the left (interference vehicle B) gradually approaches and partially occludes target A (the occlusion area increases from 10% to 50%), causing the score map output by the tracking model to have multiple peaks (target A corresponds to the primary peak, and interference vehicle B corresponds to the secondary peak). The core process verifies the multi-peak processing and loss-recapture mechanism. Detailed implementation steps:
[0074] Initialization and calibration for the first five frames follow the same logic as in Example 2: After the user selects Target A, the initialization module inputs its position information into the tracking model and Kalman filter. For the first five frames, the inference results for each frame (a clear detection frame when Target A is unobstructed) are continuously fed into the Kalman filter to calibrate the motion state model (e.g., determining Target A's average speed and linear motion preference).
[0075] At the beginning of occlusion in frame 8, a multimodal state is triggered. Interference vehicle B begins to approach target A, creating a 10% occlusion. Two significant peaks appear in the model's output score map. The system triggers the multimodal state judgment submodule of the calculation and judgment module, which performs a three-step filtering method: 1) Confidence filtering: This extracts the primary peak (target A, confidence level C1) and the secondary peak (interference vehicle B, confidence level C2). Because C2 is greater than 40% of C1 (C2 = 0.6 × C1), the secondary peak is retained. 2) Distance filtering: This calculates the pixel spacing (peak-to-peak distance) between the two peaks. If the result is greater than 1 / 4 of the score map size (to avoid interference from closely spaced peaks) and less than 1 / 8 (no fusion required), the secondary peak is retained. 3) Inter-peak valley depth filtering: This generates a rectangular region with the two peaks as the diagonal line and extracts the minimum score map value (valley depth D) within the region. Because D is less than 30% of C1 (D = 0.25 × C1), the secondary peak is not filtered out. After three filtering steps, the secondary peak still exists, and the system determines a multimodal state.
[0076] In the 10th frame, the multi-peak processing tracking score is calculated. The occlusion area increases to 30%, and the multi-peak state continues. The system performs multi-peak decision-making: the target frame coordinates of the main peak (target A) and the secondary peak (interference vehicle B) are extracted, and the IOU is calculated with the Kalman prediction results of the previous frame (s iou 1.s iou 2). Determine the mapping area of the Kalman prediction result on the score map (the possible location of the target based on the motion state prediction) and extract the intersection area between the two peaks and the mapping area: the main peak intersection area is 20×15 pixels (m=20, n=15), and the secondary peak intersection area is 10×8 pixels (m=10, n=8). Calculate the tracking score according to the formula: Main peak tracking score = (s iou 1 / (20×15))×ΣΣ(A1(x+i,y+j) / C1), where A1 is the confidence value in the intersection area, and the average confidence is obtained by cumulative calculation; the secondary peak tracking score = (s iou 2 / (10×8))×ΣΣ(A2(x+i,y+j) / C1), the calculation logic is the same as above. Set the score threshold. If the main peak tracking score is higher than the threshold and the secondary peak score is lower than the threshold, the system selects the target box corresponding to the main peak as the tracking result.
[0077] In the 15th frame, deep occlusion causes a drop in score map confidence, with the occlusion reaching 50%. The detection confidence of target A decreases significantly: the IoU between the inference result and the Kalman prediction falls below the threshold, and the maximum score map value falls between the high and low confidence thresholds. Multi-peak state determination shows that the secondary peak still exists (using the same filtering conditions as before), but the tracking score of the primary peak decreases slightly due to occlusion (still above the threshold), and the system continues to track the primary peak. Simultaneously, the Kalman filter performs an auxiliary prediction of target A's position based on historical motion states, ensuring that the tracking box always covers the visible portion of target A.
[0078] In the 18th frame, tracking loss is triggered due to complete occlusion. Interference vehicle B completely occludes target A, and the maximum scoremap output by the model falls below the low confidence threshold. The system determines tracking loss, triggering the tracking loss processing module: The re_init operation is executed. Based on the features of target A (such as outline and local texture) in the 17th frame (the last visible frame), the target feature library of the tracking model is reinitialized, strengthening the recognition weight of the unique features of target A. The search area is expanded: the original detection area is upsampled by a factor of two, expanding the monitoring range to cover the area where target A may appear (based on the Kalman predicted motion trajectory). The threshold is temporarily adjusted: the high and low confidence thresholds are increased by 10%, for example, the original high threshold of 0.8 is increased to 0.88, to reduce false detections during recapture and avoid misidentifying interfering vehicle B as target A.
[0079] In the 22nd frame, target A is successfully recaptured, exiting the occluded area and partially becoming visible again. The tracking model detects target A within the expanded search area, and the maximum score map value exceeds the increased high confidence threshold. The IoU (interference under union) between this detection bounding box and the Kalman filter prediction from the previous frame is calculated. If the result exceeds the threshold, the system's recapture logic resumes tracking and outputs the inference result for the current frame. Simultaneously, the high and low confidence thresholds are adjusted back to their original values, restoring the search area to its normal size. The Kalman filter updates its state based on the newly input inference result, correcting for prediction errors caused by occlusion.
[0080] Subsequent tracking stabilizes and resumes between frames 25 and 60, with Target A completely clear of occlusion and stable driving. The IOU between the inference result and the Kalman prediction remains above the threshold, and the system adopts the inference result and updates the filter state. Multimodal state analysis indicates no valid secondary peaks (interfering vehicle B has moved away), and tracking reverts to non-multimodal logic until Target A exits the monitoring area in the upper right corner of the left image.
[0081] This embodiment focuses on verifying the effectiveness of the three-step method for multi-peak judgment, tracking score quantification decision, loss and recapture mechanism, and multi-peak sub-module functions. In the multi-peak state, triple filtering is used to ensure that valid secondary peaks are not mistakenly deleted, and the tracking score is quantified by fusing IOU and confidence. After loss, the strategy of dynamically adjusting the search range and threshold significantly improves the recapture success rate. In summary, both embodiments are based on the scene of two white cars in the upper right area of the left figure, and the tracking logic in unobstructed and obstructed multi-peak scenes is presented in detail. Through the detailed description of the initialization calibration, dynamic threshold judgment, multi-peak filtering, loss and recapture, etc., the anti-interference ability of the present invention in similar vehicle interference and occlusion scenarios is fully verified, and the practicality and reliability are fully reflected.
[0082] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.
Claims
1. A method for processing anti-occlusion in tracking, characterized in that: The following steps are involved: After the user selects and clicks the tracking target, the tracking model and Kalman filter are initialized with the location information of the tracking target; In the tracking state of the first five frames, the inference results obtained by post-processing the model output are input into the Kalman filter and its state is updated; After 5 frames, the inference result of each frame and the Kalman filter prediction result of the previous frame are used to calculate the IOU result, and the high confidence threshold, low confidence threshold and IOU threshold are set at the same time; when the IOU result is higher than the IOU threshold, the tracking result is the inference result of each frame; When the IOU result is lower than the IOU threshold, it is processed according to the relationship between the maximum value of the score map obtained by the model output and the high confidence threshold and the low confidence threshold; If tracking is determined to be lost, the tracking state of the previous frame is obtained, and the model is re_initialized based on the target features of the previous frame. At the same time, the search area is double-upsampled for inference, and the above determination process is repeated to recapture the target. During the recapture process, the high confidence threshold and the low confidence threshold are temporarily increased by 10%. If it is still determined to be lost, the Kalman filter prediction result is used as the tracking result. If the state persists for 20 frames, the tracking state is exited.
2. The method according to claim 1, characterized in that When the IOU result is lower than the IOU threshold and the maximum value of the score map output by the model is higher than the high confidence threshold, the tracking result is the inference result of each frame.
3. The method according to claim 1, characterized in that When the IOU result is lower than the IOU threshold and the maximum value of the score map obtained by the model output is lower than the high confidence threshold and higher than the low confidence threshold, it is determined whether it is a multi-peak state. If it is not a multi-peak state, the minimum circumscribed rectangle of the inference result and the Kalman filter prediction result of the previous frame is used as the tracking result.
4. The method according to claim 3, characterized in that When it is determined to be a multi-peak state, the corresponding target frame coordinates of the main peak and the remaining secondary peaks are obtained respectively, and the IOU calculation is performed with the Kalman filter prediction results of the previous frame. At the same time, the mapping area of the Kalman filter prediction results on the score map is obtained, and the intersection area of each peak and the mapping area is obtained. The tracking score is obtained based on the IOU result of each peak and the intersection area. The calculation formula is as follows: ; Among them, m and n are the width and height of the intersection area respectively, A is the intersection area of the score map, and x and y are the upper left coordinates of the intersection area. is the iou result for each peak, and P is the peak confidence of the main peak.
5. The method according to claim 4, characterized in that Set a score threshold. When the highest tracking score is higher than the score threshold, the target frame coordinates with the highest tracking score are used as the tracking result. If there is no target frame coordinate higher than the score threshold, the tracking is considered lost.
6. The method according to claim 1, characterized in that When the IOU result is lower than the IOU threshold and the maximum value of the score map output by the model is lower than the low confidence threshold, tracking is judged to be lost.
7. The method according to claim 1, characterized in that The method for determining the multi-peak state includes the following steps: On the score map, find peak points that are significantly higher than those in the neighborhood. When multiple peak points appear, obtain the peak confidence of the main peak and each secondary peak, and filter out secondary peaks that are less than 40% of the peak confidence of the main peak. Calculate the peak distances between the remaining secondary peaks and the main peak, filter out the secondary peaks whose peak distances to the main peak are less than one-quarter of the score map size, and merge the secondary peaks whose peak distances to the main peak are less than one-eighth of the score map size; Obtain the rectangular coordinates based on the main peak and each secondary peak as the diagonal line, and obtain the minimum value of the score map within the rectangular coordinates as the inter-peak valley depth. If the inter-peak valley depth is greater than 30% of the confidence level of the main peak, filter out this peak. If secondary peaks still exist after three filtrations, it is determined to be a multi-peak state.
8. The method according to claim 7, characterized in that The formula for fusing sub-peaks is: ,in, is the main peak area, The main peak position, is the secondary peak area, It is the secondary peak position.
9. A processing system for anti-occlusion in tracking, characterized in that: include: The initialization module is used to initialize the tracking model and Kalman filter with the position information of the tracking target after the user selects and clicks the tracking target; The Kalman filter update module is used to input the inference results obtained by post-processing the model output into the Kalman filter under the tracking state of the first five frames and update its state; The calculation and judgment module is used to calculate the IOU of each frame's inference result and the Kalman filter prediction result of the previous frame after 5 frames to obtain the IOU result, set the high confidence threshold, the low confidence threshold and the IOU threshold at the same time, and judge the tracking result based on the relationship between the IOU result and the maximum value of the score map and each threshold; The tracking loss processing module is used to obtain the tracking status of the previous frame when it is determined that the tracking is lost, re_init the model based on the target features of the previous frame, and at the same time, perform two-fold upsampling on the search area for inference, and repeat the above determination process to recapture the target. During the recapture process, the high confidence threshold and the low confidence threshold are temporarily increased by 10%. If it is still determined to be lost, the Kalman filter prediction result is used as the tracking result. If the state persists for 20 frames, the tracking state is exited.
10. The system according to claim 9, characterized in that The calculation and judgment module further includes a multi-modal state judgment submodule for judging the multi-modal state.
Citation Information
Patent Citations
Moving target anti-shielding re-tracking method based on correlation filtering
CN118537365A
Object detection device and object detection method
US20200097756A1