Ship multi-target tracking method based on online spatial-temporal feature association
By adopting the online spatiotemporal feature association method in ship multi-object tracking, combined with enhanced observation Kalman filtering and spatiotemporal feature model, the problems of tracking accuracy and robustness in complex marine environments are solved, and a more efficient multi-objective tracking effect is achieved.
Patent Information
- Application Number
- CN202510141223.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
AI Technical Summary
In complex marine environments, multi-target tracking of ships faces challenges such as occlusion, nonlinear motion, observation platform shaking and diverse appearance characteristics, resulting in insufficient tracking accuracy and robustness.
A multi-objective tracking method based on online spatiotemporal feature association is adopted. By enhancing the combination of the observation Kalman filter module and spatiotemporal feature model, the target's spatiotemporal feature response is detected, and a similarity matching matrix is constructed. The Hungarian algorithm is optimized to achieve online adjustment and improve tracking accuracy.
It effectively solves problems such as occlusion, nonlinear motion and platform shaking, significantly improves the accuracy and robustness of multi-objective tracking, and provides more accurate and reliable ship monitoring support.
Smart Images

Figure CN120070502A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of computer vision target tracking, and in particular relates to a ship multi-target tracking method based on online spatiotemporal feature association. Background Art
[0002] With the rapid development of artificial intelligence and computer vision technology, multi-target tracking has become one of the key technologies in the fields of intelligent monitoring, unmanned driving and marine traffic management. In the field of ship monitoring, multi-target tracking technology is widely used in ship traffic safety monitoring, channel planning, illegal intrusion detection and other scenarios, providing important support for improving marine safety and traffic efficiency.
[0003] In a complex marine environment, multi-target tracking of ships faces the following major challenges: 1. Occlusion problem: Ships may be occluded when approaching or intersecting, resulting in loss of detection frames or mismatching. 2. Non-linear motion: Under the influence of wind, waves and currents, the motion trajectory of ships often has strong nonlinear characteristics, which makes it difficult to predict the target position. 3. Shaking of the observation platform: UAVs, surveillance cameras or other observation platforms may produce large movements in the marine environment, further exacerbating the difficulty of tracking the target. 4. Diverse appearance features: There are large differences in the appearance features of ships such as size, shape, and color, which significantly increases the difficulty of target detection and feature association.
[0004] In order to solve the above problems, multi-target tracking methods based on spatiotemporal feature association have gradually become a research hotspot in recent years. By combining the spatial position characteristics and temporal motion characteristics of the target, this type of method can not only effectively deal with the target occlusion problem, but also capture the dynamic behavior of the target, significantly improving the accuracy and robustness of multi-target tracking. Summary of the invention
[0005] Purpose of the invention: The present invention proposes a ship multi-target tracking method based on online spatiotemporal feature association. First, an enhanced observation Kalman filter module is developed and a spatiotemporal feature model is established. Then, the spatiotemporal feature model is used to detect the characteristic response of the target and a similarity matching matrix is constructed. Finally, the Hungarian algorithm is used to optimize the similarity matching matrix to achieve online adjustment and improve tracking accuracy. The ship multi-target tracking method based on online spatiotemporal feature association can not only process complex marine environments in real time, but also effectively solve problems such as occlusion, nonlinear motion and platform shaking in traditional methods, providing more accurate and reliable technical support for ship monitoring systems.
[0006] Technical solution: The present invention proposes a multi-target tracking method based on online spatiotemporal feature association, comprising the following steps:
[0007] Step 1: Obtain ship images, preprocess the initial images, and obtain the original data set;
[0008] Step 2: Perform object detection on the training samples and calibrate the time coordinates and spatial coordinates of the detection results;
[0009] Step 3: Eliminate the detection results with confidence lower than the set threshold, and introduce Occlusion-Aware Non-Maximum Suppression (ONMS) to retain more occluded detections, increase the chance of association with the trajectory, and indirectly reduce the confusion caused by missed detections;
[0010] Step 4: Develop an enhanced observation Kalman filter module to predict the position of the target in the current frame, establish a time feature model and a spatial feature model of the target respectively, fuse the time feature model and the spatial feature model of the tracked target to obtain the spatio-temporal feature model of the target, acquire the time feature and spatial feature of the target, and fuse them to generate spatio-temporal features for characterizing the dynamic behavior of the target;
[0011] Step 5: In the current frame, use the spatio-temporal feature model to detect the spatio-temporal feature response of the target, associate the spatio-temporal feature response with the spatio-temporal object feature model of the tracked object, and fuse them to obtain a similarity matching matrix;
[0012] Step 6: Use the Hungarian algorithm to optimize the similarity matching matrix, determine the best match between the current frame target and the historical trajectory, and update the parameters of the spatio-temporal feature model of the target to achieve online adjustment and improve the tracking accuracy.
[0013] Furthermore, the specific method in Step 1 is as follows:
[0014] Step 1.1: Obtain ship images from multiple angles and different scenarios to ensure the diversity and representativeness of the data;
[0015] Step 1.2: Manually annotate the collected ship images, marking the bounding boxes and key feature points of the ships;
[0016] Step 1.3: Perform preprocessing on the images such as size adjustment, denoising, brightness and contrast optimization, and remove low-quality and duplicate images. Finally, construct a training set, a validation set, and a test set.
[0017] Furthermore, the specific method in Step 2 is as follows:
[0018] Step 2.1: Use Repvit as the backbone network. First, adopt a single and deeper downsampling layer to increase the network depth and reduce the information loss caused by the resolution reduction; secondly, use a simple classifier composed of a global average pooling layer and a linear layer; finally, select a stage ratio of 1:1:7:1, and then increase the network depth to 2:2:14:2; use SlideLoss as the loss function, and adaptively learn the positive sample threshold parameter and negative sample threshold parameter μ. The calculation formula is as follows:
[0019]
[0020] Among them, x represents an unknown number and is used to limit the μ interval.
[0021] Step 2.2: Train the constructed model to obtain an object detection model through multiple rounds of optimization.
[0022] Step 2.3: Perform object detection on the ship image through the model and calibrate the time coordinates and space coordinates of the detection results to ensure the consistency of the position and state of each object in the time and space dimensions.
[0023] Furthermore, the specific method in step 3 is as follows:
[0024] Step 3.1: Set a confidence threshold and delete all images with a confidence lower than the threshold in the detection results to ensure the accuracy of subsequent multi-object tracking.
[0025] Step 3.2: Introduce occlusion-aware non-maximum suppression ONMS to retain more occluded regions. The formula is as follows:
[0026] S 1 = S·(1 - α·O i,j )
[0027] Where: S 1 is the adjusted confidence score, S is the original confidence score, α is the occlusion adjustment factor that controls the weight of the occlusion effect and is between [0, 1], and O i,j is the occlusion degree between box i and box j, represented by IOU(A, B).
[0028] Furthermore, the specific method in step 4 is as follows:
[0029] Step 4.1: Construct an enhanced observation Kalman filter module.
[0030] First, when establishing the association between adjacent frames, design a Fuse-distance metric function that fuses IOU distance, object category, and size information to determine the optimal match between the detection box and the Kalman filter prediction box. Second, introduce a Gaussian cascade matching module to solve the problem of frequent ID switching caused by the violent shaking of the ship.
[0031] The formula of the Fuse-distance metric function is as follows:
[0032] D Fuse = ω 1 ·D IOU + ω 2 ·D class + ω 3 ·Dsize
[0033] Among them: D Fuse represents the comprehensive metric value of Fuse - distance, D IOU represents the Intersection over Union (IOU) distance, which is used to measure the overlapping degree of two bounding boxes, D class represents the category difference metric, which is used to judge whether two bounding boxes belong to the same category, D size represents the size difference metric, which is used to compare the area difference of two bounding boxes, ω 1 , ω 2 , ω 3 represents the weight coefficient, which is used to adjust the contribution of each feature to the final metric value;
[0034] Step 4.2: Construct the time - feature model M T ; For the time feature, it is represented by the Histogram of Optical Flow (HOF), which is used to describe the magnitude and direction of the motion speed of the tracked object;
[0035] Step 4.3: Construct the space - feature model M s ; For the space feature, according to the rectangular image of the tracked object in the training samples, the discriminative projection vector of the target region is obtained through the Incremental Linear Discriminant Analysis (ILDA) algorithm, and then the vector is weighted with the color histogram of the region to obtain the space feature of the target;
[0036] Step 4.4: Fuse the space - feature model M s and the time - feature model M T to obtain the spatio - temporal feature model M of the target. The space - feature model M s is constructed by ILDA and color histogram, and the time - feature model M T describes the motion speed and direction through the Histogram of Optical Flow. The spatio - temporal feature model M generated by fusing the two is used to characterize the dynamic behavior of the target.
[0037] Furthermore, the specific method in step 5 is as follows:
[0038] Step 5.1: In the data association stage, adopt the spatio - temporal data association framework; First, online detect the feature response of the tracked object in the current frame and extract the spatio - temporal features;
[0039] Step 5.2: Calculate the similarity of the space and time features respectively;
[0040] The calculation of the space - feature similarity is based on the cosine similarity of the feature vectors, which measures the consistency of the detected bounding box and the target bounding box in terms of space features. The formula is:
[0041]
[0042] Among them, For detecting responses and tracking target T i of the time feature similarity For the detection response of the spatial feature vector For the tracking target T i of the spatial feature vector respectively represent and the vector norms of;
[0043] The time feature similarity calculation is based on the Gaussian distribution, measuring the similarity of the detection box and the target box in motion features. The formula is:
[0044]
[0045] where For the detection response and tracking target T i of the time feature similarity, δ is the standard deviation of the Gaussian function, used to control the width of the time feature distribution, For the detection response and tracking target T i of the time feature distance;
[0046] Step 5.3: Fuse the spatial feature similarity and the time feature similarity into a comprehensive similarity measure. The formula is S combined is the fused comprehensive similarity, ε 1 and ε 2 are weight coefficients, adjusting the influence of spatial and time features on the matching;
[0047] Step 5.4: Construct a similarity matching matrix, and fill the comprehensive similarity between all detection responses and the target feature model into the matrix:
[0048]
[0049] G is the similarity matching matrix, is the m-th detection response, T n is the n-th target.
[0050] Furthermore, the specific method in step 6 is:
[0051] Step 6.1: Based on the constructed similarity matching matrix M, the Hungarian algorithm finds the best match between the detection response and the historical trajectory by maximizing the similarity or minimizing the cost; each match considers both spatial and time features simultaneously, and then compares the optimal match result with the set threshold. If the match result is higher than the threshold, it indicates a reliable association. If it is lower than the threshold, it means the target is lost or there is a mis-match;
[0052] Step 6.2: For the successfully matched targets, update their spatio-temporal feature model parameters, including spatial feature vectors and temporal feature vectors. The parameter update is based on the latest feature values of the detection responses, enabling the model to dynamically adapt to the changes of the targets.
[0053] Beneficial effects:
[0054] 1. By combining the spatial feature model and the temporal feature model of the target, the present invention constructs a spatio-temporal feature model to effectively capture the dynamic behavior of the target. Using the enhanced observation Kalman filter module to predict the target position and combining temporal feature modeling methods such as the histogram of optical flow, the accuracy of multi-target tracking is significantly improved. Especially in complex environments, the tracking effect on non-linear moving targets is better than traditional methods.
[0055] 2. Adopting an occlusion perception mechanism, by fusing multi-dimensional features (such as target appearance, size, motion trajectory, etc.), it can still maintain a high detection and tracking accuracy rate when the target is partially or completely occluded, thus avoiding target loss and false matching.
[0056] 3. The present invention introduces a Gaussian cascade matching module and constructs a Fuse-distance metric function between adjacent frames to optimize the matching process between the detection box and the prediction box, significantly reducing the frequent ID switching phenomenon caused by the shaking of the ship observation platform or the rapid movement of the target.
[0057] 4. This method adopts an online association optimization technology. By calculating the similarity matching matrix in real time and using the Hungarian algorithm for optimization, it can quickly complete target matching in a dynamic scene, meeting the high real-time requirements of the ship monitoring scene. Description of the drawings
[0058] Figure 1 is the overall flowchart of the present invention;
[0059] Figure 2 is the specific algorithm flowchart of the present invention;
[0060] Figure 3 is the schematic diagram of the Repvit structure;
[0061] Figure 4 is the ship tracking effect diagram before algorithm improvement;
[0062] Figure 5 is the ship tracking effect diagram after algorithm improvement. Detailed implementation manners
[0063] The present invention will be further clarified below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, those skilled in the art's various equivalent modifications of the present invention all fall within the scope defined by the appended claims of this application.
[0064] The present invention discloses a method for multi-target tracking of ships based on online spatio-temporal feature association. First, a dataset image is obtained and object detection is performed on the image samples. The coordinates in the time dimension and the coordinates in the space dimension of the detection results are calibrated, the detection results with a confidence level less than a set value are eliminated, and occlusion-aware non-maximum suppression is introduced to retain more occluded detections, thereby increasing the chance of association with the trajectory and indirectly reducing the confusion caused by missed detections. Secondly, an enhanced observation Kalman filter module is developed to predict the position of the tracking target in the current frame. Then, a feature model of the target is constructed in the time dimension and the space dimension respectively, and the time feature model and the space feature model of the tracking target are fused to obtain a spatio-temporal feature model of the tracking target. Finally, the spatio-temporal feature response of the object in the current frame is detected online, and the spatio-temporal feature response is associated with the spatio-temporal object feature model of the tracking object. By calculating the similarity metric matching matrix obtained by fusion, the Hungarian algorithm is used to solve the optimal association pairs between the historical trajectories of the tracking objects and the detection responses, and the parameters of the object spatio-temporal feature model are updated. Specifically, it includes the following steps:
[0065] Step 1: Obtain ship images and preprocess the initial images to obtain the original dataset.
[0066] Step 1.1: Obtain ship images from multiple angles and different scenarios to ensure the diversity and representativeness of the data.
[0067] Step 1.2: Manually annotate the collected ship images to mark the bounding boxes and key feature points of the ships.
[0068] Step 1.3: Preprocess the images, such as resizing, denoising, optimizing brightness and contrast, and remove low-quality and duplicate images. Finally, construct a training set, a validation set, and a test set.
[0069] Step 2: Perform object detection on the training samples and calibrate the time coordinates and space coordinates of the detection results.
[0070] Step 2.1: Use Repvit as the backbone network. First, adopt a single and deeper downsampling layer to increase the network depth and reduce the information loss caused by the resolution reduction. Second, use a simple classifier composed of a global average pooling layer and a linear layer to reduce the latency. Finally, select a better stage ratio of 1:1:7:1, and then increase the network depth to 2:2:14:2, thus achieving a deeper layout, improving the detection accuracy and reducing the latency. Use SlideLoss as the loss function, which can adaptively learn the positive sample threshold parameter and the negative sample threshold parameter μ. Setting a higher weight near μ will increase the relative loss of difficult samples and focus more attention on the misclassified examples of difficult samples. The calculation formula is as follows:
[0071] Step 2.2: Train the constructed model to obtain a target detection model with higher accuracy through multiple rounds of optimization.
[0072] Step 2.3: Perform target detection on the ship images through the model and calibrate the time coordinates and space coordinates of the detection results to ensure the consistency of the position and state of each target in the time and space dimensions.
[0073] Step 3: Eliminate the detection results with confidence lower than the set threshold, and introduce Occlusion-aware Non-Maximum Suppression (ONMS) to retain more occluded detections, thereby increasing the chance of association with the trajectory and indirectly reducing the confusion caused by missed detections.
[0074] Step 3.1: Set the confidence threshold to 0.8, and delete all images with detection result confidence lower than 0.8 to ensure the accuracy of subsequent multi-object tracking.
[0075] Step 3.2: Introduce Occlusion-aware Non-Maximum Suppression (ONMS) to retain more occluded regions. The formula is as follows:
[0076] S 1 = S·(1 - α·O i,j )
[0077] Where: S 1 is the adjusted confidence score. S is the original confidence score. α is the occlusion adjustment factor, which controls the weight of the occlusion effect and is usually between [0, 1]. O i,j is the degree of occlusion between box i and box j, usually calculated using IOU(A, B) or other features.
[0078] Step 4: Develop an enhanced observation Kalman filter module to predict the position of the target in the current frame, and respectively establish the time feature model and space feature model of the target, and construct the spatio-temporal feature model of the target to characterize the dynamic behavior of the target.
[0079] Step 4.1: Construct an enhanced observation Kalman filter module. Specifically, first, when establishing the association between adjacent frames, a Fus e -distanc e metric function that fuses IOU distance, object category, and size information is designed to determine the optimal match between the detection box and the Kalman filter prediction box. Secondly, a Gaussian cascade matching module is introduced to effectively solve the problem of frequent ID switching caused by the violent shaking of the ship. To further address the challenges posed by non-linear object motion, camera motion compensation is combined with the Kalman filter method based on the observation center, significantly reducing the estimation error of the Kalman filter. The formula for the Fuse-distance metric function is as follows: D Fuse = ω 1 ·D IOU + ω 2 ·D class + ω 3 ·D size Where: D Fuse represents the comprehensive metric value of Fuse-distance. D IOU represents the intersection over union (IOU) distance, which is used to measure the overlap degree of two boxes. D class represents the category difference metric, which determines whether two boxes belong to the same category. D size represents the size difference metric, which is used to compare the area difference between two boxes. ω 1 , ω 2 , ω 3 represent the weight coefficients, which are used to adjust the contribution of each feature to the final metric value.
[0080] Step 4.2: Construct a temporal feature model of the target. For temporal features, the histogram of optical flow (HOF) is used to describe the magnitude and direction of the motion speed of the tracking target.
[0081] Step 4.3: Construct a spatial feature model. For spatial features, based on the rectangular image of the tracking target in the training samples, the discriminant projection vector of the target region is obtained through the incremental linear discriminant analysis (ILDA) algorithm, and then this vector is weighted with the color histogram of the region to obtain the spatial features of the target.
[0082] Step 4.4: Fuse the spatial feature model M s and the temporal feature model M T to obtain the spatio-temporal feature model M of the target. The spatial feature model M s is constructed by ILDA and color histogram, and the temporal feature model M T describes the motion speed and direction through the histogram of optical flow. The spatio-temporal feature model M generated by fusing the two is used to characterize the dynamic behavior of the target.
[0083] Step 5: In the current frame, use the spatio-temporal feature model to detect the feature response of the target and construct a similarity matching matrix.
[0084] Step 5.1: In the data association stage, adopt a spatio-temporal data association framework. First, online detect the feature response of the tracked target in the current frame and extract its spatio-temporal features.
[0085] Step 5.2: Calculate the similarities of the spatial and temporal features respectively. The spatial feature similarity calculation is based on the cosine similarity of the feature vectors, which measures the consistency of the detection box and the target box in terms of spatial features. The formula is:
[0086]
[0087] For the detection response and the tracked target T i of the temporal feature similarity. For the detection response of the spatial feature vector. For the tracked target T i of the spatial feature vector. respectively represent and of the vector norms.
[0088] The temporal feature similarity calculation is based on the Gaussian distribution, which measures the similarity of the detection box and the target box in terms of motion features. The formula is:
[0089]
[0090] For the detection response and the tracked target T i of the temporal feature similarity. δ is the standard deviation of the Gaussian function, which is used to control the width of the temporal feature distribution. For the detection response and the tracked target T i of the temporal feature distance.
[0091] Step 5.3: Fuse the spatial feature similarity and the temporal feature similarity into a comprehensive similarity metric.
[0092] The formula is S combined is the fused comprehensive similarity. ω 1 and ω 2 are weight coefficients, which adjust the influence of spatial and temporal features on the matching.
[0093] Step 5.4: Construct a similarity matching matrix and fill the comprehensive similarity between all detection responses and the target feature model into the matrix:
[0094]
[0095] Step 6: Use the Hungarian algorithm to optimize the similarity matching matrix, determine the best match between the current frame target and the historical trajectory, and update the spatio-temporal feature model parameters of the target to achieve online adjustment and improve the tracking accuracy.
[0096] Step 6.1: Based on the constructed similarity matching matrix M, the Hungarian algorithm finds the best match between the detection response and the historical trajectory by maximizing the similarity (or minimizing the cost). Each match takes into account both spatial and temporal features to ensure the accuracy of the association. Subsequently, the optimal matching result is compared with a set threshold: if the matching result is higher than the threshold, it indicates a reliable association. If it is lower than the threshold, it may indicate target loss or mis-matching.
[0097] Step 6.2: For the successfully matched targets, update their spatio-temporal feature model parameters (including spatial feature vectors and temporal feature vectors). The parameter update is usually based on the latest feature values of the detection response, enabling the model to dynamically adapt to the changes of the target.
[0098] The multi-target tracking method for ships based on online spatio-temporal feature association effectively reduces the defect of frequent identity exchange caused by target occlusion and overcomes the problem of trajectory prediction error due to ship overlap. To more intuitively verify the effectiveness of the algorithm improvement, the same frame of the video is selected for comparison to verify the original target tracking algorithm ( Figure 4 ) and the target tracking algorithm based on online spatio-temporal feature association ( Figure 5 ) for visual comparison. As can be seen from the figure, the target tracking algorithm based on online spatio-temporal feature association can accurately track the target ship, and the number of id switches is also reduced, which is more conducive to practical applications.
[0099] The above embodiments are only used to illustrate the technical concept and characteristics of the present invention, and their purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A multi-target tracking method based on online spatiotemporal feature association, characterized in that: The following steps are involved: Step 1: Obtain ship images, preprocess the initial images, and obtain the original data set; Step 2: Perform target detection on the training samples and calibrate the time coordinates and spatial coordinates of the detection results; Step 3: Eliminate detection results with confidence levels lower than a set threshold, and introduce occlusion-aware non-maximum suppression (ONMS) to retain more occluded detections, increase the chances of association with trajectories, and indirectly reduce confusion caused by missed detections. Step 4: Develop an enhanced observation Kalman filter module to predict the position of the target in the current frame, and establish the time feature model and space feature model of the target respectively. Then, fuse the time feature model and space feature model of the tracked target to obtain the time and space feature model of the target, obtain the time feature and space feature of the target, and fuse them to generate the time and space feature to characterize the dynamic behavior of the target. Step 5: In the current frame, the spatiotemporal feature model is used to detect the spatiotemporal feature response of the target, and the spatiotemporal feature response is associated with the spatiotemporal object feature model of the tracked object, and the similarity matching matrix is obtained by fusion; Step 6: Use the Hungarian algorithm to optimize the similarity matching matrix, determine the best match between the current frame target and the historical trajectory, and update the target's spatiotemporal feature model parameters to achieve online adjustment and improve tracking accuracy.
2. The multi-target tracking method based on online spatiotemporal feature association according to claim 1, characterized in that: The specific method in step 1 is: Step 1.1: Obtain ship images from multiple angles and different scenes to ensure data diversity and representativeness; Step 1.2: Manually annotate the collected ship images and mark the boundary box and key feature points of the ship; Step 1.3: Perform preprocessing on the images, such as resizing, denoising, brightness and contrast optimization, and remove low-quality and duplicate images. Finally, construct the training set, validation set, and test set.
3. The multi-target tracking method based on online spatiotemporal feature association according to claim 1, characterized in that: The specific method in step 2 is: Step 2.1: Using Repvit as the backbone network, first use a single and deeper downsampling layer to increase the network depth and reduce the information loss due to resolution reduction; secondly, use a simple classifier consisting of a global average pooling layer and a linear layer; finally, choose a stage ratio of 1:1:7:1, and then increase the network depth to 2:2:14:2; use SlideLoss as the loss function, and adaptively learn the positive sample threshold parameter and the negative sample threshold parameter μ, the calculation formula is as follows: Among them, x represents an unknown number, which is used to limit the μ interval. Step 2.2: Train the constructed model and obtain the target detection model through multiple rounds of optimization; Step 2.3: Use the model to detect the ship image and calibrate the time coordinates and space coordinates of the detection results to ensure that the position and state of each target are consistent in time and space dimensions.
4. The multi-target tracking method based on online spatiotemporal feature association according to claim 1, characterized in that: The specific method in step 3 is: Step 3.1: Set the confidence threshold and delete all images whose detection results have confidence values lower than the threshold to ensure the accuracy of subsequent multi-target tracking; Step 3.2: Introduce occlusion-aware non-maximum suppression (ONMS) to retain more occluded areas. The formula is as follows: S 1 =S·(1-α·O i,j ) Where: S 1 is the adjusted confidence score, S is the original confidence score, α is the occlusion adjustment factor, which controls the weight of the occlusion effect, between [0,1], O i,j It is the degree of occlusion between box i and box j, expressed as IOU(A,B).
5. The multi-target tracking method based on online spatiotemporal feature association according to claim 1, characterized in that: The specific method in step 4 is: Step 4.1: Construct the enhanced observation Kalman filter module; Firstly, when establishing the association between adjacent frames, a fuse-distance metric function is designed to integrate IOU distance, object category and size information to determine the optimal match between the detection frame and the Kalman filter prediction frame. Secondly, a Gaussian cascade matching module is introduced to solve the frequent ID switching problem caused by the violent shaking of the ship. The fuse-distance metric function formula is as follows: D Fuse =ω1·D IOU +ω2·D class +ω3·D size Where: D Fuse Represents the comprehensive metric value of Fuse-distance, D IOU It represents the intersection-over-union (IOU) distance, which is used to measure the degree of overlap between two boxes. class Represents the category difference metric, judging whether two boxes belong to the same category, D size Represents the size difference metric, which is used to compare the area difference between two boxes. ω1, ω2, ω3 represent weight coefficients, which are used to adjust the contribution of each feature to the final metric value. Step 4.2: Construct the temporal feature model M of the target T ; For the time feature, the optical flow histogram HOF is used to describe the size and direction of the moving speed of the tracking target; Step 4.3: Construct spatial feature model M s ; For spatial features, according to the rectangular image of the tracked target in the training sample, the discriminative projection vector of the target area is obtained through the incremental linear discriminant analysis ILDA algorithm, and then the vector is weighted with the color histogram of the area to obtain the spatial features of the target; Step 4.4: Transform the spatial feature model M s and time characteristic model M T Fusion, to obtain the target's spatiotemporal feature model M and spatial feature model M s The temporal feature model M is constructed by ILDA and color histogram. T The optical flow histogram is used to describe the motion speed and direction, and the spatiotemporal feature model M generated by the fusion of the two is used to characterize the dynamic behavior of the target.
6. The multi-target tracking method based on online spatiotemporal feature association according to claim 5, characterized in that: The specific method in step 5 is: Step 5.1: In the data association stage, a spatiotemporal data association framework is adopted; first, the feature response of the tracked target in the current frame is detected online, and the spatiotemporal features are extracted; Step 5.2: Calculate the similarity of spatial and temporal features respectively; The spatial feature similarity calculation is based on the cosine similarity of the feature vector, which measures the consistency of the detection box and the target box in spatial features. The formula is: in, To detect response and tracking target T i The similarity of time features, To detect the response The spatial eigenvector of To track the target T i The spatial eigenvector of Respectively and The vector norm of ; The temporal feature similarity calculation is based on Gaussian distribution, which measures the similarity between the detection frame and the target frame in terms of motion features. The formula is: in, To detect response and tracking target T i δ is the standard deviation of the Gaussian function, which is used to control the width of the temporal feature distribution. To detect the response and tracking target T i The time characteristic distance of Step 5.3: Combine spatial feature similarity and temporal feature similarity into a comprehensive similarity measure, the formula is S combined is the comprehensive similarity after fusion, ε1 and ε2 are weight coefficients, which adjust the influence of spatial and temporal features on matching; Step 5.4: Construct a similarity matching matrix and fill the matrix with the comprehensive similarity between all detection responses and the target feature model: G is the similarity matching matrix, is the mth detection response, T n is the nth target.
7. The multi-target tracking method based on online spatiotemporal feature association according to claim 1, characterized in that: The specific method in step 6 is: Step 6.1: Based on the constructed similarity matching matrix M, the Hungarian algorithm finds the best match between the detection response and the historical trajectory by maximizing the similarity or minimizing the cost; each match considers both spatial and temporal features, and then compares the best match result with the set threshold. If the match result is higher than the threshold, it means that the association is reliable. If it is lower than the threshold, it means that the target is lost or mismatched. Step 6.2: For the successfully matched target, update its spatiotemporal feature model parameters, including spatial feature vectors and temporal feature vectors. The parameter update is based on the latest feature values of the detection response, so that the model can dynamically adapt to changes in the target.