An adaptive association method for sea surface multi-target tracking
By extracting global features and generating fusion weights in multi-target tracking on the sea surface, the matching accuracy and stability issues of target tracking in complex marine environments are solved. This enables dynamic adjustment of the adaptive association strategy, thereby improving the accuracy and stability of multi-target tracking on the sea surface.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
- Filing Date
- 2026-05-26
- Publication Date
- 2026-06-23
AI Technical Summary
Existing multi-target tracking methods are prone to matching errors in complex and ever-changing marine environments, leading to target loss or misidentification. Furthermore, their tracking accuracy and stability decrease when faced with complex interference, making it difficult to meet the stable tracking requirements in complex dynamic scenarios.
Multiple global features are extracted from the current frame image to construct a global scene feature vector, which is then input into a pre-trained risk prediction model to generate target identity change risk and trajectory breakage risk. Based on these risks, the fusion weights of multi-branch association costs are dynamically generated, and weighted fusion calculations are performed to achieve target matching.
By adaptively adjusting the association strategy, the matching accuracy and trajectory continuity of multi-target tracking on the sea surface are improved, the cross-scenario generalization ability is enhanced, and the accuracy of fixed weight strategies in dynamic and complex environments and the overall tracking failure caused by limited local information vision are avoided.
Smart Images

Figure CN122265853A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target tracking technology, and in particular to an adaptive association method for multi-target tracking on the sea surface. Background Technology
[0002] Multi-target tracking technology is mainly used to determine the motion trajectory of targets in a continuous video sequence.
[0003] Existing multi-target tracking methods typically employ local adaptive weight fusion techniques to associate and match detected targets. However, in complex and dynamic scenarios (such as fog blurring target features, strong reflections affecting positioning accuracy, or severe shaking of the monitoring platform), these methods are prone to matching errors during target association, leading to target loss or incorrect identification. Furthermore, when encountering complex interference, existing methods are susceptible to overall tracking failure, resulting in a significant decrease in tracking matching accuracy and stability. This makes it difficult to meet the stable tracking requirements in complex dynamic scenarios and presents a problem of poor scene generalization ability. Summary of the Invention
[0004] This invention provides an adaptive association method for tracking multiple targets on the sea surface, in order to overcome the deficiencies in the existing technology.
[0005] This invention provides an adaptive association method for multi-target tracking on the sea surface, comprising the following steps: Extract multiple global features from the current frame image, and combine the multiple global features to construct a global scene feature vector; The global scene feature vector is input into a pre-trained risk prediction model to obtain the target identity change risk and trajectory break risk output by the risk prediction model. Obtain the target detection result and the prediction result of the previous trajectory of the current frame image, and calculate the association cost of the target detection result and the prediction result under multiple different feature dimensions to obtain the multi-branch association cost; Based on the target identity change risk and the trajectory breakage risk, generate the fusion weight corresponding to the multi-branch association cost; The multi-branch association cost is weighted and fused according to the fusion weight to obtain the final association cost; Based on the final association cost, the target detection result is matched with the prediction result to obtain the target matching result.
[0006] The present invention also provides an adaptive correlation device for multi-target tracking on the sea surface, comprising the following modules: The extraction module is used to extract multiple global features of the current frame image and combine the multiple global features to construct a global scene feature vector; The prediction module is used to input the global scene feature vector into a pre-trained risk prediction model to obtain the target identity change risk and trajectory break risk output by the risk prediction model. The calculation module is used to obtain the target detection result and the prediction result of the previous trajectory of the current frame image, and calculate the association cost of the target detection result and the prediction result under multiple different feature dimensions to obtain the multi-branch association cost; The generation module is used to generate the fusion weights corresponding to the multi-branch association cost based on the target identity change risk and the trajectory breakage risk. The fusion module is used to perform weighted fusion calculation on the multi-branch association cost according to the fusion weight to obtain the final association cost; The matching module is used to match the target detection result with the prediction result based on the final association cost to obtain the target matching result.
[0007] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the adaptive association method for multi-target tracking on the sea surface as described above.
[0008] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the adaptive association method for multi-target tracking on the sea surface as described above.
[0009] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the adaptive association method for multi-target tracking on the sea surface as described above.
[0010] This invention provides an adaptive association method for multi-target tracking on the sea surface. It extracts multiple global features from the current frame image and constructs a global scene feature vector. This global scene feature vector is input into a pre-trained risk prediction model to obtain target identity change risk and trajectory breakage risk. Then, based on the target identity change risk and trajectory breakage risk, it dynamically generates fusion weights corresponding to multi-branch association costs. After weighted fusion calculation of the multi-branch association costs according to the fusion weights, target matching is performed, achieving adaptive dynamic adjustment of the association matching strategy in multi-target tracking on the sea surface. Since the fusion weights are generated based on the target identity change risk and trajectory breakage risk inferred from the global features extracted from the current frame image through the risk prediction model, the contribution ratio of each association cost branch can be adaptively adjusted with changes in the overall scene state. This avoids the problem of decreased association matching accuracy caused by the inability to change weights in dynamic and complex marine environments when using a fixed-weight strategy. It also avoids the problem of overall tracking failure when facing global interference or unreasonable global strategy adjustments when facing local interference when adjusting weights only based on local information due to a lack of global scene perception. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating the adaptive association method for multi-target tracking on the sea surface provided by the present invention.
[0013] Figure 2 This is a schematic diagram of the overall framework of the adaptive association method for multi-target tracking on the sea surface provided in this embodiment of the invention.
[0014] Figure 3 This is a schematic diagram of the adaptive correlation device for multi-target tracking on the sea surface provided by the present invention.
[0015] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0017] Multi-object tracking technology is mainly used to determine the motion trajectory of targets in continuous video sequences. Existing multi-object tracking technologies mainly revolve around fusion strategies based on associated cues, and can be divided into two main categories: fixed-weight fusion and locally adaptive weight fusion. Fixed-weight fusion methods construct a cost matrix by fusing associated cues such as motion and appearance using preset fixed weights to achieve trajectory matching with detection boxes. Locally adaptive weight fusion methods, on the other hand, aim to overcome the rigidity of fixed weights by dynamically adjusting the weights based on local information in a single frame.
[0018] However, when faced with the complex and ever-changing marine environment, the relevant technologies cannot adapt to its dynamic characteristics and suffer from the following core defects: First, the static nature of the fixed-weight strategy means it completely lacks dynamic adaptability. On the one hand, marine scene interference is significantly dynamic (such as a ship changing from stable sailing to violent rolling), which can cause a sharp drop in the reliability of motion cues. Fixed weights can lead to a significant decline in the accuracy of association matching. On the other hand, it cannot take into account the differentiated needs of various marine sub-scenes such as ports, open seas, and rainy / foggy weather for cues, resulting in extremely poor generalization ability. Second, the limited field of vision of local adaptation means that existing methods rely only on local information such as the confidence of a single detection box and the state of a single trajectory for fine-tuning, lacking global scene perception capabilities. When faced with global interference such as heavy fog or strong glare, it cannot trigger an effective global cue switching strategy, which can easily lead to overall tracking failure. When faced with local problems such as temporary occlusion of a single target, it is easily misjudged as a global problem, resulting in suboptimal adjustment strategies and seriously affecting tracking stability. Finally, the effectiveness of the association mechanism is limited. Existing mechanisms mainly rely on manual design, and their impact on the final tracking performance and adaptability in complex dynamic scenes remains severely insufficient.
[0019] To address this issue, this invention provides an adaptive association method for multi-target tracking on the sea surface. The method extracts multiple global features from the current frame image and combines these features to construct a global scene feature vector. This global scene feature vector is then input into a pre-trained risk prediction model to obtain the target identity jump risk and trajectory breakage risk output by the model. The method acquires the target detection result and the predicted trajectory of the previous frame image, and calculates the association cost between the target detection result and the predicted result under multiple different feature dimensions, obtaining a multi-branch association cost. Based on the target identity jump risk and trajectory breakage risk, fusion weights corresponding to the multi-branch association costs are generated. The multi-branch association costs are then weighted and fused according to the fusion weights to obtain the final association cost. Based on the final association cost, the target detection result and the predicted result are matched to obtain the target matching result. This method breaks through the limitations of local information perspective and the rigidity of static weights, endowing the system with global perception and dynamic adaptability to complex marine environments. This effectively overcomes dynamic interference such as strong winds and waves, rain, fog, and glare, significantly improving the matching accuracy, trajectory continuity, and cross-scene generalization ability in the process of multi-target tracking on the sea surface. This method can be applied to marine observation platforms equipped with cameras, including but not limited to drones, unmanned surface vessels, or shore-based monitoring stations. The implementer of this method can be a computing device deployed on or connected to the aforementioned observation platform.
[0020] in, Figure 1 This is a flowchart illustrating the adaptive association method for multi-target tracking on the sea surface provided by the present invention, as shown below. Figure 1 As shown, the method includes steps 110, 120, 130, 140, 150 and 160.
[0021] Step 110: Extract multiple global features from the current frame image and combine them to construct a global scene feature vector.
[0022] During multi-target tracking on the sea surface, the marine observation platform continuously acquires video streams, obtaining image data frame by frame from these streams. The frame currently being processed is recorded as the current frame image.
[0023] Here, global features refer to statistical or descriptive features extracted from the entire scope of the current frame image that reflect the overall state of the scene corresponding to that frame image. Unlike local features that only focus on the state of a single detected target, global features are oriented towards the overall performance of all detected targets and the environment in the current frame image. Global features can cover multiple dimensions such as motion dimension, target distribution dimension, and environmental interference dimension. Among them, global features in the motion dimension are used to characterize the degree of overall image change introduced by the motion of the marine observation platform itself in the current frame image; global features in the target distribution dimension are used to characterize the overall scale distribution and detection reliability distribution of detected targets in the current frame image; global features in the environmental interference dimension are used to characterize the overall impact of environmental factors such as wave reflection and fog obscuring on target detection and appearance recognition in the current frame image.
[0024] In one optional implementation, the number of global features can be 2, 3, 4, 5, 6, or more. The specific categories of global features can be flexibly set according to the actual application scenario and target tracking requirements. For example, in a sea surface monitoring scenario, global features may include six features: camera motion intensity, average detection box size, detection confidence variance, proportion of small targets, appearance feature clarity, and proportion of wave reflective areas. In other application scenarios, global features may include only some of the above features, or may include other feature types such as target density and average inter-frame target displacement.
[0025] After extracting multiple global features, these features are combined in a preset order to construct a global scene feature vector. The global scene feature vector can be a fixed-dimensional numerical vector, with its dimension equal to the number of extracted global features. For example, when six global features are extracted, the global scene feature vector is a six-dimensional vector.
[0026] In one alternative implementation, before constructing the global scene feature vector, each global feature can be normalized to map the values of each global feature to a unified range of values, thereby eliminating the dimensional differences between different global features.
[0027] Step 120: Input the global scene feature vector into the pre-trained risk prediction model to obtain the target identity change risk and trajectory break risk output by the risk prediction model.
[0028] Specifically, a risk prediction model refers to a pre-trained computational model capable of inferring the tracking risk level of a current frame image based on the input global scene feature vector. This risk prediction model can be implemented using various architectures, including but not limited to multilayer perceptrons, recurrent neural networks, gradient boosting decision trees, or other model architectures capable of regression prediction. The core function of the risk prediction model is to establish a mapping relationship between the global scene feature vector and tracking risk, that is, to predict the degree of risk that may be encountered when performing target association matching under the current scene conditions, based on the overall scene status of the current frame image. The risk prediction model outputs two risk indicators: target identity change risk and trajectory breakage risk.
[0029] Target identity switching risk refers to the degree of risk that, during the association matching process of the current frame image, the same actual target will be incorrectly associated with different tracking identities, or different actual targets will be incorrectly associated with the same tracking identity. The target identity switching risk can be a continuous probability value. A higher target identity switching risk value indicates a higher probability of incorrect identity switching occurring during the association matching process under the current scene conditions.
[0030] Trajectory breakage risk refers to the degree of risk that the continuously tracked target trajectory will be interrupted during the association matching process of the current frame image due to failure to successfully match the corresponding detected target. The trajectory breakage risk can also be a continuous probability value. The higher the value of the trajectory breakage risk, the higher the probability that the existing trajectory has been interrupted or fragmented under the current scene conditions.
[0031] In one alternative implementation, the range of values for both the target identity change risk and the trajectory break risk can be limited to the range [0, 1].
[0032] Regarding the training process of the risk prediction model, supervisory signals can be constructed during the training phase to guide the learning of model parameters. Specifically, the training label for target identity change risk can be determined based on the ratio of the number of times the same ground truth trajectory corresponds to different tracking identity identifiers between adjacent frames to the total number of ground truth trajectories in adjacent frames. The training label for trajectory breakage risk can be determined based on the ratio of the number of times a ground truth trajectory transitions from a continuous unmatched state to a rematched state to the total number of ground truth trajectories in the current frame. Through the above-mentioned construction of supervisory signals, the risk prediction model can learn the intrinsic correlation between the global scene feature vector and the actual tracking risk.
[0033] For example, in a port scenario, when the marine observation platform experiences severe shaking and there is extensive reflection on the sea surface, the values of camera motion intensity and the proportion of wave reflection area in the global scene feature vector are relatively large. Based on this, the risk prediction model outputs a higher risk of target identity change and trajectory breakage. In a navigation environment with clear weather and calm sea, the values of each dimension of the global scene feature vector are relatively stable, and based on this, the risk prediction model outputs a lower risk of target identity change and trajectory breakage.
[0034] Step 130: Obtain the target detection result and the prediction result of the previous trajectory of the current frame image, and calculate the association cost of the target detection result and the prediction result under multiple different feature dimensions to obtain the multi-branch association cost.
[0035] Specifically, the object detection result refers to the set of information about all detected objects obtained after performing object detection on the current frame image. The object detection result can be obtained by processing the current frame image using an object detection model. The information for each detected object typically includes the bounding box coordinates, bounding box width and height, detection confidence score, and appearance features. The number of detected objects included in the object detection result depends on the number of objects that actually exist in the current frame image and have been successfully detected.
[0036] The preceding trajectory refers to the motion trajectory of all tracked targets that has been established and maintained in video frames preceding the current frame. Each preceding trajectory records the historical motion state information of the tracked target, including but not limited to position coordinates, motion velocity, bounding box size, and historical appearance features.
[0037] The prediction result of the preceding trajectory refers to the result obtained by predicting the position of each preceding trajectory in the current frame image based on the historical motion state information recorded in the preceding trajectory and using a motion state prediction method. Kalman filtering can be used for motion state prediction to obtain the prediction result. The number of prediction results is the same as the number of preceding trajectories, with one prediction result corresponding to each preceding trajectory.
[0038] After obtaining the target detection results and the prediction results of the preceding trajectory, this embodiment calculates the association cost between the two under multiple different feature dimensions.
[0039] Here, association cost refers to a numerical metric used to quantify the degree of matching between a single target detection result and the prediction result of a single preceding trajectory. A smaller association cost indicates a higher degree of matching between the target detection result and the prediction result of the preceding trajectory. Feature dimension refers to different metric angles used to measure the degree of matching. Association costs under different feature dimensions evaluate the matching relationship between the target detection result and the prediction result of the preceding trajectory from different information levels.
[0040] Multi-branch association cost refers to the set of association costs calculated separately for multiple different feature dimensions. Each feature dimension corresponds to a branch, and each branch independently calculates its own association cost.
[0041] In one optional implementation, the feature dimensions may include at least two of the following: spatial dimension, appearance dimension, motion dimension, and reliability dimension. The association cost in the spatial dimension quantifies the degree of spatial overlap between the bounding box of the detected object and the bounding box of the predicted result of the preceding trajectory. The association cost in the appearance dimension quantifies the similarity between the appearance features of the detected object and the historical appearance features of the preceding trajectory. The association cost in the motion dimension quantifies the degree of spatial overlap between the predicted result and the detected object after motion compensation. The association cost in the reliability dimension quantifies the detection reliability of the detected object itself.
[0042] For each of the aforementioned feature dimensions, a corresponding association cost is calculated between each detected target in the target detection result and each predicted position in the prediction result of the preceding trajectory, thus forming a cost matrix. The number of rows in this cost matrix represents the number of preceding trajectories, and the number of columns represents the number of detected targets in the current frame image. Multiple feature dimensions correspond to multiple cost matrices, and these cost matrices together constitute the multi-branch association cost.
[0043] It should be noted that there is no strict order restriction between the operations of obtaining the target detection results and the prediction results of the preceding trajectory in step 130, and the operations of extracting global features in step 110 and performing risk prediction in step 120. In actual implementation, step 130 can be executed in parallel with steps 110 and 120, or it can be executed sequentially according to the order of the step numbers.
[0044] Step 140: Based on the target identity change risk and trajectory breakage risk, generate the fusion weights corresponding to the multi-branch association costs.
[0045] Here, the fusion weight refers to the weight coefficient assigned to each branch in the multi-branch association cost. The fusion weight is used to determine the contribution ratio of each branch when weighting and fusioning the association costs of each branch in subsequent steps. If the multi-branch association cost contains several branches, then the same number of fusion weights are generated accordingly, with each fusion weight corresponding to one branch.
[0046] The core of the fusion weight generation process lies in dynamically determining the importance of each branch in the final association cost calculation based on the specific values of the target identity jump risk and trajectory breakage risk in the current frame image. When the values of the target identity jump risk and trajectory breakage risk change, the fusion weights corresponding to each branch also change accordingly.
[0047] In one optional implementation, the sum of the values of each fusion weight can be normalized to 1. In another optional implementation, the values of each fusion weight may not be subject to normalization constraints.
[0048] For example, when the target identity change risk is high while the trajectory breakage risk is low, it indicates that the current scenario is more prone to erroneous identity switching during association matching. In this case, the fusion weight of the branch corresponding to the appearance dimension can be increased to leverage the discriminative power of appearance features to reduce the probability of erroneous identity switching, while the fusion weight of the branch corresponding to the spatial dimension can be decreased. Similarly, when the target identity change risk is low while the trajectory breakage risk is high, it indicates that the existing trajectory is more likely to be interrupted in the current scenario. In this case, the fusion weight of the branch corresponding to the reliability dimension can be increased to maintain trajectory continuity using high-reliability detection results. Furthermore, when both risk indicators are low, it indicates that the current scenario is relatively stable. In this case, the fusion weight of the branch corresponding to the spatial dimension can be increased because spatial overlap is the most direct and efficient matching criterion in a stable scenario.
[0049] Step 150: Perform weighted fusion calculation on the multi-branch association cost according to the fusion weight to obtain the final association cost.
[0050] Specifically, the final association cost refers to the comprehensive association cost obtained by weighted fusion of the association costs of each branch in the multi-branch association cost. The final association cost can be represented in the form of a cost matrix, where the number of rows is the number of preceding trajectories and the number of columns is the number of detected targets in the current frame image.
[0051] Specifically, the final association cost between the prediction result of the i-th preceding trajectory and the j-th detected target can be calculated as follows: multiply the cost value of each branch in the multi-branch association cost at the i-th row and j-th column by the fusion weight of the corresponding branch, and then add the results of each product to obtain the final association cost value at that position.
[0052] The final association cost matrix integrates matching information from various feature dimensions, and the contribution ratio of each dimension is dynamically determined by step 140 based on the risk status of the current frame image. Therefore, the final association cost can adaptively adjust the matching strategy according to different scene conditions, which not only leverages the complementary advantages of multi-dimensional association clues but also avoids the limitations of fixed weight allocation in complex scenes.
[0053] Step 160: Based on the final association cost, match the target detection result with the prediction result to obtain the target matching result.
[0054] Specifically, the matching operation refers to the process of finding a one-to-one correspondence between all target detection results and all predicted results of preceding trajectories based on the final association cost matrix, which minimizes the total global association cost.
[0055] In an alternative implementation, an association cost threshold can be set during the matching operation. When the final association cost between a target detection result and a predicted result of a preceding trajectory exceeds the association cost threshold, the pairing is considered invalid, even if the two are paired in the current matching scheme. In one example, the association cost threshold can be set to 0.5.
[0056] The target matching results contain the following three types of information: The first type is the set of successfully matched pairs, which are pairs where a one-to-one correspondence is successfully established between the target detection result and the prediction result of the preceding trajectory; the second type is the set of unmatched preceding trajectories, which are preceding trajectories of the detected target that could not be found in the current frame image; and the third type is the set of unmatched target detection results, which are detected targets of the preceding trajectory that could not be found in the preceding trajectory.
[0057] For example, suppose that 5 targets are detected in the current frame image, and the preceding trajectories contain 4 active trajectories. After the matching operation, if 3 targets are successfully matched with 3 preceding trajectories respectively, then the set of successfully matched pairs contains 3 pairs, the set of unmatched preceding trajectories contains 1 preceding trajectory, and the set of unmatched target detection results contains 2 targets.
[0058] In the subsequent processing flow, corresponding operations can be performed according to the three types of information of the target matching result: for a successfully matched pair, the state information of the preceding trajectory can be updated using the corresponding target detection result; for a failed match of the preceding trajectory, the position can be completed using a prediction method, or the preceding trajectory can be terminated after exceeding the preset conditions; for a failed target detection result, a new tracking trajectory can be initialized accordingly.
[0059] The adaptive association method for multi-target tracking on the sea surface provided in this embodiment extracts multiple global features from the current frame image and constructs a global scene feature vector. This global scene feature vector is then input into a pre-trained risk prediction model to obtain target identity change risk and trajectory breakage risk. Based on these risks, fusion weights corresponding to multi-branch association costs are dynamically generated. After weighted fusion calculation of the multi-branch association costs according to the fusion weights, target matching is performed, achieving adaptive dynamic adjustment of the association matching strategy in multi-target tracking on the sea surface. Since the fusion weights are generated based on the target identity change risk and trajectory breakage risk inferred from the global features extracted from the current frame image through the risk prediction model, the contribution ratio of each association cost branch can be adaptively adjusted with changes in the overall scene state. This avoids the problem of decreased association matching accuracy caused by the inability to change weights in a fixed-weight strategy when facing dynamic and complex marine environments. It also avoids the problem of overall tracking failure or unreasonable global strategy adjustments when facing local interference due to a lack of global scene perception when adjusting weights based solely on local information.
[0060] Considering that in multi-target tracking scenarios on the sea surface, the interference factors affecting target association and matching are widespread and independent, including image jitter introduced by the movement of the marine observation platform itself, differences in the scale distribution of sea surface targets, fluctuations in the reliability of detector output, and environmental interference such as wave reflections, a single-dimensional global feature is insufficient to comprehensively characterize the overall scene state of the current frame image. Therefore, this embodiment selects representative features from the motion dimension, target distribution dimension, and environmental interference dimension to extract six global features from the current frame image: camera motion intensity, average detection box size, detection confidence variance, proportion of small targets, clarity of appearance features, and proportion of wave reflection areas, in order to achieve a multi-angle comprehensive characterization of the scene state of the current frame image.
[0061] Specifically, the extraction of multiple global features from the current frame image includes: The camera motion intensity, average detection box scale, detection confidence variance, small target proportion, appearance feature sharpness, and wave reflection area proportion of the current frame image are extracted as multiple global features.
[0062] ①Camera motion intensity Camera motion intensity is used to quantify the degree of overall image displacement and deformation caused by the movement of the marine observation platform itself in the current frame. In maritime monitoring scenarios, the marine observation platform is prone to swaying due to waves, resulting in overall image translation, rotation, or scaling changes between adjacent frames. The higher the value of camera motion intensity, the more severe the influence of the platform's own movement on the current frame image, and the lower the reliability of target association based on spatial location.
[0063] As an optional implementation, the camera motion intensity extraction step includes: First, extracting the motion change of feature points between the current frame and the previous frame. Specifically, several pixels with significant texture features can be selected as feature points in the previous frame. Then, a sparse optical flow algorithm is used to calculate the corresponding positions of these feature points in the current frame, thereby obtaining the displacement vector of each feature point between two adjacent frames. This displacement vector is the motion change of the feature points. During the process of obtaining the motion change of feature points, a random sampling consensus algorithm can be used to remove outliers that are inconsistent with the overall motion pattern, thereby improving the accuracy of subsequent calculations.
[0064] After obtaining the motion changes of the feature points after outlier removal, a spatial transformation matrix is constructed based on these motion changes. The spatial transformation matrix is used to describe the overall geometric transformation relationship between two adjacent frames.
[0065] In one alternative implementation, the spatial transformation matrix It can be expressed in the form of an affine transformation matrix, which is a 2x3 matrix, denoted as: ; in, and Characterizing scale variation components, and Characterizing rotational and shear change components, and Characterizes the translational change component.
[0066] After constructing the spatial transformation matrix, the camera motion intensity is determined by calculating the changes in the elements of the spatial transformation matrix. Specifically, the absolute values of the differences between each element of the spatial transformation matrix and its corresponding element of the identity matrix are summed to obtain the camera motion intensity. The calculation formula is as follows: ; in, The first element in the spatial transformation matrix Line number The element values of the column, For the elements of the identity matrix, when hour, The value is 1, when hour, The value is 0. When there is no camera motion between two adjacent frames, the spatial transformation matrix degenerates into a unit transformation, and the camera motion intensity is... The value of is 0; the more intense the camera motion, the greater the deviation of each element in the spatial transformation matrix from the corresponding element of the identity matrix, and the greater the camera motion intensity. The larger the value, the better.
[0067] ② Average detection frame size The average detection box scale is used to quantify the overall size of all detected targets in the current frame image. In maritime surveillance scenarios, the target detection box scale is closely related to the distance of the target from the maritime observation platform, and is also affected by the focal length and shooting height of the camera equipment on the maritime observation platform. The smaller the average detection box scale value, the farther the target is in the current frame image or the smaller the target size, which may reduce the discriminative power of spatial dimension association costs.
[0068] As an optional implementation, the average detection box scale is characterized by the average area of the bounding boxes of all detected objects in the current frame image. Average Detection Box Scale The calculation formula is: ; in, This represents the total number of detected targets in the current frame image. For the first The width of the bounding box of each detected target. For the first The height of the bounding box of each detected target.
[0069] ③ Detect confidence level variance The detection confidence variance quantifies the dispersion of the detection confidence values for all detected targets in the current frame image, reflecting whether the detection reliability of the target detection model is consistent for each detected target in the current frame image. When the detection confidence variance is large, it indicates that the detection reliability varies significantly between different detected targets, and some targets may have a high risk of false detection or false negative detection.
[0070] As an optional implementation, let the current frame image be the first... The detection confidence level of each target is The mean detection confidence score of all detected targets in the current frame image is Then the test confidence variance The calculation formula is: ; in, Let be the mean confidence score for the detection in frame t. This represents the total number of targets to be detected.
[0071] ④ Percentage of small goals The small target percentage quantifies the proportion of smaller detected targets in the current frame of the image out of all detected targets. In wide-format scenes such as aerial photography of the sea, distant targets appear extremely small in the image, which can easily lead to decreased positioning accuracy and blurred appearance features. The higher the small target percentage value, the more small targets in the current frame that require special attention.
[0072] As an optional implementation, detection targets with an area smaller than a preset area threshold are defined as small targets. This area threshold can be set to one-third of the average scale of targets in the training dataset. Let the area threshold be... The proportion of small goals The calculation formula is: ; in, This is an indicator function that takes the value 1 when the condition within the parentheses is true and 0 when the condition is false.
[0073] ⑤ Appearance (Re-Identification, REID) feature clarity Appearance feature sharpness is used to quantify the discriminative power of appearance features of all detected targets in the current frame image. Appearance features are high-dimensional vector representations extracted from the image regions of each detected target by an appearance feature extraction model. When appearance feature sharpness is high, it indicates that the appearance differences between the detected targets are more obvious, and the discriminative power of appearance dimension association cost is stronger; when appearance feature sharpness is low, such as in foggy or low-light conditions, the appearance features between the detected targets tend to be similar, and the reliability of appearance dimension association cost decreases.
[0074] As an optional implementation, appearance feature sharpness is characterized by the normalized variance of the appearance feature vectors of all detected targets in the current frame image. Let the appearance feature matrix of all detected targets in the current frame image be... Its dimension is the total number of detected targets multiplied by the dimension of the appearance feature vector, then the appearance feature clarity is... The calculation formula is: ; in, The appearance feature matrix (dimensions) of all detected targets in frame t. ), The mean of the column variances of the matrix. The maximum feature variance in the dataset.
[0075] ⑥ Percentage of reflective area from sea waves After extracting the clarity of appearance features, global features reflecting environmental interference are further extracted. The proportion of wave reflection area is used to quantify the area ratio of high-brightness areas formed by wave reflection in the current frame image. In sea surface monitoring scenarios, wave reflection can cause overexposure in local areas of the image, affecting the target detector's positioning accuracy and appearance feature extraction effect in that area. The larger the value of the wave reflection area proportion, the more severe the reflection interference in the current frame image.
[0076] As an optional implementation, the step of extracting the proportion of reflective areas of ocean waves includes: segmenting the current frame image using a set grayscale threshold to extract the reflective areas. Specifically, firstly, the global grayscale mean of the current frame image is calculated, and 1.5 times this global grayscale mean is set as the grayscale threshold. Then, iterate through each pixel in the current frame image, and assign pixels with gray values greater than the gray value threshold. Pixels marked as reflective pixels are considered reflective pixels. All areas marked as reflective pixels constitute the reflective region. After extracting the reflective region, the ratio of the reflective region's pixel area to the total pixel area of the current frame image is calculated to obtain the proportion of the wave's reflective region. The calculation formula is as follows: ; in, This represents the grayscale value of the pixel at coordinates (x, y) in the current frame image. This is the grayscale threshold. The width of the current frame image in pixels. The height of the current frame image in pixels. This is an indicator function that takes the value 1 when the condition within the parentheses is true and 0 when the condition is false.
[0077] After extracting the above six global features, the camera motion intensity, average detection box scale, detection confidence variance, small target proportion, appearance feature sharpness, and wave reflection area proportion are combined in a fixed order into a six-dimensional global scene feature vector, denoted as . .
[0078] In multi-target tracking scenarios on the sea surface, camera motion intensity and the proportion of wave reflection areas characterize the scene state of the current frame image from the dimensions of motion and environmental interference, respectively. The extraction accuracy of these two directly affects the quality of the subsequent global scene feature vector and the prediction accuracy of the risk prediction model. Specifically, camera motion intensity needs to accurately reflect the overall geometric deformation of the image caused by the movement of the marine observation platform itself, while the proportion of wave reflection areas needs to accurately reflect the range of interference between sea surface reflection and image quality. Therefore, this embodiment designs corresponding extraction methods for the physical characteristics of both.
[0079] Specifically, the steps for extracting the camera motion intensity and the proportion of the wave reflection area include: Extract the motion change of feature points between the current frame image and the previous frame image, construct a spatial transformation matrix based on the motion change of feature points, and determine the camera motion intensity by calculating the change of elements in the spatial transformation matrix. The current frame image is segmented by setting a grayscale threshold, the reflective area is extracted, and the ratio of the pixel area of the reflective area to the total pixel area of the current frame image is calculated to obtain the proportion of the reflective area of the waves.
[0080] In maritime surveillance scenarios, offshore observation platforms are prone to rigid body motions such as translation, rotation, and scaling due to factors like waves and wind. These motions result in a holistic spatial geometric transformation between adjacent image frames. Accurately quantifying the degree of this holistic spatial geometric transformation can provide risk prediction models with crucial information about the reliability of motion cues in the current frame. Therefore, this embodiment extracts the motion changes of feature points between adjacent frames and constructs a spatial transformation matrix accordingly. Finally, the camera motion intensity is determined by calculating the changes in the elements within the spatial transformation matrix.
[0081] Specifically, the extraction of camera motion intensity includes the following three steps.
[0082] The first step is to extract the feature point motion change between the current frame and the previous frame. Feature point motion change refers to the displacement vector of several pixels with significant texture features selected in the previous frame in the current frame. As an optional implementation, several corner points or pixels with significant texture can be detected in the previous frame as feature points. Then, a sparse optical flow algorithm is used to track the corresponding positions of these feature points in the current frame, thereby obtaining the displacement vector of each feature point between two adjacent frames. This displacement vector is the feature point motion change. After obtaining the feature point motion change, due to the motion of the sea surface target itself and the presence of detection noise, the displacement vectors of some feature points may not conform to the overall camera motion pattern. Therefore, a random sampling consensus algorithm can be further used to filter the feature point motion change, removing outliers that are inconsistent with the overall motion pattern and retaining inliers that conform to the overall motion pattern, thereby improving the accuracy of subsequent spatial transformation matrix construction.
[0083] The second step is to construct a spatial transformation matrix based on the changes in feature point motion. A spatial transformation matrix is a mathematical matrix used to describe the overall spatial geometric transformation relationship between two adjacent image frames. Based on the changes in feature point motion after outlier removal, the least squares method is used to solve for the values of each element of the affine transformation matrix, thus obtaining the spatial transformation matrix.
[0084] The third step is to determine the camera motion intensity by calculating the changes in the elements of the spatial transformation matrix. The change in an element refers to the deviation between each element in the spatial transformation matrix and its corresponding element in the identity matrix. When there is no camera motion between two adjacent frames, the spatial transformation matrix degenerates into the identity transformation matrix, at which point the changes in each element are zero. The more intense the camera motion, the greater the deviation of each element in the spatial transformation matrix from its corresponding element in the identity matrix.
[0085] For example, when the marine observation platform is in a stable sailing state, the changes between adjacent frames are minimal, the spatial transformation matrix is close to the unit transformation matrix, and the value of camera motion intensity is close to 0; however, when the marine observation platform encounters large waves and violently rocks, there are obvious translations and rotations between adjacent frames, the elements in the spatial transformation matrix deviate significantly from the corresponding elements of the unit matrix, and the value of camera motion intensity increases significantly.
[0086] After obtaining the camera motion intensity, considering that wave reflection is a common and widespread interference factor in the marine environment, causing large areas of high brightness in the image, severely affecting the detection accuracy and appearance feature extraction of targets within these areas, and consequently impacting the reliability of association matching, it is necessary to further extract the proportion of wave reflection areas to quantify the influence range of this interference factor.
[0087] The extraction of the proportion of reflective area of ocean waves includes the following two steps.
[0088] The first step is to segment the current frame image using a set grayscale threshold to extract the reflective region. The grayscale threshold is the boundary line between reflective and non-reflective pixels. As an optional implementation, the grayscale threshold can be set to 1.5 times the global grayscale mean of the current frame image. This setting adaptively adapts to images under different lighting conditions; the grayscale threshold is increased in brighter scenes and decreased in weaker scenes. After determining the grayscale threshold, each pixel in the current frame image is traversed, and pixels with grayscale values greater than the grayscale threshold are marked as reflective pixels. The area formed by all reflective pixels is the reflective region.
[0089] The second step is to calculate the ratio of the pixel area of the reflective region to the total pixel area of the current frame image, thus obtaining the proportion of the wave reflective region. The value of the wave reflective region proportion ranges from 0 to 1; a larger value indicates that the area affected by wave reflection in the current frame image is more extensive.
[0090] For example, in a sea surface monitoring scenario on a sunny day with a large solar altitude angle, the sunlight reflected from the sea surface forms a large area of high-brightness light spots in the image. At this time, the area of the reflected area extracted after image segmentation is large, and the value of the wave reflected area is high. In a cloudy or nighttime scenario, the sea surface reflection phenomenon is not obvious, the area of the reflected area is very small, and the value of the wave reflected area is close to 0.
[0091] Considering the high real-time requirements of multi-target tracking scenarios at sea, and the typically limited computing resources of maritime observation platforms, risk prediction models need to minimize computational complexity and parameter size while ensuring prediction accuracy. A three-layer cascaded architecture, comprising a feature input layer, a nonlinear hidden layer, and a probability output layer, is adopted. This architecture enables end-to-end mapping from the global scene feature vector to dual-risk probabilities with minimal model parameters. While meeting real-time tracking requirements, it effectively captures the nonlinear correlation between features of various dimensions in the global scene feature vector and tracking risks.
[0092] In this embodiment, the risk prediction model includes a feature input layer, a nonlinear hidden layer, and a probability output layer connected in sequence. The feature input layer receives the global scene feature vector and passes it to subsequent layers. The nonlinear hidden layer performs nonlinear feature transformation on the input data to extract deep correlation patterns between features of different dimensions in the global scene feature vector. The probability output layer converts the feature mapping result output by the nonlinear hidden layer into a continuous risk probability with values limited to a set probability range. The inference process of the risk prediction model is explained step by step below.
[0093] The first step is to input the global scene feature vector into the feature input layer to obtain the input layer feature vector.
[0094] The feature input layer is the first layer in the risk prediction model that directly interfaces with external data. Its function is to receive the global scene feature vector and pass it as the input layer feature vector to the nonlinear hidden layer. The input dimension of the feature input layer is the same as the dimension of the global scene feature vector. When the global scene feature vector is a 6-dimensional vector, the input dimension of the feature input layer is 6. The input layer feature vector is numerically identical to the global scene feature vector, and it serves as the input data for the subsequent nonlinear hidden layer.
[0095] As an optional implementation, before the feature input layer receives the global scene feature vector, the values of each dimension in the global scene feature vector can be normalized and preprocessed to map the values of each dimension to a unified range between 0 and 1, so as to eliminate the influence of the difference in the units of different dimensions on the subsequent layer calculations, thereby enabling the nonlinear hidden layer to perform feature transformation more stably.
[0096] The second step involves inputting the input layer feature vector into a nonlinear hidden layer, which then performs linear transformations and nonlinear activation on the input layer feature vector to obtain the hidden layer feature mapping result.
[0097] After obtaining the input layer feature vector, considering that the mapping relationship between the features of each dimension in the global scene feature vector and the risk of target identity change and trajectory breakage is usually non-linear—for example, the camera motion intensity has a small impact on risk in the low value range but a sharp increase in risk in the high value range—it is difficult to accurately model such non-linear relationships through linear transformation alone. Therefore, it is necessary to perform linear transformation and non-linear activation processing on the input layer feature vector through a non-linear hidden layer to extract the deep non-linear feature patterns contained in the global scene feature vector.
[0098] The nonlinear hidden layer contains several neurons. Each neuron performs a linear transformation on the input layer's feature vector before processing it through a nonlinear activation function. The number of neurons in the nonlinear hidden layer determines its feature representation capacity. As an optional implementation, the nonlinear hidden layer can contain 16 neurons. This number strikes a balance between feature representation capacity and computational efficiency, effectively capturing the cross-correlation information between the dimensions of the 6-dimensional global scene feature vector without causing overfitting or increased inference latency due to too many parameters.
[0099] The process by which a nonlinear hidden layer processes the input layer's feature vector can be represented as follows: ; in, This is the weight matrix. For bias vectors, This represents the feature mapping result of the hidden layer. (Non-linear activation function) Its function is to introduce nonlinear characteristics into the result of linear transformation, enabling the risk prediction model to fit the complex nonlinear mapping relationship between the global scene feature vector and the tracking risk.
[0100] The third step is to input the hidden layer feature mapping results into the probability output layer, which processes the hidden layer feature mapping results and restricts the processing results to a set probability value range through a continuous mapping function, outputting the target identity change risk and trajectory break risk representing continuous probability values.
[0101] After obtaining the hidden layer feature mapping results, considering that the risk of target identity change and trajectory breakage needs to be output in the form of continuous probability values, and the values need to be restricted within the set probability value range to ensure physical rationality, the hidden layer feature mapping results need to be further processed through the probability output layer.
[0102] The probabilistic output layer contains two neurons, corresponding to the output target identity change risk and trajectory breakage risk, respectively. The probabilistic output layer first performs a linear transformation on the hidden layer feature mapping results, and then uses a continuous mapping function to map the result of the linear transformation to a set probability value range.
[0103] Here, a continuous mapping function refers to a monotonically continuous function that can map any real-valued input to a predetermined probability value interval. The predetermined probability value interval can be a closed interval between 0 and 1. As an optional implementation, the continuous mapping function can be the Sigmoid function. The Sigmoid function smoothly maps input values from negative infinity to positive infinity to output values between zero and one, and the output values asymptotically converge at zero and one, making it suitable for representing probabilities. In another optional implementation, other functions with similar mapping properties can also be used as the continuous mapping function; this embodiment does not specifically limit this approach.
[0104] The processing of the probability output layer can be represented as follows: ; in This is the output layer weight matrix. As the output layer bias vector, the Sigmoid function ensures that the output is in the [0,1] interval. This indicates a risk of the target identity changing. This indicates the risk of trajectory breakage. This represents the feature mapping result of the hidden layer.
[0105] when When the value of approaches 1, it indicates that the risk prediction model determines that the probability of an incorrect identity switching occurs during the association matching process of the current frame image is extremely high; when When the value of approaches 0, it indicates that the probability of an incorrect identity switching occurs is extremely low.
[0106] The training process for risk prediction models requires constructing supervisory signals to guide the model parameters. , , and Learning. Training labels for target identity shift risk. The construction method is as follows: ; in, For the first Frame and the -1 is the number of different tracking identifiers corresponding to the same ground truth trajectory in a frame matching pair. Let be the number of true trajectories in the (t-1)th frame. represents the number of true trajectories in frame t, max(·) is the maximum value operation, and min(·) is the minimum value operation, used to limit the label value to between 0 and 1.
[0107] Training labels for trajectory breakage risk The construction method is as follows: ; in, This represents the number of times the true trajectory transitions from a continuous mismatch state to a rematch state, which is the number of times the pseudo-true trajectory transitions from "continuous mismatch (≥2 frames)" to "rematch". denoted as the number of true trajectories in frame t.
[0108] In one optional implementation, the total number of parameters in the risk prediction model can be controlled to within 1,000, and the inference time per frame can be controlled to within 1 ms, which can meet the computational resource constraints of real-time tracking by the marine observation platform.
[0109] For example, when the six dimensions of the global scene feature vector are: camera motion intensity 0.8, average detection box size 1200, detection confidence variance 0.05, small target proportion 0.6, appearance feature clarity 0.3, and wave reflection area proportion 0.4, the global scene feature vector is normalized and input into the feature input layer to obtain the input layer feature vector. After the input layer feature vector undergoes linear transformation and nonlinear activation processing in the nonlinear hidden layer, a 16-dimensional hidden layer feature mapping result is obtained. After the hidden layer feature mapping result is processed by the probability output layer, the output target identity jump risk is 0.72 and trajectory breakage risk is 0.65, indicating that the probability of identity identification error switching and trajectory interruption is relatively high during the association matching process under the current scene conditions.
[0110] Considering that the target identity change risk and trajectory break risk output by the risk prediction model are both continuous probability values between zero and one, directly generating fusion weights based on these continuous probability values through simple linear mapping or threshold judgment can easily lead to drastic weight jumps near the critical point. This results in excessive fluctuations in fusion weights between adjacent frames, thus affecting the stability of the association matching. Therefore, this embodiment introduces a fuzzy logic reasoning mechanism. First, the continuous probability values are converted into discrete fuzzy variables. Then, a preset logic mapping rule base is used to map the level state combinations of the discrete fuzzy variables into fusion weights. This achieves a smooth and continuous transition from risk probability to fusion weights, avoiding the adverse effects of weight oscillations on tracking stability.
[0111] In this embodiment, the process of generating the fusion weights includes the following three steps.
[0112] The first step is to transform the target identity change risk and trajectory break risk with continuous probability values into a first discrete fuzzy variable corresponding to the target identity change risk and a second discrete fuzzy variable corresponding to the trajectory break risk, based on a preset membership function. Both the first discrete fuzzy variable and the second discrete fuzzy variable contain multiple risk level states.
[0113] Since the continuous probability values of target identity change risk and trajectory break risk are inherently ambiguous, meaning that the same probability value may belong to multiple risk levels at the same time, it is necessary to quantify this ambiguity through a membership function so that subsequent logical reasoning can take into account the transitional characteristics of probability values between multiple risk levels, rather than simply defining rigid boundaries.
[0114] Here, the membership function refers to a predefined mathematical function used to map a continuous probability value to its membership value at various risk level states. The membership value ranges from zero to one, representing the degree to which the probability value belongs to a certain risk level state. The first discrete fuzzy variable refers to the set of membership values for the target identity change risk at various risk level states after being transformed by the membership function. The second discrete fuzzy variable refers to the set of membership values for the trajectory break risk at various risk level states after being transformed by the membership function. Risk level states refer to the discrete level categories formed after classifying the degree of risk.
[0115] As an optional implementation, the risk level can be divided into three levels: low risk, medium risk, and high risk. The membership function can be in the form of a triangular membership function, which is simple to calculate and has a smooth transition, making it suitable for deployment on marine observation platforms with limited computing resources. The triangular membership functions corresponding to each risk level are defined as follows.
[0116] The membership function for a low-risk state is: Applicable to ; in, This represents the risk probability value, specifically the value of the risk of the target's identity changing or trajectory breaking. Risk probability value Membership value in a low-risk state. When hour, It decreases monotonically from 1 to 0.
[0117] The membership function for the medium-risk state is: ; in, This represents the risk probability value. Risk probability value The membership value in a medium-risk state. hour, Monotonically increasing from 0 to 1; when hour, It decreases monotonically from 1 to 0.
[0118] The membership function for a high-risk state is: , ; in, This represents the risk probability value. Risk probability value Membership value in a high-risk state. When hour, It monotonically increases from 0 to 1.
[0119] By substituting the target identity change risk into each of the three membership functions mentioned above, the membership values of the target identity change risk in low-risk, medium-risk, and high-risk states are obtained. These three membership values constitute the first discrete fuzzy variable. Similarly, the trajectory breakage risk is substituting into each membership function to obtain the membership values of the trajectory breakage risk in low-risk, medium-risk, and high-risk states. These three membership values constitute the second discrete fuzzy variable.
[0120] In one alternative implementation, the number of risk level states is not limited to three; it can be divided into two, four, or more levels depending on the actual scenario. The membership function is also not limited to a triangular membership function; trapezoidal membership functions, Gaussian membership functions, or other suitable membership function forms can also be used.
[0121] The second step is to obtain logical mapping rules by matching the risk level state of the first discrete fuzzy variable and the risk level state of the second discrete fuzzy variable from the logical mapping rule base.
[0122] After obtaining the first and second discrete fuzzy variables, considering that the target identity change risk and trajectory breakage risk reflect different types of tracking risks, and that different combinations of their state levels correspond to different scenario feature patterns, the optimal contribution ratio of each branch in the multi-branch association cost varies under each scenario feature pattern. Therefore, it is necessary to match the logical mapping rules corresponding to the current state level combination from a pre-built logical mapping rule library based on the level state combinations formed by the risk level states of the first and second discrete fuzzy variables.
[0123] Here, the level-state combination refers to the binary combination formed by the current risk level state of the first discrete fuzzy variable and the current risk level state of the second discrete fuzzy variable. The logical mapping rule base refers to a pre-constructed set of rules containing multiple logical mapping rules. Each logical mapping rule defines the fusion weight allocation scheme that each branch in the multi-branch association cost should obtain under a specific level-state combination condition. The construction of the logical mapping rules is based on prior knowledge of the impact of various disturbances on the reliability of different association clues in a multi-target tracking scenario on the sea surface.
[0124] As an optional implementation, the logical mapping rule base may contain four core logical mapping rules, each with the same priority, covering typical level state combinations of the first discrete fuzzy variable and the second discrete fuzzy variable between high-risk and low-risk states.
[0125] When matching logical mapping rules, the risk level states with the highest membership values for both the first and second discrete fuzzy variables are first determined, and these two are combined to form the level state combination for the current frame image. Then, a logical mapping rule matching this level state combination is searched in the logical mapping rule library. When both the first and second discrete fuzzy variables have non-zero membership values across multiple risk level states, weighted interpolation can be performed based on the membership values of each triggered logical mapping rule, thus achieving a smooth transition between different logical mapping rules.
[0126] The third step is to output the fusion weights corresponding to the multi-branch association costs according to the logical mapping rules.
[0127] After obtaining logical mapping rules from the logical mapping rule base, the fusion weights corresponding to each branch in the multi-branch association cost are output according to the fusion weight allocation scheme defined in the logical mapping rules. The number of output fusion weights is the same as the number of branches in the multi-branch association cost, and there is a one-to-one correspondence between each fusion weight and each branch.
[0128] As an optional implementation, when the multi-branch association cost contains four branches, the output fusion weight is a four-dimensional vector. ,in The fusion weights are the costs associated with motion calibration. The fusion weights are the spatial correlation costs. The fusion weights corresponding to the appearance association cost. The fusion weights are the confidence-related costs, and the sum of all fusion weights is 1.
[0129] When multiple logical mapping rules are triggered simultaneously, the final output fusion weight is the result of weighted average of the fusion weight allocation schemes corresponding to each triggered logical mapping rule according to their respective trigger strengths, thereby achieving a smooth and continuous transition of fusion weight between different level state combinations.
[0130] For example, suppose the target identity jump risk in the current frame image is 0.8 and the trajectory breakage risk is 0.2. After the first step of transformation, the first discrete fuzzy variable has the highest membership value in the high-risk state, and the second discrete fuzzy variable has the highest membership value in the low-risk state. After the second step of matching, the matched logical mapping rule indicates that under the condition that the target identity jump risk is in the high-risk state and the trajectory breakage risk is in the low-risk state, the fusion weight corresponding to the motion calibration association cost should be increased and the fusion weight corresponding to the confidence association cost should be decreased. After the third step of output, the obtained fusion weight is... = Among them, the motion calibration association cost received the highest fusion weight of 0.5, while the confidence association cost received the lowest fusion weight of 0.1.
[0131] In multi-target tracking scenarios at sea, target identity switching risk and trajectory breakage risk reflect two fundamentally different tracking failure modes. The former focuses on the problem of erroneous identity switching, while the latter focuses on the problem of trajectory continuity interruption. Different types of association cost branches play different roles in dealing with these two failure modes. For example, appearance association cost is advantageous in distinguishing adjacent targets with similar appearances, effectively reducing the probability of erroneous identity switching; confidence association cost is advantageous in screening high-reliability detection results, effectively maintaining trajectory continuity.
[0132] Therefore, it is necessary to adjust the fusion weight allocation scheme of each branch association cost in a targeted manner according to the different risk level combinations of target identity change risk and trajectory break risk, so that the most effective association cost branch under the current scenario conditions gets a higher fusion weight, while the fusion weight of the association cost branch with lower reliability under the current scenario conditions is reduced accordingly.
[0133] In this embodiment, multiple risk level states include low-risk and high-risk states. The logical mapping rule base includes the following four logical mapping rules, each with the same priority, covering four typical level state combinations of the first and second discrete fuzzy variables between low-risk and high-risk states. The content and design basis of each logical mapping rule are explained below.
[0134] Logical mapping rule 1: If the first discrete fuzzy variable is in a high-risk state and the second discrete fuzzy variable is in a high-risk state, then increase the fusion weight of appearance association cost and decrease the fusion weight of spatial association cost.
[0135] When the first discrete fuzzy variable is in a high-risk state, it indicates a high probability that the target's identity will be incorrectly switched in the current frame image. When the second discrete fuzzy variable is also in a high-risk state, it indicates a high probability that the existing trajectory in the current frame image will be interrupted. This combination of high-risk levels typically corresponds to complex scenarios in seascapes where there is both severe image jitter and densely intersecting targets. Under these conditions, the reliability of spatial location-based association costs is significantly reduced due to position prediction errors caused by image jitter. If a high fusion weight is still assigned to spatial association costs, it will further exacerbate the problems of incorrect identity switching and trajectory interruption. In contrast, appearance association costs are matched based on the visual appearance features of the target and are not directly affected by spatial location prediction errors. Under these conditions, they can more effectively distinguish the identities of different targets. Therefore, this logical mapping rule increases the fusion weight of appearance association costs to a higher level while decreasing the fusion weight of spatial association costs to a lower level.
[0136] As an optional implementation, the fusion weight allocation scheme corresponding to this logical mapping rule is as follows: The four values correspond to the fusion weights of motion calibration association cost, spatial association cost, appearance association cost, and confidence association cost, respectively. Appearance association cost received the highest fusion weight of 0.5, while spatial association cost received the lowest fusion weight of 0.1.
[0137] Logical mapping rule 2: If the first discrete fuzzy variable is in a high-risk state and the second discrete fuzzy variable is in a low-risk state, then increase the fusion weight of motion calibration association cost and decrease the fusion weight of confidence association cost.
[0138] After obtaining the weight allocation scheme for the dual high-risk scenarios covered by the first logical mapping rule, another typical scenario is considered: when the first discrete fuzzy variable is in a high-risk state and the second discrete fuzzy variable is in a low-risk state, it indicates that the probability of erroneous switching of the target identity in the current frame image is high, but the existing trajectory remains continuous. This combination of levels of state usually corresponds to the situation where the marine observation platform undergoes significant motion but the overall target detection results are stable, such as when the marine observation platform encounters a brief lateral sway during stable navigation. In this situation, the spatial positional offset caused by the motion of the marine observation platform is the main cause of erroneous identity switching. The motion calibration association cost, by performing positional compensation on the bounding box of the prediction result based on camera motion parameters, can effectively eliminate the positional offset error introduced by the motion of the marine observation platform, thereby reducing the probability of erroneous identity switching. At the same time, since the risk of trajectory breakage is low, it indicates that the overall reliability of the detection results is already at a good level, and there is no need to further increase the fusion weight of the confidence association cost. Therefore, this logical mapping rule increases the fusion weight of the motion calibration association cost to a higher level, while decreasing the fusion weight of the confidence association cost to a lower level.
[0139] As an optional implementation, the fusion weight allocation scheme corresponding to this logical mapping rule is as follows: Among them, the motion calibration association cost received the highest fusion weight of 0.5, while the confidence association cost received the lowest fusion weight of 0.1.
[0140] Logical mapping rule three: If the first discrete fuzzy variable is in a low-risk state and the second discrete fuzzy variable is in a high-risk state, then increase the fusion weight of the confidence association cost and decrease the fusion weight of the appearance association cost.
[0141] After obtaining the weight allocation scheme for the level state combinations covered by the first two logical mapping rules, another typical scenario is considered: when the first discrete fuzzy variable is in a low-risk state and the second discrete fuzzy variable is in a high-risk state, it indicates that the target identity in the current frame image is relatively well maintained, but the probability of the existing trajectory being interrupted is high. This level state combination usually corresponds to the situation in a sea scene where heavy fog, rain fog, or low light cause the overall detection confidence to decrease. In this situation, the main reason for the trajectory interruption is that the target detection model's detection results for some targets are not reliable enough, causing some detection results to fail to meet the matching conditions. The confidence association cost, by quantifying the reliability of the detection results themselves, can prioritize the selection of highly reliable detection results for matching, thereby reducing the probability of matching failure caused by low-reliability detection results and effectively maintaining the continuity of the trajectory. At the same time, since the risk of target identity jump is low, it indicates that the distinction between targets in the current scene is already good, and there is no need to rely on the appearance association cost to enhance the identity distinction capability. Therefore, this logical mapping rule increases the fusion weight of the confidence association cost to a higher level, while reducing the fusion weight of the appearance association cost to a lower level.
[0142] As an optional implementation, the fusion weight allocation scheme corresponding to this logical mapping rule is as follows: Among them, the confidence association cost received the highest fusion weight of 0.5, while the appearance association cost received the lowest fusion weight of 0.1.
[0143] Logical mapping rule four: If the first discrete fuzzy variable is in a low-risk state and the second discrete fuzzy variable is in a low-risk state, then increase the fusion weight of the spatial correlation cost.
[0144] After obtaining the weight allocation scheme for the level state combinations covering at least one high-risk state covered by the first three logical mapping rules, considering that when both the first and second discrete fuzzy variables are low-risk states, it indicates that the scene corresponding to the current frame image is generally stable, and the target identity is correctly identified and the trajectory continuity is at a good level. This level state combination usually corresponds to ideal conditions such as clear weather, calm sea surface, and stable movement of the offshore observation platform. Under these conditions, all types of association cost branches have good reliability, and the spatial association cost is calculated based on the spatial position overlap of the bounding box, which has the most direct physical meaning and the lowest computational complexity, making it the most efficient and reliable matching basis in stable scenarios. Therefore, this logical mapping rule raises the fusion weight of the spatial association cost to a higher level, while maintaining the basic fusion weights for other branches.
[0145] As an optional implementation, the fusion weight allocation scheme corresponding to this logical mapping rule is as follows: Among them, the spatial correlation cost received the highest fusion weight of 0.5.
[0146] It should be noted that the specific values of the fusion weights for each branch in the above four logical mapping rules are merely exemplary values. In practical applications, these values can be adjusted according to the specific characteristics of the sea surface monitoring scenario and the tracking performance optimization goals. Furthermore, the number of rules in the logical mapping rule base is not limited to four. Rules covering combinations of medium-risk states and other levels of states can be added as needed to further refine the granularity of the fusion weight adjustment.
[0147] In multi-target tracking scenarios at sea, the matching relationship between target detection results and predicted trajectories can be measured from multiple different information levels. A single feature dimension's association cost is insufficient to maintain reliable discriminative power across all scenarios. For example, when a maritime observation platform experiences severe shaking, the reliability of the spatial location-based association cost decreases; when dense fog occurs, the reliability of the appearance-based association cost decreases. Therefore, it is necessary to construct an association cost branch system covering multiple feature dimensions such as space, appearance, motion, and reliability. This allows the association costs of each branch to complement each other under different scenario conditions, providing rich and diverse matching information sources for subsequent fusion weight adjustment and weighted fusion calculation.
[0148] In this embodiment, the multi-branch association cost includes four branches: spatial association cost, appearance association cost, motion calibration association cost, and confidence association cost. Each branch independently quantifies the degree of matching between the target detection result and the predicted result of the preceding trajectory from four different feature dimensions: spatial location overlap, target appearance consistency, spatial matching degree after motion compensation, and the reliability of the detection result itself. The smaller the value of each branch association cost, the higher the degree of matching between the target detection result and the predicted result of the preceding trajectory under that feature dimension. The calculation method of each branch association cost is explained below.
[0149] ① Calculation of spatial association cost Spatial association cost quantifies the degree of spatial overlap between the bounding box of the detected target and the bounding box of the predicted trajectory. Spatial overlap is the most basic and intuitive metric in target association matching. Its core assumption is that the spatial position of the same target changes only slightly between adjacent frames; therefore, detection and prediction results with high spatial overlap are more likely to correspond to the same target. In scenarios with calm sea surfaces and stable movement of the offshore observation platform, spatial association cost exhibits high reliability and discriminative power.
[0150] As an optional implementation, the spatial association cost is calculated based on the intersection-union ratio (IUR) of the bounding boxes of the target detection results and the bounding boxes of the predicted results of the preceding trajectories. The IUR is the ratio of the intersection area to the union area of two bounding boxes, ranging from zero to one; a higher value indicates a higher degree of spatial overlap between the two bounding boxes. Spatial Association Cost The calculation formula is: ; in, Let be the bounding box of the prediction result of the i-th preceding trajectory. Let the bounding box of the j-th object detection result be... The intersection-union ratio between the two bounding boxes. Let $\frac{i}{j}$ be the spatial association cost between the predicted result of the i-th preceding trajectory and the detected result of the j-th target. When the two bounding boxes completely overlap, the crossover ratio (CRI) is 1 and the spatial association cost is 0, indicating a perfect match; when the two bounding boxes do not overlap at all, the CRI is 0 and the spatial association cost is 1, indicating a complete mismatch.
[0151] ② Calculation of appearance-related costs After obtaining the spatial association cost, considering that multiple targets may be spatially close in a sea surface scene, relying solely on spatial overlap may lead to mismatches between different targets. The appearance association cost establishes a matching relationship by comparing the visual appearance features of the targets, and can effectively distinguish between targets when they are spatially close by utilizing their appearance differences, thus complementing the spatial association cost.
[0152] The appearance association cost is calculated based on the vector similarity between the current appearance features of the target detection result and the historical appearance feature mean of the preceding trajectory. The current appearance feature refers to the high-dimensional feature vector extracted from the image region corresponding to the j-th target detection result in the current frame image by the appearance feature extraction model. The historical appearance feature mean refers to the weighted average of the cumulative updated appearance feature vectors of the i-th preceding trajectory in historical frames.
[0153] As an optional implementation, the update of the historical appearance feature mean can adopt an exponentially weighted moving average mechanism. After each current trajectory successfully matches the target detection result, the current appearance feature is weighted and fused with the historical appearance feature mean according to a preset update coefficient. The update coefficient γ can be set to 0.9, that is, the retention ratio of historical information in the historical appearance feature mean is 0.9, and the integration ratio of new information in the current frame is 0.1.
[0154] Vector similarity can be measured using cosine similarity. Appearance association cost. The calculation formula is: ; in, Let be the mean vector of historical appearance features of the i-th preceding trajectory. Let j be the current appearance feature vector of the j-th target detection result. Let be the appearance association cost between the i-th preceding trajectory and the j-th target detection result. The cosine similarity value ranges from 0 to 1, with a larger value indicating greater similarity between the two appearance feature vectors. The appearance association cost ranges from zero to one, with a smaller value indicating a higher degree of appearance matching.
[0155] ③ Calculation of motion calibration associated costs After obtaining the appearance association cost, considering that in a maritime monitoring scenario, the movement of the maritime observation platform itself will cause a global shift in the positions of all targets in the image, this global shift will introduce systematic errors into the spatial location-based association cost. The motion calibration association cost first performs positional motion compensation based on camera motion parameters on the bounding boxes of the predicted results of the preceding trajectory to eliminate the systematic positional shift introduced by the movement of the maritime observation platform itself. Then, it calculates the spatial overlap between the compensated predicted bounding boxes and the bounding boxes of the target detection results, thereby obtaining more accurate motion dimension matching information.
[0156] Specifically, the calculation of motion calibration association cost consists of two steps. First, positional motion compensation is applied to the bounding box of the predicted result of the preceding trajectory based on the camera motion intensity calibration parameters of the current frame image. The camera motion intensity calibration parameters are the spatial transformation matrix constructed in the global feature extraction step. By transforming the bounding box of the predicted result of the preceding trajectory using the spatial transformation matrix, the compensated predicted bounding box is obtained. This compensated predicted bounding box eliminates the overall positional offset introduced by the motion of the maritime observation platform. Second, the motion calibration association cost is calculated based on the spatial overlap between the compensated predicted bounding box and the bounding box of the target detection result.
[0157] As an optional implementation, motion calibration associated cost The calculation formula is: ; in, Let be the bounding box of the prediction result of the i-th preceding trajectory. The spatial transformation matrix Predicted bounding box after position and motion compensation. Let the bounding box of the j-th object detection result be... The intersection-union ratio (IUU) is the ratio between the compensated predicted bounding box and the bounding box of the target detection result. The adjustment coefficient is used to control the sensitivity of the conversion from crossover ratio to cost value; exp(·) is the natural exponential function. The motion calibration association cost between the predicted result of the i-th preceding trajectory and the detection result of the j-th target. Adjustment coefficient. It can be set to 0.1. The value of motion calibration association cost ranges from 0 to 1. The higher the spatial overlap between the compensated predicted bounding box and the bounding box of the target detection result, the larger the intersection-union ratio, and the smaller the value of motion calibration association cost.
[0158] ④ Calculation of confidence level association cost After obtaining the motion calibration association cost, considering that in harsh sea environments such as rain, fog, and low light, the target detection model may have high uncertainty in the detection results of some targets, and if low-reliability detection results are directly used in association matching, it is easy to lead to incorrect matching or trajectory interruption. The association costs of the three branches mentioned above all focus on the relative matching relationship between the target detection results and the prediction results of the preceding trajectory, without measuring the reliability of the detection results themselves. The confidence association cost, by quantifying the reliability of the target detection results themselves, provides an additional source of information on the detection quality for association matching, and can guide the matching process to prioritize the selection of high-reliability detection results in harsh environments.
[0159] The confidence association cost is calculated based on the self-confidence value of the target detection result and the global detection confidence dispersion of the current frame image. The self-confidence value refers to the degree of certainty of the target detection model that the detection result is a real target, and its value ranges from 0 to 1. The detection confidence dispersion is the detection confidence variance calculated in the global feature extraction step, which is used to correct the impact of fluctuations in global detection quality on the reliability assessment of individual detection results.
[0160] As an optional implementation method, confidence-related costs The calculation formula is: ; in, Let be the self-confidence value of the detection result of the j-th target. The global detection confidence variance for the current frame image. Let $\frac{ ... Convert the confidence score into a cost form; the higher the confidence score, the lower the cost. (Denominator) This is used to correct the dispersion of global detection confidence. When the variance of global detection confidence is large, it indicates that the reliability of each detection result in the current frame image is different. Increasing the denominator value reduces the overall confidence association cost, avoiding overly strict penalties on the reliability of individual detection results in scenarios with large fluctuations in detection quality.
[0161] After calculating the association costs of the four branches mentioned above, the spatial association cost matrix, appearance association cost matrix, motion calibration association cost matrix, and confidence association cost matrix together constitute the multi-branch association cost, providing matching information under four different feature dimensions for the subsequent fusion weighted fusion calculation.
[0162] As an optional implementation, the final association cost between the i-th preceding trajectory and the j-th target detection result in the final association cost matrix is... The fusion formula is: ; in, and These are the fusion weights corresponding to the spatial association cost, appearance association cost, motion calibration association cost, and confidence association cost in the current t-th frame image, respectively.
[0163] In multi-target tracking scenarios at sea, factors such as wave reflection, rain and fog obscuring the surface, and small, distant targets can lead to false detections in target detection models. These false detections typically have low confidence levels, and directly incorporating them into the matching process would increase the probability of incorrect matches, interfere with the correct association of existing trajectories, and potentially cause incorrect switching of target identifiers or abnormal trajectory interruptions. Therefore, before performing the matching operation, target detection results need to be pre-filtered based on a confidence cost threshold to remove low-reliability detections, thereby improving the overall quality of the detection results participating in the matching. After filtering, the optimal matching operation is performed, and new trajectories are initialized for detection results that were not associated during the matching process, thus forming a complete closed loop for matching and trajectory management.
[0164] In this embodiment, the process of obtaining the target matching result includes the following three steps.
[0165] The first step is to remove low-reliability detection results from the target detection results based on the set confidence cost threshold, and obtain the filtered target detection results.
[0166] In maritime surveillance scenarios, target detection models often misclassify non-target areas as targets and assign them low detection confidence when processing image regions containing wave reflections, water surface reflections, or rain / fog noise. Including these low-reliability detections in subsequent matching processes not only increases the size of the final association cost matrix, thus increasing computational overhead, but also introduces a large number of invalid matching candidates, reducing the accuracy of the matching results. Therefore, by setting a confidence cost threshold before matching and using this threshold to eliminate low-reliability detections, false detection interference can be effectively reduced, improving the efficiency and accuracy of subsequent matching operations.
[0167] Here, the confidence cost threshold refers to a pre-set cutoff value for the confidence-related cost used to determine whether a target detection result has sufficient reliability. When the confidence-related cost of a target detection result exceeds this threshold, it indicates that the detection reliability of the target detection result is insufficient and it should be discarded. Low-reliability detection results refer to target detection results whose confidence-related cost exceeds the confidence cost threshold. Filtered target detection results refer to the set of target detection results remaining after removing all low-reliability detection results from the target detection results.
[0168] As an optional implementation, the confidence cost threshold can be set to 0.3. This threshold means that when the confidence association cost of a target detection result is greater than 0.3, the target detection result is considered a low-reliability detection result and is discarded. In one example, assuming the current frame image contains 10 target detection results, after filtering with the confidence cost threshold, 2 target detection results have a confidence association cost exceeding 0.3 and are therefore discarded, leaving 8 target detection results as the filtered target detection results.
[0169] In one alternative implementation, the specific value of the confidence cost threshold can be adjusted according to the false detection rate and tracking performance requirements in the actual application scenario. In harsh sea environments with high false detection rates, the confidence cost threshold can be appropriately reduced to enhance filtering effectiveness; in stable environments with good detection quality, the confidence cost threshold can be appropriately increased to retain more effective detection results.
[0170] The second step is to solve for the optimal match between the filtered target detection results and the prediction results based on the final association cost, and obtain the target matching result.
[0171] After obtaining the filtered target detection results, considering that the filtered target detection results have eliminated low-reliability detection results, the detection results participating in the matching all have high detection reliability. On this basis, performing the optimal matching operation can obtain more accurate matching results.
[0172] This step, based on the column vectors in the final association cost matrix corresponding to the filtered target detection results, solves for the globally optimal one-to-one matching scheme between the filtered target detection results and the predicted results of the preceding trajectories. Specifically, it extracts a submatrix from the final association cost matrix, with row indices corresponding to the predicted results of all preceding trajectories and column indices corresponding to the filtered target detection results. Based on this submatrix, it solves for the one-to-one assignment scheme that minimizes the total global association cost.
[0173] As an alternative implementation, the optimal matching can be solved using the Hungarian algorithm. The Hungarian algorithm is a classic bipartite graph optimal matching algorithm that can find the optimal assignment scheme that minimizes the total cost in polynomial time complexity. During the execution of the Hungarian algorithm, an association cost threshold can be set. If the final association cost between a filtered target detection result and a predicted result of a preceding trajectory exceeds this threshold, even if the two are paired in the Hungarian algorithm's solution, the pairing is considered invalid and split. The association cost threshold can be set to 0.5.
[0174] The target matching results contain the following three types of information: The first type is the set of successfully matched pairs. The first category consists of pairs where a one-to-one correspondence is successfully established between the filtered target detection results and the predicted results of the preceding trajectories, and the final association cost does not exceed the association cost threshold; the second category consists of sets of preceding trajectories that did not successfully match. The first category consists of targets whose preceding trajectories could not be found in the filtered target detection results; the second category consists of the set of filtered target detection results that failed to match. This refers to the filtered target detection result where no associated preceding trajectory was found in the preceding trajectory.
[0175] For example, suppose the filtered target detection results contain 8 detected targets, and the predicted results of the preceding trajectories contain 6 active trajectories. After solving using the Hungarian algorithm, 5 detected targets successfully match 5 preceding trajectories respectively, and the final association cost value of each match does not exceed 0.5. Then, the set of successfully matched pairs contains 5 pairs, the set of unmatched preceding trajectories contains 1 preceding trajectory, and the set of unmatched filtered target detection results contains 3 detected targets.
[0176] The third step is to extract the filtered target detection results that did not match successfully, and initialize a new tracking trajectory based on the extracted filtered target detection results.
[0177] After obtaining the target matching results, considering that the unsuccessfully matched filtered target detection results may correspond to newly appearing targets in the current frame image, such as newly entered vessels or newly surfaced objects, and given that these detection results have already passed the confidence cost threshold filtering in the first step and have high detection reliability, it is necessary to initialize new tracking trajectories for them to ensure that newly appearing targets can be incorporated into the tracking system in a timely manner.
[0178] Specifically, a third type of information is extracted from the target matching results, namely the set of filtered target detection results that failed to match. For each filtered target detection result in the set, a new tracking trajectory is initialized. The initialization process includes: assigning a unique identifier to the new tracking trajectory; using the bounding box coordinates and size information of the filtered target detection result as the initial motion state of the new tracking trajectory; and using the current appearance features of the filtered target detection result as the average of the initial historical appearance features of the new tracking trajectory. After initialization, the new tracking trajectory is added to the trajectory pool for management and participates in association matching as a preceding trajectory in the processing of subsequent frames.
[0179] As an optional implementation, newly initialized tracking trajectories, after being included in the trajectory pool, can be subject to a confirmation period. During this period, the trajectory must successfully match the target detection result a preset number of times consecutively before it can be confirmed as a valid trajectory. If the confirmation conditions are not met within the confirmation period, the trajectory is removed from the trajectory pool. This mechanism can further suppress the generation of false trajectories caused by occasional false detections.
[0180] For example, in the example above, the set of filtered target detection results that did not successfully match. It contains three detection targets. New tracking trajectories are initialized for each of these three targets. Each new trajectory is given a unique identification number and its motion state and appearance features are initialized with its corresponding detection result information. It is then included in the trajectory pool and will participate in the association matching as the preceding trajectory in the processing of the next frame.
[0181] Considering that the target matching results include both successfully matched and unmatched preceding trajectories, different processing operations need to be performed for each case to maintain the accuracy and timeliness of the trajectory state information in the tracking system. For successfully matched preceding trajectories, their motion state and historical appearance features need to be updated using the latest target detection results to accurately reflect the latest state of the target in the current frame. For unmatched preceding trajectories, their positions in the current frame need to be completed using prediction methods to maintain trajectory continuity and avoid premature termination of target tracking due to brief detection omissions. Simultaneously, for preceding trajectories that have failed to match for an extended period, termination conditions need to be set to release system resources and prevent invalid trajectories from interfering with subsequent matching processes.
[0182] In this embodiment, after obtaining the target matching result, the following three processing steps are also included.
[0183] The first step is to update the motion state and historical appearance features of the preceding trajectory based on the corresponding target detection results for a successfully matched preceding trajectory.
[0184] The purpose of updating the state of a successfully matched preceding trajectory is to integrate the latest observation information into the historical record of the preceding trajectory. This allows the motion state and appearance features of the preceding trajectory to continuously track the actual changes in the target, providing accurate basic data for motion state prediction and association cost calculation in the next frame. If this update operation is not performed, the state information recorded in the preceding trajectory will gradually lag behind the actual state of the target, leading to the accumulation of biases in motion state prediction and a decrease in the accuracy of appearance feature matching in subsequent frames.
[0185] Specifically, for each pair in the set M of successfully matched pairs in the target matching results, perform the following update operation.
[0186] The first step is to update the motion state of the preceding trajectory. As an optional implementation, a Kalman filter-based state update step can be used. Specifically, the bounding box position coordinates and size information of the successfully matched target detection results are used as the observations of the current frame. Combined with the predicted state and prediction covariance matrix obtained in the Kalman filter prediction step, the observations and predictions are optimally fused using Kalman gain calculation to obtain the updated motion state vector of the i-th preceding trajectory in the current frame. The updated motion state vector contains information such as position coordinates, motion velocity, and bounding box size, and its accuracy is better than the results obtained using only the predictions or only the observations.
[0187] The second step is to update the historical appearance features of preceding trajectories. As an optional implementation, the update of historical appearance features employs an exponentially weighted moving average mechanism. Specifically, the average historical appearance features of the i-th preceding trajectory are weighted and fused with the current appearance feature vector of the successfully matched target detection result according to a preset update coefficient to obtain the updated average historical appearance features. The update formula is: ; in, Let be the mean vector of historical appearance features of the i-th preceding trajectory. Let j be the current appearance feature vector of the j-th target detection result. These are update coefficients, used to control the fusion ratio between historical information and new information in the current frame. Update coefficients It can be set to 0.9, meaning that 90% of the historical appearance feature information is retained and 10% of the current frame's appearance feature information is incorporated during each update. This update mechanism allows the average historical appearance feature to smoothly follow the gradual change in the target appearance, while maintaining the ability to remember historical appearance information and avoiding drastic changes in the average historical appearance feature due to fluctuations in the appearance of a single frame.
[0188] For example, assuming that the third preceding trajectory successfully matches the fifth target detection result, Kalman filtering is used to fuse the bounding box information of the fifth target detection result with the predicted state of the third preceding trajectory to obtain the updated motion state vector. At the same time, the current appearance feature vector of the fifth target detection result is weighted and fused with the historical appearance feature mean of the third preceding trajectory according to the update coefficient of 0.9 to obtain the updated historical appearance feature mean.
[0189] The second step is to use a Gaussian probability prediction model and combine it with the acceleration estimate of the previous time step to predict and complete the position of the previous trajectory in the current frame for the unmatched previous trajectory.
[0190] After updating the state of a successfully matched preceding trajectory, it's important to consider that there might be unmatched preceding trajectories in the target matching results. These preceding trajectories failed to find associated target detection results in the current frame. The reasons for this unmatch might include the target being briefly obscured by waves, the target detection model missing the target in the current frame, or the target being rejected in the pre-filtering stage due to insufficient confidence. In these cases, the target still exists in the scene; it's just temporarily missed during the detection process in the current frame. If tracking of a preceding trajectory is terminated due to a single frame's detection omission, it will lead to unnecessary trajectory interruption. When the target is re-detected in subsequent frames, a new identity needs to be assigned, causing an incorrect switching of the target's identity. Therefore, for unmatched preceding trajectories, a prediction method is needed to complete their position in the current frame to maintain trajectory continuity.
[0191] Here, the Gaussian probability prediction model refers to a mathematical model that predicts the target's position in the current frame by using motion state information from previous time steps, based on the probability distribution assumption of the target's motion state. The Gaussian probability prediction model assumes that the target's motion state follows a Gaussian distribution over a short period, meaning that changes in the target's position can be approximated using velocity and acceleration. The acceleration estimate from previous time steps refers to the target's acceleration value at a previous time step, calculated based on historical motion state information recorded in the previous trajectory.
[0192] As an optional implementation, the formula for predicting and completing the position of the unmatched preceding trajectory in the current frame is as follows: ; in, The predicted completion position of the i-th unmatched preceding trajectory in the current frame. This represents the position and state of the preceding trajectory in the previous frame. This is the velocity estimate of the preceding trajectory in the previous frame. The acceleration is an estimate from the preceding time step. This is the time interval between adjacent frames. When the video stream's frame rate is 25 frames per second, The value is 0.04 seconds.
[0193] After completing the prediction completion, the prediction completion position will be... The position state of the unmatched preceding trajectory in the current frame is recorded, and this preceding trajectory will continue to participate in the association matching during the processing of the next frame. At the same time, the counter for the consecutive frames in which the preceding trajectory has been unmatched is incremented by one.
[0194] The third step is to terminate the tracking and updating of the preceding trajectory if the number of consecutive frames in which the preceding trajectory is in an unmatched state exceeds the set threshold.
[0195] After performing prediction completion on a failed preceding trajectory, the accuracy of the prediction completion operation gradually decreases with the increase of consecutive unmatched frames, because prediction completion relies entirely on extrapolation of the preceding motion state and cannot obtain new observation information for correction. When a preceding trajectory fails to match for several consecutive frames, on the one hand, the deviation between the predicted position of the preceding trajectory and the actual position of the target may have accumulated to a non-negligible level; on the other hand, the target may have actually left the monitoring screen or been permanently occluded. In the above situation, continuing to maintain the preceding trajectory not only fails to provide effective tracking information but also consumes system resources and increases the computational overhead of correlation matching in subsequent frames. Therefore, it is necessary to set a consecutive frame threshold. When the number of consecutive frames in which a preceding trajectory is in a failed matching state exceeds this threshold, the tracking update of the preceding trajectory is terminated, and it is removed from the trajectory pool.
[0196] Here, the threshold is a pre-determined maximum number of frames that a preceding trajectory is allowed to remain in a state of unsuccessful matching. As an optional implementation, this threshold can be set to 5 frames. That is, when a preceding trajectory fails to match for 5 consecutive frames, it is determined that the target is no longer within the current monitoring range or can no longer be effectively tracked, and the tracking and updating of that preceding trajectory is terminated.
[0197] In one alternative implementation, the threshold can be adjusted based on the frame rate of the video stream and the target motion characteristics. In high-frame-rate video streams, because the time interval between adjacent frames is short, brief occlusions typically only affect a few frames, and the threshold can be set to a smaller value. In low-frame-rate video streams or scenarios where the target motion is slow, the threshold can be appropriately increased to tolerate longer periods of missed detection.
[0198] The specific operations for terminating tracking updates include: marking the preceding trajectory as terminated, ceasing motion state prediction and appearance feature update operations on it, and removing it from the pool of active trajectories participating in association matching. Historical information of the terminated preceding trajectory can be retained in the system's history for subsequent tracking result statistics and analysis.
[0199] In multi-target tracking scenarios at sea, the cameras on maritime observation platforms typically possess high imaging resolution; for example, a single frame of a 4K image can have over 8 million pixels. In such high-resolution images, distant sea targets occupy a very small pixel area. If the entire high-resolution image is directly input into a target detection model, the model usually needs to scale the input image to a fixed, lower resolution to meet the model's input size requirements. This scaling operation further compresses or even loses pixel information for small targets, severely reducing the model's sensitivity and accuracy in detecting and locating them. Therefore, it is necessary to segment the high-resolution image before target detection, dividing it into multiple smaller sub-images and feeding each sub-image into the target detection model. This allows for effective detection of targets across the entire image while maintaining the integrity of the target pixel information in each sub-image.
[0200] In this embodiment, before extracting multiple global features of the current frame image, the following three steps are also included.
[0201] The first step is to acquire the original video stream of the sea surface observation and extract single-frame high-resolution images from the video stream.
[0202] Acquiring the raw video stream of sea surface observations is the initial step in the multi-target sea surface tracking process, and the video stream is the data source for all subsequent processing steps. The raw video stream of sea surface observations refers to the continuous sequence of images captured by the camera equipment mounted on the marine observation platform during sea surface monitoring. The marine observation platform may include drones, unmanned surface vessels, or shore-based monitoring stations, and the camera equipment may include visible light cameras, infrared cameras, or multispectral cameras. This embodiment does not limit the specific types of marine observation platforms and camera equipment.
[0203] As an optional implementation, the frame rate of the video stream can be set to 25 frames per second, meaning the camera captures 25 frames per second. After acquiring the video stream, each frame is extracted sequentially from the video stream according to its frame number. The frame currently being processed is the single-frame high-resolution image. The resolution of the single-frame high-resolution image is determined by the imaging parameters of the camera device and can be 4K resolution, 1080P resolution, or other resolutions. Its single-frame pixel size is denoted as W (width pixels) × H (height pixels).
[0204] In one alternative implementation, when the video stream resolution is 1080P or lower, the proportion of the pixel area occupied by the target in a single frame image to the total image area is relatively large. In this case, subsequent segmentation operations can be omitted, and the single frame image can be directly input into the target detection model for processing. When the video stream resolution is 4K or higher, the need for detecting small-sized targets becomes more prominent, and subsequent segmentation operations are required.
[0205] The second step is to split the single-frame high-resolution image into multiple sub-images with a set overlap rate.
[0206] After extracting a single high-resolution image from the video stream, directly scaling the entire high-resolution image to the fixed input size of the object detection model would result in severe loss of pixel information for small targets. Furthermore, if a simple, uniform, non-overlapping segmentation method is used, targets at the segmentation boundaries would be truncated into two incomplete segments, causing the object detection model to fail to correctly identify the target or produce duplicate, incomplete detection results. Therefore, a segmentation strategy with a set overlap rate is required, ensuring a certain width of overlap between adjacent sub-images and guaranteeing that targets at the segmentation boundaries are completely contained within at least one sub-image.
[0207] Here, the overlap rate is defined as the ratio of the area of the overlapping region between two adjacent sub-images to the area of a single sub-image. The higher the overlap rate, the higher the probability that the target located at the cutting boundary is completely included, but it will also increase the total number of sub-images, thus increasing the computational cost.
[0208] As an optional implementation, a single-frame high-resolution image can be split into sliding window sub-images according to a preset sub-image size. The sub-image size can be set to 1280×1280, which matches the standard input size of the object detection model, allowing the sub-images to be directly input into the object detection model without additional scaling. The overlap rate can be set to 10%, meaning that adjacent sub-images overlap by 128 pixels in both the horizontal and vertical directions.
[0209] Specifically, using a sub-image size of 1280×1280 and an overlap width of 128 pixels as parameters, starting from the top left corner of a single-frame high-resolution image, a sliding window-style cropping process is performed along the horizontal and vertical directions with a step size of 1280-128=1152 pixels, resulting in a 1280×1280 sub-image each time. When the sliding window reaches the right or bottom boundary of the single-frame high-resolution image, the last column or last row of sub-images can maintain their complete sub-image size by moving back to the left or up, or by padding to fill in any missing parts. After the above splitting operation, the single-frame high-resolution image is divided into multiple sub-images, each of which covers the entire area of the single-frame high-resolution image, and adjacent sub-images have overlapping areas with a set overlap rate.
[0210] In one alternative implementation, the specific value of the overlap rate can be adjusted based on the typical size of the target and the probability of the target being truncated at the segment boundary. When small targets are relatively dense in the sea scene, the overlap rate can be appropriately increased to reduce the risk of missing boundary targets; when the target size is large and the density is low, the overlap rate can be appropriately decreased to reduce computational overhead.
[0211] The third step is to input multiple sub-images into the object detection model to obtain the object detection results.
[0212] After obtaining multiple sub-images with a set overlap rate, considering that the resolution of each sub-image matches the standard input size of the object detection model, the multiple sub-images can be input into the object detection model one by one or in parallel for object detection processing.
[0213] Here, the object detection model refers to a pre-trained deep learning model capable of locating and classifying objects in an input image. After inferring for each sub-image, the object detection model outputs information such as the bounding box coordinates, bounding box width and height, and detection confidence of each detected object in that sub-image.
[0214] As an alternative implementation, the target detection model can employ a single-stage target detection network fine-tuned using a sea surface target dataset. This network takes a sub-image as input and directly outputs the detection information of each target in that sub-image, exhibiting high inference speed and making it suitable for sea surface monitoring scenarios with high real-time requirements.
[0215] After outputting the detection results for each sub-image, the detection results of each sub-image need to be coordinate-mapped according to the spatial position relationship of each sub-image in the single-frame high-resolution image. This transforms the bounding box coordinates of the targets in each sub-image from the local coordinate system of the sub-image to the global coordinate system of the single-frame high-resolution image. Since there are overlapping areas between adjacent sub-images, the same target may be detected repeatedly in multiple sub-images. Therefore, after coordinate mapping, a non-maximum suppression operation needs to be performed to merge and remove duplicate detection results in the overlapping areas. After coordinate mapping and non-maximum suppression, a complete target detection result within the single-frame high-resolution image is obtained. The information for each detected target in this result includes the bounding box center coordinates, bounding box width and height, and detection confidence score in the global coordinate system, denoted as . ,in and Let be the coordinates of the center of the bounding box of the i-th detected target. and For the width and height of the bounding box, To test the confidence level.
[0216] The target detection result is the target detection result of the current frame image used in subsequent steps to extract global features, calculate association costs, and perform association matching.
[0217] Based on any of the above embodiments, the specific implementation methods for obtaining the prediction results of the preceding trajectory will be described in detail below.
[0218] In multi-target tracking on the sea surface, to calculate the multi-branch association cost between the target detection result of the current frame and the previous trajectory, it is necessary to predict the expected state of the previous trajectory in the current frame beforehand. Since sea surface targets usually maintain a certain motion inertia between adjacent frames, introducing a kinematic prediction model to predict the historical motion state of the previous trajectory can effectively narrow the target search and matching range and provide an accurate spatial position comparison benchmark for the subsequent calculation of spatial association cost and motion calibration association cost.
[0219] In this embodiment, the position of the preceding trajectory in the current frame image is predicted using Kalman filtering. Kalman filtering is an algorithm based on the state equation of a linear system that optimally estimates the system state using system input and output observation data, effectively filtering observation noise caused by the sea surface environment.
[0220] Specifically, the process of predicting the preceding trajectory using Kalman filtering includes setting a state vector and performing iterative calculations using the state equation and the observation equation.
[0221] First, construct the state vector of the preceding trajectory. As an optional implementation, the state vector of frame t is denoted as... It contains kinematic information such as the target's center position, velocity, and bounding box size, specifically represented as: ; in, and Let x and y be the x and y coordinates of the center point of the bounding box of the preceding trajectory in frame t, respectively. and These represent the horizontal and vertical velocities of the preceding trajectory at frame t, respectively. and Let be the width and height of the bounding box of the preceding trajectory in frame t, respectively, and let T denote the transpose of the matrix.
[0222] Secondly, the state of the preceding trajectory is forward- deduced based on the Kalman filter state equation. The Kalman filter state equation describes the evolution of the motion state of the preceding trajectory from frame t-1 to frame t, and its calculation formula is as follows: ; in, Let be the state vector of frame t. Let A be the state vector of the preceding trajectory in frame t-1 (i.e., the previous frame), and let B be the state transition matrix and B be the control matrix. To control the quantity, This is process noise.
[0223] Meanwhile, the Kalman filter observation equation is used to describe the relationship between the system state and the actual observation data, and its calculation formula is as follows: ; in, Let be the observation vector of frame t, which represents the position and size of the bounding box actually output by the object detection model. This is the observation matrix, used to map the state space to the observation space. For observation noise, it is used to characterize the detection and measurement errors introduced by the camera equipment itself or by changes in sea surface illumination.
[0224] In the above equations, process noise With observation noise All are assumed to follow a Gaussian distribution with a mean of zero, specifically expressed as follows: and , where Q is the process noise covariance matrix and R is the observation noise covariance matrix.
[0225] After obtaining the state vector of each preceding trajectory predicted by Kalman filtering, the position coordinates and size information in the state vector are extracted to obtain the predicted bounding box of the preceding trajectory in the current frame. For all m active preceding trajectories, the final output trajectory prediction result set is denoted as . It is represented as: ; in, to These represent the predicted bounding boxes of the first to m preceding trajectories in frame t. This set of trajectory prediction results serves as the input basis for subsequent multi-branch association cost matching calculations.
[0226] Figure 2 This is a schematic diagram of the overall framework of the adaptive association method for multi-target tracking on the sea surface provided in this embodiment of the invention, as shown below. Figure 2 As shown, the framework's operation flow is as follows: First, the system receives the input image frame sequence. For the current frame image, it extracts global features that can characterize multiple dimensions such as "camera motion", "target size", "appearance similarity" and "confidence", such as camera motion intensity, average detection box size, detection confidence variance, and appearance feature clarity. Then, it combines the multiple global features through scene analysis to construct a global scene feature vector.
[0227] Then, the global scene feature vector is input into the pre-trained risk prediction model to obtain the target identity change risk and trajectory break risk directly output by the risk prediction model.
[0228] Next, based on the predicted target identity change risk and trajectory break risk, fuzzy logic weight mapping is performed to dynamically generate the fusion weights corresponding to the multi-branch association costs.
[0229] Subsequently, the target detection result and the prediction result of the previous trajectory of the current frame image are obtained, and multi-clue parallel computing is performed to calculate the association cost of the target detection result and the prediction result under multiple different feature dimensions, so as to obtain the multi-branch association cost. The multi-branch association cost specifically includes spatial association cost (IoU cost), appearance association cost (ReID cost), motion calibration association cost (OCM cost), and confidence association cost (Conf cost).
[0230] Finally, the multi-branch association cost is weighted and fused according to the dynamically generated fusion weight to construct the final association cost matrix. Based on the final association cost matrix, the target detection result and the prediction result are optimally matched, and the target matching result is finally output, thereby effectively realizing stable multi-target tracking in complex marine environments.
[0231] Based on any of the above embodiments, in order to further verify the effectiveness and advantages of the adaptive association method for multi-target tracking on the sea surface provided by the present invention, the following detailed description is provided in conjunction with experimental data and comparison results of three specific publicly available datasets, and the corresponding experimental data tables are retained.
[0232] In this embodiment, the method provided by the present invention is compared with relevant mainstream multi-target tracking algorithms. The baseline algorithms for comparison include SORT, DeepSORT, ByteTrack, OC-SORT, and HybridSORT. The three typical complex scene datasets selected for the experiment are: the Marine Obstacle Detection Dataset (MODS) for severe near-shore shaking scenes in ports, the SeaDronesSee dataset for wide-format small-target aerial photography of the sea surface, and the USVTrack dataset for rain, fog, and nighttime scenes in inland waterways. To ensure the fairness of the comparison, all experimental results are evaluated based on a unified YOLOv11m target detection model and tracking evaluation tool (TrackingEvaluation, TrackEval).
[0233] In the comparative analysis of the experimental results, the core tracking metrics used included: Higher Order Tracking Accuracy (HOTA), Detection Accuracy (DetA), Association Accuracy (AssA), Detection Recall (DetRe), Detection Precision (DetPr), Association Recall (AssRe), Association Precision (AssPr), and Localization Accuracy (LocA).
[0234] First, we analyze the experimental results on the MODS dataset, which depicts severe shaking in near-shore port environments. The core challenge of this dataset lies in the severe shaking of the ship's hull, which easily leads to distortion of target motion characteristics and missed detection of small targets. The experimental results of the method provided in this invention and various baseline algorithms on this dataset are shown in Table 1.
[0235] Table 1 Comparison of Experimental Results for MODS Dataset Experimental results show that the method provided in this invention achieves a HOTA index of 15.937, a 12.8% improvement compared to the baseline method HybridSORT; and an AssA index of 27.129, a 33.0% improvement compared to the baseline method OC-SORT. This significant improvement is primarily attributed to the fact that, in scenarios with severe shaking, the method of this invention effectively reduces misclassifications caused by image shake through the dynamic synergy of confidence association cost and motion calibration association cost. Specifically, while maintaining the basic fusion weight of the confidence association cost (e.g., 0.15) to filter high-confidence target detection results, the fusion weight corresponding to the motion calibration association cost is dynamically increased (e.g., increased to 1.2 times the basic value) to compensate for motion distortion, thereby effectively reducing misclassifications caused by image shake. Furthermore, the method provided in this invention achieves a DetRe index of 12.830, a 32.2% improvement compared to HybridSORT, fully demonstrating the good suppression effect of introducing the confidence association cost branch on the problem of missed detection of small targets.
[0236] Secondly, the experimental results of the SeaDronesSee dataset, a wide-frame aerial photography dataset depicting small targets on the sea surface, are analyzed. This dataset features single-frame images with a resolution of 4K wideframe, and the distribution of small targets is extremely dense, requiring very high positioning accuracy. In this experiment, the method provided in this invention employs a single-frame high-resolution image segmentation strategy, that is, splitting the original image into multiple sub-images with a uniform resolution of 1280 pixels by 1280 pixels. The experimental results of the method provided in this invention and various baseline algorithms on this dataset are shown in Table 2.
[0237] Table 2 Comparison of Experimental Results for the SeaDronesSee Dataset Experimental results show that the method provided in this invention achieves a HOTA index of 65.126, a 4.5% improvement over the DeepSORT algorithm; and a LocA index of 86.217, a 2.8% improvement over the HybridSORT algorithm. In this scenario, by increasing the fusion weight corresponding to the confidence association cost (e.g., increasing its weight to 1.3 times the base value in small target clustering scenarios), the filtering of high-confidence small targets is strengthened, effectively improving the localization accuracy. Simultaneously, by increasing the fusion weight of the appearance association cost (e.g., increasing it to 1.4 times the base value), the appearance differentiation between targets is strengthened, effectively solving the misjudgment problem caused by the spatial overlap of small targets. Furthermore, the method provided in this invention achieves a DetPr index of 85.852, a 7.0% improvement over HybridSORT, further verifying the optimization effect of the confidence association cost on detection accuracy.
[0238] Finally, the experimental results on the USVTrack dataset, depicting rainy, foggy, and nighttime scenes in inland waterways, are analyzed. This dataset includes harsh low-light environments such as rain, fog, and nighttime, which can easily cause drastic fluctuations in the detection confidence of target detection models and severe distortion of target appearance features. The experimental results of the method provided in this invention and various baseline algorithms on this dataset are shown in Table 3.
[0239] Table 3 Comparison of experimental results for the USVTrack dataset Experimental results show that the method provided in this invention achieves a HOTA index of 71.412, which is basically on par with the optimal baseline algorithm OC-SORT, but achieves an AssA index of 63.018, an improvement of 0.6% over OC-SORT. More importantly, in nighttime and rain / fog scenarios, the method provided in this invention reduces the number of target identity jumps by 11.2% compared to OC-SORT. This is mainly attributed to the fact that in scenarios with clustered low-confidence detection results, this embodiment dynamically increases the fusion weight corresponding to the confidence association cost (e.g., to 1.8), thereby effectively filtering out high-quality target detection results and greatly suppressing false detection interference caused by rain and fog. At the same time, the method achieves an AssRe index of 65.923, an improvement of 2.9% over OC-SORT, which fully demonstrates that by dynamically weighting and fusing the confidence association cost with the motion calibration association cost and the appearance association cost, the overall target association matching recall rate is significantly optimized.
[0240] The adaptive association device for multi-target tracking on the sea surface provided by the present invention will be described below. The adaptive association device for multi-target tracking on the sea surface described below can be referred to in correspondence with the adaptive association method for multi-target tracking on the sea surface described above.
[0241] Figure 3 This is a schematic diagram of the adaptive correlation device for multi-target tracking on the sea surface provided by the present invention, as shown below. Figure 3 As shown, the device includes: The extraction module 310 is used to extract multiple global features of the current frame image and combine the multiple global features to construct a global scene feature vector; Prediction module 320 is used to input the global scene feature vector into the pre-trained risk prediction model to obtain the target identity change risk and trajectory break risk output by the risk prediction model; The calculation module 330 is used to obtain the target detection result and the prediction result of the previous trajectory of the current frame image, and calculate the association cost of the target detection result and the prediction result under multiple different feature dimensions to obtain the multi-branch association cost. The generation module 340 is used to generate fusion weights corresponding to multi-branch association costs based on the target identity change risk and trajectory breakage risk. The fusion module 350 is used to perform weighted fusion calculation on the multi-branch association cost according to the fusion weight to obtain the final association cost; The matching module 360 is used to match the target detection results with the prediction results based on the final association cost to obtain the target matching results.
[0242] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440. The processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions from the memory 430 to execute an adaptive association method for multi-target tracking on the sea surface.
[0243] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0244] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the adaptive association method for multi-target tracking on the sea surface provided by the above methods.
[0245] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the adaptive association method for multi-target tracking on the sea surface provided by the methods described above.
[0246] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0247] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0248] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive association method for multi-target tracking on the sea surface, characterized in that, include: Extract multiple global features from the current frame image, and combine the multiple global features to construct a global scene feature vector; The global scene feature vector is input into a pre-trained risk prediction model to obtain the target identity change risk and trajectory break risk output by the risk prediction model. Obtain the target detection result and the prediction result of the previous trajectory of the current frame image, and calculate the association cost of the target detection result and the prediction result under multiple different feature dimensions to obtain the multi-branch association cost; Based on the target identity change risk and the trajectory breakage risk, generate the fusion weight corresponding to the multi-branch association cost; The multi-branch association cost is weighted and fused according to the fusion weight to obtain the final association cost; Based on the final association cost, the target detection result is matched with the prediction result to obtain the target matching result; The extraction of multiple global features from the current frame image includes: The camera motion intensity, average detection box scale, detection confidence variance, small target proportion, appearance feature sharpness, and wave reflection area proportion of the current frame image are extracted as multiple global features.
2. The adaptive association method for sea surface multi-target tracking according to claim 1, characterized in that, The steps for extracting the camera motion intensity and the proportion of the wave reflection area include: Extract the motion change of feature points between the current frame image and the previous frame image, construct a spatial transformation matrix based on the motion change of feature points, and determine the camera motion intensity by calculating the change of elements in the spatial transformation matrix. The current frame image is segmented by setting a grayscale threshold, the reflective area is extracted, and the ratio of the pixel area of the reflective area to the total pixel area of the current frame image is calculated to obtain the proportion of the reflective area of the waves.
3. The adaptive association method for multi-target tracking over the sea surface according to any one of claims 1 to 2, characterized in that, The risk prediction model includes a feature input layer, a nonlinear hidden layer, and a probability output layer connected in sequence. The step of inputting the global scene feature vector into a pre-trained risk prediction model to obtain the target identity change risk and trajectory break risk output by the risk prediction model includes: The global scene feature vector is input into the feature input layer to obtain the input layer feature vector; The input layer feature vector is input into the nonlinear hidden layer, and the nonlinear hidden layer performs linear transformation and nonlinear activation processing on the input layer feature vector to obtain the hidden layer feature mapping result. The hidden layer feature mapping result is input to the probability output layer, which processes the hidden layer feature mapping result and restricts the processing result to a set probability value range through a continuous mapping function, outputting the target identity change risk and the trajectory break risk representing continuous probability values.
4. The adaptive association method for multi-target tracking over the sea surface according to any one of claims 1 to 2, characterized in that, The step of generating the fusion weights corresponding to the multi-branch association cost based on the target identity change risk and the trajectory breakage risk includes: Based on a preset membership function, the target identity change risk and the trajectory break risk with continuous probability values are respectively transformed into a first discrete fuzzy variable corresponding to the target identity change risk and a second discrete fuzzy variable corresponding to the trajectory break risk. Both the first discrete fuzzy variable and the second discrete fuzzy variable contain multiple risk level states. Based on the risk level state combination formed by the risk level state of the first discrete fuzzy variable and the risk level state of the second discrete fuzzy variable, a logical mapping rule is obtained by matching from the logical mapping rule base. Based on the logical mapping rules, output the fusion weights corresponding to the multi-branch association costs.
5. The adaptive association method for multi-target tracking on the sea surface according to claim 4, characterized in that, The multiple risk level states include low-risk states and high-risk states; The logical mapping rule base includes: If the first discrete fuzzy variable is the high-risk state and the second discrete fuzzy variable is the high-risk state, then increase the fusion weight of appearance association cost and decrease the fusion weight of spatial association cost. If the first discrete fuzzy variable is the high-risk state and the second discrete fuzzy variable is the low-risk state, then increase the fusion weight of motion calibration association cost and decrease the fusion weight of confidence association cost. If the first discrete fuzzy variable is the low-risk state and the second discrete fuzzy variable is the high-risk state, then increase the fusion weight of the confidence association cost and decrease the fusion weight of the appearance association cost. If both the first discrete fuzzy variable and the second discrete fuzzy variable represent the low-risk state, then the fusion weight of the spatial correlation cost is increased.
6. The adaptive association method for multi-target tracking on the sea surface according to any one of claims 1 to 2, characterized in that, The multi-branch association cost includes spatial association cost, appearance association cost, motion calibration association cost, and confidence association cost; The calculation of the association cost between the target detection result and the prediction result under multiple different feature dimensions yields a multi-branch association cost, including: The spatial association cost is calculated based on the spatial overlap between the bounding box of the target detection result and the bounding box of the prediction result. The appearance association cost is calculated based on the vector similarity between the current appearance features of the target detection result and the average historical appearance features of the preceding trajectory. The bounding box of the prediction result is compensated for positional motion based on the camera motion intensity calibration parameters of the current frame image, and the motion calibration association cost is calculated based on the spatial overlap between the compensated predicted bounding box and the bounding box of the target detection result. Based on the self-confidence value of the target detection result and the global detection confidence dispersion of the current frame image, the confidence association cost used to quantify the reliability of the detection result is calculated.
7. The adaptive association method for multi-target tracking on the sea surface according to any one of claims 1 to 2, characterized in that, The step of matching the target detection result with the prediction result based on the final association cost to obtain the target matching result includes: Based on the set confidence cost threshold, low-reliability detection results in the target detection results are removed to obtain filtered target detection results; Based on the final association cost, the optimal match between the filtered target detection result and the prediction result is solved to obtain the target matching result; Extract the filtered target detection results that did not match successfully, and initialize a new tracking trajectory based on the extracted filtered target detection results.
8. The adaptive association method for multi-target tracking on the sea surface according to any one of claims 1 to 2, characterized in that, After matching the target detection result with the prediction result based on the final association cost to obtain the target matching result, the method further includes: For a successfully matched preceding trajectory, the motion state and historical appearance features of the preceding trajectory are updated based on the corresponding associated target detection results. For unmatched preceding trajectories, a Gaussian probability prediction model is used in conjunction with the acceleration estimate of the preceding time step to predict and complete the position of the preceding trajectory in the current frame. If the number of consecutive frames in which the preceding trajectory is in a failed matching state exceeds a set threshold, the tracking and updating of the preceding trajectory will be terminated.
9. The adaptive association method for multi-target tracking on the sea surface according to any one of claims 1 to 2, characterized in that, Before extracting multiple global features from the current frame image, the process also includes: Acquire the video stream of the original sea surface observation and extract a single high-resolution image from the video stream; The single-frame high-resolution image is split into multiple sub-images with a set overlap rate; The multiple sub-images are input into the target detection model respectively to obtain the target detection result.