Underwater multi-target tracking method based on HDS-mamba and adaptive variational bayes
Patent Information
- Application Number
- CN202610778580.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-25
AI Technical Summary
基于交互多模型(IMM)的方法虽可处理多模型切换,但模型集合需人工预设且无法覆盖所有机动模式
1、本申请通过构建HDS-Mamba网络,利用观测重构网络专门处理声呐量测域的非线性坐标变换与阵列畸变,利用状态残差网络专门处理目标动力学域的运动模型失配,两个网络处理不同物理量,通过物理引导与数据驱动的深度结合,从量测域和状态域两个维度同时修正跟踪误差,从而在目标发生复杂机动时仍能保持高精度的状态估计,显著提升了强机动与恶劣噪声耦合下的跟踪鲁棒性。
Smart Images

Figure CN122815437A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field, specifically to an active sonar underwater multi-target tracking method based on HDS-Mamba and adaptive variational Bayes. Background Technology
[0002] Underwater environmental perception is a core component of maritime security and defense as well as underwater resource exploration. Due to the severe attenuation of visible light, active sonar detection is the primary means of underwater detection. Although multi-target tracking technology has made significant progress in radar and vision fields, it faces many key challenges in underwater sonar applications.
[0003] First, underwater targets are susceptible to ocean currents and maneuvers, and their motion often undergoes nonlinear abrupt changes, leading to mismatches in prior physical models such as constant velocity and acceleration. Furthermore, the inherent nonlinear coordinate transformations and array distortions of active underwater sonar detection make it difficult to explicitly and analytically express measurement models. Therefore, addressing the uncertainties in system dynamics and observation models during active underwater tracking is the primary challenge for achieving robust tracking. Simultaneously, influenced by multipath effects, reverberation, and unknown interference in the underwater acoustic channel, the statistical characteristics of measurement noise are often unknown and change drastically over time. Traditional methods based on fixed prior measurement noise often lead to subsequent tracking inaccuracies or even divergence. Therefore, handling non-stationary measurement noise during active underwater tracking is another crucial issue.
[0004] In existing technologies, various attempts have been made to address the aforementioned problems. For example, the traditional Extended Kalman Filter (EKF) method uses a first-order Taylor expansion to linearize the nonlinear system, but its accuracy drops sharply under conditions of strong nonlinearity and model mismatch. While Interactive Multiple Model (IMM) methods can handle multi-model switching, the model set needs to be manually pre-set and cannot cover all maneuvering modes. In recent years, deep learning-based target tracking methods (such as Transformer-based visual trackers) have made significant progress in computer vision, but these methods are mainly designed for RGB video frames, relying on rich appearance and texture features, making them difficult to directly transfer to sonar polar coordinate measurement scenarios. Moreover, while existing variational Bayesian adaptive filtering methods can estimate noise parameters online, they typically employ no-information priors or fixed empirical priors, resulting in slow convergence or even divergence when noise changes drastically.
[0005] Therefore, there is an urgent need for an underwater multi-target tracking method that can improve the tracking robustness of complex underwater sonar maneuvering targets. Summary of the Invention
[0006] The purpose of this application is to address the shortcomings of the existing technology and provide an active sonar underwater multi-target tracking method based on HDS-Mamba and adaptive variational Bayes.
[0007] An active sonar underwater multi-target tracking method based on HDS-Mamba and adaptive variational Bayes includes: Step S1: Construct an HDS-Mamba network, which includes an observation reconstruction network and a state residual network; The observation reconstruction network is used to receive measurement sequences and output reconstructed measurements in Cartesian space and their corresponding heteroscedastic measurement noise covariance matrix. The state residual network is used to receive the sliding history posterior state sequence and output the dynamic residual correction and the corresponding process noise covariance matrix. Step S2: Extract the sliding history posterior state sequence corresponding to each track from the maintained set of active tracks. The state residual network is then input into the state residual network, and based on the state residual network, a dynamic residual correction for the trajectory is output. and the predicted process noise covariance matrix The maintained active track set is the set of tracks that have been successfully matched with detection boxes and pre-stored by the system; the sliding history is a posterior state sequence. This is a historical posterior state sequence with a window length of L extracted from the historical state cache sequence of each active track; Step S3: Utilize the aforementioned dynamic residual correction amount Perform a Kalman filter one-step prediction to obtain the prior predicted state of each track at time k. ; Step S4: Extract the historical polar coordinate sequence corresponding to the detection box matched by each track from the maintained active track set, and append the polar coordinate measurement matched at the current time k to the end of the sequence. A sliding measurement sequence with a window length of L is constructed and input into the observation reconstruction network. Based on the observation reconstruction network, the reconstructed measurement of the Cartesian space at time k is output. and its corresponding heteroscedasticity measurement noise covariance matrix ; Step S5: Measure the noise covariance matrix using the heteroscedasticity. Dynamically anchor the inverse Wishart prior distribution and initialize the degrees of freedom parameters for Bayesian inference. and inverse scaling matrix ;in, The hyperparameter d is used to adjust the prior confidence of the network, where d is the dimension of the measurement vector. ; Step S6: Under the IW prior constraints using the aforementioned degree-of-freedom parameters and inverse scaling matrix, first... and reconstruction measurement Through iterative updates, the optimal state estimate at time k is finally obtained. and the true noise covariance matrix ; Step S7: Estimate the optimal state at time k. Store the historical posterior state sequence of the corresponding track and update the tracking trajectory of the track; finally, output the tracking trajectories of all active tracks to complete the underwater multi-target tracking at the current moment.
[0008] Furthermore, as described above, the maintained set of active tracks is obtained as follows: Step 1: Acquire continuous sonar images at time k using active sonar detection, and extract the set of detection boxes from the sonar images using a target detection head. And based on the confidence threshold, the detection box set is... Divided into high-resolution detection sets and low-resolution detection set ; Step 2: Based on the process noise covariance matrix Calculate the prediction error covariance matrix of trajectory j at time k. ; Step 3: Based on the prior predicted state Prediction error covariance matrix and the original measurement sequence The i-th measurement Calculate the Mahalanobis distance between the active track j and the target i, and construct the cost matrix accordingly. ; The i-th measurement By collecting the detection boxes The result is obtained by projecting onto a Cartesian coordinate system; Step 4: Utilize the constructed cost matrix High-resolution detection set With active flight paths Perform the first round of Hungarian matching and output the first set of successfully matched tracks. Subsequently, the unmatched tracks from the first round were combined with the low-resolution detection set. Perform a second matching and output the second set of successfully matched tracks. ; set the first track With the second set of tracks Merge and finally output the set of successfully matched tracks. and the set of successfully matched tracks As a set of active tracks maintained by the system.
[0009] Furthermore, in the method described above, the cost matrix in step 3... The construction includes: The squared Mahalanobis distance between track j and measurement i is calculated as the cost element. ;
[0010] Among them, information vector Information covariance , For measurement observation matrix, For associated gated noise; Determine the chi-square distribution threshold based on the confidence level of the detection box. , will satisfy If the pairing is deemed unreasonable, its cost is set to infinity, while the remaining pairs retain the original squared Mahalanobis distance, thus obtaining the cost matrix. ;
[0011] in, For the corresponding confidence level The chi-square distribution threshold.
[0012] Furthermore, in the method described above, the observation reconstruction network and the state residual network employ an isomorphic general network skeleton, which includes: An input embedding multilayer perceptron is used to map the input sequence to a high-dimensional latent space; the output of the input embedding multilayer perceptron is connected to the main processing branch and the residual branch, respectively. The main processing branch consists of the following connected components: layer normalization, one-dimensional causal convolutional layer, SiLU activation function, and N stacked Mamba modules. Each Mamba module contains a Selective State Space Model (SSM) for dynamically generating state transition parameters and performing long-range temporal dependency modeling. The residual branch directly transmits the input to the output of the multilayer perceptron to the output of the last Mamba module. An adder is used to add the output of the main processing branch to the output of the residual branch, forming a residual connection; The summed feature sequence takes the hidden state of the last time step and is input to the parallel first linear projection head and Softplus activation function head respectively; the first linear projection head is used to output the expected vector, and the Softplus activation function head is used to output the diagonal elements of the diagonal covariance matrix; When the general network skeleton is instantiated as an observation reconstruction stream, the input sequence is a sliding window sequence formed by splicing the historical polar coordinate measurement sequence with the polar coordinate measurement matched at the current time, the expectation vector is the reconstruction measurement in Cartesian space, and the diagonal covariance matrix is the heteroscedastic measurement noise covariance matrix. When the general network skeleton is instantiated as a state residual flow, the input sequence is a sliding window sequence composed of historical posterior state sequences, the expectation vector is the dynamic residual correction amount, and the diagonal covariance matrix is the process noise covariance matrix.
[0013] Furthermore, according to the method described above, the loss function of the general network skeleton is:
[0014]
[0015]
[0016]
[0017]
[0018] in, These are the hyperparameters that balance the weights of each loss term. The gating parameter for state branches is used when the history sequence length is insufficient. Take 0 if the value is less than 0, otherwise take 1. The true Cartesian coordinate state at time k. This represents the true Cartesian coordinate state at time k-1. To observe the reconstruction measurement at time k of the reconstruction network output; To observe the reconstruction measurement at time k-1 of the reconstruction network output, Here, is the heteroscedastic measurement noise covariance matrix; F is the physical reference state transition matrix; This is the correction amount for the dynamic residual output by the state residual network.
[0019] Further, in the method described above, step 4 includes: First association: linking high-resolution detection sets With all active tracks Based on cost matrix Perform Hungarian matching; successfully matched track-detection box pairs constitute the first track set. It also outputs the set of unmatched tracks. and unmatched high-resolution detection boxes ; Second association: linking low-resolution detection sets With the set of unmatched tracks A second matching process is performed, this time using a chi-squared gate threshold. The successful matching results are denoted as the second track set. Merging, but ultimately not matching low-scoring boxes It was considered background noise and was removed. For all tracks that were successfully matched in the first two attempts The matched raw measurement data will be sent to the observation Mamba branch for reconstruction, and subsequent variational Bayesian filtering updates will be performed. For unmatched high-resolution detection boxes... For high-resolution isolated detection boxes, the system initializes them as new tracks; however, for tracks that still do not match after two association attempts... Mark it as missing if it is consecutive If a frame fails to rematch, it is permanently deleted from memory.
[0020] Further, in the method described above, step S6 includes: After determining the prior distribution of Mamba-guided noise, it is necessary to solve for the joint posterior distribution of state and noise. Using Mean-Field variational theory, it is approximately decomposed into a product of mutually independent marginal distributions:
[0021] in, This represents the target's state vector at time k; Indicates the measurement noise covariance; Within the variational Bayesian framework, this makes the approximate distribution... The closest distribution to the true posterior distribution This is equivalent to maximizing the lower bound of evidence. According to Bayes' theorem, the joint probability can be expanded into the product of the observation likelihood and the prior:
[0022] in, This represents the pseudo-measure vector of the Mamba observation network reconstructed stream output used for posterior update at time k; This represents the set of cumulative historical measurement sequences from the initial time to time k-1.
[0023] By taking the derivative of the maximizing evidence lower bound using variational calculus and setting the derivative to zero, we obtain the general form of the optimal solution for the approximate distribution: for any factor The logarithm of the optimal solution is equal to the expectation of the logarithm of the joint probability over all other variables. Therefore, for the measurement noise covariance... The best approximate posterior The following fixed-point equations must be satisfied:
[0024] Will contain The term is treated as constant absorption, which further simplifies to a recursive fixed-point equation:
[0025] Gaussian observation likelihood Reverse Saudi Prior with Mamba module boots Substituting into the above equation and utilizing the trace-cycle invariant property of the matrix... have to:
[0026] in, d represents the predictive degrees of freedom of the inverse Wieshart prior distribution; d represents the dimension of the measurement feature space. Let represent the predictive inverse scaling matrix of the inverse Wissaud prior distribution; Represents the target state The residual second-order expectation term incorporates the bias of the reconstructed measurement and the uncertainty of the current state estimate:
[0027] in, It is the outer product of the pseudo-measurement information vector. It is the covariance of the state estimation error mapped to the measurement space; Observation and deduction The form is completely isomorphic to the logarithmic form of the standard inverse Wissaud distribution; the iteration variable is set as... initial value The true posterior is approximated through alternating iterations of the following two steps; Post-state update: utilizing the first Noise estimation Perform Kalman measurement updates:
[0028] in, Let represent the one-step prior state error covariance matrix of the i-th variational iteration of the target state at time k; This represents the transpose of the system observation matrix; This represents the time-varying measure-noise covariance estimation matrix obtained in the (i-1)th variational iteration; This represents the target posterior state estimate vector after the i-th iteration update at time k; This represents the one-step prior state prediction of the target at time k; This represents the Kalman gain matrix at time k during the i-th iteration; The one-step posterior state error covariance matrix represents the estimate of the target state at time k in the i-th variational iteration. Noise Posterior Update: Calculate the second moment using the latest state estimate Update the posterior parameters of the IW distribution:
[0029] in, Let represent the predicted degrees of freedom of the inverse Wissaud posterior distribution after the i-th variational iteration update at time k; Let represent the predicted degrees of freedom of the inverse Wissaud prior distribution at time k; Let represent the inverse scaling matrix of the inverse Wissaud posterior distribution after the i-th variational iteration update at time k; Let represent the inverse scaling matrix of the inverse Wissaud prior distribution at time k; The expected noise covariance for the next iteration is then derived:
[0030] Reaching the maximum number of iterations Then, take the final result. As The final posterior state estimate of the target at time t is output.
[0031] Beneficial effects: 1. This application constructs an HDS-Mamba network, which utilizes an observation reconstruction network to specifically handle the nonlinear coordinate transformation and array distortion in the sonar measurement domain, and a state residual network to specifically handle the motion model mismatch in the target dynamic domain. The two networks handle different physical quantities, and through the deep integration of physical guidance and data-driven approaches, the tracking error is corrected simultaneously from both the measurement domain and the state domain. This allows for high-precision state estimation even when the target undergoes complex maneuvers, significantly improving the tracking robustness under the coupling of strong maneuvers and severe noise.
[0032] 2. This invention overcomes the shortcomings of traditional variational filtering, which suffers from slow convergence due to the lack of informational priors, by dynamically anchoring the inverse Wishart prior distribution using the heteroscedastic measurement noise covariance matrix. In the rapidly changing underwater acoustic environment, it forms an adaptive closed-loop estimation of measurement noise, ensuring the long-term tracking stability of the system under harsh underwater acoustic conditions such as multipath interference and reverberation. It also realizes adaptive filtering of non-stationary acoustic noise, guaranteeing the stability of long-term tracking.
[0033] 3. This invention provides a data foundation for hierarchical matching by dividing the detection boxes into high-scoring and low-scoring detection sets according to their confidence levels. Then, it employs a cascaded Hungarian algorithm to first match the high-scoring detection set with active tracks, and then match the tracks that did not match in the first round with the low-scoring detection set in the second round. This effectively retrieves low-confidence but real targets, reduces track breaks, significantly reduces the false negative rate, and ensures track continuity, effectively solving the problem of track breaks for weak targets in low signal-to-noise ratio environments. Attached Figure Description
[0034] Figure 1 This is one of the flowcharts of the active sonar underwater multi-target tracking method based on HDS-Mamba and adaptive variational Bayes provided in this application; Figure 2 The second schematic diagram of the active sonar underwater multi-target tracking method based on HDS-Mamba and adaptive variational Bayes provided in this application; Figure 3 This is a schematic diagram of the HDS-Mamba network structure provided in this application; Figure 4 This is a flowchart illustrating the hierarchical data association process in this application; Figure 5 This is a flowchart of the variational Bayesian adaptive update process in this application; Figure 6 This is a comparison diagram of two-dimensional trajectories under multi-target tracking in the embodiments of this application; Figure 7 This is a comparison curve of the root mean square error of the multi-target position based on 100 Monte Carlo simulations in an embodiment of this application. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0036] This invention provides an active sonar underwater multi-target tracking method and system based on a heterogeneous decoupled sonar Mamba network (HDS-Mamba) and measurement noise adaptive variational Bayesian. Taking sonar polar coordinate measurement sequences and target historical state sequences as inputs, this invention uses a functionally heterogeneous decoupled dual-stream architecture (observation reconstruction stream + state residual stream) to jointly drive online variational inference of the measurement noise covariance R. It aims to systematically solve problems such as model uncertainty, measurement noise non-stationarity, and difficulty in correlating multiple targets under low signal-to-noise ratio conditions during active underwater sonar tracking, thereby achieving robust tracking of underwater targets.
[0037] Figure 1 This is one of the flowcharts illustrating the active sonar underwater multi-target tracking method based on HDS-Mamba and adaptive variational Bayes provided in this application, as shown below. Figure 1 As shown, it includes: Step S1: Construct an HDS-Mamba network, which includes an observation reconstruction network and a state residual network. The observation reconstruction network is used to receive measurement sequences and output reconstructed measurements in Cartesian space and their corresponding heteroscedastic measurement noise covariance matrix. The state residual network is used to receive sliding history posterior state sequences and output dynamic residual corrections and their corresponding process noise covariance matrix. Step S2: Extract the sliding history posterior state sequence corresponding to each track from the maintained set of active tracks. The state residual network is then input into the state residual network, and based on the state residual network, a dynamic residual correction for the trajectory is output. and the predicted process noise covariance matrix The maintained active track set is the set of tracks that have been successfully matched with detection boxes and pre-stored by the system; the sliding history is a posterior state sequence. This is a historical posterior state sequence with a window length of L extracted from the historical state cache sequence of each active track; Step S3: Utilize the aforementioned dynamic residual correction amount Perform a Kalman filter one-step prediction to obtain the prior predicted state of each track at time k. ;in, Let be the prior predicted state at time k that incorporates network nonlinear compensation, and F be the physical baseline dynamic state transition matrix; This is a gating parameter; when the historical sequence length is insufficient... Select 0, otherwise select 1; Step S4: Extract the historical polar coordinate sequence corresponding to the detection box matched by each track from the maintained active track set, and append the polar coordinate measurement matched at the current time k to the end of the sequence. A sliding measurement sequence with a window length of L is constructed and input into the observation reconstruction network. Based on the observation reconstruction network, the reconstructed measurement of the Cartesian space at time k is output. and its corresponding heteroscedasticity measurement noise covariance matrix ; Step S5: Measure the noise covariance matrix using the heteroscedasticity. Dynamically anchor the inverse Wishart prior distribution and initialize the degrees of freedom parameters for Bayesian inference. and inverse scaling matrix ;in, The hyperparameter d is used to adjust the prior confidence of the network, where d is the dimension of the measurement vector. ; Step S6: Under the IW prior constraints using the aforementioned degree-of-freedom parameters and inverse scaling matrix, first... and reconstruction measurement Through iterative updates, the optimal state estimate at time k is finally obtained. and the true noise covariance matrix ; Step S7: Estimate the optimal state at time k. Store the historical posterior state sequence of the corresponding track and update the tracking trajectory of the track; finally, output the tracking trajectories of all active tracks to complete the underwater multi-target tracking at the current moment.
[0038] The aforementioned set of maintained active tracks is obtained as follows: Step 1: Acquire continuous sonar images at time k using active sonar detection, and extract the set of detection boxes from the sonar images using a target detection head. And based on the confidence threshold, the detection box set is... Divided into high-resolution detection sets and low-resolution detection set ; Step 2: Based on the process noise covariance matrix Calculate the prediction error covariance matrix of trajectory j at time k. ; This is the uncertainty covariance matrix corresponding to the prior predicted state; The error covariance at the previous time step; This is a gating parameter; when the historical sequence length is insufficient... Take 0 if the value is less than 0, otherwise take 1. Step 3: Based on the prior predicted state Prediction error covariance matrix and the original measurement sequence The i-th measurement Calculate the Mahalanobis distance between the active track j and the target i, and construct the cost matrix accordingly. The i-th measurement By collecting the detection boxes The result is obtained by projecting onto a Cartesian coordinate system; The cost matrix The construction includes: The squared Mahalanobis distance between track j and measurement i is calculated as the cost element. ;
[0039] Among them, information vector Information covariance , For measurement observation matrix, For associated gated noise; Determine the chi-square distribution threshold based on the confidence level of the detection box. , will satisfy If the pairing is deemed unreasonable, its cost is set to infinity, while the remaining pairs retain the original squared Mahalanobis distance, thus obtaining the cost matrix. ;
[0040] in, For the corresponding confidence level The chi-square distribution threshold.
[0041] Step 4: Utilize the constructed cost matrix High-resolution detection set With active flight paths Perform the first round of Hungarian matching and output the first set of successfully matched tracks. Subsequently, the unmatched tracks from the first round were combined with the low-resolution detection set. Perform a second matching and output the second set of successfully matched tracks. ; set the first track With the second set of tracks Merge and finally output the set of successfully matched tracks. and the set of successfully matched tracks As a set of active tracks maintained by the system.
[0042] Figure 2 The second schematic diagram of the active sonar underwater multi-target tracking method based on HDS-Mamba and adaptive variational Bayes provided for this application (the relevant vector space is defined as follows): ),like Figure 2As shown, the method provided by the present invention will be described in general below, and specifically includes the following steps: Step 1: Acquire continuous sonar images at time k using active sonar detection, where k is the time index of the current time step. This provides the necessary temporal context information for subsequent temporal modeling of the HDS-Mamba network, enabling the system to capture the temporal evolution of target motion.
[0043] Step 2: Use the target detection head to detect the sonar image at time k obtained in Step 1, and output the set of detection boxes at the current time. Where M represents the total number of targets detected in the sonar image. For the i-th detection box, the polar coordinate parameter vector in the sonar image at time k. ,in, The radial distance to the target. The target azimuth angle, and These represent the width and height of the detection box in the sonar image, respectively. To detect the confidence score, this step transforms the original high-dimensional sonar image into structured polar coordinate detection box parameters, including radial distance, azimuth angle, detection box size, and confidence score. This achieves a dimensionality reduction mapping from sensory data to measurement information. Preserving the confidence score information provides a crucial basis for the hierarchical association strategy in subsequent step 3, enabling the system to effectively distinguish reliable detections from suspected clutter under low signal-to-noise ratio conditions. This avoids the irreversible loss of weak target information caused by a single hard threshold filtering in traditional methods.
[0044] Step 3: Based on the set confidence threshold Total set of detection boxes Divided into high-resolution detection sets and low-resolution detection set In underwater active sonar detection, the echo signal strength of distant and weak targets is low, resulting in generally low detection confidence. This hierarchical strategy divides the detection boxes into high-resolution and low-resolution detection sets, rather than simply discarding low-confidence detection boxes, thus preserving information about these potentially valid targets and providing a complete data foundation for the cascaded matching in step 10. The system maintains a set of active tracks. Where N is the total number of currently active tracks, and each track Store its historical state sequence and historical measurement sequences The historical state sequence is used to learn the nonlinear motion pattern of the target through the state residual network, while the historical measurement sequence is used to learn the nonlinear transformation and noise characteristics of sonar observations through the observation reconstruction network. At the initial time k=1, Empty set, high-resolution detection set Each detection box is initialized as a newly created track awaiting confirmation, and is added after a new track is successfully associated. When the steady-state cycle k>1, The trajectory lifecycle management module, which was maintained and updated in step 14 of the previous time step, automatically appends the latest estimated state to its historical queue after each round of filtering. This allows for direct invocation in step 5 of the next time step. This track history caching mechanism provides a temporal context for the HDS-Mamba state residual network, enabling the network to make accurate maneuver predictions based on the target's historical motion patterns. This reflects the efficient reuse of historical information in the "memory-prediction-update" closed-loop design of the tracking system.
[0045] Step 4: Collect all detection boxes Projecting onto a Cartesian coordinate system, the original measurement sequence at time k is obtained. .in, This represents the approximate physical measurement of the i-th detection box in the Cartesian coordinate system at time k, and its calculation formula is: Raw sonar measurements are in polar coordinates, while target motion modeling is typically performed in Cartesian space.
[0046] This step unifies the measurement space and state space through coordinate projection, providing physically consistent observation input for subsequent Kalman filtering state updates. Compared to the Jacobian matrix accumulation error problem caused by direct nonlinear tracking in the polar coordinate domain, this projection strategy allows subsequent linear motion models to be directly applied to target state estimation, effectively reducing system modeling complexity and linearization error.
[0047] Step 5: For the set of active tracks in the tracking system Every active track that already exists in China Extract a sliding window sequence of length L from its historical state cache sequence to obtain the historical posterior state sequence. And use it as the input to the state residual network. .in, Right now Figure 2 In Where the subscript kL:k-1 represents the historical sequence from time kL to time k-1, This represents the state estimate of the j-th track at time k-1, output by the variational Bayesian filter in step 13 and stored in the queue. x and y represent the two-dimensional spatial coordinates of the target in the Cartesian coordinate system, and vx and vy represent the two-dimensional velocity of the target in the corresponding coordinate axis directions.
[0048] This step extracts a state sequence sliding window of length L from the track history cache as input to the state residual network, providing the Mamba selective state-space model with the complete temporal context required to model the target motion trend. For new tracks with fewer than L historical frames, zero padding to length L is used, and the gating parameters of the network residual output are also adjusted. Setting the value to 0 degrades the system's prediction logic to a pure physical CV model prediction, thereby ensuring stable initialization of the system during periods of limited temporal information. This design effectively avoids state divergence caused by unreliable residuals in the network output when temporal information is scarce, ensuring a smooth transition of the new trajectory from initialization to steady-state tracking.
[0049] Step 6: Extract the historical posterior state sequence The input is fed into the HDS-Mamba state residual network, which is modeled through normalization layers, one-dimensional causal convolutional layers, and N layers of Mamba selective state space modules. The output is the dynamic residual correction for the trajectory. and the predicted process noise covariance matrix That is, the expression is .in, It is a high-order nonlinear dynamic residual extracted by a neural network to compensate for the mismatch in the physical constant velocity model caused by the target's large maneuvers; This is the diagonal covariance matrix that reflects the noise during the current flight path.
[0050] Underwater targets are influenced by ocean currents and autonomous maneuvers, often experiencing nonlinear abrupt changes in their motion state. Traditional constant-velocity physical models cannot accurately describe such complex maneuvers. This step extracts high-order nonlinear dynamic features from historical state sequences using an HDS-Mamba state residual network. This data-driven approach compensates for the mismatch error in the physical CV model, while dynamically estimating the process noise covariance Q to reflect the magnitude of model uncertainty at the current moment. Compared to the IMM method, which relies on manually pre-set multiple model sets, this network can adaptively learn arbitrary maneuvering patterns without human intervention. Furthermore, Mamba's linear time complexity provides a significant computational efficiency advantage for practical underwater platform deployment.
[0051] Step 7: Utilize dynamic residual corrections Perform a Kalman filter prediction step to calculate the prior predicted state of track j at time k. . Let k be the prior predicted state at time k that incorporates network nonlinear compensation; F is the physical baseline dynamic state transition matrix.
[0052] This application employs a constant-velocity CV model. This step integrates the state transition and dynamic residual corrections from the physical baseline CV model. This mechanism achieves an organic synergy between physical priors and data-driven approaches. It preserves the reliability of the physical model during steady-state cruise motion, avoiding uncontrollable extrapolation errors generated outside the training distribution in purely data-driven methods; and it compensates for nonlinear deviations at maneuver moments through neural networks, making up for the inherent defect of slow response of pure physical models to complex maneuvers.
[0053] Step 8: Combine the process noise covariance matrix output by the state residual network. Calculate the prediction error covariance matrix of trajectory j at time k. . This is the uncertainty covariance matrix corresponding to the prior predicted state; This represents the error covariance at the previous time step.
[0054] In traditional Kalman filtering methods, the process noise covariance Q is usually set to a fixed empirical value, which fails to reflect the increased prediction uncertainty caused by target maneuvers. This leads to a severe mismatch between the predicted covariance and the actual error when the underwater target undergoes sudden maneuvers. This step utilizes a dynamically estimated Q value from a neural network, enabling the predicted covariance to adaptively reflect the target's maneuverability. This provides a more accurate measure for subsequent Mahalanobis distance calculations, effectively distinguishing between predicted offsets caused by maneuvers and true mismatches during the data association stage.
[0055] Step 9: Use the prior predicted state obtained by performing a one-step prediction using Kalman filtering. Step 8 calculates the prediction error covariance matrix. and the original measurement sequence obtained in step 4. The i-th measurement Calculate the Mahalanobis distance between active track j and target i to construct a multi-target data association cost matrix. Define the cost matrix. The element corresponding to track j and measurement i is the square of the Mahalanobis distance. Its calculation formula is Information vector Information covariance .in, For measurement observation matrix, This is related to gated noise.
[0056] In this application, the covariance matrix of the sonar polar coordinate measurement error is defined as the covariance matrix propagated to Cartesian space via the Jacobian matrix. After calculating the Mahalanobis distance for all "track-detection box" pairs, a chi-square distribution threshold corresponding to the target detection confidence is introduced. As a gating mechanism, remove This addresses the issue of illogical matching combinations. In underwater sonar scenarios, the measurement accuracy of radial distance and azimuth angle often differs significantly. Constructing an association cost matrix using Mahalanobis distance allows for directional normalization of the deviation between state prediction and measurement using the information covariance matrix, thus considering the uncertainty caused by the anisotropy of measurement accuracy. Furthermore, a chi-square distribution threshold associated with detection confidence is introduced as a gating mechanism, eliminating erroneous matching pairs that significantly deviate from the predicted distribution from a statistical hypothesis testing perspective. In dense underwater clutter environments, this mechanism effectively controls the probability of false associations within a preset statistical significance level, providing a reliable cost input for subsequent Hungarian optimal allocation.
[0057] Step 10: Transform the association problem into a bipartite graph optimal allocation problem and solve it using the Hungarian algorithm. Based on the constructed cost matrix, allocate the high-scoring detection set... With active flight paths Perform the first round of Hungarian matching and output the set of successfully matched pairs. Subsequently, the unmatched tracks from the first round were combined with the low-resolution detection set. A second matching process is performed using a more stringent chi-square gating threshold, outputting the set of successfully matched matches. The final successful "track-detection box" matching results. Each element represents a successfully matched (j,i) pair. This cascaded matching strategy uses two rounds of matching. In the first round, high-resolution detection sets are matched with active tracks to ensure high-reliability association. In the second round, tracks not matched in the first round are matched with low-resolution detection sets to capture faint targets partially obscured by clutter or in acoustic shadow areas. This fully utilizes the potential effective information in each low-confidence detection frame, rather than simply discarding it as noise. Under low signal-to-noise ratio conditions underwater, this strategy can significantly reduce the frequency of track breaks caused by missed detections, effectively ensuring the trajectory continuity and identity consistency of multi-target tracking.
[0058] Step 11: For the set of successfully matched tracks For each track, extract its historical polar coordinate sequence of successfully associated data, and append the latest polar coordinate measurement matched at time k to the end of the sequence. The updated sliding measurement sequence was constructed. .
[0059] This step temporally concatenates historically successfully correlated polar coordinate measurement sequences with the latest matched measurement at the current moment to construct a temporally complete sliding observation sequence. This sequence preserves the original distribution characteristics of the measurements in the original polar coordinate domain and the temporal correlation of measurement noise, enabling the observation reconstruction network to directly learn the temporal evolution pattern of sonar measurements from the original measurement domain. Compared to methods that only use single-frame measurements for reconstruction, the temporal sequence input allows the network to distinguish between real target measurements and transient clutter using inter-frame consistency, thereby outputting more reliable reconstructed pseudo-measurements and more accurate uncertainty estimates.
[0060] Step 12: Combine the observation sequence constructed above The observation reconstruction network input to HDS-Mamba outputs the reconstructed Cartesian space measurement after sequence modeling. and the corresponding heteroscedasticity measurement noise covariance matrix That is, the expression is Underwater sonar detection suffers from nonlinear mapping and array distortion, making explicit analytical expression of the measurement model difficult. Traditional methods employ first-order Taylor linearization, which introduces significant truncation errors in long-range, large-azimuth scenarios. This step directly learns the end-to-end nonlinear mapping from polar coordinate measurement sequences to Cartesian reconstructed measurements using the HDS-Mamba observation reconstruction network, while simultaneously outputting the noise heteroscedasticity matrix. The reliability of current observations is quantified on a sample-by-sample basis. Compared to traditional fixed measurement noise models, this heteroscedastic output enables the system to perceive differences in measurement quality at different times and for different targets. When the target is severely disturbed, the network outputs a larger R value to reduce the weight of that measurement in subsequent filtering updates, thereby preventing the poorly measured data from affecting the accuracy of state estimation.
[0061] Step 13: Reconstruct the measurement noise covariance matrix of the network output from observations. Dynamically anchor the inverse Wishart (IW) prior distribution and initialize the degrees of freedom parameters for Bayesian inference. and inverse scaling matrix and. Among them, The hyperparameter d is used to adjust the prior confidence of the network. Traditional variational Bayesian adaptive filtering usually uses no-information priors or directly uses the posterior from the previous time step as the prior for the current time step. When the statistical characteristics of underwater acoustic noise change drastically, it faces the risk of slow convergence or even divergence.
[0062] This step dynamically anchors the parameters of the inverse Wishart prior distribution by observing the heteroscedasticity covariance matrix of the network output, and flexibly adjusts the confidence level of the network prior through hyperparameters. This mechanism ensures that the prior distribution includes both noisy prediction information from the deep learning network and retains the adaptive adjustment space of the variational Bayesian framework, allowing the model to be finely corrected based on actual measurement residuals in subsequent iterations.
[0063] Step 14: Integrate the prior predicted state at time k obtained in Step 7. Reconstruction measurement in step 12 Under the IW prior constraints set in step 13, iterative updates are performed using the variational Bayesian algorithm. Within each iteration, Kalman state parameter updates and noise parameter R updates are performed alternately until convergence, and the optimal state estimate at time k is jointly estimated. and the true noise covariance matrix Save the tracking trajectory. and will Add it to the historical state queue of the track for cyclical recall in the next moment.
[0064] Under the pre-defined Mamba network-guided prior constraints, this step iteratively updates the posterior parameters of the state and the measurement noise using a variational Bayesian algorithm until convergence, jointly estimating the time-varying measurement noise covariance and the target posterior state online. Even under the harsh conditions of drastic time-varying measurement noise statistical characteristics caused by multipath effects, reverberation, and unknown interference in the underwater acoustic channel, this mechanism ensures that the filtering estimation maintains high accuracy throughout, effectively overcoming the covariance collapse and filter divergence problems that occur in traditional fixed noise models during sudden noise changes.
[0065] Step 15: Based on the association results from Step 10, perform track lifecycle management, initialize unmatched high-value components as new tracks and add them to the active list, and deregister historical tracks that have not matched for multiple consecutive frames. This mechanism works in conjunction with the hierarchical detection strategy in Step 3 and the cascaded matching strategy in Step 10 to form a complete lifecycle closed loop from target appearance and continuous tracking to target disappearance. For underwater multi-target tracking scenarios, this lifecycle management mechanism ensures that the system can promptly capture newly emerging targets to avoid missed tracking, while preventing the infinite accumulation of false tracks from reducing computational efficiency and association accuracy, thus maintaining the high efficiency and stability of the tracking system in the long-term operation of the dynamic underwater environment.
[0066] Example: Step S1: Acquire and preprocess underwater sonar image sequences In underwater active sonar detection scenarios, acquire continuous sonar images. At any given time, the front-end object detection head outputs a set of detection boxes. Each detection box is parameterized as , To detect confidence level. Given a confidence threshold. ,Will Divided into high-resolution detection sets and low-resolution detection set The polar coordinate detection box is then roughly projected onto a Cartesian coordinate system to obtain... .
[0067] Step S2: Extract the spatiotemporal features of the sonar using the HDS-Mamba (Heterogeneous Decoupled Sonar Mamba) network.
[0068] Figure 3 This is a schematic diagram of the HDS-Mamba network structure provided in this application, as shown below. Figure 3 As shown, this invention utilizes the HDS-Mamba network (Heterogeneous Decoupled Sonar Mamba) to extract sonar spatiotemporal features. The HDS-Mamba architecture of this invention consists of two independent network branches with completely different functions and physical meanings: an observation reconstruction network (first stream) and a state residual network (second stream). Unlike traditional symmetric isomorphic structures that perform forward and inverse dual-path temporal coding of the same sequence, the dual-stream architecture of this invention employs a heterogeneous decoupled design: the observation reconstruction stream is specifically used to handle the nonlinear coordinate transformations and uncertainties in the sonar measurement domain, while the state residual stream is specifically used to model the motion model mismatch residuals in the target state domain.
[0069] Both branches of HDS-Mamba employ a homogeneous, general network architecture at their underlying backbone, aiming to reduce the computational overhead of underwater platforms through parameter sharing or structural reuse. However, the input physical quantities of the two streams, the meaning of features extracted by the network, the output content, and the subsequently associated Kalman noise matrices are completely independent and do not interfere with each other. Let the generalized input sequence of the network be... ,in The time window length, For the input dimension. When this general architecture is applied to the observation reconstruction flow (first flow), the generalized sequence... Instantiated as a historical primitive polar coordinate measurement sequence of length L Its input feature dimension d=5; when this general architecture is applied to state residual streams (second streams), the generalized sequence Instantiated as a system history state estimation sequence of length L Its input feature dimension d=4. The specific process of this general architecture is as follows.
[0070] Because underwater active sonar observations are often accompanied by high-frequency clutter, directly inputting low-dimensional coordinate vectors into the state space can easily induce gradient oscillations. Therefore, the input sequence is... First, a multilayer perceptron (MLP) is used to map from the low-dimensional space to the high-dimensional latent space. After mapping, the feature sequence is divided into two parallel branches in the data flow: the residual branch directly points to the addition operator after the Mamba module, serving as an identity mapping for deep networks to alleviate the gradient vanishing problem; the main processing branch is input to the core unit of the Mamba module after amplitude adjustment by the normalization layer.
[0071] Upon entering the Mamba module, the data flow is further decoupled into two highly asymmetric sub-paths to separate feature information from control gating. On the main branch, data first undergoes linear projection to increase the number of channels, then enters a one-dimensional causal convolution (1DCausalConv) to extract local spatiotemporal features, and is further enhanced nonlinearly using the SiLU activation function. The one-dimensional causal convolution strictly adheres to temporal causality, ensuring that information from future frames is not leaked during local feature aggregation. Subsequently, the data enters the core Selective State-Space Model (Selective SSM) to perform long-range temporal dependency modeling. Unlike traditional linear time-invariant systems with fixed parameters, this Selective State-Space Model introduces a selective scanning mechanism. In the... Layer, for time step Input hidden state The network dynamically generates state transition parameters through three independent linear layers. , With discretization step size :
[0072] Subsequently, the continuous system was discretized using the zero-order hold (ZOH) principle, resulting in... and And perform efficient sequence state recursion:
[0073] Synchronized with the main branch, the gated branch data is linearly projected onto another path and activated by the SiLU function to generate a gated weight vector with the same feature dimensions as the main branch. The outputs of the two sub-paths are multiplied element-wise, and the merged features are finally processed by a linear projection of the output to complete the internal processing of the Mamba module.
[0074] The features output from the Mamba module are summed with the aforementioned residual branch embedding features at the addition operator. Since this architecture employs causal modeling logic, the system only retains the hidden feature vector from the last time step of the sequence (t=L). , which serves as a representation of the goal at the current moment.
[0075] Finally, the truncated feature vectors The data is simultaneously input into two parallel linear prediction heads to achieve the final mapping from latent space to physical space parameters. The calculation formula is as follows:
[0076] In the formula, W and b are the weights and biases of the corresponding linear mapping layer. Expected vector The output is directly obtained through a linear mapping; while the expected variance... The diagonal elements of the diagonal covariance matrix are output by using the Softplus activation function. Mathematically, the positive definiteness of the variance is strictly guaranteed.
[0077] The above outputs will be assigned different statistical meanings depending on the application stream. In the observation reconstruction stream, the output is instantiated as the Cartesian pseudomeasure and the heteroscedasticity measurement noise covariance. In the state residual flow, the output is instantiated as the dynamic residual correction and the process noise covariance. .
[0078] The observation reconstruction network (first-line) uses raw measurement sequences containing past L frames. As input, the Mamba module combines one-dimensional causal convolution with a selective scanning mechanism to output reconstructed measurements. and the corresponding covariance matrix , denoted as:
[0079] State residual networks use historical posterior state sequences Input, output state residual correction amount and the predicted process noise covariance , denoted as:
[0080] To ensure that the HDS-Mamba (Heterogeneous Decoupled Sonar Mamba) network can accurately extract features, it is necessary to perform end-to-end joint training on the network using an offline dataset before conducting online multi-target tracking. This invention addresses the pain points of underwater sonar tracking by specifically designing a multi-target hybrid loss function, as detailed below.
[0081] First, to avoid the model falling into a degenerate state with high variance in the early stages of training, a mean squared error function is used. The reconstructed measurements output by the constrained observation reconstruction network approximate the true Cartesian coordinates of the target. The loss function directly supervises the expected head of the observation reconstruction network (first-order). This provides the most basic spatial location anchoring for the network, and the specific formula is as follows:
[0082] Secondly, through the negative log-likelihood loss function The uncertainty of observations is learned. This loss function, combined with supervised observations, reconstructs the network's expected head. With variance head It is for Figure 3 The heteroscedasticity matrix output by the Softplus activation function Apply core constraints. When the expected header output... As a reconstruction measurement When the output has a large error compared to the true value, this loss will drive the variance head to adaptively amplify the covariance value of the output. This not only verifies... Figure 3 The introduction of Softplus ensures the structural rationality of positive variance definiteness, and further enables the network to accurately assess the uncertainty of its own output:
[0083] Next, based on the assumption of the continuity of the target motion, a first-order difference smoothing loss is introduced. .because Figure 3 The Mamba module in the code has a time-series memory function. This function effectively suppresses non-physical jitter in the reconstructed trajectory caused by high-frequency sonar noise by constraining the first derivative of the expected output value in adjacent time steps to be consistent with the actual motion trend.
[0084] Finally, to ensure that the observation reconstruction network can correctly learn the target's motion trend, a dynamic prediction loss is introduced. This loss function is specifically designed to supervise the expected prediction head of the state residual network (second flow). In the formula... That is, the output value of the corresponding state flow expected head. By adding the predictions from the pure physics-based CV model to the network's predictions and then applying a truth constraint, the system is forced to... Figure 3 The state-flow network in the model only learns complex maneuver residuals that cannot be explained by the physical model, achieving a deep integration of data-driven and physical priors at the network structure output:
[0085] In summary, the final total loss function of the network is defined as:
[0086] in, These are the hyperparameters that balance the weights of each loss term. The gating parameter for state branches is used when the history sequence length is insufficient. The gate selects 0 if the input is positive and 1 otherwise. This gating mechanism effectively prevents errors during the training and inference phases. Figure 3 The system diverges due to the forced output of residuals when the state flow is in a state with scarce timing information.
[0087] Step S3: Perform hierarchical data association guided by sonar measurements like Figure 5 As shown, all existing tracks must be obtained before performing association matching. The prior distribution at the current time. For the ... A historical track, its historical posterior sequence Input to the state residual network to obtain the dynamic residual. With process noise Subsequently, the physical reference transfer matrix was used. Perform one-step state prediction:
[0088] in, This is a gating parameter; when the historical sequence length is insufficient... It takes 0, otherwise it takes 1, and is used to control the switching between new and historical tracks.
[0089] To quantify the detection frame With predicted position of the flight path To assess spatial similarity between them, we construct the cost matrix using Mahalanobis distance. Compared to Euclidean distance, Mahalanobis distance is more effective at measuring uncertainty in the process of state prediction and measurement transformation.
[0090] Among them, information vector Information covariance .in, For measurement observation matrix, To mitigate noise, this paper defines the covariance matrix of the sonar polar coordinate measurement error propagated to Cartesian space via the Jacobian matrix. Based on this, a chi-square gating mechanism is introduced to eliminate obviously unreasonable matching pairs:
[0091] In the formula, For the corresponding confidence level The chi-square distribution threshold.
[0092] After constructing the cost matrix, we transform the association problem into a classic bipartite graph optimal allocation problem and solve it using a cascaded Hungarian algorithm.
[0093] First association. High-resolution detection set. With all activity tracks Based on the above cost matrix Perform a Hungarian matching. Successfully matched pairs form a set. It also outputs the set of unmatched tracks. and unmatched high-resolution detection boxes .
[0094] Second association. In complex underwater environments, targets affected by acoustic shadowing or severe multipath interference often only generate low-confidence detection boxes. Directly discarding these boxes would lead to frequent trajectory breakage. Thanks to the powerful nonlinear motion prediction capabilities of the state Mamba branch, unmatched tracks... It still has highly reliable predicted locations. Therefore, we will use the low-resolution detection set. and A second matching is performed. To avoid introducing background clutter, a more stringent chi-square gate threshold is used for this association (…). A successful match is recorded as... The low-scoring boxes that ultimately did not match It was considered background noise and was removed.
[0095] For all tracks that were successfully matched in the first two attempts The matched raw measurement data will be sent to the observation Mamba branch for reconstruction, and subsequent variational Bayesian filtering updates will be performed. For For high-resolution isolated detection boxes, the system initializes them as new tracks. Tracks that still do not match after two association attempts... Mark it as missing if it is consecutive If a frame fails to be rematched, it is permanently removed from memory to free up computing resources.
[0096] Step S4: Adaptive Kalman Filter Update Based on Variational Bayes Unlike traditional variational filtering algorithms that rely on blind initialization using empirical values from the previous time step, this paper utilizes the heteroscedastic uncertainty of the observed Mamba network output. To dynamically anchor the current prior distribution. This is to ensure the expected value of the prior distribution is... We introduce the adjustment of hyperparameters. As degrees of freedom, the degree-of-freedom parameters are derived by inversely using the expectation formula. and inverse scaling matrix The initialization equation is as follows:
[0097] In the formula, To measure the dimension of the vector, The hyperparameter for adjusting the Mamba prediction confidence, representing the prior confidence in the Mamba prediction results, needs to satisfy the following conditions: Larger This makes the prior distribution extremely steep, and the system is more inclined to believe the variance extracted by Mamba; smaller This gives the system a larger adaptive adjustment space in subsequent variational iteration updates.
[0098] Having determined the prior distribution of Mamba guidance, we need to solve for the joint posterior distribution of the state and noise. Since this joint distribution is difficult to solve analytically directly, Mean-Field variational theory is used to approximate its decomposition into a product of independent marginal distributions:
[0099] in, This represents the target's state vector at time k; This represents the measurement noise covariance.
[0100] Within the variational Bayesian framework, this makes the approximate distribution... The closest distribution to the true posterior distribution This is equivalent to maximizing the lower bound of evidence (ELBO). According to Bayes' theorem, the joint probability can be expanded as the product of the observation likelihood and the prior:
[0101] in, This represents the pseudo-measure vector of the Mamba observation network reconstructed stream output used for posterior update at time k; This represents the set of cumulative historical measurement sequences from the initial time to time k-1.
[0102] By taking the derivative of ELBO using variational calculus and setting the derivative to zero, we can obtain the general form of the optimal solution for the approximate distribution: for any factor The logarithm of the optimal solution is equal to the expectation of the logarithm of the joint probability over all other variables. Therefore, for the measurement noise covariance... The best approximate posterior The following fixed-point equations must be satisfied:
[0103] Because we only focus on The distribution can contain The term is treated as constant absorption, which further simplifies to a recursive fixed-point equation:
[0104] Gaussian observation likelihood With Mamba-guided Reverse Saudi A priori Substituting into the above equation and utilizing the trace-cycle invariant property of the matrix... have to:
[0105] in, d represents the predictive degrees of freedom of the inverse Wieshart prior distribution; d represents the dimension of the measurement feature space. Let represent the predictive inverse scaling matrix of the inverse Wissaud prior distribution; Represents the target state The residual second-order expectation term incorporates the bias of the reconstructed measurement and the uncertainty of the current state estimate:
[0106] in, It is the square of the residual. It is the uncertainty of the estimate; Observation and deduction The form is completely isomorphic to the logarithmic form of the standard inverse Wissaud distribution. Through parameter identification, this theoretical derivation provides us with rigorous parameter update rules. In the actual algorithm implementation, the iteration variable is set as... initial value The true posterior is approximated through alternating iterations of the following two steps.
[0107] Post-state update: utilizing the first Noise estimation Perform Kalman measurement updates:
[0108] in, Let represent the one-step prior state error covariance matrix of the i-th variational iteration of the target state at time k; This represents the transpose of the system observation matrix; This represents the time-varying measure-noise covariance estimation matrix obtained in the (i-1)th variational iteration; This represents the target posterior state estimate vector after the i-th iteration update at time k; This represents the one-step prior state prediction of the target at time k; This represents the Kalman gain matrix at time k during the i-th iteration; The one-step posterior state error covariance matrix represents the estimate of the target state at time k in the i-th variational iteration. Noise Posterior Update: Calculate the second moment using the latest state estimate Update the posterior parameters of the IW distribution:
[0109] in, Let represent the predicted degrees of freedom of the inverse Wissaud posterior distribution after the i-th variational iteration update at time k; Let represent the predicted degrees of freedom of the inverse Wissaud prior distribution at time k; Let represent the inverse scaling matrix of the inverse Wissaud posterior distribution after the i-th variational iteration update at time k; Let represent the inverse scaling matrix of the inverse Wissaud prior distribution at time k; The expected noise covariance for the next iteration is then derived:
[0110] Reaching the maximum number of iterations Then, take the final result. As Estimate the optimal state at time 1 and push it into the historical state queue. In this process, a complete closed-loop reasoning is formed, and the entire process is as follows: Figure 5 As shown.
[0111] Figure 6 This is a comparison diagram of two-dimensional trajectories under multi-target tracking in the embodiments of this application, such as... Figure 6 As shown, to visually demonstrate the algorithm's anti-maneuverability performance, a typical target trajectory involving complex S-shaped maneuvers is selected for magnified display (the purple circle in the figure represents the starting point, and the purple pentagram represents the ending point). When the target undergoes complex maneuvers accompanied by non-stationary noise, traditional methods (standard EKF and traditional VB-EKF) are limited by rigid physical models and lag priors, resulting in severe trajectory out-of-bounds errors at curves and violent oscillations in noise-induced jump zones. In contrast, the method of this invention (HDS-Mamba-VB), through high-frequency residual compensation of the state flow and precise adaptive prior guidance of the observation flow, effectively eliminates the rigid amplitude limiting set by traditional algorithms to prevent divergence. Its tracking trajectory (red line) is consistently smooth from beginning to end and closely follows the true trajectory (black line), significantly improving the tracking robustness under the coupling of strong maneuvers and severe noise.
[0112] Figure 7 This is a comparison curve of the root mean square error (RMSE) of multi-target positions based on 100 Monte Carlo simulations in this application embodiment; as shown... Figure 7The graph shown is a comparison of the root mean square error (RMSE) of position based on 100 independent Monte Carlo simulations. This graph effectively eliminates the randomness of single-shot noise and objectively reflects the statistical performance of the algorithm. In the first 50 stationary time steps, the traditional algorithm experiences error spikes due to model mismatch, while the method of this invention consistently maintains extremely low error levels. After the 50th step, when noise intensity such as multipath / reverberation changes drastically, the traditional method experiences covariance collapse due to error attribution fallacy, resulting in a steep divergence in the RMSE curve. However, the root mean square error of the coordinates of the method of this invention (HDS-VBKF) remains significantly stable (the red line is close to the horizontal axis), which statistically rigorously demonstrates that the invention has superior long-term tracking stability under harsh underwater acoustic conditions.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An active sonar underwater multi-target tracking method based on HDS-Mamba and adaptive variational Bayes, characterized in that, include: Step S1: Construct an HDS-Mamba network, which includes an observation reconstruction network and a state residual network; The observation reconstruction network is used to receive measurement sequences and output reconstructed measurements in Cartesian space and their corresponding heteroscedastic measurement noise covariance matrix. The state residual network is used to receive the sliding history posterior state sequence and output the dynamic residual correction amount and the corresponding process noise covariance matrix. Step S2: Extract the sliding history posterior state sequence corresponding to each track from the maintained set of active tracks. The state residual network is then input into the state residual network, and based on the state residual network, a dynamic residual correction for the trajectory is output. and the predicted process noise covariance matrix The maintained active track set is the set of tracks that have been successfully matched with detection boxes and pre-stored by the system; the sliding history is a posterior state sequence. This is a historical posterior state sequence with a window length of L extracted from the historical state cache sequence of each active track; Step S3: Utilize the aforementioned dynamic residual correction amount Perform a Kalman filter one-step prediction to obtain the prior predicted state of each track at time k. ; Step S4: Extract the historical polar coordinate sequence corresponding to the detection box matched by each track from the maintained active track set, and append the polar coordinate measurement matched at the current time k to the end of the sequence. A sliding measurement sequence with a window length of L is constructed and input into the observation reconstruction network. Based on the observation reconstruction network, the reconstructed measurement of the Cartesian space at time k is output. and its corresponding heteroscedasticity measurement noise covariance matrix ; Step S5: Measure the noise covariance matrix using the heteroscedasticity. Dynamically anchor the inverse Wishart prior distribution and initialize the degrees of freedom parameters for Bayesian inference. and inverse scaling matrix ;in, The hyperparameter d is used to adjust the prior confidence of the network, where d is the dimension of the measurement vector. ; Step S6: Under the IW prior constraints using the aforementioned degree-of-freedom parameters and inverse scaling matrix, first predict the prior state at time k. and reconstruction measurement Through iterative updates, the optimal state estimate at time k is finally obtained. and the true noise covariance matrix ; Step S7: Estimate the optimal state at time k. Store the historical posterior state sequence of the corresponding track and update the tracking trajectory of the track; finally, output the tracking trajectories of all active tracks to complete the underwater multi-target tracking at the current moment.
2. The method according to claim 1, characterized in that, The maintained set of active tracks is obtained as follows: Step 1: Acquire continuous sonar images at time k using active sonar detection, and extract the set of detection boxes from the sonar images using a target detection head. ; And based on the confidence threshold, the detection box set is... Divided into high-resolution detection sets and low-resolution detection set ; Step 2: Based on the process noise covariance matrix Calculate the prediction error covariance matrix of trajectory j at time k. ; Step 3: Based on the prior predicted state Prediction error covariance matrix and the original measurement sequence The i-th measurement Calculate the Mahalanobis distance between the active track j and the target i, and construct the cost matrix accordingly. ; The i-th measurement By collecting the detection boxes The result is obtained by projecting onto a Cartesian coordinate system; Step 4: Utilize the constructed cost matrix High-resolution detection set With active flight paths Perform the first round of Hungarian matching and output the first set of successfully matched tracks. ; Subsequently, the unmatched tracks from the first round were combined with the low-resolution detection set. Perform a second matching and output the second set of successfully matched tracks. ; set the first track With the second set of tracks Merge and finally output the set of successfully matched tracks. and the set of successfully matched tracks As a set of active tracks maintained by the system.
3. The method according to claim 2, characterized in that, The cost matrix mentioned in step 3 The construction includes: The squared Mahalanobis distance between track j and measurement i is calculated as the cost element. ; Among them, information vector Information covariance , For measurement observation matrix, For associated gated noise; Determine the chi-square distribution threshold based on the confidence level of the detection box. , will satisfy If the pairing is deemed unreasonable, its cost is set to infinity, while the remaining pairs retain the original squared Mahalanobis distance, thus obtaining the cost matrix. ; in, For the corresponding confidence level The chi-square distribution threshold.
4. The method according to any one of claims 1-3, characterized in that, The observation reconstruction network and the state residual network adopt an isomorphic general network skeleton, which includes: An input embedding multilayer perceptron is used to map the input sequence to a high-dimensional latent space; the output of the input embedding multilayer perceptron is connected to the main processing branch and the residual branch, respectively. The main processing branch consists of the following connected components: layer normalization, one-dimensional causal convolutional layer, SiLU activation function, and N stacked Mamba modules. Each Mamba module contains a Selective State Space Model (SSM) for dynamically generating state transition parameters and performing long-range temporal dependency modeling. The residual branch directly transmits the input to the output of the multilayer perceptron to the output of the last Mamba module. An adder is used to add the output of the main processing branch to the output of the residual branch, forming a residual connection; The summed feature sequence takes the hidden state of the last time step and is input to the parallel first linear projection head and Softplus activation function head respectively; the first linear projection head is used to output the expected vector, and the Softplus activation function head is used to output the diagonal elements of the diagonal covariance matrix; When the general network skeleton is instantiated as an observation reconstruction stream, the input sequence is a sliding window sequence formed by splicing the historical polar coordinate measurement sequence with the polar coordinate measurement matched at the current time, the expectation vector is the reconstruction measurement in Cartesian space, and the diagonal covariance matrix is the heteroscedastic measurement noise covariance matrix. When the general network skeleton is instantiated as a state residual flow, the input sequence is a sliding window sequence composed of historical posterior state sequences, the expectation vector is the dynamic residual correction amount, and the diagonal covariance matrix is the process noise covariance matrix.
5. The method according to claim 4, characterized in that, The loss function of the general network skeleton is: in, These are the hyperparameters that balance the weights of each loss term. The gating parameter for state branches is used when the history sequence length is insufficient. Take 0 if the value is less than 0, otherwise take 1. The true Cartesian coordinate state at time k. This represents the true Cartesian coordinate state at time k-1. To observe the reconstruction measurement at time k of the reconstruction network output; To observe the reconstruction measurement at time k-1 of the reconstruction network output, Here, is the heteroscedastic measurement noise covariance matrix; F is the physical reference state transition matrix; This is the correction amount for the dynamic residual output by the state residual network.
6. The method according to claim 4, characterized in that, Step 4 includes: First association: linking high-resolution detection sets With all active tracks Based on cost matrix Perform Hungarian matching; successfully matched track-detection box pairs constitute the first track set. It also outputs the set of unmatched tracks. and unmatched high-resolution detection boxes ; Second association: linking low-resolution detection sets With the set of unmatched tracks A second matching process is performed, this time using a chi-squared gate threshold. The successful matching results are denoted as the second track set. Merging, but ultimately not matching low-scoring boxes It was considered background noise and was removed. For all tracks that were successfully matched in the first two attempts The matched raw measurement data will be sent to the observation Mamba branch for reconstruction, and subsequent variational Bayesian filtering updates will be performed. For unmatched high-resolution detection boxes... For high-resolution isolated detection boxes, the system initializes them as new tracks; however, for tracks that still do not match after two association attempts... Mark it as missing if it is consecutive If a frame fails to rematch, it is permanently deleted from memory.
7. The method according to claim 4, characterized in that, Step S6 includes: After determining the prior distribution of Mamba-guided noise, it is necessary to solve for the joint posterior distribution of state and noise. Using Mean-Field variational theory, it is approximately decomposed into a product of mutually independent marginal distributions: in, This represents the target's state vector at time k; Indicates the measurement noise covariance; Within the variational Bayesian framework, this makes the approximate distribution... The closest distribution to the true posterior distribution This is equivalent to maximizing the lower bound of evidence. According to Bayes' theorem, the joint probability can be expanded into the product of the observation likelihood and the prior: in, This represents the pseudo-measure vector of the Mamba observation network reconstructed stream output used for posterior update at time k; This represents the set of cumulative historical measurement sequences from the initial time to time k-1. By taking the derivative of the maximizing evidence lower bound using variational calculus and setting the derivative to zero, we obtain the general form of the optimal solution for the approximate distribution: for any factor The logarithm of the optimal solution is equal to the expectation of the logarithm of the joint probability over all other variables. Therefore, for the measurement noise covariance... The best approximate posterior The following fixed-point equations must be satisfied: Will include The term is treated as constant absorption, which further simplifies to a recursive fixed-point equation: Gaussian observation likelihood Reverse Saudi A priori with Mamba module boots Substituting into the above equation and utilizing the trace-cycle invariant property of the matrix... have to: in, d represents the predictive degrees of freedom of the inverse Wieshart prior distribution; d represents the dimension of the measurement feature space. Let represent the predictive inverse scaling matrix of the inverse Wissaud prior distribution; Represents the target state The residual second-order expectation term incorporates the bias of the reconstructed measurement and the uncertainty of the current state estimate: in, It is the outer product of the pseudo-measurement innovation vector. It is the covariance of the state estimation error mapped to the measurement space; Observation and deduction The form is completely isomorphic to the logarithmic form of the standard inverse Wissaud distribution; the iteration variable is set as... initial value The true posterior is approximated through alternating iterations of the following two steps; Post-state update: utilizing the first Noise estimation Perform Kalman measurement updates: in, Let represent the one-step prior state error covariance matrix of the i-th variational iteration of the target state at time k; This represents the transpose of the system observation matrix; This represents the time-varying measure-noise covariance estimation matrix obtained in the (i-1)th variational iteration; This represents the target posterior state estimate vector after the i-th iteration update at time k; This represents the one-step prior state prediction of the target at time k; This represents the Kalman gain matrix at time k during the i-th iteration; The one-step posterior state error covariance matrix represents the estimate of the target state at time k in the i-th variational iteration. Noise Posterior Update: Calculate the second moment using the latest state estimate Update the posterior parameters of the IW distribution: in, Let represent the predicted degrees of freedom of the inverse Wissaud posterior distribution after the i-th variational iteration update at time k; Let represent the predicted degrees of freedom of the inverse Wissaud prior distribution at time k; Let represent the inverse scaling matrix of the inverse Wissaud posterior distribution after the i-th variational iteration update at time k; Let represent the inverse scaling matrix of the inverse Wissaud prior distribution at time k; The expected noise covariance for the next iteration is then derived: Reaching the maximum number of iterations Then, take the final result. As The final posterior state estimate of the target at time step 1 is output.