Vector array high-precision direction estimation method based on wideband signal covariance matrix focusing

CN122592324APending Publication Date: 2026-08-18HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610818979.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]本发明是为了解决现有矢量阵列宽带信号方位估计方法中,多频带独立估计后加权综合易导致目标混叠与方位偏移,且在低信噪比和阵列误差环境下估计精度低、鲁棒性差的问题,现提供一种基于宽带信号协方差矩阵聚焦的矢量阵列高精度方位估计方法

Benefits of technology

[0012]与现有技术相比,本发明具有以下有益效果:通过神经网络实现多频带协方差矩阵的智能聚合,避免了传统加权综合导致的目标混叠与方位偏移,显著提升了宽带信号的方位估计精度;所提出的AMSTUNet网络嵌入混合注意力软阈值子网络,能够自适应抑制非均匀噪声,在低信噪比及阵列误差条件下展现出优异的鲁棒性;通过将可微矢量ESPRIT算法嵌入网络,构建了端到端训练框架,无需真实聚合协方差矩阵标签即可完成网络优化;仿真与实测数据验证表明,本发明在小角间距分辨、多目标估计及实际水声环境中的性能均显著优于传统方法,为水下矢量阵列宽带高精度测向提供了具有实际应用价值的技术途径。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122592324A_ABST
    Figure CN122592324A_ABST
Patent Text Reader

Abstract

The application discloses a vector array high-precision azimuth estimation method based on broadband signal covariance matrix focusing, and belongs to the technical field of underwater acoustic signal processing. The method is aimed at the problems that the existing method is prone to target aliasing and azimuth deviation after independent estimation and weighted combination of multiple frequency bands, and has low estimation accuracy and poor robustness in a low signal-to-noise ratio and array error environment. The method comprises the following steps: constructing a covariance matrix of multi-band sound pressure and velocity signals to form a real input feature map; inputting the real input feature map into a pre-trained AMSTUNet network, wherein the network adopts a U-shaped encoder-decoder structure and embeds a mixed attention soft threshold sub-network, is used for adaptively suppressing non-uniform noise and extracting features, and obtains an aggregated covariance matrix feature map; performing complex reconstruction, Hermite symmetry and normalization processing on the aggregated covariance matrix feature map to obtain an aggregated covariance matrix; extracting a sound pressure and velocity signal subspace from the aggregated covariance matrix, and solving a wave direction of arrival estimation by using a rotation invariant relationship. The application is suitable for underwater acoustic azimuth estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of underwater acoustic signal processing technology. Background Technology

[0002] In the field of underwater acoustic detection, traditional acoustic pressure arrays are limited by the design requirements of "large aperture + multiple elements," which restricts their application on miniaturized mobile platforms. Vector sensors, on the other hand, can simultaneously measure scalar sound pressure and vector particle velocity in the sound field, achieving equivalent beam performance with smaller apertures. This provides a key approach for high-precision DOA estimation on limited platforms.

[0003] Early DOA estimation methods primarily relied on beamforming, which, while applicable to vector arrays, limited multi-target estimation accuracy due to the Rayleigh criterion. Subspace-based algorithms (such as MUSIC and ESPRIT) improved accuracy, but their performance degraded sharply in low signal-to-noise ratio underwater acoustic environments. Sparse representation algorithms (such as sparse Bayesian learning, ...) While SVD (Simplified Vector Array Determination) has the potential for high accuracy, it suffers from problems such as high computational cost, strong parameter dependence, and weak environmental adaptability. In recent years, deep learning methods have been introduced into DOA estimation, but their application on vector arrays is still in its early stages, and model adaptability and structure optimization require further in-depth research.

[0004] In addition, broadband signals can improve direction finding accuracy through frequency diversity gain, but existing processing procedures mostly adopt the "frequency division-independent estimation-weighted synthesis" method, which has fundamental limitations such as high computational complexity, easy aliasing of multi-target azimuth information and false estimation. Summary of the Invention

[0005] This invention addresses the problems in existing vector array broadband signal azimuth estimation methods, such as target aliasing and azimuth shift caused by weighted synthesis after independent estimation of multiple frequency bands, and low estimation accuracy and poor robustness under low signal-to-noise ratio and array error environments. It provides a high-precision vector array azimuth estimation method based on focusing the broadband signal covariance matrix.

[0006] The high-precision azimuth estimation method for vector arrays based on wideband signal covariance matrix focusing, as described in this invention, includes:

[0007] Step 1: A uniform linear array of vector hydrophones is used to receive the broadband signal in the far field to be measured, and multi-band sound pressure and vibration velocity signals are obtained to construct the covariance matrix for each band; the covariance matrix is ​​then used to construct a real-valued input feature map.

[0008] Step 2: Input the input feature map into the pre-trained AMSTUNet network to obtain the aggregated covariance matrix feature map;

[0009] The AMSTUNet network adopts a U-shaped encoder-decoder structure, with a hybrid attention soft thresholding subnetwork embedded in both the encoder and decoder to adaptively suppress non-uniform noise in the covariance matrix and extract features; and aggregates the input covariance matrix feature maps.

[0010] Step 3: Perform complex reconstruction, Hermitian symmetry, diagonal loading, and normalization on the feature map of the aggregated covariance matrix in sequence to obtain an aggregated covariance matrix with positive definiteness and Hermitian properties.

[0011] Step 4: Extract the signal subspace of sound pressure and vibration velocity from the aggregated covariance matrix, use the rotationally invariant relationship between sound pressure and vibration velocity to solve the trigonometric function estimate of the azimuth angle, and then calculate the direction of arrival estimation result.

[0012] Compared with existing technologies, this invention has the following advantages: It achieves intelligent aggregation of multi-band covariance matrices through neural networks, avoiding target aliasing and azimuth shift caused by traditional weighted synthesis, and significantly improving the azimuth estimation accuracy of broadband signals; the proposed AMSTUNet network, embedded with a hybrid attention soft thresholding subnetwork, can adaptively suppress non-uniform noise and exhibits excellent robustness under low signal-to-noise ratio and array error conditions; by embedding the differentiable vector ESPRIT algorithm into the network, an end-to-end training framework is constructed, allowing network optimization to be completed without actual aggregation of covariance matrix labels; simulation and experimental data verification show that this invention significantly outperforms traditional methods in small angular spacing resolution, multi-target estimation, and performance in actual underwater acoustic environments, providing a practically valuable technical approach for broadband high-precision direction finding of underwater vector arrays. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the array model;

[0014] Figure 2 Block diagram of the broadband signal vector array DOA estimation algorithm (AMSTU_ESPRIT)

[0015] Figure 3 This is a diagram of the architecture of the Attention Soft Threshold U-Shaped Network (AMSTUNet).

[0016] Figure 4 This is a diagram of the Attention Soft Thresholding (AMST) subnetwork architecture;

[0017] Figure 5The graphs show the changes in DOA and error of the algorithm based on the attention soft threshold U-shaped network (AMSTUNet) over time at a fixed angular interval; where (a) is the change in DOA of the algorithm based on the attention soft threshold U-shaped network (AMSTUNet) over time at a fixed angular interval; and (b) is the change in error of the algorithm based on the attention soft threshold U-shaped network (AMSTUNet) over time at a fixed angular interval.

[0018] Figure 6 The graphs show the changes in estimated DOA and error over time for the U-shaped symmetric network (UNet) algorithm at a fixed angular interval; where (a) is the graph showing the change in estimated DOA over time for the U-shaped symmetric network (UNet) algorithm at a fixed angular interval; and (b) is the graph showing the change in algorithm estimation error over time for the U-shaped symmetric network (UNet) algorithm at a fixed angular interval.

[0019] Figure 7 The graphs show the changes in DOA and error of the CNN-based algorithm over time at a fixed angular interval; where (a) shows the change in DOA of the CNN-based algorithm over time at a fixed angular interval; and (b) shows the change in error of the CNN-based algorithm over time at a fixed angular interval.

[0020] Figure 8 The graphs show the changes in DOA and error of the rotation-invariant subspace algorithm (ESPRIT) over time at fixed angular intervals; where (a) is the graph showing the change in DOA of the rotation-invariant subspace algorithm (ESPRIT) over time at fixed angular intervals; and (b) is the graph showing the change in the error of the rotation-invariant subspace algorithm (ESPRIT) over time at fixed angular intervals.

[0021] Figure 9 The graphs show the changes in DOA and error of the Multiple Signal Classification (MUSIC) algorithm over time at a fixed angular interval; where (a) is the graph showing the change in DOA of the MUSIC algorithm over time at a fixed angular interval; and (b) is the graph showing the change in the error of the MUSIC algorithm over time at a fixed angular interval.

[0022] Figure 10 For fixed angle intervals Norm singular value decomposition algorithm ( -SVD) estimates the DOA and error over time; where (a) is the result of fixed angular intervals. Norm singular value decomposition algorithm ( (a) SVD estimation of DOA over time; (b) plot of DOA at fixed angular intervals. Norm singular value decomposition algorithm ( -SVD) estimation error over time;

[0023] Figure 11 This is a symmetrical layout signal orientation distribution diagram;

[0024] Figure 12 Comparison of RMSE and accuracy of different algorithms as a function of angular spacing; (a) is a comparison of RMSE of different algorithms as a function of angular spacing; (b) is a comparison of accuracy of different algorithms as a function of angular spacing.

[0025] Figure 13 Comparison of RMSE and accuracy of different algorithms as a function of SNR; where (a) is a comparison of RMSE and SNR of different algorithms; and (b) is a comparison of accuracy and SNR of different algorithms.

[0026] Figure 14 Estimate RMSE and accuracy coefficients as a function of error for different algorithms A comparison chart of the changes; where (a) shows the RMSE estimated by different algorithms as a function of the error coefficient. (b) Comparison chart of changes; (c) shows the estimation accuracy of different algorithms as a function of the error coefficient. Comparison chart of changes;

[0027] Figure 15 This is a test situation diagram;

[0028] Figure 16 The actual data processing results based on the Attention Soft Thresholding U-Shaped Network (AMSTUNet) algorithm—time-location history diagram;

[0029] Figure 17 The actual data processing results of the Multiple Signal Classification (MUSIC) algorithm are shown below; (a) is the actual data processing result of the Multiple Signal Classification (MUSIC) algorithm—time-location history diagram; (b) is the actual data processing result of the Multiple Signal Classification (MUSIC) algorithm—spatial spectrum diagram.

[0030] Figure 18 for Norm singular value decomposition algorithm ( The actual data processing results of -SVD; where (a) is Norm singular value decomposition algorithm ( (a) Actual data processing results of SVD—Time-location history plot; (b) is Norm singular value decomposition algorithm ( The actual data processing result of -SVD—spatial spectrum; Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0032] Specific implementation method one: Refer to Figures 1 to 4 This embodiment describes a high-precision azimuth estimation method for vector arrays based on wideband signal covariance matrix focusing. The method includes:

[0033] Step 1: A uniform linear array of vector hydrophones is used to receive the broadband signal in the far field to be measured, and multi-band sound pressure and vibration velocity signals are obtained to construct the covariance matrix for each band; the covariance matrix is ​​then used to construct a real-valued input feature map.

[0034] Step 2: Input the input feature map into the pre-trained AMSTUNet network to obtain the aggregated covariance matrix feature map;

[0035] The AMSTUNet network adopts a U-shaped encoder-decoder structure, with a hybrid attention soft thresholding subnetwork embedded in both the encoder and decoder to adaptively suppress non-uniform noise in the covariance matrix and extract features; and aggregates the input covariance matrix feature maps.

[0036] Step 3: Perform complex reconstruction, Hermitian symmetry, diagonal loading, and normalization on the feature map of the aggregated covariance matrix in sequence to obtain an aggregated covariance matrix with positive definiteness and Hermitian properties.

[0037] Step 4: Extract the signal subspace of sound pressure and vibration velocity from the aggregated covariance matrix, use the rotationally invariant relationship between sound pressure and vibration velocity to solve the trigonometric function estimate of the azimuth angle, and then calculate the direction of arrival estimation result.

[0038] Furthermore, in this invention, the method for constructing the input feature map in real form using the covariance matrix in step 1 is as follows:

[0039] The received wideband signal of the far field to be measured is subjected to Fourier transform and decomposed into P narrowband components. The covariance matrix corresponding to each frequency band is calculated. The covariance matrices of each frequency band are concatenated along the feature dimension, and then global normalization is performed. The real part and imaginary part of the complex matrix are concatenated along the feature channel dimension to obtain the input feature map.

[0040] ;

[0041] ;

[0042] ;

[0043] in, Indicates along the first dimension (index) ) performs stacking operations; operators and These are used to extract the real and imaginary components of a complex matrix, respectively. Here is the covariance matrix for each frequency band, and N is the number of channels in the vector hydrophone array. For array receiving data matrix, Let be the covariance matrix when the noise is stationary white noise.

[0044] Furthermore, in this invention, the training process of the AMSTUNet network in step 2 is as follows:

[0045] Step 21: Use a uniform linear array of vector hydrophones to receive broadband signals with known azimuth labels, obtain multi-band sound pressure and vibration velocity signals, construct the covariance matrix of each band, construct the input feature map in real form, and use the trigonometric function value of the azimuth angle of the incoming wave direction as the azimuth label to form a training dataset.

[0046] Step 22: Input the input feature map into the AMSTUNet network to be trained. The AMSTUNet network to be trained adopts a U-shaped encoder-decoder structure. Both the encoder and the decoder are embedded with a hybrid attention soft thresholding subnetwork, which is used to adaptively suppress non-uniform noise in the covariance matrix and extract features, and output the aggregated covariance matrix feature map.

[0047] Step 23: Perform complex reconstruction, Hermitian symmetry, diagonal loading and normalization on the feature map of the aggregated covariance matrix in sequence to obtain the aggregated covariance matrix; extract the signal subspace from the aggregated covariance matrix using eigenvalue decomposition, and solve the trigonometric function estimate of the azimuth angle of the incoming wave direction by combining the rotation invariance between sound pressure and vibration velocity and between the two subarrays.

[0048] Step 24: Construct a loss function from the error between the estimated trigonometric function and the directional label in Step 21. Use the calculation result of the loss function to update the weight parameters of the AMSTUNet network through backpropagation. Repeat steps 22 to 24 until the AMSTUNet network converges to obtain the trained AMSTUNet network.

[0049] The loss function includes permutation minimization of root mean square error, trigonometric identity constraint loss, and sound pressure-velocity consistency constraint loss.

[0050] Furthermore, in this invention, in step 24, the loss function is:

[0051] ;

[0052] in,

[0053] ;

[0054] ;

[0055] ;

[0056] Where i and j represent the feature category index and the source number index, respectively. This represents the true label value matrix, with the first row being the cos1 label, the second row being the cos2 label, and the third row being the sin label. Represents the network estimation matrix. azimuth of the incoming wave direction to the target The estimated value, The cosine value of the azimuth angle of the j-th source obtained by the relationship between sound pressure levels (i.e., the rotation invariance of the sound pressure subspace of adjacent array elements); This represents the estimated cosine value of the azimuth angle of the j-th source obtained by the relationship between sound pressure and x-axis vibration velocity; The estimated sine value of the j-th source azimuth angle obtained by the relationship between sound pressure and y-axis vibration velocity. Let K represent the dataset, and K be the number of information sources. For index The set of all permutations, for each permutation and combination The predicted vector after permutation is defined as:

[0057] ;

[0058] in, Let π(k) represent the estimated value of the i-th type of feature and the π(k)-th source after rearranging by π, where π(k) represents... The mapping corresponding to the kth permutation in the algorithm is used to align the column order of the predicted values ​​with the column order of the true values.

[0059] Furthermore, in this invention, in step 3, the aggregated covariance matrix is:

[0060] ;

[0061] in, For the aggregated covariance matrix, This is the output matrix after regularization; This indicates the search for the Frobenius norm. This represents a complex matrix with a batch dimension of 1 and a matrix size of 3N×3N.

[0062] Furthermore, in this invention, the regularized output matrix :

[0063] ;

[0064] in, The output matrix after Hermitian symmetry For white noise, It is an identity matrix with the same dimensions as R;

[0065] ;

[0066] in, This is a complex matrix obtained by complex reconstruction of the features of the aggregated covariance matrix.

[0067] Furthermore, in this invention, the complex matrix obtained by complex reconstruction of the aggregated covariance matrix features is... :

[0068] ;

[0069] in, This indicates that the network output matrix will be used. The complex matrix, composed of real and imaginary parts, is uniformly divided along the second dimension and has a dimension of 1×6N×3N, where N is the number of array elements and 3N corresponds to the number of channels in the vector array (each hydrophone contains three components: sound pressure, x-velocity, and y-velocity). This means taking all elements in the first dimension of Y, the elements in the second dimension with indices from 1 to 3N, and all elements in the third dimension as the real part of the complex matrix; This means taking all elements in the first dimension of Y, the elements in the second dimension with indices from 3N+1 to 6N, and all elements in the third dimension as the imaginary part of the complex matrix, where j represents the imaginary unit.

[0070] Furthermore, in this invention, in step 4, the estimated value of the trigonometric function is:

[0071] ;

[0072] ;

[0073] in, The azimuth angle of the incoming wave direction to the target. For signal wavelength, The element spacing of a uniform linear array, This represents the diagonal element value obtained by eigenvalue decomposition of the rotation-invariant matrix of the sound pressure subspace (a complex number whose phase includes azimuth information). This represents the argument (phase value) of a complex number. This represents the diagonal element values ​​(complex numbers, whose real part corresponds to the azimuth cosine and the imaginary part corresponds to the azimuth sine) obtained by performing eigenvalue decomposition on the rotation-invariant matrix of the vibrational velocity subspace.

[0074] Over the past decade, sparse signal representation theory has spurred the development of various high-precision DOA estimation algorithms, such as Sparse Bayesian Learning (SBL), L1-SVD, and Orthogonal Matching Pursuit (OMP). These algorithms offer significant accuracy potential but suffer from drawbacks including high computational cost, strong parameter dependence, and poor environmental adaptability. In recent years, deep learning methods, leveraging their data-driven advantages, have significantly improved estimation accuracy in complex environments such as low signal-to-noise ratios and array mismatches. However, research on vector sensor arrays remains in its early stages, and model adaptability needs further refinement. Furthermore, while frequency diversity gain can improve direction-finding accuracy for broadband signals, the traditional "frequency division-independent estimation-weighted synthesis" approach is not only computationally intensive but also prone to aliasing and spurious estimations in multi-target scenarios, limiting the upper limit of broadband performance.

[0075] To address the aforementioned problems, this invention proposes a high-precision DOA estimation algorithm (AMSTU_ESPRIT) that combines a UNet network improved with a CBAM soft thresholding function sub-network (AMSTUNet) with the ESPRIT method. The network part of this algorithm filters out redundant noise while extracting useful features, completing the aggregation of multi-band covariance matrices. Then, using the vector ESPRIT algorithm, the aggregated matrix output by the network is directly transformed into obtainable azimuth information. Since the ESPRIT algorithm itself is differentiable, it can directly achieve high-precision and robust DOA estimation through end-to-end training. The main contributions are summarized as follows:

[0076] (1) An end-to-end DOA estimation framework was designed. This framework completes the multi-band covariance matrix aggregation process through a neural network and embeds the vector ESPRIT algorithm to construct the backpropagation path from azimuth information to network parameters. Through network training, the problems of target aliasing and azimuth offset caused by the weighted synthesis of multi-band results are effectively solved, and the azimuth estimation accuracy is significantly improved.

[0077] (2) An AMSTUNet structure is proposed. An AMST subnetwork is constructed by improving the existing soft thresholding network using CBAM, which enables it to adapt to the uneven distribution of noise levels in the covariance matrix. It is then fused with UNet, which enhances the noise suppression capability while effectively preserving the ability to represent weak signals and small features, thereby improving the estimation performance of the model in complex scenarios.

[0078] (3) The DOA estimation performance of the model under the condition of a four-vector uniform linear matrix was analyzed through experimental simulation. The simulation results show that the present invention has significant advantages in various scenarios such as small spacing, low signal-to-noise ratio, and array error. We also further proved the effectiveness and superiority of the proposed method by comparing the DOA estimation time-location history (BTR) plots of measured data.

[0079] Combination Figure 2 To illustrate, consider a... The spacing is A uniform linear array (ULA) composed of vector hydrophones receives... A broadband signal from the far field The frequency range of the signal is The number of quick shots is And array spacing The sound pressure level output by the array With vibrational velocity components , It can be modeled as:

[0080] (1)

[0081] in, For array manifold matrix, and These characterize the directional coupling relationship between vibration velocity and sound pressure, respectively. , and The sound pressure level output by the array With vibrational velocity components , The corresponding noise component.

[0082] The array receiving data matrix can be integrated into Each column of this matrix corresponds to a complete array observation vector at a snapshot time.

[0083] When the noise is stationary white noise, its covariance matrix is ​​expressed as:

[0084] (2)

[0085] in, The target signal covariance matrix; For noise power, It is the normalized noise covariance matrix of the vector hydrophone array.

[0086] In the data preprocessing stage, the received signal is first subjected to a Fourier transform to decompose it into P narrowband signals, and the covariance matrix corresponding to each frequency band is calculated. Then, all covariance matrices are concatenated along the feature dimension. To further improve numerical stability, global normalization is performed on the concatenated data. Finally, the complex matrix is ​​converted to real number form by concatenating the real and imaginary parts along the height dimension for input into the neural network for subsequent processing. The formula is as follows:

[0087] (3)

[0088] (4)

[0089] (5)

[0090] in, Indicates along the first dimension (index) Stacking operations are performed. Operators and These are used to extract the real and imaginary components of a complex matrix, respectively. Let P represent a real matrix of size P×3N×3N.

[0091] The main neural network architecture is based on UNet, and an improved soft thresholding algorithm is introduced to suppress noise interference and improve model robustness. This network completes the mapping from a high-dimensional extended covariance matrix to a single-frequency aggregation matrix. Furthermore, covariance matrixization is performed to ensure that the network output satisfies the Hermitian symmetry and positive definiteness of the covariance matrix. It is worth noting that we cannot directly obtain the true aggregation covariance matrix as a reference label, making it difficult to train the network by comparing the neural network output with the true value. This constitutes a major obstacle to achieving supervised learning. We consider embedding the vector ESPRIT algorithm into the network's iteration process to perform azimuth estimation on the aggregation covariance matrix output by the neural network, obtaining signal source azimuth parameters whose true values ​​can be known in advance.

[0092] Unlike methods such as MUSIC that calculate spatial spectra, the ESPRIT algorithm is inherently differentiable, allowing it to successfully complete the backpropagation process of neural network training. While root-MUSIC is also differentiable, it cannot effectively utilize the combined sound pressure and vibration velocity information from vector hydrophones, cannot resolve port / ship ambiguity issues, and struggles to achieve omnidirectional direction finding. Therefore, we chose the ESPRIT method, which combines differentiability with applicability to vector arrays. In forward propagation, it receives the covariance matrix estimated by the network front-end and calculates the direction-of-arrival (DOA) estimate. In backpropagation, its differentiability is used to construct a gradient calculation path, feeding back the error between the DOA estimate and the true value to the front-end network in the form of a gradient. Thus, we constructed an end-to-end trainable network model and completed DOA estimation for a broadband signal vector array.

[0093] To address the needs of this network model, a dataset pattern is proposed to obtain a dataset consisting of L pairs of observations and their corresponding azimuth trigonometric function values, denoted as:

[0094] (6)

[0095] in, From equation (5),

[0096] (7)

[0097] The label set is used for network training. During dataset construction, the label set... The known signal orientation Obtained through direct calculation. During network training, It is obtained from the sound pressure relationship between different array elements. It is obtained from the relationship between the sound pressure level of the same element and the sound velocity along the x-axis. It is obtained from the relationship between the sound pressure of the same element and the sound velocity on the y-axis.

[0098] AMSTUNet architecture:

[0099] like Figure 4 The AMSTUNet architecture diagram described above uses UNet as its core framework. UNet is a symmetrical U-shaped encoder-decoder structure. To better suit the tasks described in this paper, it has been modified: the double convolutional blocks within the same layer have been replaced with AMST subnetworks.

[0100] AMSTUNet architecture such as Figure 3As shown. In the encoder section (left), the AMST sub-network (blue line) is first used to increase the number of feature channels and initially filter out noise. Then, a max-pooling layer with a stride of 2 is used for downsampling (orange line), reducing the spatial size of the feature map. This allows neurons to integrate information from a wider region, thereby capturing more global and abstract high-level semantic information. In the decoder section (right), the high-level information obtained from upsampling (green line) is concatenated with the high-resolution detail information transmitted through skip connections (gray dashed lines), and then passed through the AMST sub-network.

[0101] The feature maps are then fused (blue line) and redundant information is suppressed again. Transposed convolutions are used to perform self-learning upsampling (green line), thereby reconstructing the spatial dimensions of the feature maps layer by layer. Finally, a single-channel feature matrix is ​​obtained through convolutional layers and ReLU nonlinear transformation.

[0102] The purpose of the AMST subnetwork is to attempt to suppress redundant features using a soft thresholding function. A soft thresholding function is introduced into the neural network to construct a Deep Residual Shrinking Network (DRSN). DRSN, as a one-dimensional convolutional network, has each channel corresponding to a segment of vibration signal. It adaptively learns thresholds for each signal segment through internal subnetworks and uses linear transformation layers to suppress noise-related features. However, the input to the network described in this invention is not a one-dimensional signal, but a covariance matrix obtained from signal transformation. The noise distribution within the covariance matrix is ​​uneven. Directly using the DRSN network structure can easily lead to insufficient filtering of redundant information or incorrect removal of effective information. Therefore, this invention introduces a Convolutional Attention (CBAM) module into the soft thresholding module. By fusing channel and spatial attention, the network can adaptively assign differentiated thresholds to different feature locations, thereby achieving accurate and differentiated suppression of redundant features.

[0103] Sub-network structure such as Figure 4 As shown, the module's foundation consists of two convolutional layers, each containing batch normalization (BN), a ReLU activation function, and a convolutional layer. This part doubles the number of feature channels in the UNet encoder and performs cross-layer feature fusion in the decoder. Thresholding The noise level at each location in the feature map is used to quantize the value, and its value must be non-negative. Therefore, after the convolution operation, the absolute value of the feature map needs to be taken to obtain the non-negative feature intensity used to calculate the threshold. Complete threshold. From characteristic intensity Channel attention threshold coefficient Spatial attention threshold coefficient Multiplying them together, we get:

[0104] (8)

[0105] in, , , , .

[0106] The process for generating the channel attention threshold coefficient is as follows: Figure 4 As shown in blue. First, global average pooling (GAP) and global max pooling (GMP) are performed on the feature map along the spatial dimension to obtain two C-dimensional vectors describing the global statistical information of each channel. Then, the two vectors are each passed through a shared fully connected layer for dimensionality reduction. The outputs are added together and then passed through a second fully connected layer to restore the number of channels C, thus obtaining the channel attention threshold coefficients. .

[0107] The process for generating spatial attention threshold coefficients is as follows: Figure 4 As shown in green. First, global average pooling and global max pooling are performed on the input features along the channel dimension, resulting in two spatial dimensions. The feature maps are then concatenated along the channel dimension to form a dual-channel feature. This feature is then fused and transformed through two consecutive convolutional layers, ultimately generating a single-channel threshold coefficient matrix with the same spatial dimension as the input. .

[0108] Both attention modules employ a hierarchical activation design of "ReLU + Sigmoid": ReLU is placed in the first layer to introduce nonlinearity and alleviate gradient vanishing, while Sigmoid is placed in the last layer to normalize the output coefficients to the (0,1) interval. The resulting hybrid threshold coefficients... , With characteristic strength Multiplication ensures adaptive thresholding The values ​​are non-negative and have a controllable range. By performing a soft thresholding operation according to the following formula, adaptive noise suppression can be achieved while ensuring numerical stability.

[0109] (9)

[0110] Where x is the input feature of the soft thresholding function, and y is the output feature after passing through the soft thresholding function.

[0111] In summary, the AMSTUNet model architecture used in this invention is based on the UNet architecture and optimized with the introduction of an AMST subnetwork. The UNet architecture extracts semantic information through the encoder and transmits the spatial details captured by the encoder to the decoder through skip connections, achieving efficient fusion of the two. Semantic information helps the network accurately distinguish between signals and interference, while spatial location information provides high-fidelity spatial structure data, helping the network to more accurately characterize the array manifold. The AMST subnetwork performs soft thresholding on layer-by-layer features, enabling the model to more accurately identify and filter noise in the covariance matrix. This network model lays a crucial foundation for achieving high-precision DOA estimation in this invention.

[0112] Covariance matrix processing and vector ESPRIT algorithm

[0113] This section describes how to make the aggregate matrix output by the network conform to the mathematical properties of the covariance matrix, and then use the vector ESPRIT algorithm to obtain the label estimate or DOA estimate result.

[0114] First, output the network matrix Divide the complex matrix into two parts along the second dimension, using them as the real and imaginary parts respectively, and reconstruct the corresponding complex representation:

[0115] (10)

[0116] Secondly, Hermitian symmetry is performed to ensure that the eigenvalues ​​of the matrix are all real numbers:

[0117] (11)

[0118] Next, regularization is performed using diagonal loading, which is equivalent to introducing power of Weak white noise to avoid matrix ill-conditioning:

[0119] (12)

[0120] Finally, the matrix is ​​normalized:

[0121] (13)

[0122] Through the above steps, a covariance matrix with good mathematical properties is obtained. .

[0123] Therefore, the ESPRIT algorithm can be further combined to estimate the direction-of-arrival (DOA) parameters. The basic ESPRIT algorithm estimates DOA based on the rotation invariance between subarrays, meaning that the displacement vectors between corresponding elements in two subarrays remain consistent. The formula is:

[0124] (14)

[0125] (15)

[0126] in, This is the rotation matrix between the two subarrays; This represents the sound pressure signal vector received by the 1st to N-1th array elements (subarray 1) at time t. This represents the sound pressure signal vector received by the 2nd to Nth array elements (subarray 2) at time t. Represents the source vector. Represents the noise vector. Let be the rotation factor corresponding to the k-th source, where j is the imaginary unit and f is the signal frequency. The time delay difference of the received signal between corresponding array elements of the two subarrays at the k-th received signal angle;

[0127] The relationship between sound pressure and vibration velocity components, unique to vector array signals, exhibits similar rotational invariance, making... ,at this time:

[0128] (16)

[0129] (17)

[0130] in, This is the rotation matrix between the sound pressure and vibration velocity components. It is a diagonal matrix, with the diagonal elements being the cosine values ​​of the azimuth angles of each signal source. This is a diagonal matrix, with diagonal elements representing the sine values ​​of the azimuth angles of each signal source. The sound pressure signal can be converted into a vibration velocity signal using a rotation matrix that depends only on the target's azimuth, reflecting the fundamental idea of ​​the ESPRIT algorithm.

[0131] Considering the noise-free case, since the signal subspace is the same as the subspace spanned by the vector array manifold matrix, that is... Therefore, there must exist a non-singular matrix. Make:

[0132] (18)

[0133] Among them, the subspaces corresponding to the two subarrays are and Extract all sound pressure components and x and y velocity components to construct the signal subspace. and , Vibration velocity The corresponding subspace is The subspace corresponding to the complex vibration velocity signal This represents the rotation-invariant matrix of the sound pressure subspace. This represents the rotationally invariant matrix between sound pressure and vibration velocity.

[0134] Obviously, the matrix and Since they are similar matrices, they have the same eigenvalues. Therefore, by... Eigenvalue decomposition yields diagonal element values Thus, the orientation of the target can be calculated using equations (19) and (20).

[0135] (19)

[0136] (20)

[0137] in, To achieve this by rotating the invariant matrix of the sound pressure subspace The diagonal element values ​​(complex numbers) obtained by eigenvalue decomposition. To achieve this by using the rotationally invariant matrix between sound pressure and vibration velocity The diagonal element values ​​obtained by eigenvalue decomposition.

[0138] Loss function and backpropagation:

[0139] Similar to standard regression task loss metrics, the quality of the algorithm's estimation is evaluated by comparing the error between the DOA (Direction of Arrival) values ​​estimated by the network model and the true values. However, since it's impossible to restrict the order of the predicted orientation angles to be the same as the labels, we attempt to calculate the root mean square error (RMSE) of all permutations and combinations of the predicted values ​​with the true labels, and select the minimum error as the final loss function error value. Formula:

[0140] (twenty one)

[0141] Where i and j represent the row and column indices, respectively, and K is the number of information sources. For index The set of all permutations, for each permutation and combination The predicted vector after permutation is defined as .

[0142] Furthermore, by leveraging the characteristics of the labels themselves, the update direction of the network is further restricted, enabling it to reach a stable convergence state as quickly as possible. One approach is to utilize the properties of trigonometric functions, specifically obtained through the relationship between sound pressure and x-axis vibration velocity. The relationship between sound pressure and y-axis vibration velocity was obtained. The space should meet Secondly, it utilizes the combined information of sound pressure and sound pressure velocity, that is, it obtains the information through the relationship between sound pressure levels. The results obtained from sound pressure and x-axis vibration velocity The space should meet Based on the above characteristics, we define two other error functions, as follows:

[0143] (twenty two)

[0144] (twenty three)

[0145] because and for eigenvalues The real and imaginary parts are inherently related, so no permutation or combination is required.

[0146] For dataset The final loss function is the average of the loss functions for all samples. Given the dataset... All features are trigonometric function values ​​defined in the interval [-1, 1], which inherently possess dimensionless and uniform scale properties. Therefore, the calculation results of the three loss functions are directly comparable and can be linearly added. The final network training error function is shown below:

[0147] (twenty four)

[0148] After calculating the error between the two, the training loss of the network can be obtained, thus realizing the backpropagation process from the output of the vector ESPRIT algorithm, through the covariance matrix processing stage, to the parameters of each layer of the network model. Figure 2 As shown by the red dashed line, after completing the backpropagation process described above, the network model can continuously optimize its parameters through iterative training. Once training is complete, the model can effectively learn and output a low-noise aggregated covariance matrix directly from the extended covariance matrix of the input broadband signal, thus providing reliable input for the subsequent vector ESPRIT algorithm and realizing the complete DOA estimation process.

[0149] Simulation experiments / performance evaluation:

[0150] Simulation environment settings:

[0151] 1) Model and Training Strategy: The simulation considers a 4-element vector uniform linear array with an element spacing of 0.375 times the wavelength. Two independent far-field broadband signals are received simultaneously, with a sampling time of 6 seconds and a sampling frequency of 6000Hz. The broadband signal frequency ranges from 500Hz to 2000Hz, divided into 76 sub-bands in 20Hz increments. Two random integers with an angular spacing greater than 3° are randomly generated within the range [0, 359] degrees as the signal's direction of arrival. This minimum spacing constraint aims to avoid inherent resolution difficulties caused by signals being too close together. The noise is assumed to be Gaussian white noise, with a signal-to-noise ratio (SNR) ranging from -20dB to 20dB, totaling 41 sets. 500 sets of received data are generated under each SNR condition, forming a dataset of 500 * 41 = 20500 data points. To enhance the model's ability to handle low signal-to-noise ratio (SNR) environments, we added additional data with SNRs ranging from -20dB to -5dB, generating 550 sets of data for each SNR, ultimately resulting in an augmented dataset containing 20,500 + 16 * 500 = 28,500 sets of data. In this dataset, the training and validation sets were divided in a 9:1 ratio. The initial learning rate for network training was 0.01, with each batch consisting of 32 data points, and a total of 300 iterations. Every 40 epochs, the learning rate was increased by 0.5 times.

[0152] 2) Performance Metrics: This invention uses the root mean square error (RMSE) and accuracy of the estimation results to measure the DOA estimation performance of the algorithm. For Monte Carlo simulation experiments, the formula for calculating the RMSE is as follows:

[0153] (25)

[0154] Where M represents the total number of Monte Carlo experiments, and K represents the total number of signal sources in each experiment; This represents the estimated angle for the k-th real source in the m-th experiment. This represents the true angle of the k-th information source.

[0155] (26)

[0156] It is a defined function for calculating circumferential error, used to avoid error distortion caused by jumps across the 0 point.

[0157] Accuracy reflects the percentage of successful estimates out of M trials, and is calculated using the following formula:

[0158] (27)

[0159] in, This represents the root mean square error of the m-th experiment. The indicator function is used; if the root mean square error is less than 1, then the estimation is considered a successful estimation.

[0160] (28)

[0161] 3) Algorithm Comparison: To comprehensively evaluate the algorithm's performance, we selected five high-precision DOA estimation methods for comparison. First, among neural network methods, we chose the standard UNet to specifically evaluate the denoising effect of the proposed soft thresholding module. This UNet architecture is similar to... Figure 3 The simulation is identical, but without the AMST module. Considering the widespread use of U-shaped deep neural networks (Unet), we include it in the comparison. The CNN network used in the simulation consists of five stacked convolutional layers, each followed by batch normalization and ReLU activation. Both comparison methods replace the network part (i.e., the AMSTUNet module) in the algorithm of this invention with standard UNet and CNN respectively, while keeping the rest unchanged; therefore, their full names are UNet_ESPRIT and CNN_ESPRIT, respectively. For simplicity, the network names AMSTUNet, UNet, and CNN will be used as abbreviations for the corresponding methods below. Secondly, in traditional subspace algorithms, a comparison is made with the classic MUSIC algorithm and the ESPRIT algorithm. The latter is the same as the vector ESPRIT algorithm used in the backend of the algorithm of this invention, thus directly highlighting the performance gain brought by the front-end neural network. Finally, sparse representation high-precision algorithms, represented by SVD, are introduced as a higher-level performance reference to demonstrate the advancement of the broadband covariance matrix aggregation processing strategy adopted in this invention.

[0162] Simulation results:

[0163] 1) Generalization ability to untrained angles

[0164] Since the directions of arrival (ROAs) of the data used in the training set are all integers, and it is difficult to traverse all possible angle values ​​under a single signal-to-noise ratio (SNR), we designed the following test to evaluate the algorithm's angle generalization ability: SNR of -5 dB, two signals spaced... The direction of arrival of signal 1 By iterating through all integer angles from 0 to 359, the direction of arrival of signal 2 is... In the tests, all algorithms were evaluated using the same set of data simulated under the above conditions to eliminate the influence of background noise randomness on the results. The signal history and estimation error of each algorithm are shown below. Figures 5-10 As shown in the figure; theta1 and theta2 are related to... and Correspondingly, the root mean square error (RMSE) and accuracy are listed in Table 1.

[0165] Table 1 RMSE and accuracy of different algorithms

[0166]

[0167] Observation Figures 5 to 10 It can be seen that the estimated signal trajectory is highly consistent with the true signal trajectory, and it is overall smooth without significant jumps. The estimation error basically remains within 0.35°, and the error of Signal 2 does not increase significantly compared with that of Signal 1. The above results indicate that the algorithm can effectively complete the DOA estimation task and has excellent generalization ability for untrained angles. From the perspective of overall performance, the root mean square error (RMSE) of all methods is sorted from small to large as: AMSTUNet < UNet < CNN < SVD sparse algorithm < vector ESPRIT < MUSIC. In terms of accuracy: AMSTUNet = UNet = CNN > SVD sparse algorithm > MUSIC > vector ESPRIT. This shows that the introduced neural network can significantly improve the accuracy and stability of estimation, and AMSTUNet has the best performance.

[0168] Longitudinal observation of the estimation error diagrams of different algorithms reveals that centered around the 146 s and 326 s, the estimation errors of each algorithm increase to varying degrees. At this time, Angle 1 is located at or nearby, and the corresponding Angle 2 is located at or nearby, as shown below Figure 11 shown. Further simulation of the situation with different angular spacings also shows that if the azimuths of the two signals are symmetric about the x-axis, the estimation accuracy will decrease significantly. The main reason for the performance degradation is that when the incident directions of the two signals are symmetric about the x-axis, their steering vectors are highly correlated, resulting in the degradation of the eigenstructure of the received signal covariance matrix, and thus making it difficult for the algorithm to effectively distinguish these two directions. In such a symmetric layout, the sensitivity of the algorithm to noise will be exacerbated, and even a small perturbation may lead to a sharp drop in the estimation accuracy. In contrast, the neural network method adopted in the present invention can effectively suppress noise interference through a data-driven denoising mechanism, alleviate the estimation degradation problem caused by the correlation of steering vectors, and thus significantly improve the estimation stability and accuracy in symmetric scenarios.

[0169] 2) Generalization ability for angular spacing

[0170] Next, the DOA estimation errors and accuracies of each algorithm at different angular spacings will be tested to systematically evaluate their angle resolution capabilities. The angular spacing in the training dataset is randomly selected, and during the testing process, the angular spacing between the two signals The angle is increased from 0° to 180°. Since the azimuth angle has circular symmetry, the angle spacing within this range can cover all possible relative orientations of the signal in space. The formula for calculating the angle spacing is shown in equation (26). For each set angle spacing, test data is generated under a signal-to-noise ratio of -5 dB. Similar to the simulation above, the direction of arrival of signal 1 is set... By iterating through all integer angles from 0° to 359°, the direction of arrival of signal 2 is... Subsequently, the simulation results of the 360 ​​angle combinations generated were statistically averaged. This traversal averaging process aims to eliminate random biases introduced by specific angle combinations, thereby ensuring the generalizability and statistical robustness of the evaluation conclusions.

[0171] from Figure 12 As can be seen, with the increase of angular spacing, the estimation error (RMSE) of all algorithms gradually decreases, while the accuracy increases synchronously. This trend is consistent with the theoretical expectation that "DOA estimation performance improves with the increase of signal angular separation." The AMSTUNet algorithm proposed in this invention performs best in both RMSE and accuracy. For the extreme case of an angular spacing of 5°, the RMSE estimated by the algorithm in this invention is only 0.4134, while traditional algorithms cannot effectively distinguish between two signals under this condition. With the increase of angular spacing, the performance of AMSTUNet further improves: when the angular spacing exceeds 20°, its estimation accuracy reaches 100%, and the RMSE converges to below 0.2°; at 180°, the RMSE reaches a minimum of 0.07546°. Under the same conditions, the estimated RMSE of CNN and UNet networks is greater than 0.1°.

[0172] Analysis shows that, under conditions of small angular spacing, skip connections can effectively optimize the network's ability to capture fine-grained features, which is the key reason why UNet has a more significant performance advantage over CNN. Furthermore, the added hybrid attention soft thresholding module can more effectively restore key information in the original signal, which makes the AMSTUNet algorithm ultimately exhibit better estimation accuracy.

[0173] 3) Generalization ability regarding signal-to-noise ratio

[0174] To systematically evaluate the algorithm's performance boundaries, we set up simulation experiments with different signal-to-noise ratios (SNRs). The SNR range was set from -20 dB to 20 dB, covering the entire range from the critical state of extremely low SNR to good channel conditions, to comprehensively analyze the algorithm's performance trend with SNR. The signal orientation generation strategy for the test set was consistent with that for the training set, but continuous orientation values ​​were used instead of integer angles to cover a wider orientation space. To obtain statistically reliable results, 5000 Monte Carlo experiments were independently conducted under each SNR condition. The algorithm performance changes with SNR as follows: Figure 13As shown

[0175] Analyzing the performance of various algorithms under different signal-to-noise ratios (SNRs) reveals that when the SNR is below 15 dB, the neural network-based DOA estimation algorithm significantly outperforms traditional methods. However, when the SNR exceeds 15 dB, the estimation accuracy of traditional methods gradually surpasses that of neural network methods. This phenomenon is because neural networks, as data-driven function approximators, inherently suffer from biases introduced by finite training samples and model structure during the estimation process; while traditional methods, under Gaussian noise with sufficient SNR, can achieve asymptotically unbiased parameter estimation, thus exhibiting superior theoretical limit performance.

[0176] Overall, the AMSTUNet algorithm proposed in this invention achieves the best comprehensive performance: in a favorable environment with high signal-to-noise ratio (≥0dB), the algorithm's RMSE is less than 0.1°, while the accuracy remains above 99.9%, demonstrating extremely high precision and reliability; in a harsh environment with low signal-to-noise ratio (0dB~-15dB), its estimation error only slowly increases to 0.373°, while still maintaining a high accuracy of 96.9%; in extreme cases (-20dB), the algorithm can still provide effective estimation with an estimation error of 1.063° and an accuracy of 56.62%.

[0177] This is mainly due to the optimization of two key aspects of the algorithm proposed in this invention: First, it can effectively fuse information from multi-band covariance matrices, avoiding the errors caused by weighted averaging of multi-subband results in traditional methods; Second, the AMSTUNet module in the algorithm, which introduces a learnable soft thresholding module oriented towards the covariance matrix based on the UNet architecture, can adaptively filter noise and retain key signal features, exhibiting superior nonlinear expression and feature selection capabilities.

[0178] 4) Generalization ability to array errors:

[0179] Non-ideal conditions such as element position errors, channel amplitude and phase errors, and element mutual coupling effects can severely disrupt the array manifold, thereby significantly impairing algorithm performance. To evaluate the algorithm's generalization ability to array errors, we construct a parameterized error model: the position error follows a parameterized model with a mean of zero and a standard deviation of 0. The channel gain follows a normal distribution; the channel gain fluctuates based on unity gain. The standard deviation of the phase offset is The mutual coupling error is modeled using an exponential decay model, and the coupling coefficient between adjacent array elements is... And this is extended to the coupling relationship between channels in the vector array through a weight matrix. This is the error severity coefficient, and the range is designed to cover the entire spectrum from minor deviations to critical performance failure.

[0180] Based on this error model, we generated a new dataset considering array error and retrained and tested three neural network models on this dataset. It covers data samples with varying degrees of random array error at each signal-to-noise ratio level. During test set generation, a fixed random number seed for the array error was used to simulate the robustness of the algorithm under array error present at a real-world time. Since no prior information relating to array error and steering vector is utilized, the above end-to-end training and testing strategy can be directly extended to other array error scenarios. Furthermore, to avoid the influence of angular resolution on the results, the angular distance between the two signals is constrained to be greater than 20° to more realistically reflect the algorithm's adaptability to array error.

[0181] Test results are as follows Figure 14 As shown: AMSTUNet, UNet, and CNN are the performance of models retrained on new datasets; while AMSTUNet' is the original model trained on an ideal dataset; ESPRIT, MUSIC, and The SVD algorithm is tested based on an ideal array manifold and does not utilize prior knowledge of array errors for correction.

[0182] When the array has no errors, its response function matches the ideal parameterized model, thus all methods achieve optimal performance in DOA estimation. As array errors increase, the estimation error of the ESPRIT algorithm rises almost linearly, due to its direct reliance on numerical estimation of the array manifold. The MUSIC algorithm, benefiting from its spectral peak search mechanism, is relatively less sensitive to errors. The -SVD method achieves regularization through sparsity constraints. Therefore, the estimation errors of both methods show a more gradual trend with increasing array error compared to ESPRIT.

[0183] In contrast, neural network-based methods exhibit stronger robustness to array errors. Even the original model AMSTUNet' trained on an ideal dataset shows lower estimation errors than traditional methods such as ESPRIT and MUSIC. Furthermore, retraining with a new dataset incorporating array errors further enhances the generalization ability of the network model, with the proposed AMSTUNet demonstrating the best performance. Specifically, as the error severity coefficient ρ increases from 0 to 1, AMSTUNet's estimation error only slightly increases from 0.1485° to 0.5468°, while maintaining a stable estimation accuracy above 88.8%. This result fully validates the algorithm's superior robustness to array errors. This is because its data-driven nature allows it to directly learn and extract robust deep features from array data containing errors, thus maintaining high-precision DOA estimation over a considerably wide error range.

[0184] Actual data processing:

[0185] The algorithm's performance was evaluated using real-world acoustic data. The test site was located in Qiandao Lake, China, and the test conditions were as follows: Figure 15 As shown, a 4-element vector uniform linear array with an element spacing of 0.5m was fixedly installed on an underwater maneuvering platform. A sound source towed by the test vessel was located in the far field, continuously emitting a broadband random signal ranging from 1000Hz to 5000Hz. Both the sound source and the maneuvering platform were at an underwater depth of 15 meters. To verify the omnidirectional estimation performance of the algorithm, the two sound sources were positioned on opposite sides of the array. 496 snapshots were processed as one frame, with 50 frames per second, for a total of 120 seconds of data.

[0186] To verify the performance of the algorithm presented in this paper, MUSIC and The algorithm is compared with _SVD. This choice is based on the following considerations: Due to factors such as sonar deployment, significant array errors exist in the experiment, causing the measured steering vector to deviate severely from the theoretical value. This renders the ESPRIT algorithm ineffective because it cannot directly compensate for the actual steering vector, while MUSIC and L1_SVD can directly utilize the measured steering vector to complete the azimuth estimation. Furthermore, the algorithm in this paper can use the measured steering vector to generate simulation data for network training, compensating for ESPRIT's inability to compensate for steering vectors, thus enabling DOA estimation of real data. Therefore, selecting these two methods for comparison allows for a more reasonable evaluation of the estimation performance of the algorithm in this paper.

[0187] Data processing results are as follows Figures 16 to 18 As shown. Because the AMSTUNet_ESPRIT method does not obtain results through spectral peak search by calculating the spatial spectrum, but directly outputs the DOA estimation value, there is no second spatial spectrum. The MUSIC method is affected by spectral peak aliasing, resulting in both sound sources moving closer to the center to varying degrees. Furthermore, the estimation error is larger at locations where the signal azimuth distance between the beginning and end is small. Overall, it is difficult to accurately reflect the azimuth trajectory of the sound sources. The SVD algorithm can obtain a clear target azimuth trajectory, but it exhibits significant fluctuations. The spatial spectrum shows that this algorithm has a narrow main lobe width, indicating high angular resolution. However, during the weighted superposition of broadband signals, small estimation deviations in each sub-band due to noise or array errors are easily amplified during weighting, affecting the accuracy of the final azimuth estimation and making it difficult to achieve theoretically optimal performance. The AMSTUNet_ESPRIT method, on the other hand, largely avoids these problems and achieves higher estimation accuracy. Figure 16As shown in the azimuth time history graph, the trajectory estimated by this method is relatively smooth and continuous, and can accurately reflect the azimuth trajectory of the target, which fully verifies the effectiveness of the method in real data.

[0188] While the invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features described herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other described embodiments.

Claims

1. A high-precision azimuth estimation method for vector arrays based on focusing of broadband signal covariance matrix, characterized in that, The method includes: Step 1: A uniform linear array of vector hydrophones is used to receive the broadband signal in the far field to be measured, and multi-band sound pressure and vibration velocity signals are obtained to construct the covariance matrix for each band; the covariance matrix is ​​then used to construct a real-valued input feature map. Step 2: Input the input feature map into the pre-trained AMSTUNet network to obtain the aggregated covariance matrix feature map; The AMSTUNet network adopts a U-shaped encoder-decoder structure, with a hybrid attention soft thresholding subnetwork embedded in both the encoder and decoder to adaptively suppress non-uniform noise in the covariance matrix and extract features; and aggregates the input covariance matrix feature maps. Step 3: Perform complex reconstruction, Hermitian symmetry, diagonal loading, and normalization on the feature map of the aggregated covariance matrix in sequence to obtain an aggregated covariance matrix with positive definiteness and Hermitian properties. Step 4: Extract the signal subspace of sound pressure and vibration velocity from the aggregated covariance matrix, use the rotationally invariant relationship between sound pressure and vibration velocity to solve the trigonometric function estimate of the azimuth angle, and then calculate the direction of arrival estimation result.

2. The high-precision azimuth estimation method for vector arrays based on wideband signal covariance matrix focusing according to claim 1, characterized in that, In step 1, the method for constructing the input feature map in real form using the covariance matrix is ​​as follows: The received wideband signal in the far field to be measured is subjected to Fourier transform and decomposed into P narrowband components. The covariance matrix corresponding to each frequency band is calculated. The covariance matrices of each frequency band are concatenated along the feature dimension, and then global normalization is performed. The real and imaginary parts of the complex matrix are concatenated along the feature channel dimension to obtain the input feature map.

3. The high-precision azimuth estimation method for vector arrays based on wideband signal covariance matrix focusing as described in claim 1, characterized in that, In step 2, the training process of the AMSTUNet network is as follows: Step 21: Use a uniform linear array of vector hydrophones to receive broadband signals with known azimuth labels, obtain multi-band sound pressure and vibration velocity signals, construct the covariance matrix of each band, construct the input feature map in real form, and use the trigonometric function value of the azimuth angle of the incoming wave direction as the azimuth label to form a training dataset. Step 22: Input the input feature map into the AMSTUNet network to be trained. The AMSTUNet network to be trained adopts a U-shaped encoder-decoder structure. Both the encoder and the decoder are embedded with a hybrid attention soft thresholding subnetwork, which is used to adaptively suppress non-uniform noise in the covariance matrix and extract features, and output the aggregated covariance matrix feature map. Step 23: Perform complex reconstruction, Hermitian symmetry, diagonal loading and normalization on the feature map of the aggregated covariance matrix in sequence to obtain the aggregated covariance matrix; extract the signal subspace from the aggregated covariance matrix using eigenvalue decomposition, and solve the trigonometric function estimate of the azimuth angle of the incoming wave direction by combining the rotation invariance between sound pressure and vibration velocity and between the two subarrays. Step 24: Construct a loss function from the error between the estimated trigonometric function and the directional label in Step 21. Use the calculation result of the loss function to update the weight parameters of the AMSTUNet network through backpropagation. Repeat steps 22 to 24 until the AMSTUNet network converges to obtain the trained AMSTUNet network. The loss function includes permutation minimization of root mean square error, trigonometric identity constraint loss, and sound pressure-velocity consistency constraint loss.

4. The high-precision azimuth estimation method for vector arrays based on wideband signal covariance matrix focusing according to claim 3, characterized in that, In step 24, the loss function is: ; in, ; ; ; Where i and j represent the feature category index and the source number index, respectively. This represents the true label value matrix, with the first row being the cos1 label, the second row being the cos2 label, and the third row being the sin label. Represents the network estimation matrix. azimuth of the incoming wave direction to the target The estimated value, The cosine value of the j-th source azimuth angle is estimated by the relationship between sound pressure levels; This represents the estimated cosine value of the azimuth angle of the j-th source obtained by the relationship between sound pressure and x-axis vibration velocity; The estimated sine value of the j-th source azimuth angle obtained by the relationship between sound pressure and y-axis vibration velocity. Let K represent the dataset, and K be the number of information sources. For index The set of all permutations, for each permutation and combination The predicted vector after permutation is defined as: ; in, Let π(k) represent the estimated value of the i-th type of feature and the π(k)-th source after rearranging by π, where π(k) represents... The source sequence number mapping corresponding to the kth permutation position is used to align the column order of the predicted value with the column order of the true value.

5. The high-precision azimuth estimation method for vector arrays based on wideband signal covariance matrix focusing according to claim 1, characterized in that, In step 3, the aggregated covariance matrix is: ; in, For the aggregated covariance matrix, This is the output matrix after regularization; This indicates the search for the Frobenius norm. This represents a complex matrix with a batch dimension of 1 and a matrix size of 3N×3N.

6. The high-precision azimuth estimation method for vector arrays based on wideband signal covariance matrix focusing according to claim 1, characterized in that, Regularized output matrix : ; in, The output matrix after Hermitian symmetry For white noise, It is an identity matrix with the same dimensions as R; ; in, This is a complex matrix obtained by complex reconstruction of the features of the aggregated covariance matrix.

7. The high-precision azimuth estimation method for vector arrays based on wideband signal covariance matrix focusing according to claim 1, characterized in that, The complex matrix obtained by complex reconstruction of the aggregated covariance matrix features : ; in, This indicates that the network output matrix will be used. A complex matrix composed of real and imaginary parts, uniformly divided along the second dimension, has a dimension of 1×6N×3N, where N is the number of array elements and 3N corresponds to the number of channels in the vector array. This means taking all elements in the first dimension of Y, the elements in the second dimension with indices from 1 to 3N, and all elements in the third dimension as the real part of the complex matrix; This means taking all elements in the first dimension of Y, the elements in the second dimension with indices from 3N+1 to 6N, and all elements in the third dimension as the imaginary part of the complex matrix, where j represents the imaginary unit.

8. The high-precision azimuth estimation method for vector arrays based on wideband signal covariance matrix focusing according to claim 1, characterized in that, In step 4, the estimated values ​​of the trigonometric functions are: ; ; in, The azimuth angle of the incoming wave direction to the target. For signal wavelength, The element spacing of a uniform linear array. This represents the diagonal element values ​​obtained by eigenvalue decomposition of the rotation-invariant matrix of the sound pressure subspace. Indicates the argument of a complex number. This represents the diagonal element values ​​obtained by performing eigenvalue decomposition on the rotationally invariant matrix of the vibrational velocity subspace.