A high-resolution direction-of-arrival estimation method based on deep unfolding network
By mapping the ISTA iterative process to a fixed-layer neural network through deep unfolding, and combining gradient descent and residual connections, the performance degradation problem of traditional methods under low signal-to-noise ratio and limited snapshot conditions is solved, and efficient high-resolution direction-of-arrival estimation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN ENG UNIV
- Filing Date
- 2026-02-28
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional high-resolution direction-of-arrival estimation methods suffer from performance degradation under low signal-to-noise ratio and limited snapshot conditions, and have high computational complexity, making it difficult to meet the requirements of real-time processing and energy efficiency balance.
A high-resolution orientation estimation method based on deep unfolded networks is adopted. The iterative shrinking threshold algorithm ISTA is mapped to a neural network with a fixed number of layers through the dCv-ISTA-Net network. The gradient descent module, the near-end mapping module and the residual connection are combined. The iterative process is adaptively optimized by the convolutional autoencoder, and the position priority loss function is introduced to improve the estimation accuracy.
It significantly improves estimation performance under low signal-to-noise ratio and limited snapshot conditions, reduces computational complexity, meets real-time processing requirements, and maintains high estimation accuracy in complex environments.
Smart Images

Figure CN122131228A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of direction-of-arrival estimation technology, specifically a high-resolution orientation estimation method based on deep unfolded networks. Background Technology
[0002] Direction of Arrival (DOA) estimation is one of the core tasks of array signal processing, with wide applications in radar, sonar, and wireless communication. Currently, with the increasing frequency of human activities in the ocean and the continuous development of underwater vibration reduction and noise reduction technologies, underwater acoustic DOA estimation faces the following challenges: On the one hand, under non-ideal conditions such as low signal-to-noise ratio caused by strong marine environmental noise and a small number of effective snapshots due to target maneuvering, there is an urgent need for DOA estimation algorithms that can maintain high estimation accuracy even in complex environments; on the other hand, due to the strict power consumption and computational resource limitations typically imposed on underwater detection platforms, many existing high-resolution DOA methods suffer from high computational complexity, making it difficult to meet the requirements of real-time processing and energy efficiency balance.
[0003] Traditional high-resolution DOA estimation methods mainly include subspace-based methods, sparse reconstruction-based methods, and deconvolutional beamforming methods. However, subspace-based methods suffer from severe performance degradation under low signal-to-noise ratio and limited snapshot conditions; while sparse reconstruction algorithms and deconvolutional beamforming methods perform well in terms of estimation accuracy and spatial resolution, their computational complexity is high. Deep unfolded networks, which integrate physical models and data-driven hybrid methods, combine interpretability with the strong fitting capabilities of deep learning, and have become an important research direction in the field of DOA estimation. Summary of the Invention
[0004] The purpose of this invention is to address the problem that traditional high-resolution direction-of-arrival estimation methods suffer significant performance degradation under conditions of low signal-to-noise ratio and limited snapshots, and to provide a high-resolution orientation estimation method based on deep unfolded networks.
[0005] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0006] A high-resolution orientation estimation method based on deep unfolded networks includes the following steps:
[0007] The array received signal is obtained using a uniform linear array, and the CBF spatial spectrum vector corresponding to the array received signal is obtained. Then, the CBF spatial spectrum vector is input into the dCv-ISTA-Net network to obtain the output signal power distribution vector, i.e., the azimuth estimation result.
[0008] The dCv-ISTA-Net network includes a gradient descent module and a proximal mapping module. The gradient descent module is used to perform the gradient descent mechanism in the ISTA algorithm on the input to obtain the intermediate result after gradient descent. The proximal mapping module is used to map the intermediate result after gradient descent to a high-dimensional space through a convolutional layer, and perform feature selection and noise suppression through an encoder and decoder. Finally, the high-dimensional features are restored to the original dimension through a convolutional layer. The encoder and decoder are cascaded with ReLU activation functions and are cascaded across layers through residual connections.
[0009] Furthermore, the intermediate result after gradient descent is represented as follows:
[0010] ;
[0011] in, For a pre-defined dictionary matrix, This is an intermediate result. For the model number The output of the layer, For the first The learnable step size parameter of the layer, This is the CBF spatial spectrum.
[0012] Furthermore, the proximal mapping module is represented as follows:
[0013] ;
[0014] in, For learnable proximal mapping operators, The convolution kernel is The number of channels is convolutional layers, For decoder, For activation function, For encoder, The convolution kernel is A convolutional layer with 1 channel. This represents the number of channels in the convolutional layer.
[0015] Furthermore, the residual connection is represented as:
[0016] .
[0017] Furthermore, the loss function of the dCv-ISTA-Net network is expressed as:
[0018] ;
[0019] ;
[0020] in, The number of training samples in each batch. To improve the loss function, For the first The model output results for each sample. For the first The true label of each sample This refers to the temperature parameter.
[0021] Furthermore, the dCv-ISTA-Net network is a pre-trained dCv-ISTA-Net network, and the training process is as follows:
[0022] Step 1: Based on the spatial region of interest and the preset grid, discretize the far-field spatial region to obtain overcomplete spatial discrete sampling points;
[0023] Step 2: Use a uniform linear array to acquire the array received signal, and use conventional beamforming to obtain the CBF azimuth spectrum of the array received signal under spatial overcompleteness at each discrete sampling point;
[0024] Step 3: Obtain the CBF spatial spectrum vector based on the CBF orientation spectrum of all discrete sampling points under overcomplete spatial conditions;
[0025] Step 4: Based on the CBF spatial spectral vector, obtain the signal power distribution vector, expressed as:
[0026] ,
[0027] in, For CBF space spectral vectors, For a pre-defined dictionary matrix, The signal power distribution vector. This refers to the noise component in the azimuth spectrum;
[0028] Step 5: Construct a training set using the CBF spatial spectral vector and the signal power distribution vector, and train the dCv-ISTA-Net network using the training set to obtain the trained dCv-ISTA-Net network.
[0029] Furthermore, the CBF azimuth spectrum of the array received signal in the desired beamforming direction is expressed as follows:
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] ;
[0035] in, For the first The power of each signal, for A uniform linear array in the desired direction of beamforming The CBF weighted vector, , The number of beam grid points, For weighted vectors, For the spacing between array elements, For array to receive signals, for Quantity, The incident azimuth angle is... For the number of array elements, , For the number of snapshots, For the far-field array manifold matrix, For far-field signal vectors, For noise vectors, It is the symbol for imaginary numbers. The center frequency of the signal. For the speed of sound, for The first moment Signals incident from one direction, For the expected operation, For the first The far-field steering vector of a signal. .
[0036] Furthermore, the CBF azimuth spectrum under the overcomplete spatial domain is represented as follows:
[0037] ;
[0038] in, The number of grid points, For the first Each grid point in the direction On the signal power, For the first The guide vector of each grid point.
[0039] Furthermore, the learnable parameters of the dCv-ISTA-Net network include the gradient descent step size for each layer. and in the near-end mapping module and encoder and decoder .
[0040] Furthermore, the total number of parameters of the dCv-ISTA-Net network is:
[0041] ;
[0042] in, The number of channels in the convolutional layer. This represents the number of cascaded layers in the network.
[0043] The beneficial effects of this invention are:
[0044] This application maps the Iterative Shrink Thresholding (ISTA) algorithm to a fixed-layer neural network using deep unrolling techniques, thereby inheriting interpretability while introducing data-driven learning capabilities. Through a convolutional autoencoder with residual connections, the iterative process is adaptively optimized, and a loss function focuses the network on the accurate reconstruction of the signal support set. This addresses the problem of significant performance degradation of estimation methods under low signal-to-noise ratio and limited snapshot conditions. The technical solution of this application significantly improves estimation performance. Attached Figure Description
[0045] Figure 1 This is a schematic diagram of an array receiver model;
[0046] Figure 2 Here is a diagram of the dCv-ISTA-Net model structure;
[0047] Figure 3 This is a schematic diagram of the spatial spectrum output of each network layer under a single objective.
[0048] Figure 4 This is a schematic diagram of the spatial spectrum for testing the generalization capability of the number of information sources;
[0049] Figure 5 A schematic diagram of the spatial spectrum of each algorithm;
[0050] Figure 6 A schematic diagram illustrating the RMSE of DOA estimation for each algorithm under different signal-to-noise ratios;
[0051] Figure 7 A schematic diagram illustrating the RMSE of DOA estimation for each algorithm under different snapshot numbers;
[0052] Figure 8 A schematic diagram showing the resolution probability of each algorithm at different angular intervals;
[0053] Figure 9 A schematic diagram showing the resolution probability of each algorithm under different signal-to-noise ratios;
[0054] Figure 10 Time-azimuth history of the CBF algorithm;
[0055] Figure 11Time-orientation history of the dCv-RL algorithm;
[0056] Figure 12 Time-orientation history diagram of the MUSIC algorithm;
[0057] Figure 13 Time-azimuth history diagram of the SBL algorithm;
[0058] Figure 14 Time-location history diagram of the SPICE algorithm;
[0059] Figure 15 The time-orientation history of the dCv-ISTA-Net algorithm. Detailed Implementation
[0060] It should be noted that, where there is no conflict, the various embodiments disclosed in this application can be combined with each other.
[0061] Specific Implementation Method 1: The high-resolution orientation estimation method based on deep unfolded networks described in this implementation method specifically includes the following:
[0062] 1. Signal Model
[0063] 1.1 Array Model
[0064] consider A few uncorrelated far-field narrowband signals From the direction ( The number of incident elements is The spacing between array elements is A uniform linear array (ULA). Using the leftmost element as the reference element, a schematic diagram of the array receiving signals is shown below. Figure 1 As shown.
[0065] Array in The received signal vector at time t can be represented as:
[0066] (1)
[0067] in, for Time of the first The received data of each array element; for The first moment Signals incident from all directions; for Time of the first Additive noise received by each array element; For the first Each element receives the signal. The time delay relative to the reference array element; It is the symbol for imaginary numbers. The center frequency of the signal. The speed of sound.
[0068] Define the entire array in The received signal vector at time is Far-field signal vector noise vector The above formula can be further expressed as:
[0069] (2)
[0070] in, For the number of snapshots, For the far-field array manifold matrix, the first... The far-field steering vector of each signal The expression is:
[0071] (3)
[0072] Assuming that the signal and noise are independent of each other, the covariance matrix of the array received data can be expressed as:
[0073] (4)
[0074] in, This represents the expectation operation. and These represent the signal covariance matrix and the noise covariance matrix, respectively.
[0075] 1.2 Deconvolution Beamforming Model
[0076] To distinguish between the desired direction of beamforming and the direction of arrival. This application uses the symbol to represent the desired direction of beamforming. Indicates. For A uniform linear array, in the direction The CBF weighted vector is:
[0077] (5)
[0078] Assuming the noise is isotropic white Gaussian noise, and that the signals are uncorrelated with each other and with the noise, then the array receives the following signals: In direction The CBF azimuth spectrum can be represented as:
[0079] (6)
[0080] in, For the first The power of each signal, This represents the noise component in the azimuth spectrum.
[0081] The far-field spatial domain is discretized into... grid points , direction The signal power on is ,but The positions of the non-zero elements correspond to the azimuth of the actual signal, i.e. CBF azimuth spectrum under complete airspace conditions It can be represented as:
[0082] (7)
[0083] As shown in equation (7), the power distribution of the signal in the spatial domain has the following linear relationship with the CBF azimuth spectrum:
[0084] (8)
[0085] CBF spatial spectral vector The signal power distribution vector is:
[0086] Beam diagram dictionary It can be represented as:
[0087] (9)
[0088] Because the signal power is constant and non-negative, that is , need to Apply nonnegativity constraints. Also, due to the dictionary matrix... It is usually pathological, therefore it needs to be introduced into the objective function. The regularization term improves the stability of the solution, resulting in the optimization problem as follows:
[0089] (10)
[0090] The regularization term can be viewed as a regularization term for the spatial signal power. The sparse constraints can be used to suppress noise and spurious peaks by utilizing the sparse priors of the source distribution. The optimization problem of Equation (10) can be solved by ISTA, and the process is shown in Table 1.
[0091] Table 1: ISTA Solution Flow
[0092]
[0093] In Table 1, the step size is usually taken. Lipschitz constant The reciprocal, , To obtain the eigenvalue function. Proximal mapping function. The expression is:
[0094] (11)
[0095] 2. A Direction Estimation Method Based on Deep Unfolded Networks
[0096] 2.1 Network Structure
[0097] While ISTA has the advantages of not requiring matrix inversion and having low single-step computational complexity, it still has the following limitations:
[0098] (1) ISTA requires a large number of iterations to converge and the overall convergence is slow;
[0099] (2) The regularization term of this algorithm needs to be set in advance and lacks an adaptive adjustment mechanism;
[0100] (3) The regularization factor of this algorithm is affected by factors such as noise statistics and measurement matrix, making it difficult to select.
[0101] To address the aforementioned issues, this application proposes a deconvolutional beamforming method, dCv-ISTA-Net, based on a deep unfolded network. This method maps the ISTA iterative process described in Table 1 to a neural network with a fixed number of layers, models a suitable proximal mapping function through a convolutional autoencoder, and introduces a residual connection structure to achieve cross-layer gradient propagation.
[0102] The structure of dCv-ISTA-Net is as follows Figure 2 As shown, the network uses the CBF spatial spectrum The input is spatial power distribution, and the output is spatial power distribution. The network is... It consists of several cascaded layers, each corresponding to one iteration update of ISTA. The step size parameter, regularization term, etc. are all learned adaptively through the neural network.
[0103] dCv-ISTA-Net can be divided into the following three main modules:
[0104] (1) Gradient Descent Module: This module corresponds to the gradient descent mechanism in the ISTA algorithm. Model 1 The mathematical expression for implementing gradient descent in layers is:
[0105] (12)
[0106] in, For a pre-defined dictionary matrix, For intermediate parameters, For the model number The output of the layer, For the first The learnable step size parameter of the layer.
[0107] (2) Proximal mapping module: This module replaces manually designed regularization terms with convolutional autoencoders. Its processing flow is as follows:
[0108] I. The intermediate results after gradient descent Through convolutional layers Mapping to higher-dimensional space , The number of channels in the convolutional layer is set to [number] in this application. .
[0109] II. Encoders using symmetrical structures and decoder This enables feature selection and noise suppression. and Cascading ReLU activation functions between them can preserve non-negative signal components while suppressing noise interference. Among them, the encoder... The receptive field is expanded using two cascaded, progressively larger convolutional kernels, while the decoder... Using encoder The symmetrical structural design, with progressively smaller convolutional kernels, allows the network to capture broader contextual information and effectively improve its feature learning capabilities.
[0110] III. Through convolutional layers From high-dimensional features Restored to the original signal dimension .
[0111] No. The structure of the layer near-end mapping module is as follows: Figure 2 The mathematical expression shown within the dashed box is:
[0112] (13)
[0113] (3) Residual module
[0114] Inspired by ResNet, a residual connection is introduced between the encoding and decoding parts, and its mathematical expression is:
[0115] (14)
[0116] Residual connections allow gradients to propagate directly back to shallower layers, effectively solving the gradient vanishing problem in deep networks. Simultaneously, cross-layer connections enable the network to efficiently utilize lower-layer features, preserving original signal information and accelerating convergence.
[0117] The learnable parameters of dCv-ISTA-Net include the gradient descent step size for each layer. and convolutional layers in the near-end mapping module and encoder and decoder The parameter set:
[0118] ;
[0119] Optimization is performed using gradient backpropagation. The total number of parameters in the network is... It is independent of the dimension of the input.
[0120] 2. Loss Function
[0121] The Mean-Square Error (MSE) loss function has certain limitations in DOA estimation tasks. Its overemphasis on the precise recovery of amplitude information leads to insufficient accuracy in source location. To address this issue, this application proposes a location-priority loss function, enabling the network to focus more on the recovery of non-zero locations, thereby improving DOA estimation accuracy.
[0122] First, we introduce the hyperbolic tangent transform function with temperature parameters, namely:
[0123] (15)
[0124] in, It is the hyperbolic tangent function. For the temperature parameter, this application sets it as follows: . Functions are often used for approximation in sparse reconstruction algorithms. Norm, satisfying:
[0125] (16)
[0126] From this property, we can know that the temperature parameter control Sensitivity to amplitude. Specifically, for smaller amplitudes... , Able to have a larger amplitude Mapping to values approaching 1, with smaller magnitudes The mapping is to values that approach 0.
[0127] Considering that the training samples contain additive noise, which inevitably leads to errors in the preset label amplitude values, the following fidelity-constrained loss function is constructed:
[0128] (17)
[0129] in, The output of the model, This loss function uses hyperbolic tangent transform to suppress the proportion of amplitude error in the gradient, forcing the model to focus more on the accurate recovery of the signal support set location, thus enhancing the model's robustness to noise.
[0130] In batch training, the loss function The expression is shown in equation (18):
[0131] (18)
[0132] in, The number of training samples in each batch, For the first The model output for each sample. This corresponds to the actual label.
[0133] 3. Numerical Simulation Analysis
[0134] All simulations in this section were performed on the same computer, which has an Intel(R) Core(TM) i9-13900HX CPU, an NVIDIA GeForce RTX4060 Laptop GPU, and 32GB of RAM. All simulations were implemented in Python.
[0135] This section uses simulation data to construct the dataset. Consider... A uniform linear array of elements, with an element spacing of half a wavelength (the center frequency of the signal). ), speed of sound Far-field airspace According to equal sine intervals Divided into From discrete grid points, we obtain a spatial grid point set. The simulation generates 10,000 training samples, each containing 5 independent far-field narrowband signals with their incident directions randomly distributed within a spatial grid. The signal-to-noise ratio (SNR) of each signal is randomly distributed between -10 and 5 dB, and the number of snapshots is fixed at 100. Due to the large dynamic range of the spatial spectrum, to eliminate the influence of outliers and accelerate the convergence process of model training, the generated input CBF spatial spectrum and labels are further normalized to obtain the final training sample set. The number of cascaded layers in dCv-ISTA-Net is set to... The Xavier initialization method was used to initialize the convolutional layer parameters, and the AdamW optimizer was selected for training (learning rate was 100%). (Batch size is 64).
[0136] To evaluate the performance of the proposed algorithm, the trained dCv-ISTA-Net was compared with traditional DOA estimation methods such as CBF, Richardson-Lucy deconvolution (dCv-RL), MUSIC, Sparse Bayesian learning (SBL), and Sparse Iterative Covariance Estimation (SPICE). Specifically, dcv-RL used the results of 10 iterations, while the other iterative algorithms had an upper limit of 200 iterations.
[0137] 3.1 Feasibility Analysis
[0138] Under single-target conditions with a signal-to-noise ratio of 5dB and a snapshot count of 100, Figure 3 The estimation results and intermediate layer outputs of dCv-ISTA-Net are presented. Figure 3 It can be seen that as the number of network layers increases, the background level of the azimuth spectrum gradually decreases, and the spectral peaks tend to become sharper. This result indicates that each network layer of dCv-ISTA-Net is equivalent to one ISTA iteration, demonstrating interpretability. To further test the algorithm's adaptability to unknown sources, under the conditions of a signal-to-noise ratio of 5dB and a snapshot count of 100, Figure 4 The estimation results of the model for incident signals with K=10 sources from different orientations are presented. As shown in the figure, dCv-ISTA-Net successfully estimated the true orientations of 10 incident signals without encountering K>5 samples during the training phase. Thanks to inheriting the prior structure and domain knowledge of the ISTA algorithm, rather than relying entirely on a data-driven learning mechanism, dCv-ISTA-Net exhibits good generalization ability.
[0139] Consider two uncorrelated far-field signals with an incident azimuth interval of 1 / 2. The signal-to-noise ratio is 5dB and the number of snapshots is 100. Figure 5 The spatial spectral estimation results of various algorithms are presented, with the right figure showing a magnified view of the peak values in the left figure. CBF and dCv-RL fail to separate the two targets due to their excessively wide main lobes. While MUSIC, SBL, and SPICE can separate the two targets, their spectral peak widths are relatively wide. This is because MUSIC suffers from decreased subspace orthogonality under noise and limited snapshots, while SBL and SPICE exhibit incomplete iterative convergence, limiting the quality of their sparse solutions. In contrast, dCv-ISTA-Net exhibits the lowest background noise level and the sharpest spectral peak, validating its high-resolution advantage over model-driven methods.
[0140] To evaluate the computational complexity of each algorithm, Table 2 shows the average computation time for each algorithm running 100 times on the same CPU under the above conditions. The results indicate that the sparse algorithms SBL and SPICE suffer from limited real-time performance due to multiple iterations and matrix inversion. dCv-ISTA-Net transforms the iterative algorithm into a fixed-layer neural network using the deep expansion concept and leverages the advantages of offline training and online inference in deep learning, shifting the complexity to the training phase and achieving high real-time performance during inference.
[0141] Table 2: Average computation time for each algorithm
[0142]
[0143] 3.2 Statistical Performance Analysis
[0144] This section evaluates the DOA estimation accuracy and spatial resolution of each algorithm through Monte Carlo method experiments, with 500 Monte Carlo trials conducted for each scenario. The root mean square error (RMSE) is used as the evaluation metric for DOA estimation accuracy. Each far-field sound source is located at an angle. Incident, definition:
[0145] (19)
[0146] in, This indicates the total number of Monte Carlo runs. Indicates the first The Monte Carlo trial The true location of a far-field target. Its estimated value.
[0147] To measure the angular resolution capability of each algorithm, the probability of successfully resolving two targets is chosen as the metric. Consider two far-field sound sources... and For incident light, the condition for successful resolution of two targets is defined as follows:
[0148] (20)
[0149] When the condition in the above formula is true, it is considered that the algorithm is in the first... In this Monte Carlo trial, two targets were successfully distinguished, and the effective resolution angle was defined as the angle with a probability of successful resolution greater than 100%. The angle.
[0150] To verify the effectiveness of the position-first loss function, this section compares dCv-ISTA-Net (denoted as dCv-ISTA-Net1) trained using the position-first loss function with dCv-ISTA-Net (denoted as dCv-ISTA-Net2) trained using the original MSE loss function.
[0151] Assume two targets are incident on ULA from different angles, with the same signal-to-noise ratio, both changing from -15dB to 5dB, and 100 snapshots. Figure 6 The RMSE of DOA estimation for each algorithm under different signal-to-noise ratios (SNRs) is shown. Simulation results show that the proposed method has the highest DOA estimation accuracy in the SNR range below -5 dB. This is attributed to the powerful nonlinear feature extraction capability of deep learning, which enables it to effectively extract the orientation information of weak targets from a noisy background. However, at higher SNRs, the RMSE of this method is slightly higher than that of MUSIC, SBL, and SPICE. This is because, at high SNRs, traditional subspace-based or sparse reconstruction-based methods approach the Cramer-Rao Lower Bound (CRLB), and the inherent function approximation error of neural networks limits their performance at high SNRs. Furthermore, training with a position-first loss function reduces the overall RMSE of the model. Furthermore, the performance improvement is mainly concentrated in the low signal-to-noise ratio range, proving that the proposed loss function can enhance the robustness of the model under low signal-to-noise ratio conditions.
[0152] Assuming two targets are incident on ULA from different angles, with a signal-to-noise ratio of -5dB for both, the number of snapshots changes from 10 to 200. Figure 7 The RMSE of DOA estimation for each algorithm is shown under different snapshot numbers. As the number of snapshots increases, the MUSIC algorithm shows a significant decrease in RMSE due to the more accurate estimation of the sample covariance matrix. Since the covariance domain algorithm SPICE can fully utilize the accumulated power of multiple snapshots compared to the element domain algorithm SBL, it achieves higher DOA estimation accuracy when the number of snapshots is greater than 50. Under low snapshot conditions, the SPICE algorithm's performance deteriorates significantly due to the increased error in covariance matrix estimation. In contrast, dCv-ISTA-Net, although trained with only 100 samples, still maintains excellent estimation accuracy across all snapshot numbers, thanks to its inherited robustness to snapshots from deconvolutional beamforming.
[0153] To quantify the angular resolution of the evaluation algorithm, Figure 8 The results show that, with a fixed signal-to-noise ratio of 0dB and a snapshot count of 100, the probability of resolving two targets using each algorithm varies with the angular interval ( ). The trend of change. When the target interval At the same time, both dCv-ISTA-Net methods can still efficiently distinguish between two targets, significantly outperforming traditional high-resolution algorithms. Figure 9 This further demonstrates the fixed angle interval. The resolution probabilities of each algorithm under different signal-to-noise ratios (SNRs) are shown. dCv-ISTA-Net1 can effectively distinguish dual targets even with an SNR of -10dB, while MUSIC, SPICE, and SBL require an SNR above 0dB to achieve target resolution at equal intervals. This is because the MUSIC algorithm suffers from subspace leakage under low SNR and limited snapshots, leading to a decrease in resolution; while sparse reconstruction algorithms theoretically possess super-resolution capabilities, their performance upper limit is constrained by the finite isometric property; dCv-ISTA-Net, through its powerful data-driven learning capabilities, overcomes the dependence of traditional high-resolution algorithms on ideal model assumptions, thus exhibiting superior performance in neighbor target resolution and weak signal detection.
[0154] 4. Experimental Data Analysis
[0155] This section uses measured data from a specific sea area to verify the practicality of the proposed algorithm. A uniform horizontal linear array of 30 elements was placed on the seabed with an element spacing of 2m. Multiple targets, including surface vessels, were present near the experimental area. In data processing, the data was divided into frames every 5 seconds, and the 320Hz frequency signal was processed. The dCv-ISTA-Net network model used is the same as in the previous section, trained using simulation data and an improved loss function.
[0156] Figures 10 to 15 The time-azimuth history diagrams obtained by each algorithm from processing the measured data are shown. In the experiment, the trajectories of the two target ships with higher energy were close together and... The targets converge nearby. Limited by resolution, traditional algorithms can barely distinguish neighboring targets throughout the entire observation period; only the SPICE algorithm can barely distinguish two targets within 0-100 frames. In contrast, the proposed algorithm dCv-ISTA-Net can clearly distinguish the neighboring target pair within both 0-150 frames and 350-500 frames. Two weaker neighboring targets were identified. These results demonstrate that dCv-ISTA-Net exhibits a significant advantage in orientation resolution compared to traditional DOA estimation algorithms. Furthermore, dCv-ISTA-Net demonstrates a narrower main lobe width and better sidelobe suppression compared to dCv-RL in deconvolution processing. It should be noted that for the data recorded in the experiment... Near-field interference is incorrectly fitted as two spurious far-field targets because methods such as SPICE and dCv-ISTA-Net are based on the far-field plane wave assumption, which is mismatched with the near-field model.
[0157] 5. Conclusion
[0158] This application proposes a deconvolutional beamforming method, dCv-ISTA-Net, based on a deep unfolded network. By fusing the interpretability of model-driven methods with the fitting capability of data-driven methods, it achieves efficient high-resolution DOA estimation. This method maps the iterative process of ISTA to a convolutional neural network with a fixed number of layers. Through learnable gradient step size parameters and a proximal mapping module, the iterative process is adaptively optimized, significantly improving computational efficiency and estimation performance. Furthermore, by introducing a hyperbolic tangent transform into the loss function, the method guides the model to focus on accurate reconstruction of the signal support set, effectively enhancing noise robustness. Simulation results show that the proposed method outperforms traditional high-resolution algorithms under different signal-to-noise ratios and snapshot numbers, and can meet real-time processing requirements. Experimental data processing results further validate the practicality and high-resolution advantages of dCv-ISTA-Net in real underwater acoustic environments. Future work will focus on constructing a self-supervised training method based on deep unfolded networks, utilizing unlabeled data collected in real-world scenarios to further optimize the model's adaptability in complex marine environments, reducing reliance on simulation data.
[0159] It should be noted that the specific embodiments are merely explanations and illustrations of the technical solution of the present invention and should not be used to limit the scope of protection. Any modifications made in accordance with the claims and specification of the present invention that are only partial should still fall within the protection scope of the present invention.
Claims
1. A high-resolution orientation estimation method based on deep unfolded networks, characterized in that... Includes the following steps: The array received signal is obtained using a uniform linear array, and the CBF spatial spectrum vector corresponding to the array received signal is obtained. Then, the CBF spatial spectrum vector is input into the dCv-ISTA-Net network to obtain the output signal power distribution vector, i.e., the azimuth estimation result. The dCv-ISTA-Net network includes a gradient descent module and a proximal mapping module. The gradient descent module is used to perform the gradient descent mechanism in the ISTA algorithm on the input to obtain the intermediate result after gradient descent. The proximal mapping module is used to map the intermediate result after gradient descent to a high-dimensional space through a convolutional layer, and perform feature selection and noise suppression through an encoder and decoder. Finally, the high-dimensional features are restored to the original dimension through a convolutional layer. The encoder and decoder are cascaded with ReLU activation functions and are cascaded across layers through residual connections.
2. The high-resolution orientation estimation method based on deep unfolded networks according to claim 1, characterized in that... The intermediate result after gradient descent is represented as follows: ; in, For a pre-defined dictionary matrix, This is an intermediate result. For the model number The output of the layer, For the first The learnable step size parameter of the layer, This is the CBF spatial spectrum.
3. The high-resolution orientation estimation method based on deep unfolded networks according to claim 2, characterized in that... The proximal mapping module is represented as follows: ; in, For learnable proximal mapping operators, The convolution kernel is The number of channels is convolutional layers, For decoder, For activation function, For encoder, The convolution kernel is A convolutional layer with 1 channel. This represents the number of channels in the convolutional layer.
4. The high-resolution orientation estimation method based on deep unfolded networks according to claim 3, characterized in that... The residual connection is represented as follows: 。 5. A high-resolution orientation estimation method based on deep unfolded networks according to claim 4, characterized in that... The loss function of the dCv-ISTA-Net network is expressed as follows: ; ; in, The number of training samples in each batch. To improve the loss function, For the first The model output results for each sample. For the first The true label of each sample This refers to the temperature parameter.
6. The high-resolution orientation estimation method based on deep unfolded networks according to claim 5, characterized in that... The dCv-ISTA-Net network is a pre-trained dCv-ISTA-Net network, and the training process is as follows: Step 1: Based on the spatial region of interest and the preset grid, discretize the far-field spatial region to obtain overcomplete spatial discrete sampling points; Step 2: Use a uniform linear array to acquire the array received signal, and use conventional beamforming to obtain the CBF azimuth spectrum of the array received signal under spatial overcompleteness at each discrete sampling point; Step 3: Obtain the CBF spatial spectrum vector based on the CBF orientation spectrum of all discrete sampling points under overcomplete spatial conditions; Step 4: Based on the CBF spatial spectral vector, obtain the signal power distribution vector, expressed as: , in, For CBF space spectral vectors, For a pre-defined dictionary matrix, The signal power distribution vector. This refers to the noise component in the azimuth spectrum; Step 5: Construct a training set using the CBF spatial spectral vector and the signal power distribution vector, and train the dCv-ISTA-Net network using the training set to obtain the trained dCv-ISTA-Net network.
7. A high-resolution orientation estimation method based on deep unfolded networks according to claim 6, characterized in that... The CBF azimuth spectrum of the array received signal in the desired beamforming direction is represented as follows: ; ; ; ; ; in, For the first The power of each signal, for A uniform linear array in the desired direction of beamforming The CBF weighted vector, , The number of beam grid points, For weighted vectors, For the spacing between array elements, For array to receive signals, for Quantity, The incident azimuth angle is... For the number of array elements, , For the number of snapshots, For the far-field array manifold matrix, For far-field signal vectors, For noise vectors, It is the symbol for imaginary numbers. The center frequency of the signal. For the speed of sound, for The first moment Signals incident from one direction, For the expected operation, For the first The far-field steering vector of a signal. .
8. A high-resolution orientation estimation method based on deep unfolded networks according to claim 7, characterized in that... The CBF azimuth spectrum under the overcomplete spatial domain is represented as follows: ; in, The number of grid points, For the first Each grid point in the direction On the signal power, For the first The guide vector of each grid point.
9. A high-resolution orientation estimation method based on deep unfolded networks according to claim 1, characterized in that... The learnable parameters of the dCv-ISTA-Net network include the gradient descent step size of each layer. and in the near-end mapping module and encoder and decoder .
10. A high-resolution orientation estimation method based on deep unfolded networks according to claim 1, characterized in that... The total number of parameters in the dCv-ISTA-Net network is: ; in, The number of channels in the convolutional layer. This represents the number of cascaded layers in the network.