DOA estimation method and system based on deep complex-valued convolutional attention residual network
By processing the covariance matrix of coprime arrays through a deep complex-valued convolutional attention residual network and adaptively learning DOA features, the model mismatch problem of sparse arrays is solved, and high-precision DOA estimation is achieved under conditions of low signal-to-noise ratio and few snapshots, thus improving estimation accuracy and robustness.
Patent Information
- Application Number
- CN202510512662.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Existing DOA estimation algorithms suffer from model mismatch on sparse arrays such as coprime arrays, leading to performance degradation. Furthermore, deep learning methods lack robustness under conditions of low noise ratio and few snapshots, making it difficult to achieve high-precision DOA estimation.
We employ a Deep Complex-Valued Convolutional Attention Residual Network (DC-CARN) to directly process the covariance matrix in the complex domain through complex-valued convolution, attention mechanisms, and residual connections. This allows for adaptive learning of DOA features, which are then combined with a complex-valued fully connected network for high-dimensional mapping, outputting the estimated probability distribution of DOA.
Under conditions of low signal-to-noise ratio and few snapshots, it achieves high-precision and robust DOA estimation, breaks through the assumption limitations of traditional methods, significantly improves estimation accuracy and resolution, and has good generalization performance and scalability.
Smart Images

Figure CN120670808B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of array signal processing, in particular to a DOA estimation method and system based on a deep complex convolution attention residual network. BACKGROUND
[0002] Direction of Arrival (DOA) estimation is a core problem of array signal processing, which aims to determine the spatial angle of the signal source relative to the receiving array. It is widely used in many fields such as radar detection, wireless communication, sonar positioning, radio astronomy, etc. Traditional DOA estimation algorithms, such as Multiple Signal Classification (MUSIC) algorithm, Estimation of Signal Parameters via Rotational Invariance Techniques (ESPRIT) algorithm and their various improved versions, such as Weighted MUSIC, Root-MUSIC (R-MUSIC), Weighted ESPRIT, Total Least Squares ESPRIT (TLS-ESPRIT), etc., are mostly based on regular array structures such as Uniform Linear Array (ULA) and have strong assumptions on signal and Gaussian white noise, signal incoherence, etc. And to avoid phase ambiguity, the inter-element spacing of the uniform array is usually no more than half a wavelength, which limits the array aperture and affects the resolution of DOA estimation and the degree of freedom of identifiable signal sources. Increasing the number of array elements to expand the ULA aperture will directly lead to a substantial increase in hardware cost, system complexity and power consumption, which is not feasible in many practical applications.
[0003] To overcome the above limitations, sparse arrays such as coprime arrays (CA) have emerged, which synthesize a large-aperture virtual array with fewer physical elements through non-uniform arraying and coprime subarray design. Therefore, CA has a significant advantage in increasing the array aperture and degrees of freedom relative to ULA, but its application to DOA estimation also brings new challenges. The non-uniformity of CA destroys the Vandermonde structure of the array flow pattern, and directly applying subspace decomposition-based algorithms such as MUSIC and ESPRIT will cause a serious model mismatch problem, and the performance will deteriorate sharply. Therefore, researchers have proposed virtual array processing techniques, with spatial smoothing MUSIC (SS-MUSIC) being a typical representative. It restores the full rank of the matrix by spatial smoothing the covariance matrix of the virtual array, and then applies the MUSIC algorithm. There are also studies that introduce ESPRIT, compressed sensing (CS), sparse Bayesian learning (SBL), and other theories into coprime array DOA estimation. However, these model-driven schemes have high computational complexity and insufficient robustness under non-ideal conditions such as low SNR, limited number of snapshots, coherent sources or array element position errors, channel gain / phase inconsistencies, etc. Off-grid and gridless methods can alleviate the grid mismatch problem, but often sacrifice computational efficiency or are sensitive to noise, and still require parameter tuning in practical applications, limiting their universality.
[0004] Deep learning (DL) provides a new data-driven approach for DOA estimation. Deep neural networks (DNN) have strong non-linear fitting capabilities and the advantage of automatically learning features from data. Fully connected neural networks (FCNN) and convolutional neural networks (CNN) are used to learn the mapping relationship from array received data or its covariance matrix to DOA. CNN is particularly suitable for processing the spatial structure of the covariance matrix. For the inherent complex number characteristics of array signals, complex-valued neural networks (CVNN) preserve amplitude and phase information, and exhibit better accuracy and robustness than real-valued networks in low SNR and few snapshots scenarios. However, current deep learning research for CA is still insufficient, especially the design of efficient network architectures that combine complex-valued convolution, attention mechanisms, and residual connections to address the unique model mismatch and noise sensitivity problems of CA. SUMMARY
[0005] The present application provides a new deep learning network structure, which can fully adapt to and utilize the characteristics of CA, effectively process complex array signal data, and realize high-precision and high-robustness DOA estimation method and system based on deep complex-valued convolution attention residual network for coprime array signal sources by integrating advanced network components such as complex-valued convolution, attention mechanism and residual connection.
[0006] To achieve the above purpose, the present application is realized by the following technical scheme: the DOA estimation method based on deep complex-valued convolution attention residual network (DC-CARN) provided by the present application comprises the following steps:
[0007] S1, a data preprocessing module obtains complex data, acquires data by coprime array (CoprimeArray, CA) and pre-processes to obtain complex-valued sample covariance matrix (SCM) data as the original input information for DOA estimation;
[0008] S2, a feature extraction module, a deep complex-valued convolution attention residual network (Deep Complex-Valued Convolutional Attention Residual Network, DC-CARN) is constructed to directly process the covariance matrix information in the complex domain, the deep complex-valued convolution attention residual network utilizes an initial two-dimensional complex-valued convolution layer, a complex-valued convolution block attention network and a cross-layer residual connection structure to deeply extract complex-valued features related to spatial angles;
[0009] S3, a DOA estimation module; including a multi-layer complex-valued fully connected network layer (Complex-Valued Fully Connected Neural Network, CV-FCNN), which is responsible for finally mapping the extracted high-dimensional complex-valued feature vector to the representation space directly related to the DOA estimation task, i.e. for mapping the complex-valued features to the angle domain; wherein the complex-valued fully connected network layer performs nonlinear transformation and mapping on the one-dimensional complex-valued feature vector, and maps the high-dimensional complex-valued feature vector extracted in step S2 to the representation space directly related to the DOA estimation task;
[0010] S4, an output module for converting the angle domain feature representation after completing the DOA estimation task into a user interpretable DOA estimation probability distribution result.
[0011] Preferably, in step S2, a complex-valued convolutional block attention module (CV-CBAM) is embedded in the feature extraction module, which includes channel attention and spatial attention sub-modules specially designed for complex features, and specifically receives the complex SCM data obtained in step S1. Through channel attention and spatial attention mechanisms, the above sub-module network can adaptively learn and emphasize the most important feature channels and spatial positions for the current task, suppress noise and irrelevant feature interference, and deeply extract complex feature vectors related to spatial angles through an initial two-dimensional complex convolution layer, a cascaded complex convolution block attention module, and a cross-layer residual connection structure.
[0012] Preferably, in step S1, obtaining the complex SCM data specifically includes:
[0013] The two uniform linear arrays with prime numbers of elements are arranged in a superposition manner to form a coprime array, wherein the number of elements of a first sub-array is M, and the element spacing is Nd; the number of elements of a second sub-array is N, and the element spacing is Md; M and N are prime numbers and M < N, and d is half the wavelength of the signal. The total number of elements of the coprime array is P = M + N - 1.
[0014] J far-field narrowband signals are received to form an array received signal data model:
[0015]
[0016] wherein X Ω (t) is an array received signal at time t, N Ω (t) is a Gaussian white noise, A Ω is an array flow matrix of (M + N - 1) × J dimensions, a Ω (θ j ) is an array steering vector corresponding to the jth signal;
[0017] The SCM data is calculated according to L snapshots, specifically:
[0018]
[0019] wherein, is a P × P complex matrix, L represents the number of snapshots, and (·) Η represents a conjugate transpose operation.
[0020] Preferably, in step S2, S21 is included, and in the deep complex convolution attention residual network, an initial two-dimensional complex convolution layer is used to perform preliminary local feature perception and abstraction on the input complex SCM data, specifically:
[0021] The input complex-valued SCM data is first processed by a two-dimensional complex-valued convolution layer, which has N complex-valued filter of size (q, q) in the complex domain, (q, q) represents the size of the two-dimensional complex-valued convolution kernel is q x q; the complex filter matrix W = A + iB is convolved with the complex vector h = x + iy, and the complex-valued convolution operation is realized by the following matrix form:
[0022]
[0023] Where A and B are real matrices, x and y are real vectors, R(·) represents the real part, I(·) represents the imaginary part, and * represents the real-valued convolution operation.
[0024] The output of the initial two-dimensional complex-valued convolution layer is sequentially subjected to two-dimensional complex-valued batch normalization processing and complex-valued activation function processing; wherein the complex-valued LeakyReLU activation function is applied to the real part and the imaginary part of the variable of the previous layer respectively, therefore, the output of a two-dimensional complex-valued convolution layer can be represented as:
[0025] f k (X) = CV-LeakyReLU(CV-BN(W k *X+b k ))
[0026] Where W k and b k represent the filter matrix and bias vector of the kth two-dimensional complex-valued convolution layer.
[0027] Preferably, it further comprises S22, the deep complex-valued convolution attention residual network utilizes a plurality of complex-valued convolution attention modules connected in series, sequentially connected along the signal processing direction, forming a deep feature extraction path, and the number of filter channels of the convolution layer in the subsequent complex-valued convolution attention module is greater than or equal to the number of filter channels of the convolution layer in the previous module;
[0028] Each complex-valued convolution attention module comprises a plurality of two-dimensional complex-valued convolution layers and a complex-valued convolution block attention network; the two-dimensional complex-valued convolution layer continues to operate in the complex domain, further performing nonlinear transformation and spatial feature extraction on the feature map, and the number of filters between modules is designed to gradually increase to learn feature representations of different levels and dimensions;
[0029] The channel attention module in the complex-valued convolution block attention network is used to adjust the channel weight, and the input is the complex-valued feature X obtained by the two-dimensional complex-valued convolution layer, and the complex-valued feature X is mapped to the real domain to calculate the feature matrix M = |X|; global max pooling M max and global average pooling M avg are performed on the feature map of each channel, and then M concat=[M avg M max The concatenated result is used as input to the perceptron MLP, and the complex scaling parameter P = MLP(M) is generated by passing the sigmoid activation function. concat Simultaneously, the output is split into scaling factors S applied to the real and imaginary parts. r and S i Construct the complex scaling factor S = S r +jS i Finally, the feature map is scaled by a complex number at the channel scale to obtain the enhanced features.
[0030] The spatial attention module in the complex-valued convolutional block attention network is used to adjust the spatial location weights. It performs average pooling M′ in the spatial dimension by taking the modulus M′=|X′| of the feature map enhanced by the complex-valued channel attention module. avg and maximize pooling M′ max Then, by splicing them together, we get M′. spatial =[M′ avg ,M′ max The input is then fed into a complex-valued convolutional layer to learn the weights at each location, and then passed through a sigmoid activation function to obtain the spatial attention weights W. s Finally, the feature maps are weighted to obtain the enhanced features.
[0031] Preferably, it also includes S23, where the deep complex-valued convolutional attention residual network utilizes a cross-layer residual connection structure and employs cross-layer fast connections to add the module's input features directly or after simple transformation to the module's output features, achieving cross-module feature fusion, specifically as follows:
[0032] Each complex-valued convolutional attention module has a residual connection, which directly adds the input feature x of the module to the processed output F(x) to form the final output F(x)+x; this helps with gradient propagation and training deep networks.
[0033] The final output of the feature extraction module also includes a flattening layer for structural transformation, which converts the multidimensional complex-valued feature maps from the feature extraction module into one-dimensional complex-valued feature vectors to adapt to the input requirements of subsequent fully connected layers.
[0034] Preferably, in step S3, the DOA estimation module includes multiple CV-FCNN layers and a complex-valued Dropout layer inserted between every two adjacent CV-FCNN layers;
[0035] The CV-FCNN layer uses a complex weight matrix and a complex bias vector to perform affine transformation on a one-dimensional complex-valued feature vector input, and introduces nonlinearity through a complex-valued activation function to realize high-level semantic mapping of the feature space; wherein the kth CV-FCNN layer can be expressed as:
[0036] n k = W k,k-1 h k-1 + b k
[0037] h k = CV-LeakyReLU(net k )
[0038] wherein W k,k-1 represents the (k-1)th complex-valued weight, h k-1 represents the (k-1)th input, b k represents the kth bias, n k represents the kth output, and LeakyReLU represents a complex-valued activation function;
[0039] The complex-valued Dropout layer is used as a regularization method to randomly "discard" neurons with a certain probability during training, while discarding the real part and the imaginary part, reducing the coordination between neurons, enhancing the generalization ability of the model, and preventing overfitting when the training data is limited.
[0040] Preferably, in step S4, the output module specifically comprises:
[0041] a feature conversion unit for converting the complex-valued feature vector of the complex-valued fully connected network layer into a real-valued representation; the real part and the imaginary part of the complex-valued vector are spliced to form a real-valued vector with doubled dimensions;
[0042] an output mapping layer, which is a standard real-valued fully connected layer FCNN, and the number of output neurons thereof is equal to the total number G of pre-set discrete angle grid points; this layer is responsible for linearly mapping the converted high-dimensional real-valued feature vector to the G angle dimensions;
[0043] an activation function unit adopting a Sigmoid activation function, which compresses each output value to the interval (0, 1), so that it can be interpreted as the posterior probability of the existence of a signal source corresponding to the angle grid point.
[0044] Preferably, the output module outputs a G-dimensional real-valued probability vector, i.e., an estimated angle space spectrum, and maps the feature vector to the pre-divided grid, and finally maps it to the angle distribution probability space through the Sigmoid activation function to determine the DOA estimation value; which can be specifically expressed as:
[0045]
[0046] where h' = {Real(h3), Imag(h3)} represents the real and imaginary parts of the output of the last CV-FCNN layer, and W' represents the weight matrix of the output layer.
[0047] The DOA estimation system based on the deep complex-valued convolutional attention residual network uses the above DOA estimation method, and specifically comprises:
[0048] A system input interface is configured to receive externally input complex-valued SCM data.
[0049] A feature extraction module is electrically connected or in data communication with the system input interface at an input end and connected with the input end of the flattening layer at an output end, and is configured to extract complex-valued features related to DOA from the original SCM in a deep and layer-by-layer manner.
[0050] The flattening layer converts the multi-dimensional complex-valued feature tensor output by the feature extraction module into a one-dimensional long vector form, so as to adapt to the input requirements of the subsequent fully connected layer.
[0051] A DOA estimation module is connected with the output end of the flattening layer at an input end and connected with the output module at an output end, and is configured to map the high-dimensional complex-valued feature vector extracted by the feature extraction module to a representation space directly related to the DOA estimation task.
[0052] An output module and an output interface are electrically connected or in data communication with each other, and are configured to output a DOA estimation probability spectrum.
[0053] The DOA estimation method and system based on the deep complex-valued convolutional attention residual network have the following beneficial effects:
[0054] (1) The DOA estimation method based on the deep complex-valued convolutional attention residual network combines the complex-valued network structure with the attention mechanism, completely retains the signal phase information through complex domain operation, and uses complex-valued convolutional attention to strengthen key feature extraction and residual connection to optimize gradient propagation, so that the model can still maintain stable performance under low SNR and few snapshots. In addition, the deep complex-valued convolutional attention residual network system adopts an end-to-end learning architecture, breaks through the strong hypothesis limitation of the traditional method on the signal model, can automatically learn the complex and highly nonlinear mapping relationship from SCM to DOA from data, and adaptively mines the virtual aperture characteristics of the coprime array. Finally, through the cooperative optimization of complex value processing, attention weighting and residual connection, the subtle feature differences corresponding to different DOAs can be captured and distinguished more finely, so that the system can maintain the strong expression ability of the deep network while achieving higher estimation accuracy, and significantly improve the robustness and estimation accuracy under low SNR and few snapshots.
[0055] Meanwhile, the feature extraction module combines the attention mechanism to focus on key features and the residual connection structure to optimize gradient propagation. The attention mechanism enables the model to adaptively focus on the feature information most relevant to the DOA estimation, suppressing redundancy and interference. The residual connection structure allows the construction of deeper network models, promotes the effective transfer and fusion of features, and enables deeper network structures to be stably trained and play a role, thereby improving the model's capacity and feature representation level.
[0056] (2) The deep complex-valued convolutional attention residual network system of the present invention innovatively uses complex-valued SCM data generated by CA as input, and learns the virtual aperture characteristics of coprime arrays implicitly through the deep complex-valued convolutional attention residual network, thereby obtaining resolution and degrees of freedom that exceed the physical array element limit; and the use of regularization methods such as complex-valued Dropout helps to prevent overfitting and effectively improve the generalization performance of the deep complex-valued convolutional attention residual network system on unseen data.
[0057] (3) The deep complex-valued convolutional attention residual network system of the present invention integrates a variety of advanced deep learning components such as complex-valued operations, convolutional networks, attention mechanisms, and residual networks to construct a dedicated optimized architecture for DOA estimation. At the same time, the modular design also gives the system good scalability and configurability, making it easy to adjust the network size and complexity according to specific application scenarios. Attached Figure Description
[0058] Figure 1 This is a schematic diagram of the coprime array CA in Embodiment 1.
[0059] Figure 2 This is a schematic diagram of the element positions of a 6-element coprime array with M=3 and N=4 used in Embodiment 1.
[0060] Figure 3 This is a schematic diagram of the overall architecture of the depthwise complex-valued convolutional attention residual network DC-CARN in Example 1.
[0061] Figure 4 for Figure 3 A schematic diagram of the internal structure and data processing flow of a two-dimensional complex-valued convolutional layer;
[0062] Figure 5 for Figure 3 A schematic diagram of the internal structure of the complex-valued convolutional attention module;
[0063] Figure 6 for Figure 5 A schematic diagram of the complex-valued convolutional block attention network CV-CBAM used in the paper;
[0064] Figure 7 for Figure 6Structure diagram of middle channel attention module;
[0065] Figure 8 For Figure 6 Structure diagram of middle spatial attention module;
[0066] Figure 9 For Figure 3 Or Figure 5 Principle diagram of cross-layer residual connection structure used in the middle;
[0067] Figure 10 The training and verification loss curve diagram of the DOA estimation system in the training process;
[0068] Figure 11 The DOA estimation root mean square error performance change comparison diagram of the estimation method and various existing DOA estimation algorithms under the condition of fixed snapshot number and signal-to-noise ratio;
[0069] Figure 12 The DOA estimation root mean square error performance change comparison diagram of the estimation method and various existing DOA estimation algorithms under the condition of fixed signal-to-noise ratio and snapshot number. DETAILED DESCRIPTION
[0070] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all.
[0071] Embodiment 1
[0072] The DOA estimation method based on the deep complex-valued convolution attention residual network comprises the following steps:
[0073] S1, the data preprocessing module obtains complex data, acquires data by a coprime array (CoprimeArray, CA) and pre-processes to obtain complex-valued sampling covariance matrix (SCM) data as the original input information of DOA estimation, specifically including:
[0074] As Figure 1 shown, the present application is aimed at the application scene of utilizing the coprime array to estimate the direction of arrival (DOA), the coprime array is used as a kind of sparse array, by the uniform linear array of two array element numbers being prime to each other superposition arrangement. Wherein, the array element number of the first subarray is M, the array element number of the second subarray is N, M and N are prime to each other and M < N. The array element spacing of the first subarray and the second subarray is Nd and Md respectively, d is the half wavelength of signal. The two uniform linear arrays are combined together, due to the coprime relationship of M and N, only the first array element of two subarrays will coincide, other array elements will not coincide, so there are M+N-1 array elements.
[0075] As Figure 2 shown. This embodiment adopts a 6-element P=M+N-1=6 array composed of coprime numbers M=3 and N=4, and adopts such an array to receive signals and generate complex-valued SCM data for subsequent DOA estimation. Since the spacing between the elements of the coprime array is no longer half the wavelength of the signal, the positions of the elements in the coprime array are:
[0076] Ω={Mnd}∪{Nmd}
[0077] n=0,1,…,N-1;m=0,1,…,M-1
[0078] J far-field narrowband signals are received to form an array received signal data model:
[0079]
[0080] where X Ω (t) is the array received signal at time t, N Ω (t) is Gaussian white noise, A Ω is a (M+N-1)×J dimensional array stream matrix, a Ω (θ j ) is the array steering vector corresponding to the jth signal, which is specifically expanded as:
[0081]
[0082] Ωi∈Ω,i=1,…,M+N-1
[0083] where Ωi is the actual position of each element, and the position of the first element is 0.
[0084] From the above array received signal X Ω (t), all possible incoming signal directions θ1…θ K are estimated, and the covariance matrix of the coprime array received signal is obtained, which is specifically represented as:
[0085]
[0086] where p s =[p s1 ,…,p sJ ] T is the power of each incident signal source; Rs is the covariance matrix of S(t); is the noise power; I M+N-1 is a (M+N-1)×(M+N-1) dimensional unit matrix; is an operation for statistical expectation; (·) His a conjugate transpose operation; diag(·) is a diagonal matrix formed by taking the elements in a vector as values on the diagonal of the matrix;
[0087] The SCM is calculated according to L snapshot data, specifically:
[0088]
[0089] wherein, is a P x P dimensional complex matrix, L represents the number of snapshots, (·) Η represents a conjugate transpose operation. When the receiving array has P array elements, the SCM data is a P x P dimensional complex matrix,
[0090]
[0091] S2, a feature extraction module; receives the complex-valued SCM data shown in Figure 2 The SCM is a 6 x 6 complex Hermitian matrix, a deep complex-valued convolutional attention residual network (DC-CARN) is constructed to directly process the covariance matrix information in the complex domain, the deep complex-valued convolutional attention residual network uses an initial two-dimensional complex-valued convolutional layer, a complex-valued convolutional block attention network, and a cross-layer residual connection structure to deeply extract complex-valued features related to spatial angles;
[0092] S21, in the deep complex-valued convolutional attention residual network, an initial two-dimensional complex-valued convolutional layer is used to perform preliminary local feature perception and abstraction on the input complex-valued SCM, specifically:
[0093] As shown in Figure 4 , the input complex-valued SCM is first processed by a two-dimensional complex-valued convolutional layer, the convolutional layer operates in the complex domain and uses a complex-valued convolutional kernel. It is configured with 64 filters of size (3, 3), (3, 3) representing the size of the two-dimensional complex-valued convolutional layer convolutional kernel is 3 x 3; stride (1, 1), padding (1, 1). The complex filter matrix W = A + iB is convolved with the complex vector h = x + iy, where A and B are real matrices, x and y are real vectors, and W * h = (A * x - B * y) + i(B * x + A * y) is obtained by convolving the filter W with the vector h. The complex-valued convolution operation is implemented by the following matrix form:
[0094]
[0095] wherein, A and B are real matrices, x and y are real vectors, R(·) represents the real part, I(·) represents the imaginary part, and * represents the real-valued convolution operation.
[0096] The output of the initial two-dimensional complex-valued convolution layer sequentially performs two-dimensional complex-valued batch normalization processing and complex-valued activation function processing; wherein the complex-valued LeakyReLU activation function is applied to the real part and the imaginary part of the variable of the previous layer respectively, therefore, the output of one two-dimensional complex-valued convolution layer can be expressed as:
[0097] f k (X)=CV-LeakyReLU(CV-BN(W k *X+b k ))
[0098] Wherein, W k and b k represent the filter matrix and the bias vector of the kth two-dimensional complex-valued convolution layer.
[0099] S22, the deep complex-valued convolution attention residual network utilizes a plurality of complex-valued convolution attention modules connected in series, sequentially connected along the signal processing direction, forming a deep feature extraction path, and the number of filter channels of the convolution layer in the subsequent complex-valued convolution attention module is greater than or equal to the number of filter channels of the convolution layer in the previous module;
[0100] As shown in Figure 5 , each cascaded complex-valued convolution attention module includes a plurality of two-dimensional complex-valued convolution layers and a complex-valued convolution block attention network (CV-CBAM); the two-dimensional complex-valued convolution layer continues to operate in the complex domain, further performing nonlinear transformation and spatial feature extraction on the feature map, and the number of filters and the number of output channels between modules are designed to gradually increase from 64 to 128 to 256, in order to learn feature representations of different levels and dimensions.
[0101] As shown in Figure 6 , the complex-valued convolution block attention network in the complex-valued convolution attention module is set between or after the adjacent two two-dimensional complex-valued convolution layers, and includes two sub-modules of complex-valued channel attention module and complex-valued spatial attention module, which are both designed based on the modulus information of complex features, and these sub-modules can adaptively learn and emphasize the most important feature channels and spatial positions for the current task, suppress irrelevant or noise interference information, thereby improving the discriminability and robustness of the features.
[0102] Wherein, as shown in Figure 7 , the channel attention module in the complex-valued convolution block attention network is used to adjust the channel weight, the input is the complex-valued feature X obtained by the two-dimensional complex-valued convolution layer, and the complex-valued feature X is mapped to the real number domain to calculate the feature matrix M = |X|; global maximum pooling M max and global average pooling M avg are performed on each channel feature map, and then M concat = [M avg , M maxThe concatenated result is used as input to the perceptron MLP, and the complex scaling parameter P = MLP(M) is generated by passing the sigmoid activation function. concat Simultaneously, the output is split into scaling factors S applied to the real and imaginary parts. r and S i Construct the complex scaling factor S = S r +jS i Finally, the feature map is scaled by a complex number at the channel scale to obtain the enhanced features.
[0103] like Figure 8 As shown, the spatial attention module in the complex-valued convolutional block attention network is used to adjust the spatial location weights. It takes the modulus M′=|X′| of the feature map enhanced by the complex-valued channel attention module and performs average pooling M′ along the spatial dimension. avg and maximize pooling M′ max Then, by splicing them together, we get M′. spatial =[M′ avg ,M′ max The input is then fed into a complex-valued convolutional layer to learn the weights at each location, and then passed through a sigmoid activation function to obtain the spatial attention weights W. s Finally, the feature maps are weighted to obtain the enhanced features.
[0104] S23, such as Figure 9 As shown, the deep complex-valued convolutional attention residual network utilizes a cross-layer residual connection structure and employs cross-layer shortcut connections to add the input features of a module directly or after simple transformation to the output features of the module, achieving cross-module feature fusion. Specifically, each complex-valued convolutional attention module has a residual connection, which directly adds the module's input feature x to the module's processed output F(x) to form the final output F(x)+x. This facilitates gradient propagation and training of deep networks, specifically as follows:
[0105]
[0106] As can be seen from the above formula, even if the derivative parameter decreases as the network structure deepens, the final derivative value will never be less than 1, and the final result of the gradient chain rule will not tend to 0. Therefore, the gradient vanishing phenomenon will not occur during parameter updates.
[0107] like Figure 3 As shown, the final output of the feature extraction module also includes a flattening layer for structural transformation, which converts the multidimensional complex-valued feature maps from the feature extraction module into one-dimensional complex-valued feature vectors to adapt to the input requirements of the subsequent fully connected layers.
[0108] S3、as shown, the flattened feature vector enters the direction of arrival (DOA) estimation module, which performs nonlinear transformation and mapping on the one-dimensional complex-valued feature vector, including multiple complex-valued fully connected network layers (CV-FCNN) and a complex-valued Dropout layer inserted between each adjacent two CV-FCNN layers; Figure 3
[0109] The present application uses three CV-FCNN layers, each corresponding to the input feature and the output feature, respectively. The CV-FCNN layer uses a complex weight matrix and a complex bias vector to perform affine transformation on the input one-dimensional complex-valued feature vector, and introduces nonlinearity through a complex activation function to realize high-level semantic mapping of the feature space. The CV-FCNN of the kth layer can be expressed as:
[0110] n k =W k,k-1 h k-1 +b k
[0111] h k =CV-LeakyReLU(net k )
[0112] where W k,k-1 is the complex weight of the (k-1)th layer, h k-1 is the input of the (k-1)th layer, b k is the bias of the kth layer, n k is the output of the kth layer, and LeakyReLU is a complex activation function.
[0113] The complex-valued Dropout layer is a regularization method that randomly "discards" neurons with a certain probability during training, discarding both the real and imaginary parts, reducing the coordination between neurons, enhancing the generalization ability of the model, and preventing overfitting when the training data is limited.
[0114] S4、the output module specifically includes a feature conversion unit, an output mapping layer, and an activation function unit;
[0115] The feature conversion unit converts the complex-valued feature vector of the complex-valued full connection network layer into a real-valued representation; adopts real part and imaginary part splicing: h' = {Real(h3), Imag(h3)}, splices the real part and imaginary part of the complex-valued vector to form a real-valued vector with doubled dimensions; the output mapping layer is a standard real-valued full connection layer FCNN, the number of output neurons of which is equal to the total number G of preset discrete angle grid points; the layer is responsible for linearly mapping the converted high-dimensional real-valued feature vector to G angle dimensions; the activation function unit adopts a Sigmoid activation function, and the Sigmoid function compresses each output value to the interval (0, 1), so that it can be interpreted as the posterior probability of the existence of a signal source corresponding to the angle grid point.
[0116] The output module outputs a G-dimensional real-valued probability vector, that is, an estimated angle space spectrum, and maps the feature vector to a pre-divided grid, and finally maps it to an angle distribution probability space through a Sigmoid activation function to determine the DOA estimation value; the specific representation is:
[0117]
[0118] Where h' = {Real(h3), Imag(h3)} represents the real part and imaginary part splicing of the output of the last CV-FCNN layer, and W' represents the weight matrix of the output layer.
[0119] Embodiment 2
[0120] The application is based on a deep complex-valued convolutional attention residual network DOA estimation system, and adopts the DOA estimation method in Embodiment 1, which specifically includes:
[0121] A system input interface for receiving externally transmitted complex-valued SCM data;
[0122] A feature extraction module, the input end of which is electrically connected or in data communication with the system input interface, and the output end of which is connected with the input end of the flattening layer; for deeply extracting complex-valued feature representations related to DOA from the original SCM layer by layer;
[0123] A flattening layer for converting the multi-dimensional complex-valued feature tensor output by the feature extraction module into a one-dimensional long vector form for adapting to the input requirements of the subsequent full connection layer;
[0124] A direction of arrival DOA estimation module, the input end of which is connected with the output end of the flattening layer, and the output end of which is connected with the output module; for mapping the high-dimensional complex-valued feature vector extracted by the feature extraction module to a representation space directly related to the DOA estimation task;
[0125] An output module and an output interface, the output module being electrically connected or in data communication with the output interface, for outputting the direction of arrival DOA estimation probability spectrum.
[0126] The DOA estimation system based on deep complex-valued convolutional attention residual network in the embodiment is trained. There are a total of 121 wave directions in the training set, the spatial signal angle range received by the array is selected as [-60°, 60°], and the angle interval is set as 1°. It is assumed that there are two signal sources incident on the array, that is, J = 2, the angles of the two signal sources are traversed in all combinations of 121 angles in the range of [-60°, 60°], and finally there are 7260 pairs of different angle combinations. The environment of each pair of signals is traversed in 7 different SNR levels {-20dB, -15dB, -10dB, -5dB, 0dB, 5dB, 10dB}, and there are a total of D = 7 x 7260 = 50,820 training samples. At the specified SNR level, the complex-valued SCM data is constructed from the covariance matrix data corresponding to each angle pair, the sampling snapshot number L is set as 500, 80% of the samples are used for training, and the remaining 20% of the samples are used for verification, and each sample is marked with the corresponding wave direction angle.
[0127] The parameter update and optimization of the network adopts the Adam optimizer, which can dynamically modify the learning rate of the parameters. The initial learning rate is set as 0.001, the momentum parameters β1 = 0.85, and β2 = 0.95. In addition, according to the loss decrease of the verification set, the learning rate is dynamically adjusted by a factor of 0.7 under the condition of patience for 5 epochs, that is, if the loss of the verification set does not decrease for 5 consecutive epochs, the learning rate is scaled by 0.7 times, thereby effectively avoiding the local optimal problem. In order to further alleviate the risk of overfitting, the L1 norm regularization is introduced to the FCNN layer weight, so as to ensure the sparsity and stability of the model parameters. The binary cross entropy loss (BCELoss) is used to measure the error between the model output and the real angle label probability distribution. The goal of network training is to minimize the loss function, which is represented as follows:
[0128]
[0129] wherein v represents the learnable parameters of the network model, represents and the cross entropy of y i
[0130] Figure 10 The training and validation loss curves of the DOA estimation system of the application in the training process are shown, and it can be seen from the figure that the DOA estimation system based on the deep complex-valued convolution attention residual network of the application has been effectively converged on the training set and the validation set. Although the validation loss is slightly higher than the training loss, the difference between the two is relatively stable, indicating that the model does not overfit and has strong generalization ability. The loss ratio is kept in a reasonable range, and the overall training effect is ideal. After training, the deep complex-valued convolution attention residual network-based mutual DOA estimation system receives new complex-valued SCM data input, estimates according to the DOA estimation method and Figure 3 the flow shown in embodiment 1, and finally outputs the probability spectrum of DOA estimation.
[0131] Embodiment 3
[0132] In order to prove the advantages of the deep complex-valued convolution attention residual network model DC-CARN proposed in the application in DOA estimation based on the co-prime array, as shown in Figure 10 and Figure 11 the application compares the performance of the deep complex-valued convolution attention residual DOA estimation method of embodiment 1 with the traditional estimation methods in the prior art under different signal-to-noise ratio levels and snapshot numbers.
[0133] These traditional estimation methods mainly involve some traditional model-driven methods, including MUSIC, R-MUSIC, Coprime MUSIC, DNN method and CV-CNN method, wherein the DNN is composed of four layers of FCNN. All the traditional methods are implemented in MATLAB 2023b and simulated on a computer with a 3.00GHz Intel Core i9 processor and 64GB RAM, and the network model is deployed on a NVIDIA GeForce RTX 3090 GPU based on Pytorch. The complex-valued SCM data is used with an initial learning rate lr=10-4, a batch size of 32, training for 200 epochs, and an Adam optimizer. In order to make a reasonable comparison, the angle resolution is set to be the same in different methods, and the root mean square error (RMSE) is used to evaluate the DOA estimation performance, which is defined as follows:
[0134]
[0135] wherein θ j is the true label of the jth DOA angle, is the estimated value of the jth DOA angle in the tth Monte Carlo experiment, M c is the number of Monte Carlo experiments; the central processing unit time is used to evaluate the time complexity of DOA estimation.
[0136] Figure 11 The performance curves of the DOA estimation root mean square error of the estimation method of the application and various existing DOA estimation algorithms under the condition of fixed snapshot number L = 200 and the signal-to-noise ratio from low to high -10dB to 20dB are shown in the figure, from which it can be seen that the RMSE value of the DC-CARN network system proposed in the application is significantly lower than that of all the comparison methods in the entire SNR range of the test, especially in the low SNR area SNR < 5dB, the performance advantage is more obvious. In contrast, the traditional MUSIC and Root-MUSIC methods perform sharply deteriorated at low SNR, with high RMSE values. Even compared with other deep learning methods such as CV-CNN, the DC-CARN network system of the application also exhibits lower RMSE, further proving that the application has stronger robustness and higher estimation accuracy at different noise levels.
[0137] Figure 12 The performance curves of the DOA estimation root mean square error of the estimation method of the application and various existing DOA estimation algorithms under the condition of fixed signal-to-noise ratio SNR = 0dB and the snapshot number from few to many 20 to 1000 are shown in the figure, from which it can be seen that the RMSE value of the DC-CARN network system proposed in the application remains at a low level in the entire snapshot number range of the test, which is significantly better than the comparison methods. Especially in the case of very limited snapshot number L < 100, the performance advantage of the DC-CARN network system is particularly prominent, which shows its strong ability and robustness in processing small sample data. At the same time, although the performance of all methods improves with the increase of snapshot number, the DC-CARN network system of the application always maintains the optimal or near-optimal performance level.
[0138] As Figure 11 and Figure 12 shown in the performance change curve, it is fully proved that the DC-CARN network system designed in the application has the technical advantages of excellent performance and robustness. Mainly due to the overall use of complex-valued convolutional layers, complex-valued fully connected layers and deep complex-valued network architecture, it can directly process the covariance matrix information in the complex domain, avoid the information loss caused by the separation processing of real and imaginary parts, and is more in line with the physical nature of array signal processing.
[0139] At the same time, the complex-valued convolutional block attention network CV-CBAM is embedded in the feature extraction module, which can make the network adaptively focus on the feature channels and spatial regions that contribute more to DOA estimation through channel attention and spatial attention mechanisms, and suppress the interference of noise and irrelevant features. And the residual connection structure is adopted, which effectively alleviates the gradient vanishing problem in the training of deep network, promotes the effective transmission and fusion of features, so that the deeper network structure can be stably trained and play a role.
[0140] In summary, the DOA estimation method and system based on the deep complex-valued convolution attention residual network proposed in the application, through its unique structure design and component configuration, has shown significant performance improvement compared with the prior art in simulation verification, especially in challenging conditions such as low signal-to-noise ratio and few snapshots, has higher estimation accuracy and robustness, and has strong practical value.
[0141] On the basis of the above-mentioned embodiments, the technical features involved therein and the functions and effects of the technical features in the application are described in detail to help those skilled in the art fully understand the technical solutions of the application and reproduce them.
[0142] Finally, although the present specification is described in terms of embodiments, not every embodiment contains only one independent technical solution, and the description manner of the specification is only for the sake of clarity, those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be appropriately combined to form other embodiments that those skilled in the art can understand.
Claims
1. A DOA estimation method based on deep complex-valued convolutional attention residual network, characterized in that, The steps include the following: S1, a data preprocessing module obtains complex data, acquires data with a coprime array (CA), and pre-processes to obtain complex-valued sample covariance matrix (SCM) data as the original input information for DOA estimation; In step S1, obtaining the complex-valued SCM data specifically includes: The coprime array is established by superimposed arrangement of two uniform linear arrays with array element numbers being coprime, wherein the array element number of the first subarray is M , the array element spacing is Nd ; the array element number of the second subarray is N , the array element spacing is Md ; M and N are coprime and M N , d is the half wavelength of the signal, and the total array element number of the coprime array is P = M + N -1; receive J a far-field narrowband signal, forming an array received signal data model: ; wherein is an array receiving signal at time t , is a Gaussian white noise, is a (N M + N -1) x N array stream pattern matrix, J , , is an array steering vector corresponding to the j th signal; According to L SCM data is calculated from the individual snapshot data, specifically: ; wherein is one P x P complex matrix of dimension L denotes the number of snapshots, denotes the operation of taking the conjugate transpose; S2, a feature extraction module, a deep complex-valued convolutional attention residual network (DC-CARN) is constructed to directly process the covariance matrix information in the complex domain, which utilizes an initial two-dimensional complex-valued convolutional layer, a complex-valued convolutional block attention network, and a cross-layer residual connection structure to deeply extract complex-valued features related to spatial angles; In step S2, a complex-valued convolutional block attention module (CV-CBAM) is embedded in the feature extraction module, which includes channel attention and spatial attention sub-modules specially designed for complex features. Specifically, the complex-valued SCM data obtained in step S1 is received, and through channel attention and spatial attention mechanisms, the above sub-module network can adaptively learn and emphasize the most important feature channels and spatial positions for the current task, suppress noise and irrelevant feature interference, and deeply extract complex-valued feature vectors related to spatial angles through an initial two-dimensional complex-valued convolutional layer, a cascaded complex-valued convolutional block attention module, and a cross-layer residual connection structure. S3, a DOA estimation module; including a multi-layer complex-valued fully connected neural network layer (CV-FCNN), which is responsible for finally mapping the extracted high-dimensional complex-valued feature vectors to a representation space directly related to the DOA estimation task, i.e., for mapping complex-valued features to the angle domain; wherein the complex-valued fully connected neural network layer performs nonlinear transformation and mapping on one-dimensional complex-valued feature vectors, and maps the high-dimensional complex-valued feature vectors extracted in step S2 to a representation space directly related to the DOA estimation task. S4, an output module for converting the angle domain feature representation after completing the DOA estimation task into a user interpretable DOA estimation probability distribution result.
2. The DOA estimation method based on deep complex-valued convolutional attention residual network according to claim 1, characterized in that, In step S2, S21 is included, and in the deep complex-valued convolutional attention residual network, an initial two-dimensional complex-valued convolutional layer is used to perform preliminary local feature perception and abstraction on the input complex-valued SCM data, specifically: The input complex-valued SCM data is first processed by a two-dimensional complex-valued convolution layer, which has a filter with a size of N q q q q The size of the convolution kernel of the two-dimensional complex-valued convolution layer is q q The complex filter matrix is convolved with the complex vector , and the complex-valued convolution operation is implemented by the following matrix form: ; where A and B are real matrices, and x and y are real vectors, denotes the real part, denotes the imaginary part, denotes a real-valued convolution operation; The output of the initial two-dimensional complex-valued convolutional layer is sequentially subjected to two-dimensional complex-valued batch normalization processing and complex-valued activation function processing; wherein the complex-valued LeakyReLU activation function is applied to the real part and the imaginary part of the variable of the previous layer, respectively. Therefore, the output of a two-dimensional complex-valued convolutional layer can be represented as: ; wherein, and denote the filter matrix and bias vector of the k th two-dimensional complex-valued convolutional layer.
3. The DOA estimation method based on deep complex-valued convolution attention residual network according to claim 2, characterized in that, S22, the deep complex-valued convolutional attention residual network utilizes a plurality of complex-valued convolutional attention modules connected in sequence along a signal processing direction to form a deep feature extraction path, and the number of filter channels of the convolutional layer in the subsequent complex-valued convolutional attention module is greater than or equal to the number of filter channels of the convolutional layer in the previous module; Each complex-valued convolutional attention module includes a plurality of two-dimensional complex-valued convolutional layers and a complex-valued convolutional block attention network; the two-dimensional complex-valued convolutional layers continue to operate in the complex domain, further performing nonlinear transformation and spatial feature extraction on the feature map, and the number of filters between modules is designed to gradually increase to learn feature representations at different levels and dimensions; The channel attention module in the complex convolution block attention network is used for adjusting channel weights, and the input is complex features obtained by a two-dimensional complex convolution layer , and mapping the complex features to a real number field to calculate a feature matrix ; Global max pooling is performed on the feature map of each channel and global average pooling , and then the results are spliced to obtain The spliced result is input into the perception machine MLP, and the complex scaling parameter is generated through the Sigmoid activation function ; at the same time, the output is divided into scaling factors applied to the real part and the imaginary part and , to construct a complex scaling factor ; finally, the feature map is scaled in the channel dimension to obtain the enhanced feature , represents element-level multiplication, that is, corresponding position elements are multiplied one by one. The spatial attention module in the complex convolution block attention network is used for adjusting the spatial position weight, and the enhanced feature map is taken modulo by the complex channel attention module , average pooling in the spatial dimension is performed , and maximum pooling is performed , and then splicing is performed to obtain , then input to the complex convolution layer to learn the weight of each position, and the spatial attention weight is obtained through the Sigmoid activation function , and finally the feature map is weighted to obtain the enhanced feature .
4. The DOA estimation method based on deep complex-valued convolution attention residual network according to claim 3, characterized in that, S23, the deep complex-valued convolutional attention residual network utilizes a cross-layer residual connection structure to adopt cross-layer shortcut connection to add the input features of the module to the output features of the module directly or after a simple transformation to realize cross-module feature fusion, specifically: Each complex convolution attention module is provided with a residual connection, and the input features of the module are added to the output of the module after processing x to form the final output F ( x ) + F ( x ) + x ; help gradient propagation and train deep network; The final output end of the feature extraction module is further provided with a flattening layer for structure conversion to map and convert the multi-dimensional complex-valued features from the feature extraction module into a one-dimensional complex-valued feature vector to adapt to the input requirements of the subsequent fully connected layer.
5. The DOA estimation method based on deep complex-valued convolutional attention residual network according to claim 1, characterized in that, In step S3, the DOA estimation module includes a plurality of CV-FCNN layers and a complex-valued Dropout layer inserted between each adjacent two CV-FCNN layers; The CV-FCNN layer uses a complex weight matrix and a complex bias vector to perform affine transformation on a one-dimensional complex-valued feature vector input, and introduces nonlinearity through a complex-valued activation function to realize high-level semantic mapping of the feature space; wherein the CV-FCNN of the first k layer can be represented as: ; ; wherein, is the output of the first k -1 layer complex weight, is the input of the first k -1 layer, is the output of the first k layer bias, is the output of the first k layer, and LeakyReLU is a complex activation function. The complex-valued Dropout layer serves as a regularization means to randomly "discard" neurons with a certain probability during training while discarding the real part and the imaginary part, reducing the synergistic adaptation between neurons and enhancing the generalization ability of the model to prevent overfitting when the training data is limited.
6. The DOA estimation method based on deep complex-valued convolutional attention residual network according to claim 1, characterized in that, In step S4, the output module specifically includes: A feature conversion unit for converting the complex-valued feature vector of the complex-valued fully connected network layer into a real-valued representation; the real part and the imaginary part of the complex-valued vector are spliced together to form a real-valued vector with doubled dimensions; An output mapping layer serving as a standard real-valued fully connected layer FCNN with the number of output neurons equal to the total number G of pre-set discrete angle grid points; this layer is responsible for linearly mapping the converted high-dimensional real-valued feature vector to G angle dimensions; An activation function unit adopting a Sigmoid activation function to compress each output value to the interval (0, 1) so that it can be interpreted as the posterior probability of the existence of a signal source corresponding to the angle grid point.
7. The DOA estimation method based on deep complex-valued convolutional attention residual network according to claim 6, characterized in that, The output module outputs a G-dimensional real-valued probability vector, i.e., an estimated angle space spectrum, which maps the feature vector to the pre-divided grid and finally to the angle distribution probability space through the Sigmoid activation function to determine the DOA estimation value; specifically represented as: ; wherein denotes the concatenation of the real and imaginary parts of the output of the last CV-FCNN layer, denotes the weight matrix of the output layer.
8. The DOA estimation system based on deep complex-valued convolutional attention residual network, characterized in that, The DOA estimation method of any one of claims 1-7, specifically includes: A system input interface for receiving externally incoming complex-valued SCM data; A feature extraction module having an input end electrically connected or in data communication with the system input interface and an output end connected to the input end of the flattening layer; for extracting complex-valued feature representations related to DOA from the original SCM layer by layer and deeply; A flattening layer, which converts the multi-dimensional complex-valued feature tensor output by the feature extraction module into a one-dimensional long vector form, so as to adapt to the input requirements of the subsequent fully connected layer; A DOA estimation module, which is connected with the output end of the flattening layer and is connected with an output module, is used to map the high-dimensional complex-valued feature vector extracted by the feature extraction module to a representation space directly related to the DOA estimation task; An output module and an output interface, which are electrically connected or in data communication, are used to output the DOA estimation probability spectrum.
Citation Information
Patent Citations
Joint modulation identification method based on attention mechanism and residual structure
CN115514597A
DOA (direction of arrival) estimation method based on dual-classification marking vision Transform
CN117556350A