Direction of arrival estimation method based on complex value convolution and electronic equipment

By combining complex-valued convolution and attention mechanism, the complex-valued covariance matrix features are extracted and weighted, which solves the low resolution problem of the DOA estimation algorithm under low signal-to-noise ratio and small number of snapshots, and achieves more accurate direction of arrival estimation.

CN120802167APending Publication Date: 2025-10-17应急管理部大数据中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510877719.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing DOA estimation algorithms have low resolution and accuracy in scenarios with low signal-to-noise ratio and a small number of snapshots, making it difficult to accurately estimate the direction of arrival of the signal source.

Method used

A direction of arrival estimation method based on complex-valued convolution is adopted, combining the complex-valued convolution module, the complex-valued attention module and the complex-valued fully connected module. By generating a complex-valued covariance matrix, the complex-valued two-dimensional convolution layer and the complex-valued attention mechanism are used to extract features, and then reweight and map them to achieve more accurate DOA estimation.

Benefits of technology

The accuracy and angular resolution of DOA estimation are improved, and the direction of arrival of the signal source can be accurately estimated under low signal-to-noise ratio and small number of snapshots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120802167A_ABST
    Figure CN120802167A_ABST
Patent Text Reader

Abstract

The invention discloses a direction-of-arrival estimation method based on complex value convolution and electronic equipment, and one embodiment of the direction-of-arrival estimation method based on complex value convolution comprises the following steps: generating a complex value covariance matrix based on an array receiving signal; inputting the complex covariance matrix into a direction-of-arrival estimation model to obtain a direction-of-arrival estimation value output by the direction-of-arrival estimation model; wherein the direction of arrival estimation model comprises a complex value two-dimensional convolution module, a complex value attention module and a complex value full connection module; the complex-valued two-dimensional convolution module is used for extracting features of the complex-valued covariance matrix and outputting first complex-valued features; the complex value attention module is used for re-weighting the first complex value feature based on an attention mechanism to generate a second complex value feature; and the complex value full connection module is used for flattening the input second complex value feature and generating a direction-of-arrival estimation value. Therefore, direction-of-arrival estimation with higher precision can be realized in a scene with low SNR (Signal to Noise Ratio) and a small number of snapshots.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application is in the field of array signal processing, in particular to a method for direction of arrival estimation based on complex value convolution and electronic device. BACKGROUND

[0002] Direction of Arrival (DOA) estimation is a technique for determining the angle or elevation of one or more signal sources relative to an array of sensors, the core idea of which is to use the time difference, phase difference or amplitude difference of signals received by sensors at different positions in the array, combined with signal processing algorithms, to infer the position information of the signal source.

[0003] DOA estimation algorithms have a wide range of application scenarios in tasks such as target source positioning in earthquake rescue, jammer or radiator positioning in electronic warfare, satellite communication, sound source positioning and medical imaging.

[0004] DOA estimation algorithms mainly include model-driven methods and data-driven methods. Model-driven methods mainly include conventional beamformers, subspace-based methods and compressed sensing-based methods. Conventional beamformers are the most basic DOA estimation method, the core principle of which is to adjust the weighting phase of the sensor array to coherently superimpose signals in a specific direction, thereby enhancing the signal in that direction and suppressing other direction interference, and determining the direction of arrival of the signal source through spatial scanning. This method is simple to calculate and does not require matrix inversion or eigenvalue decomposition, but its resolution is low and it cannot distinguish between signals with similar angles and can introduce false peaks to cause side interference in low SNR conditions. Subspace-based methods are mainly represented by Multiple Signal Classification (MUSIC) algorithm and Estimation of Signal Parameters via Rotational Invariance Techniques (ESPRIT) algorithm. The MUSIC algorithm uses the orthogonality of the noise subspace and the steering vector to construct a spatial spectrum, and the spectral peak corresponds to the DOA. It has high resolution and can break through the Rayleigh limit, but it is sensitive to coherent signals. The ESPRIT algorithm uses the translational invariance of the array to directly solve the DOA through rotational invariance, without the need for spectral peak search, and it does not need to perform angle scanning, so it is computationally efficient, but the array geometry needs to satisfy the translational invariance. Compressed sensing-based methods use the sparsity of signals in the spatial angle domain and combine with compressed sensing theory to accurately recover the direction of the signal source from a small amount of observation data. This method can suppress noise effects through sparsity constraints, but has high computational complexity.

[0005] The information disclosed in this Background section is only for the purpose of increasing an understanding of the general context of the present application and does not therefore constitute an acknowledgement or a form of suggestion that this information forms prior art that is already known to a person of ordinary skill in the art. SUMMARY

[0006] The purpose of the present application is to achieve higher precision DOA estimation in low SNR and fewer snapshot scenarios.

[0007] To achieve the above-mentioned purpose, in a first aspect, the present application provides a method for direction of arrival estimation based on complex-valued convolution, comprising the following steps: generating a complex-valued covariance matrix based on array received signals; inputting the complex-valued covariance matrix into a direction of arrival estimation model to obtain a direction of arrival estimation value output by the direction of arrival estimation model; wherein the direction of arrival estimation model comprises a complex-valued two-dimensional convolution module, a complex-valued attention module and a complex-valued fully connected module; the complex-valued two-dimensional convolution module is used to extract features of the complex-valued covariance matrix and output a first complex-valued feature; the complex-valued attention module is used to: based on an attention mechanism, reweight the first complex-valued feature to generate a second complex-valued feature; the complex-valued fully connected module is used to: flatten the input second complex-valued feature to obtain a third complex-valued feature; connect the real part and the imaginary part of the third complex-valued feature to obtain a connection result; map the connection result to a real space to obtain a real-valued vector, and map the real-valued vector as input to a classification grid to obtain the direction of arrival estimation value.

[0008] In an embodiment of the present application, the complex-valued two-dimensional convolution module comprises a complex-valued two-dimensional convolution layer and a complex-valued two-dimensional batch normalization layer, comprising: the weight matrix of the convolution kernel of the complex-valued two-dimensional convolution layer is represented as W = A + iB, the complex-valued element vector is represented as V = C + iD, the real part and the imaginary part of the complex number are represented as R(·) and I(·) respectively, and the real-valued convolution operation is represented as *, then the complex-valued convolution operation is represented as:

[0009] The complex-valued two-dimensional batch normalization layer performs normalization operation on the output of the complex-valued two-dimensional convolution layer, and scales the data centered on zero through the square root of the variance of two components, so that the mean of the output is 0 and the variance is 1.

[0010] In an embodiment of the present application, the complex-valued attention module comprises a complex-valued channel attention module; the complex-valued channel attention module is used to: process the real part and the imaginary part of the first complex-valued feature respectively to obtain the real part corresponding to the real part channel attention weight and the imaginary part channel attention weight.

[0011] In an embodiment of the present application, the respective processing of the real part and the imaginary part of the first complex-valued feature to obtain the real-part channel attention weight corresponding to the real part and the imaginary-part channel attention weight includes: performing global average pooling operation on the real part and the imaginary part respectively to obtain channel-level statistics; the real part or the imaginary part is sequentially subjected to channel re-scaling through a fully connected layer, a ReLU activation function and a Sigmoid activation function to generate the channel attention weight corresponding to the channel; and the attention weights of the real part and the imaginary part are extended to the original spatial dimension for channel weighting to obtain a second complex-valued feature.

[0012] In an embodiment of the present application, the complex-valued attention module includes a spatial attention module; the spatial attention module is configured to generate spatial attention weights corresponding to spatial positions, wherein the spatial attention weights are used for reweighting the first complex-valued feature.

[0013] In an embodiment of the present application, the complex-valued attention module includes a Transformer structure; the first complex-valued feature is a complex-valued embedding vector, and position encoding is added, wherein the position encoding indicates the spatial position information of the signal; the Transformer structure calculates a complex-valued dot product for each complex-valued embedding vector to generate position attention weights, wherein the position attention weights are used for reweighting the first complex-valued feature.

[0014] In an embodiment of the present application, the complex-valued fully connected module includes a plurality of complex-valued fully connected layers and a joint dropout layer corresponding to each complex-valued fully connected layer; wherein the joint dropout layer is configured to randomly assign a value of 0 to elements in a weight matrix of the complex-valued fully connected layer to perform joint dropout operation on the real part and the imaginary part.

[0015] In an embodiment of the present application, the training process of the DOA estimation model includes a pre-training stage and a fine-tuning stage; the pre-training stage is trained using general simulation data samples; and the fine-tuning stage is trained using DOA data samples.

[0016] In an embodiment of the present application, the signal model is that K narrowband independent source signals are incident on an M-element antenna array from directions θ k with Gaussian noise; the complex-valued covariance matrix is generated based on the array received signal, including: the complex-valued covariance matrix of the signal under a limited snapshot is obtained through the following formula

[0017]

[0018] wherein X(t) represents a signal vector at time t, H represents its conjugate transpose, and T represents the number of snapshots.

[0019] In a second aspect, the present application provides an electronic device, comprising: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of the first aspect.

[0020] Compared with the prior art, the method and the electronic device for direction of arrival estimation based on complex-valued convolution according to the present application fuse complex-valued convolution with attention mechanism, so that the whole network can not only learn joint features between complex numbers and global features, maintain the amplitude correlation of the signal, but also enhance key features and suppress irrelevant features, so that the network learns features that help locate the signal source and improve the DOA estimation accuracy. The present application designs complex-valued convolution, which can more naturally process the original complex signal and directly process complex input. The present application designs complex-valued attention mechanism, which can dynamically adjust the weight of each feature and make the model pay more attention to important information. The design of complex-valued convolution is more in line with the physical properties of the signal, retains the phase difference information of the signal, and improves the authenticity of modeling. The introduction of attention can utilize global context information for inter-channel modeling, improve the global modeling capability of the network, and enhance key features to help the model focus on features with stronger discrimination, thereby improving the angle resolution. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is a schematic diagram of an embodiment of the method for direction of arrival estimation based on complex-valued convolution according to the present application;

[0022] Figure 2 is a schematic diagram of an embodiment of the method for direction of arrival estimation according to the present application;

[0023] Figure 3 is a schematic diagram of an embodiment of the system for direction of arrival estimation based on complex-valued convolution according to the present application. DETAILED DESCRIPTION

[0024] The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings, but it should be understood that the scope of protection of the present application is not limited by the specific embodiments.

[0025] Unless otherwise clearly indicated, throughout the specification and claims, the term "comprise" or variations such as "comprises" or "comprising" will be understood to imply the inclusion of a stated element or group of elements but not the exclusion of any other element or group of elements.

[0026] As shown in Figure 1 , it shows the flow of an embodiment of the method for direction of arrival estimation based on complex-valued convolution according to the present application. As shown in Figure 1 , the method for direction of arrival estimation based on complex-valued convolution comprises the following steps:

[0027] Step 101, generating a complex-valued covariance matrix based on an array received signal.

[0028] Step 102, inputting the complex-valued covariance matrix into a DOA estimation model to obtain a DOA estimation value output by the DOA estimation model.

[0029] In the embodiment, the DOA estimation model comprises a complex-valued two-dimensional convolution module, a complex-valued attention module and a complex-valued full connection module.

[0030] In the embodiment, the complex-valued two-dimensional convolution module is configured to extract features of the complex-valued covariance matrix and output a first complex-valued feature.

[0031] In the embodiment, the complex-valued attention module is configured to reweight the first complex-valued feature based on an attention mechanism and generate a second complex-valued feature.

[0032] In the embodiment, the complex-valued full connection module is configured to flatten the input second complex-valued feature to obtain a third complex-valued feature, connect real and imaginary parts of the third complex-valued feature to obtain a connection result, map the connection result to a real space to obtain a real-valued vector, and map the real-valued vector as input to a classification grid to obtain the DOA estimation value.

[0033] It should be noted that the complex-valued convolution and the attention mechanism are fused so that the entire network can not only learn the joint features and global features between complex numbers, maintain the amplitude correlation of the signal, but also enhance the key features and suppress the irrelevant features, so that the network learns the features that are helpful for positioning the signal source and improves the DOA estimation accuracy. The complex-valued convolution is designed to more naturally process the original complex signal and directly process the complex input. The complex-valued attention mechanism is designed to dynamically adjust the weight of each feature, so that the model pays more attention to important information. The design of the complex-valued convolution is more consistent with the physical properties of the signal, retains the phase difference information of the signal, and improves the authenticity of modeling. The introduction of attention can utilize global context information for inter-channel modeling, improve the global modeling capability of the network, and enhance key features to help the model focus on features with stronger discrimination, thereby improving the angle resolution.

[0034] Meanwhile, the method provided by the present application can use the complex-valued convolution module to simultaneously convolve the real value and the imaginary value. The shallow network is used to extract low-level features, and the deep network is used to obtain deep features. In order to enhance the dependency relationship of the intermediate layer, a complex-valued attention mechanism is also designed. Specifically, the present application combines the advantages of complex-valued convolution and attention, can simultaneously learn the information of real value and imaginary value, and realizes more effective hierarchical feature extraction. Meanwhile, the present application can pay attention to the feature information of real value and imaginary value respectively, extract more important feature information and improve the accuracy of DOA estimation.

[0035] By contrast, (1) in a low signal-to-noise ratio environment, the signal can be submerged in noise. This will cause the covariance matrix of the received signal to be contaminated, and then the noise subspace and the signal subspace are difficult to separate, so that the MUSIC algorithm and the ESPRIT algorithm cannot correctly extract the signal subspace, resulting in spectral peak ambiguity or false peak value. (2) In the environment with a small number of snapshots, the estimated value of the covariance matrix deviates greatly from the real covariance matrix, which will cause the signal eigenvalue to be underestimated and the noise eigenvalue to be overestimated, and then the signal subspace is mixed with the noise component, causing the subspace aliasing, and finally causing the performance of the DOA estimation algorithm to be inaccurate. (3) The CNN-based method can automatically extract spatial features from the covariance matrix of the received signal without the need for artificial design of complex feature engineering. However, many CNN-based DOA estimation methods use real-value networks to split complex data into real and imaginary parts, which makes the network only learn features from the real and imaginary parts respectively, resulting in the loss of coupling information in the complex domain. In addition, the traditional CNN has excellent local modeling, but focusing on the extraction of local information will ignore the global information of the signal, which will inevitably affect the estimation effect of the global structure, resulting in inaccurate final DOA estimation results.

[0036] As an example, the network model structure of the DOA estimation model of the present application is as shown in Figure 2 The complex-valued covariance matrix of the array received signal is taken as the input of the network. After feature extraction by the complex-valued convolution layer, after the second convolution and activation, the model introduces a complex-valued channel attention module to reweight and adjust the global features of the real and imaginary channels respectively, and enhance the important feature channels. This module improves the selectivity of the features, so that the model pays more attention to important features. Subsequently, the attention-enhanced features are recombined, and the complex-valued attention module improves the selectivity and robustness of the feature representation. Subsequently, the attention-enhanced features (i.e. the second complex-valued features) are mapped back to the spectral space through a fully connected layer. This process effectively decomposes the covariance matrix into multiple orthogonal feature subspaces containing different angle resolution information.

[0037] As an example, the relevant explanations are given to the terms involved in Figure 2

[0038] Complex-valued two-dimensional convolution layer: performs complex convolution operation to extract the joint features hidden in the real and imaginary parts, and preserves the phase information in the signal.

[0039] Complex-valued channel attention module: enhances the modeling ability of the convolutional neural network for channel features.

[0040] Real part compression: compresses the spatial features of the real part into a global description value to capture the real part channel information.

[0041] ​Real part excitation: generate a weight for each channel of the real part.

[0042] Real part weighting: weight each channel of the real part in the original feature map according to the real part channel weight.

[0043] Imaginary part compression: compress the spatial features of the imaginary part into a global description value to capture the imaginary part channel information.

[0044] Imaginary part excitation: generate a weight for each channel of the imaginary part.

[0045] Imaginary part weighting: weight each channel of the imaginary part in the original feature map according to the imaginary part channel weight.

[0046] Complex-valued fully connected layer: perform nonlinear transformation, integration and mapping of the features extracted by the complex-valued convolution to the spectral space while preserving the complex structure features.

[0047] Fully connected layer: map the real number vector composed of the real part and the imaginary part of the output of the complex-valued fully connected layer to the final probability space to realize multi-label binary classification.

[0048] In some embodiments, the complex-valued two-dimensional convolution module includes a complex-valued two-dimensional convolution layer and a complex-valued two-dimensional batch normalization layer, wherein the weight matrix of the convolution kernel of the complex-valued two-dimensional convolution layer is represented as W = A + iB, the complex-valued element vector is represented as V = C + iD, the real part and the imaginary part of the complex number are represented as R(·) and I(·) respectively, the real-valued convolution operation is represented as *, and the complex-valued convolution operation is represented as:

[0049]

[0050] The complex-valued two-dimensional batch normalization layer performs normalization operation on the output of the complex-valued two-dimensional convolution layer, and scales the data centered on zero through the square root of the variance of two components, so that the mean of the output is 0 and the variance is 1.

[0051] The present implementation designs a complex-valued two-dimensional convolution (CV-Conv2d) layer to better extract features from the covariance matrix. Assuming that the weight matrix of the convolution kernel is represented as W = A + iB, the complex-valued element vector is represented as V = C + iD, the real part and the imaginary part of the complex number are represented as R(·) and I(·) respectively, and the real-valued convolution operation is represented as *, then the complex-valued convolution operation can be represented as:

[0052]

[0053] After the CV-Conv2d operation, in order to improve the training efficiency, the output of the convolution layer is normalized using a complex-valued two-dimensional batch normalization (CV-BN) layer, so that the mean of the output is 0 and the variance is 1. CV-BN2d is not simply normalized by separating the real part and the imaginary part, but is scaled by the square root of the variance of the two components to center the data around zero. In addition, a complex-valued scaled exponential linear unit (CV-SELU) is used for the real part and the imaginary part of the variable of the previous layer, respectively. The CV-SELU can be expressed as:

[0054]

[0055] CV-SELU(x) = SELU(R(x)) + iSELU(I(x))

[0056] The output O processed by the i-th CV-Conv2d layer is processed by CV-SELU, which can be expressed as:

[0057] O = CV-SELU(CV-BN(w i *O i +b i ))

[0058] In the above formula, w i represents the weight matrix of the i-th CV-Cond2d layer, b i represents the bias term of the i-th CV-Cond2d layer, and CV-BN(·) represents the complex-valued batch normalization operation.

[0059] Therefore, the signal data is essentially a complex signal that carries amplitude information and phase information. Using complex-valued convolution can preserve the complex multiplication and addition rules, truly reflect the mathematical properties of the signal, and naturally model high-level semantic features such as phase difference, thereby improving the estimation performance. The complex-valued convolution layer can better adapt to the covariance matrix of the array received signal, which is a complex-valued covariance matrix, and extract phase information from it.

[0060] In contrast, traditional DOA estimation methods based on convolutional neural networks usually use real-valued convolution, which splits the complex-valued data into real and imaginary parts for convolution. This destroys the algebraic structure of the signal data and makes it difficult to capture key information such as phase difference.

[0061] In some embodiments, the complex-valued attention module includes a complex-valued channel attention module.

[0062] The complex-valued channel attention module is configured to process a real part and an imaginary part of the first complex-valued feature respectively to obtain a real part channel attention weight corresponding to the real part and an imaginary part channel attention weight.

[0063] Introducing an attention mechanism in DOA estimation can reduce the noise part in the signal. The channel attention mechanism is a second-order channel attention mechanism that can adaptively adjust the channel interdependence in high-order statistical features, so that the network pays more attention to important parts.

[0064] The channel attention mechanism for complex values is used to redistribute channel weights in the network. The complex-valued feature is split into a real part and an imaginary part, which are sent into two SE modules respectively, and then the attention weights are fused to enhance useful feature channels and suppress redundant channels, thereby improving the sensitivity of the model to effective direction information in DOA. Especially in low SNR scenarios, some imaginary part channels have more noise and less information, and the SE attention can automatically weaken and reduce the influence of useless or interfering channels. The generalization ability of the model is improved.

[0065] In some embodiments, the processing of the real part and the imaginary part of the first complex-valued feature to obtain the real part channel attention weight corresponding to the real part and the imaginary part channel attention weight includes: performing a global average pooling operation on the real part and the imaginary part respectively to obtain channel-level statistics; the real part or the imaginary part is sequentially subjected to a fully connected layer, a ReLU activation function and a Sigmoid activation function for channel recalibration to generate channel attention weights corresponding to the channels; and the attention weights of the real part and the imaginary part are extended to the original spatial dimension for channel weighting to obtain a second complex-valued feature.

[0066] First, the real part O R and the imaginary part O I are subjected to a global average pooling operation to obtain channel-level statistics. This process can be represented as:

[0067]

[0068] In the above formula, S p takes S R or S I , which represents the channel statistics of the real part and the imaginary part, respectively. O p takes O R or O I , which represents the output of the real part and the imaginary part after complex-valued convolution. This process achieves spatial compression by averaging each channel, introduces global context information of the channel, so that the network focuses on the relationship between the channels, improves the ability of the model to distinguish important channels, and enhances the expressiveness.

[0069] Then O R or O IThe channel re-labeling is performed through the full connection layer and ReLU activation function and Sigmoid activation function in sequence, and the attention weights corresponding to the channels are generated respectively, and the process can be expressed by a formula as follows:

[0070]

[0071] In the above formula, A p , A R and A I respectively represent the attention weights generated by the real part channel and the imaginary part channel. and respectively represent the dimension reduction matrix and the dimension recovery matrix, and the dimension reduction matrix and the dimension recovery matrix are respectively used for the real part channel and the imaginary part channel, and the values are different according to the different real part and imaginary part. After the process, the model can learn the importance weight of the channel.

[0072] Finally, the attention weights of the real part and the imaginary part are extended to the original space dimension for channel weighting, so as to enhance the important channel and suppress the unimportant channel, and the process can be expressed as follows:

[0073]

[0074] In the above formula, A and respectively represent the outputs of the real part and the imaginary part after the re-weighting (i.e., the second complex value feature), and then the outputs are fused and the complex value convolution operation is continued.

[0075] Therefore, based on the attention weights independently generated for each channel, the model can adaptively enhance the response of the important channel to achieve noise suppression. The independent full connection network enables the real part and the imaginary part to have different weight generation strategies, and can more flexibly adapt to the asymmetric characteristics of the complex signal.

[0076] In some embodiments, the complex attention module comprises a spatial attention module; the spatial attention module is configured to generate spatial attention weights corresponding to spatial positions, wherein the spatial attention weights are used to re-weight the first complex value feature.

[0077] The spatial attention mechanism focuses on the importance of the spatial region of the feature map, rather than the relationship between the channels. By weighting the spatial positions, the spatial features corresponding to the signal source direction are enhanced, and the interference of the noise region is suppressed.

[0078] In some embodiments, the complex-valued attention module comprises a Transformer structure; the first complex-valued feature is a complex-valued embedding vector, and position encoding is added, wherein the position encoding indicates the spatial position information of the signal; the Transformer structure calculates a complex-valued dot product for each complex-valued embedding vector to generate a position attention weight, wherein the position attention weight is used to reweight the first complex-valued feature.

[0079] Here, the Transformer structure adopts a self-attention mechanism to model global feature dependency relationships and solve the limitations of traditional CNN local modeling.

[0080] As an example, the complex-valued covariance matrix is mapped to a complex-valued embedding vector through a complex-valued linear layer of the model, and position encoding is added to retain the spatial position information of the signal. The channel attention module is replaced to calculate a complex-valued dot product of Query, Key, and Value for each complex-valued embedding vector to generate an attention weight, capture global inter-channel and spatial dependency relationships, and maintain a pre-training and fine-tuning two-stage training process in the training strategy.

[0081] In some embodiments, the complex-valued fully connected module comprises a plurality of complex-valued fully connected layers and a joint dropout layer corresponding to each complex-valued fully connected layer; wherein the joint dropout layer is configured to randomly assign a value of 0 to elements in a weight matrix of the complex-valued fully connected layer and perform a joint dropout operation on the real part and the imaginary part.

[0082] First, the features extracted from each filter on the CV-Conv2d are flattened and connected to a series of complex-valued fully connected layers. The i-th complex-valued fully connected layer can be formulated as:

[0083] FC i = W i,i-1 (CV-SELU(FC i-1 ))+b i

[0084] In the above formula, FC i represents the output after the i-th complex-valued fully connected layer, W i,i-1 represents the weight matrix of the i-1-th layer and the i-th layer, and b i represents the corresponding bias term. After each complex-valued fully connected layer, CV-Dropout is used to prevent overfitting during training, which randomly assigns a value of 0 to a specific element in the weight matrix W i,i-1 with a probability of p. CV-Dropout performs a joint dropout operation on the real part and the imaginary part, ensuring that these components can be set to 0 at the same time. This avoids the problem of spurious phase shift caused by only dropping the real or imaginary component in traditional dropout.

[0085] After obtaining the last complex-valued fully connected layer output, the real part and the imaginary part are connected to realize the mapping of the complex-valued feature to the real space. Then, the real-valued vector is taken as the input of the fully connected layer to be mapped to the predefined grid.

[0086] In some embodiments, in the model training stage, the present application uses a two-stage training process of pre-training and fine-tuning to further improve the performance of the model.

[0087] In some embodiments, the training process of the direction of arrival estimation model includes a pre-training stage and a fine-tuning stage; wherein the pre-training stage uses general simulation data samples for training; and the fine-tuning stage uses direction of arrival data samples for training.

[0088] As an example, the model is trained using simulation data. The potential space to be estimated ranges from -60° to 60°, which is divided into L equally spaced grids at a resolution of 1°. In the training set, the directions of the two signal sources are two randomly selected values in the potential space. In order to improve the generalization ability of the model under different signal-to-noise ratios, samples with the same angle pair are repeatedly generated under a set of uniform signal-to-noise ratio levels. The spatial spectrum recovery problem is modeled as a multi-label classification task by CCAN, and binary classification is performed on each grid. CCAN only outputs 1 at the grid of the true direction, and outputs 0 at other grids.

[0089] In addition, the proposed CCAN model is pre-trained on a training set composed of expected covariance matrices. Then, fine-tuning is performed on a training set composed of sampled covariance matrices using a smaller learning rate, and the cross-entropy loss between the estimated spatial spectrum and the true label is used as the loss function.

[0090] It should be noted that, in order to solve the problems of instability, long training time and instability of traditional training, the present application divides the training process into two stages of pre-training and fine-tuning. In the pre-training stage, the general low-level or middle-level features are learned through simulation data to adapt the convolution kernel to the complex signal structure and reduce the instability in the early stage of training, such as gradient explosion or disappearance. In the fine-tuning stage, the model learns to transition from general features to specific features for DOA estimation to adapt to different signal-to-noise ratio changes, classify the direction information of the final output, and improve the accuracy of DOA estimation.

[0091] In some embodiments of the present application, the DOA estimation problem is regarded as an inverse problem of spectral recovery, and the signal model is assumed that K narrowband independent source signals from direction θ k are incident on an M-element antenna array, and the received signal is a complex-valued signal with Gaussian noise. The specific implementation process is as follows:

[0092] Step one: get the estimation of signal covariance matrix under limited snapshots by the following formula.

[0093]

[0094] In the formula, X(t) represents a signal vector at time t, H represents a conjugate transpose, and T represents the number of snapshots.

[0095] Step two: input the complex covariance matrix into the network. First, the phase feature information is extracted through the complex convolution layer, then the channel recalibration is performed through the complex channel attention module, then the feature map is mapped back to the frequency spectrum space through the complex full connection layer, finally the complex features are converted into real value probability distribution, the binary classification label of sparse spectrum is matched, and the final result is obtained.

[0096] It should be noted that in the network structure, the model is improved from two aspects of convolution and attention mechanism. The attention mechanism is used in cooperation with the complex convolution, so that the model can locally extract features and globally adjust responses, and is more suitable for modeling the directional changes of array signals. In the implementation mode, a complex convolution network is designed, and the complex convolution network is fused with the channel attention mechanism to construct a new complex channel attention network. The channel attention module is designed on the intermediate structure of the network, obtains the input from the complex convolution, and transmits the result to the subsequent structure to obtain more accurate hierarchical information features. Thus, higher precision DOA estimation can be realized in the scene of low SNR and small number of snapshots.

[0097] The experimental results prove that the DOA estimation method proposed in the application has good precision when detecting unknown signal sources. The DOA estimation precision of the application is better than that of MUSIC, ESPRIT and the method based on real value convolution CNN under the change of signal-to-noise ratio from-20dB to 20dB. In the extremely low snapshot scene, the DOA estimation precision of the application is better than that of the above methods.

[0098] Reference will now be made to Figure 3 , which shows a structural schematic diagram of an electronic device (such as a terminal device or a server) suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. The electronic device shown in the figure is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0099] As Figure 3As shown, the electronic device can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 302 or loaded into a random access memory (RAM) 303 from a storage device 308. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0100] In general, the following devices can be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 can allow the electronic device to communicate wirelessly or wired with other devices to exchange data. Although Figure 3 An electronic device having various devices is shown, but it is understood that all of the shown devices are not required to be implemented or possessed. More or less devices can be alternatively implemented or possessed.

[0101] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 309, or installed from the storage devices 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0102] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.

[0103] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0104] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and not be assembled into the electronic device.

[0105] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0106] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0107] The units described in the embodiments of the present disclosure can be implemented by hardware, software, or a combination of hardware and software. In some cases, the names of the units do not constitute a limitation on the units themselves.

[0108] The functions described in this specification can be implemented in part or in whole through one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0109] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a lined- up electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] The above description is only preferred embodiments of the present disclosure and the explanation of the applied technical principles. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and also covers other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0111] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0112] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0113] The foregoing description of specific exemplary embodiments of the application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise forms disclosed, and obviously many modifications and variations are possible in light of the above teaching. It is intended that the scope of the application be limited not with this detailed description, but rather by the claims appended hereto.

Claims

1. A method for estimating direction of arrival based on complex-valued convolution, characterized in that: The following steps are involved: generating a complex-valued covariance matrix based on the array received signal; Inputting the complex-valued covariance matrix into a direction-of-arrival estimation model to obtain a direction-of-arrival estimation value output by the direction-of-arrival estimation model; The DOA estimation model includes a complex-valued two-dimensional convolution module, a complex-valued attention module, and a complex-valued fully connected module. The complex-valued two-dimensional convolution module is used to extract the features of the complex-valued covariance matrix and output a first complex-valued feature; The complex-valued attention module is configured to: reweight the first complex-valued feature based on an attention mechanism to generate a second complex-valued feature; The complex-valued fully connected module is used to: flatten the second complex-valued feature of the input to obtain a third complex-valued feature; concatenate the real part and the imaginary part of the third complex-valued feature to obtain a concatenation result; map the concatenation result to real space to obtain a real-valued vector; and map the real-valued vector as input to a classification grid to obtain a direction of arrival estimate.

2. The method for estimating the direction of arrival based on complex-valued convolution according to claim 1, wherein: The complex-valued two-dimensional convolution module includes a complex-valued two-dimensional convolution layer and a complex-valued two-dimensional batch normalization layer, including: The weight matrix of the convolution kernel of the complex-valued two-dimensional convolution layer is expressed as W = A + iB, the complex-valued element vector is expressed as V = C + iD, the real and imaginary parts of the complex number are expressed as R(·) and I(·), respectively, and the real-valued convolution operation is expressed as *. Then the complex-valued convolution operation is expressed as: The complex-valued 2D batch normalization layer normalizes the output of the complex-valued 2D convolutional layer by scaling the zero-centered data by the square root of the variance of the two components so that the output has a mean of 0 and a variance of 1.

3. The method for estimating the direction of arrival based on complex-valued convolution according to claim 1, wherein: The complex-valued attention module includes a complex-valued channel attention module; The complex-valued channel attention module is used to: The real part and imaginary part of the first complex-valued feature are processed separately to obtain the real part channel attention weight and imaginary part channel attention weight corresponding to the real part.

4. The method for estimating the direction of arrival based on complex-valued convolution according to claim 3, wherein: The processing of the real part and the imaginary part of the first complex-valued feature to obtain the real part channel attention weight and the imaginary part channel attention weight corresponding to the real part includes: Perform global average pooling operations on the real and imaginary parts respectively to obtain channel-level statistics; The real part or imaginary part passes through the fully connected layer and the ReLU activation function and the Sigmoid activation function for channel recalibration to generate the channel attention weight of the corresponding channel respectively; The attention weights of the real and imaginary parts are extended to the original spatial dimension for channel weighting to obtain the second complex-valued feature.

5. The method for estimating the direction of arrival based on complex-valued convolution according to claim 1, wherein: The complex-valued attention module includes a spatial attention module; The spatial attention module is used to generate a spatial attention weight corresponding to a spatial position, wherein the spatial attention weight is used to reweight the first complex-valued feature.

6. The method for direction of arrival estimation based on complex-valued convolution according to claim 1, wherein: The complex-valued attention module includes a Transformer structure; The first complex-valued feature is a complex-valued embedded vector and is added with a position code, wherein the position code indicates spatial position information of the signal; The Transformer structure calculates a complex-valued dot product for each complex-valued embedding vector and generates a position attention weight, wherein the position attention weight is used to reweight the first complex-valued feature.

7. The method for estimating direction of arrival based on complex-valued convolution according to claim 1, wherein: The complex-valued fully connected module includes a plurality of complex-valued fully connected layers and a joint dropout layer corresponding to each complex-valued fully connected layer; The joint dropout layer is used to randomly assign a value of 0 to elements in the weight matrix of the complex-valued fully connected layer, and perform a joint dropout operation on the real part and the imaginary part.

8. The method for estimating direction of arrival based on complex-valued convolution according to claim 1, wherein: The training process of the DOA estimation model includes: a pre-training phase and a fine-tuning phase; wherein The pre-training stage uses general simulation data samples for training; The fine-tuning stage uses direction of arrival data samples for training.

9. The method for direction of arrival estimation based on complex-valued convolution according to claim 1, wherein: The signal model assumes that there are K narrowband independent source signals incident on an M antenna array from direction θ, and the received signal is a complex-valued signal with Gaussian noise; The generating of a complex-valued covariance matrix based on the array-received signal comprises: The complex-valued covariance matrix of the signal under finite snapshots is obtained by the following formula Where X(t) represents the signal vector at time t, H represents its conjugate transpose, and T represents the number of snapshots.

10. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.