Sparse array direction-of-arrival estimation method based on local and global self-supervised networks
Through the differential virtual array expansion and ANM self-supervised methods of local and global self-supervised networks, the performance limitation of sparse array DOA estimation under low signal-to-noise ratio and hole conditions is solved, and high-precision recovery and estimation of gridless signals are achieved.
Patent Information
- Application Number
- CN202510513492.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-25
AI Technical Summary
Existing sparse array DOA estimation methods are limited in performance under low signal-to-noise conditions or when there are holes in the array, and self-supervised neural networks rely on labeled data collection is challenging, resulting in fundamental mismatch problems.
Local and global self-supervised networks are adopted to achieve grid-free signal recovery through differential virtual array expansion, local and global feature extraction network, and atomic norm minimization of ANM self-supervised networks, and label-free self-supervised learning is performed using Transformer network and ANM loss function.
In the case of any sparse array and many holes, it has more stable performance, solves the accuracy problem of traditional methods and achieves more accurate gridless recovery and estimation.
Smart Images

Figure CN120372218A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless signal processing, and particularly relates to a method for estimating the direction of arrival of sparse arrays based on a local and global self-supervised network. Background Technique
[0002] As an important branch of array signal processing, the DOA estimation method based on sparse arrays has been widely used in the fields of communication, radar, sonar, seismic exploration, and radio astronomy and has developed rapidly. The element spacing of a non-uniform linear array (sparse array) is relatively large, which can effectively reduce the mutual coupling effect and improve the spatial resolution. At the same time, through the covariance vectorization of sparse array signals, the corresponding virtual array and virtual array measurement data can be obtained. A larger array aperture can be obtained through the virtual array, improving the degree of freedom and parameter estimation accuracy. The DOA estimation methods for sparse arrays mainly include subspace methods and compressive sensing methods. Subspace methods mainly include: a DOA estimation method that uses the spatial spectrum results of two linear sub-arrays of a co-prime array and proves the existence and uniqueness of the solution of this method; a local spectral peak search MUSIC method that uses the phase ambiguity relationship; a DOA estimation method that uses spatial smoothing technology to construct a full-rank covariance matrix of the virtual array and then realizes it through MUSIC. However, the virtual arrays generated by sparse arrays are usually non-uniform and have holes. This processing method discards discontinuous virtual elements, and the virtual array information is not fully utilized, which in turn leads to a decrease in DOA estimation accuracy. In order to avoid the performance loss of parameter estimation caused by insufficient utilization of virtual array information, compressive sensing methods that utilize all virtual array information have been successively proposed. Meshless compressive sensing methods mainly include: a method for reconstructing the virtual array Toeplitiz matrix by interpolating the difference set virtual array of a co-prime array and constructing a nuclear norm minimization problem; a DOA estimation method for virtual array interpolation based on atomic norm; a DOA estimation method based on low-rank reconstruction of the Toeplitz covariance matrix; a meshless DOA estimation method for non-uniform sparse arrays of arbitrary geometric shapes and further deriving an irregular Vandermonde decomposition method for irregular Toeplitz matrices; however, the performance of meshless signal recovery methods is still limited under low signal-to-noise ratio conditions or when there are more holes in the array.
[0003] In addition to the above traditional methods, there are also DOA estimation methods based on neural networks, mainly including: Convolutional neural networks have been applied to DOA estimation and show excellent performance under low signal-to-noise ratio conditions; A sparse prior DOA estimation method based on a deep convolutional network that learns the inverse transform from a training dataset; For any sparse array, a denoising network-based circular and non-circular signal and dynamic attention network. The above methods rely on supervised neural network training using labeled data. However, in the real world, collecting and labeling accurate data can be challenging. Researchers have begun to focus on unsupervised or self-supervised DOA estimation methods; An unsupervised pre-training method for small datasets that uses a restricted Boltzmann machine to capture the underlying features of the input data and improve weight initialization; Designing the L_1 norm minimization as a loss function to implement an unsupervised learning method, but its performance depends on the size of the grid design and suffers from the basis mismatch problem, the same as traditional grid-based compressive sensing DOA estimation methods. Summary of the Invention
[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide a sparse array direction of arrival (DOA) estimation method based on a local and global self-supervised network in view of the deficiencies of the prior art, including the following steps:
[0005] Step 1, perform signal processing: First, construct the received signal of an arbitrary sparse array and perform difference set virtual array expansion to obtain difference set virtual array measurement data; Subsequently, construct a neural network training dataset according to the measurement data;
[0006] Step 2, establish a local and global feature extraction network: Construct a local and global feature extraction network through two or more layers of gated linear units GLU (Gated Linear Unit) and a Transformer network. Input the difference set virtual array measurement data, and establish a spatial feature relationship through the local and global feature extraction network and obtain the measurement data missing at the virtual array hole positions;
[0007] Step 3, establish an atomic norm minimization ANM (Atomic Norm Minimization) self-supervised network; The atomic norm minimization ANM self-supervised network converts complex number operations into real-valued matrix operations, and through a complex Toeplitz matrix construction layer, a complex singular value decomposition SVD (Singular Value Decomposition) layer, a complex positive semidefinite constraint PSD (Positive Semidefinite) layer, and an atomic norm minimization ANM loss function, realizes gridless signal recovery while ensuring the complex structure in the network;
[0008] Step 4, Meshless Signal Recovery and Direction of Arrival (DOA) Estimation: Input the neural network training dataset into the local and global feature extraction network. Utilize the amplitude correlation relationship between array elements to establish a reasonable model to recover the missing measurement data. At the same time, embed the Atomic Norm Minimization (ANM) method into the network to achieve label-free self-supervised data recovery. Finally, use the recovered meshless data for DOA estimation.
[0009] Step 1 includes: Set an arbitrary sparse array S with K h holes, and the array aperture is M. K narrowband uncorrelated far-field sources from different directions are θ = [θ1, θ2, …, θ K , where θ K represents the Kth signal source, and the signal x S (t) received by the sparse array S at time t is expressed as:
[0010] x S (t) = As(t) + n(t) (1),
[0011] where s(t) is the signal received at time t, n(t) is independent and identically distributed zero-mean additive Gaussian noise, and A represents the array manifold.
[0012] In Step 1, the array manifold A is expressed as:
[0013] A = [a(θ1), a(θ2), … a(θ K )] (2),
[0014] where the steering vector a(θ k ) of the kth signal source is specifically expressed as:
[0015]
[0016] where the intermediate parameter d = λ / 2, λ is the wavelength, p1 represents the position of the first array element, and p(M - K h ) represents the position of the (M - K h )th array element; j represents the imaginary unit, and e represents the natural constant;
[0017] The covariance R S of the signal is expressed as:
[0018]
[0019] where E represents the expectation of the matrix, H represents the conjugate transpose of the matrix, I represents the identity matrix, p k represents the power of the kth signal source, and p n represents the noise power.
[0020] Step 1 further includes: calculating the signal covariance matrix from finite sampling
[0021]
[0022] where, T s is the number of sampling snapshots.
[0023] Step 1 further includes: vectorizing the covariance matrix to obtain the measurement data corresponding to the difference set virtual array Calculating the set of position serial numbers of the difference set virtual array corresponding to the sparse array S where, m and n respectively represent the position serial number of the m-th array element and the position serial number of the n-th array element in the sparse array S, and deleting the repeated array element position elements in the difference set virtual array position set to obtain the final difference set virtual array position set S diff ;
[0024] Removing the measurement data of the difference set virtual array corresponding to the repeated array element positions to obtain the virtual array signal corresponding to the difference set virtual array position set S diff
[0025]
[0026] where, b(θ k ) represents the steering vector corresponding to the difference set virtual array; i represents the result after removing the repeated elements after vectorizing the identity matrix I;
[0027] Virtual array signal Interpolating 0 at the places where there are holes in the virtual array element positions, and the input data set of the local and global feature extraction network consists of the real part and the imaginary part of the interpolated measurement data; represents the set of real numbers;
[0028] Collecting D training samples with different signal directions, different numbers and positions of missing array elements, and the d-th training sample Input data set Y = {Y (1) , …, Y (D)}; d takes values from 1 to D.
[0029] Step 1 further includes: using the original training data to generate the data in the loss function, that is and G, where is the vector obtained by vectorizing and interpolating the sampling covariance in formula (5) Constructed Toeplitz matrix, G represents a binary matrix, representing The data at the corresponding array interpolation positions is 0, and the data at the non-interpolation positions of the array is 1; then, the self-supervised data corresponding to the d-th training sample is expressed as Samples are collected under different signal directions and different sparse array configurations, and the input self-supervised data set is expressed as P = {P (1) , …, P (D)};
[0030] The training data consists of data pairs (Y (d) , P (d) ), and the training data set is specifically expressed as Data = {(Y (1) , P (1) ), …, (Y (D) , P (D) )}.
[0031] Step 2 includes: The local and global feature extraction network includes a local feature modeling layer and a global feature modeling layer;
[0032] The local feature modeling layer performs the following operations: The local feature modeling layer includes a fully connected feature expansion layer, and the fully connected feature expansion layer includes a complex linear connection layer, a connection layer, a convolutional layer, and an activation layer. The training data set first passes through the fully connected feature expansion layer to obtain data First, the data features of adjacent 3 array elements are grouped by a sliding window, and the grouped data is expressed as where L = 2M - 1 is the total number of the virtual array after inserting 0 at the virtual array holes; B represents the batch size of training; d f represents the expanded feature dimension; R represents the real number field;
[0033] Local feature aggregation is performed through convolution, and the convolution kernel k e is 3 * 1, and the output X s of the convolutional layer is expressed as:
[0034]
[0035] where, d g represents the feature dimension after passing through the convolutional layer;
[0036] Then, the output data of the convolutional layer is passed through more than two layers of gated linear units GLU. For the z-th layer of gated linear unit GLU, the gating threshold δ z is obtained by a linear connection layer and a sigmoid activation layer:
[0037] δ z = Sigmoid(X s W + b) (8),
[0038] Among them, W is the weight of the linear connection layer, and b is the bias of the linear connection layer;
[0039] The output X of the z-th layer gated linear unit GLU lz is expressed as:
[0040] X lz = δ z ·(X s W + b) (9),
[0041] Set the gated linear unit GLU to have Z layers, and the output data of the Z layers is expressed as X l1 ,..., X lZ , and stack X l1 ,..., X lZ to obtain the output X of the local feature modeling layer l :
[0042]
[0043] where ° means directly stacking and paralleling vectors;
[0044] The global feature modeling layer performs the following operations: establish a Transformer network, and through the Transformer network, perform parallel processing on X d and global feature modeling to capture the long-range dependencies of the virtual array measurement data features. The Transformer network includes position encoding, multi-head attention layer, and feed-forward layer;
[0045] The position encoding models the relative or absolute position of the features of the array elements. The feature encoding X d,l of the features at the l-th array element p,l is expressed as:
[0046]
[0047] where i is the index of the input feature, using sine encoding for even indices and cosine encoding for odd indices;
[0048] Model the signal feature relationships of different array element positions through the multi-head attention layer. For the h-th head attention mechanism head h it is specifically expressed as:
[0049] head h = Attention(Q h , K h , V h ) (12),
[0050] where the attention mechanism Attention is expressed as:
[0051]
[0052] Among them, Q h , K h , V h respectively represent the query, key, and value of the h-th attention mechanism, and are obtained by linearly projecting the data X p through different linear transformation layers. d h is the dimension of Q h and K h ;
[0053] Connect the H a attention results to obtain the multi-head attention MultiHead feature X a :
[0054]
[0055] Among them, head1 represents the attention result of the first head, represents the attention result of the H a -th head, is the multi-head connection matrix, d v represents the dimension of V, and d model represents the dimension expected to be obtained after multi-head splicing; Concat represents the connection operation; Q, K, and V respectively represent the query, key, and value of the multi-head attention after splicing multiple attention results;
[0056] The feed-forward layer includes two linear layers and a linear rectifier Relu activation function. The output X f of the feed-forward layer is expressed as:
[0057] X f = Relu((X a W1)W2) (15),
[0058] where, W1 and W2 represent the network weights of the two linear layers; X a represents the multi-head attention MultiHead feature;
[0059] After residual and layer normalization processing, the output of the Transformer network is denoted as X g , and then, the learned local features and global features are connected to obtain the output X gl of the local and global feature extraction network.
[0060] Step 3 includes: the ANM self-supervised network includes a complex Toeplitz reconstruction layer, a complex singular value decomposition SVD layer, a complex semi-positive definite constraint PSD layer and an ANM loss function, which is used to realize gridless data recovery;
[0061] The complex Toeplitz reconstruction layer performs the following operations:
[0062] For the output X of the local and global feature extraction networks gl ∈R B×2×M , take X gl First row X gl,1 As real data, the second row X gl,2 The data is taken as the imaginary data, and the complex Toeplitz matrix X is obtained according to the data. T The real part X T,real and the imaginary part X T,imag The structure is:
[0063]
[0064] in, <X gl,1 > M and <X gl,2 > M Respectively represent X gl The Mth value of the first row of data and the Mth value of the second row of data;
[0065] The complex singular value decomposition SVD layer performs the following operations: SVD decomposition is performed on the Toeplitz matrix. First, the complex Toeplitz matrix X is decomposed using formula (16). T The real part X T,real and the imaginary part X T,imag Construct a real-valued combinatorial matrix C:
[0066]
[0067] Then, perform singular value decomposition SVD on the matrix C:
[0068] [U C ,Σ C ,V C ]=SVD(C) (19),
[0069] Among them, U C With V C is the eigenvector, Σ C is the eigenvalue;
[0070] The real part of the eigenvector of the complex matrix U real 、V real and the imaginary part U imag 、Vimag They are respectively:
[0071] U real = U C (1:M, 1:2:2M) (20),
[0072] U imag = U C (M + 1:2M, 1:2:2M) (21),
[0073] V real = V C (1:M, 1:2:2M) (22),
[0074] V imag = V C (M + 1:2M, 1:2:2M) (23),
[0075] X T The eigenvalue Σ of is:
[0076] Σ = Σ C (1:2:2M, 1:2:2M) (24),
[0077] The complex semi - positive definite constraint PSD layer performs the following operations: For the (i, i) - th eigenvalue, define the eigenvalue threshold activation function f max (Σ):
[0078]
[0079] where ∈ is a positive threshold, generally taken as e -6 ; Σ (i,i) represents the value of the i - th row and i - th column in the eigenvalue Σ.
[0080] Perform the complex semi - positive definite constraint PSD algorithm operation to obtain the intermediate parameter X p :
[0081]
[0082] The outputs of the complex semi - positive definite constraint PSD layer are respectively the real part and the imaginary part of X p ;
[0083] Interpolate the virtual array signal along the positive axis of the difference - set virtual array. The noise - free signal y of the positive - axis virtual array after interpolation is expressed as:
[0084]
[0085] where z(θ k ) represents the steering vector corresponding to the positive axis of the virtual array;
[0086] Based on the atomic norm theory, the atomic set of the virtual array signal is defined as:
[0087]
[0088] The atomic norm minimization problem of y is constructed as:
[0089]
[0090] where, denotes taking the infimum;
[0091] The atomic norm minimization problem is equivalent to a semidefinite programming form:
[0092]
[0093] where min T(y) denotes minimization, s.t. denotes subject to, denotes the PSD constraint: Tr denotes taking the trace; denotes calculating the Frobenius norm of the matrix after subtracting from
[0094] The Toeplitz matrix constructed from the vector y is specifically expressed as:
[0095]
[0096] The matrix is expressed as:
[0097]
[0098] where, denotes taking the L-th value in the item vector, and * denotes taking the conjugate of the matrix;
[0099] In formula (30), the PSD constraint for is implemented through the complex SVD layer and the PSD constraint layer, and the convex optimization problem equivalent to the atomic norm minimization is expressed as the loss function of the self-supervised network Specifically expressed as:
[0100]
[0101] The present invention also provides an electronic device, including a processor and a memory, where the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method described above.
[0102] The present invention also provides a storage medium storing a computer program or instructions, which, when running on a computer, execute the steps of the above-described method.
[0103] The method of the present invention can be applied to the following fields:
[0104] Radar target positioning: In a radar system, DOA estimation is used to identify the direction of a target and is commonly used in object tracking, target detection, collision warning, etc. DOA estimation helps the radar system determine the position and movement trajectory of the target.
[0105] Communication positioning: DOA estimation is used in the field of communication for beamforming, interference suppression, multi-user detection, and spatial multiplexing. By accurately estimating the direction of arrival of a signal, the capacity and signal quality of a wireless communication system can be improved, especially in a multiple-input multiple-output system.
[0106] Sonar positioning: In the acoustic field, DOA estimation is used in speech positioning, sound source positioning, and multi-microphone arrays. For example, in a speech recognition system, DOA estimation can determine the position of a speaker, thereby enabling more accurate sound separation and enhancement.
[0107] Advantageous effects: Compared with the existing meshless DOA estimation traditional methods, the method proposed by the present invention can have more stable performance in any sparse array with many holes. At the same time, compared with the existing self-supervised neural network DOA estimation methods, the new scheme proposed in this paper has meshless recovery performance, solves the basis mismatch problem of traditional neural network methods, and has more accurate recovery and estimation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0109] Figure 1 It is a schematic diagram of a local and global feature extraction network.
[0110] Figure 2 It is a schematic diagram of an ANM self-supervised network.
[0111] Figure 3 It is a complete network architecture diagram.
[0112] Figure 4 It is the root mean square error simulation result of different methods under different signal-to-noise ratios.
[0113] Figure 5 It is the root mean square error simulation result of different methods under different numbers of snapshots. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0114] The embodiment of the present invention provides a sparse array direction of arrival estimation method based on a local and global self-supervised network, including the following steps:
[0115] Step 1, perform signal processing;
[0116] Consider an arbitrary sparse array S with K h holes and an array aperture of M. K narrowband uncorrelated far-field sources from different directions are θ = [θ1, θ2,..., θ K . The signal received by the array S is expressed as:
[0117] x S (t) = As(t) + n(t) (1)
[0118] where s(t) is the received signal, n(t) is independent and identically distributed zero-mean additive Gaussian noise, and A represents the array manifold, specifically expressed as:
[0119] A = [a(θ1), a(θ2),..., a(θ K )] (2)
[0120]
[0121] where d = λ / 2, and p(M - h k ) represents the position of the (M - h k )th array element.
[0122] The covariance of the signal is expressed as:
[0123]
[0124] where p k represents the signal source power, and p n represents the noise power.
[0125] Actually, the signal covariance matrix is calculated from its finite sampling:
[0126]
[0127] where T is the number of sampling snapshots.
[0128] Vectorize the covariance matrix to obtain the measurement data corresponding to the difference set virtual array For the sparse array S, the corresponding set of difference set virtual array positions is obtained by calculation, but this includes duplicate difference set virtual array elements. Delete the duplicate array element positions in the set of difference set virtual array element positions to obtain the final set of difference set virtual array positions, denoted as S diffAccording to the above rules, the difference set virtual array measurement data corresponding to the repeated array element positions is removed to obtain the difference set virtual array S diff The corresponding virtual array signal is expressed as:
[0129]
[0130] where b(θ k ) represents the steering vector corresponding to the difference set virtual array.
[0131] The difference set virtual array signal Interpolate "0" at the places where there are holes in the virtual array element positions. The input data set of the network consists of the real part and the imaginary part of the interpolated measurement data, which is expressed as and The training samples can be expressed as Collect D training samples with different signal directions, different numbers and positions of missing array elements. The input data set is expressed as Y = {Y (1) , …, Y (D)}.
[0132] Since this network takes the atomic norm minimization (ANM) problem as the self-supervised loss, there is no need to generate labeled data. The original training data is needed to generate the data in the loss function, that is and G. Then, the self-supervised data corresponding to the d-th training sample is expressed as Collect D samples under different signal directions and different sparse array configurations. The input self-supervised data set is expressed as P = {P (1) , …, P (D)}.
[0133] Therefore, the training data consists of the data pairs of (Y (d) , P (d) ). The training data set is specifically expressed as D = {(Y (1) , P (1) ), …, (Y (D) , P (D) )}.
[0134] Step 2, establish a local and global feature extraction network;
[0135] There are specific amplitude and phase relationships between the signals received by different array elements. Utilizing the correlation between these array elements helps to recover the signals at the missing positions. A parallel local-global feature extraction network is adopted to jointly model the local and global spatial features to extract the spatial relationship of the received data in order to recover the missing measurement data. The local-global feature extraction network is as shown in Figure 1 Figure
[0136] (1) Local Feature Modeling Layer
[0137] The data first passes through the feature expansion layer, and the obtained data First, the data features of adjacent 3 array elements are grouped into a group through a sliding window, and the grouped data is represented as where L = 2M - 1 is the total number of virtual arrays after inserting 0 at the virtual array holes. Local feature aggregation is performed through convolution, and the convolution kernel is 3*1, which is expressed as:
[0138]
[0139] Then, the data is passed through multiple layers of GLU to establish a more accurate local feature model. For the z-th layer of GLU, the gating threshold is obtained from the linear connection layer and the Sigmoid activation layer:
[0140] δ z = Sigmoid(X s W + b) (8)
[0141] The output of the z-th layer of GLU is expressed as:
[0142] X lz = δ z ·(X i W + b) (9)
[0143] Set the z layer of GLU, and the output data of the z layer is expressed as X l1 ,..., X lZ , and they are stacked to obtain the output of the local feature modeling layer:
[0144]
[0145] (2) Global Feature Modeling Layer
[0146] Parallel processing of X is achieved through the Transformer network d and global feature modeling to capture the long-range dependencies of the virtual array measurement data features. The Transformer network consists of positional encoding, multi-head attention, and a feed-forward layer.
[0147] The positional encoding models the relative or absolute positions of the features of the array elements. The feature encoding of the l-th array element is expressed as follows:
[0148]
[0149] where i refers to the index of the input feature, and sine encoding is used for even indices and cosine encoding is used for odd indices.
[0150] The multi - head attention layer models the relationship of signal features at different array element positions. For the h - th head, the attention mechanism is specifically expressed as:
[0151] head h =Attention(Q h ,K h ,V h ) (12)
[0152]
[0153] Among them, Q h 、K h 、V h are obtained by linearly projecting the data X p through different linear transformation layers, and d h is the dimension of Q h and K h .
[0154] Connect the attention results of H heads to obtain the multi - head attention features:
[0155] MultiHead(Q,K,V)=Concat(head1,…,head H )W O (14)
[0156] Among them, is the multi - head connection matrix.
[0157] The feed - forward layer contains two linear layers and a Relu activation function, and is expressed as:
[0158] X f =Relu((X a W1)W2) (15)
[0159] After residual and layer normalization processing, the output of the Transformer network is denoted as X g , and then, the learned local features and global features are connected to obtain the output of the local - global feature extraction network, denoted as X gl .
[0160] Step 3, ANM self - supervised network
[0161] The ANM self-supervised network uses a complex Toeplitz reconstruction layer, a complex singular value decomposition (SVD) layer, a complex positive semidefinite constraint (PSD) layer, and an ANM loss function to achieve meshless data recovery, ensuring the complex structure of the data in the network while implementing self-supervised learning. The ANM self-supervised network is as Figure 2 shown.
[0162] (1) Complex Toeplitz matrix reconstruction layer
[0163] For the output X gl ∈R B×2×M of the local-global feature extraction network, take its first row as the real part data and the second row data as the imaginary part data. According to the taken data, the real part and the imaginary part of the complex Toeplitz matrix are constructed as:
[0164]
[0165] (2) Complex SVD decomposition
[0166] Before applying the PSD constraint, the Toeplitz matrix needs to be decomposed by SVD. Real-valued arithmetic operations are used in the complex SVD layer to approximate complex operations.
[0167] First, construct a real-valued combined matrix:
[0168]
[0169] Then, decompose the matrix C by SVD:
[0170] [U C ,Σ C ,V C =SVD(C) (19)
[0171] where U C and V C are eigenvectors, and Σ C are eigenvalues.
[0172] The real part and the imaginary part of the eigenvector of the complex matrix are respectively:
[0173]
[0174]
[0175]
[0176] Since X Tis a Hermitian matrix, and its eigenvalues are real numbers, then the eigenvalues of X T are:
[0177] Σ = Σ C (1:2:2M, 1:2:2M) (24)
[0178] (3) Complex PSD constraint layer
[0179] To ensure that the constructed Toeplitz matrix satisfies the PSD constraint, its eigenvalues must be non - negative. Therefore, a complex PSD constraint layer is used to apply a threshold to the eigenvalues to ensure their non - negativity.
[0180] For the (i, i) - th eigenvalue, define the eigenvalue threshold activation function:
[0181]
[0182] where ∈ is a small positive threshold that can prevent the eigenvalues from being non - negative.
[0183] The complex - number algorithm operation of the PSD constraint layer is further expressed as:
[0184]
[0185] The outputs of the complex PSD constraint layer are the real part and the imaginary part of the above formula respectively.
[0186] (4) Interpolation ANM loss function
[0187] Interpolate the measurement data along the positive axis of the difference - set virtual array After interpolation, the noise - free signal y of the positive - axis virtual array is expressed as:
[0188]
[0189] Based on the atomic norm theory, the atomic set of the virtual - array signal is defined as:
[0190]
[0191] Then the atomic - norm minimization problem of y is constructed as:
[0192]
[0193] The above atomic - norm minimization problem is equivalent to a semidefinite programming form:
[0194]
[0195] where represents the PSD constraint:
[0196]
[0197] The vector after sampling covariance vectorization and interpolation in formula (5) Constructed, and the specific construction method is as follows:
[0198]
[0199] G is a binary matrix, indicating that the data at the corresponding array interpolation positions is 0, and the data at the array non-interpolation positions is 1.
[0200] In formula (30), for the PSD constraint is implemented through a complex SVD layer and a PSD constraint layer. The convex optimization problem equivalent to atomic norm minimization can be expressed as the loss function of a self-supervised network, specifically expressed as:
[0201]
[0202] Step 4, result evaluation;
[0203] The goal of this scheme is to improve the signal recovery performance of any sparse array under different numbers of holes and hole positions. The amplitude correlation relationship between sparse array elements can be used to establish a more reasonable model to recover the measured data. At the same time, the meshless sparse optimization method is embedded into the network to realize meshless data recovery without labels. The complete network architecture is as Figure 3 shown.
[0204] In the simulation, the aperture of the sparse array is set to 8, the number of holes is randomly set to 1, 2, 3, 4, and the hole positions are also random, obtaining any sparse array S. The simulation experiment considers the root mean square error of any sparse array under different numbers of snapshots and signal-to-noise ratios, defined as
[0205]
[0206] where C represents the Monte Carlo simulation experiment, represents the estimated value.
[0207] By comparing the RMSE performance of the proposed scheme with the traditional ANM and interpolation ANM methods under any sparse array, the performance advantages of the proposed method are demonstrated. Considering a case where half of the array elements are missing, that is, when there are four array element holes, the sparse array position set is [0, 5, 6, 7], and the difference set virtual array position set is [-7, -6, -5, -2, -1, 0, 1, 2, 5, 6, 7], and its virtual array has four holes. The directions of the source signals are from 4.2° and 20.9° respectively. The RMSE performance simulation results are as Figure 4 , Figure 5As shown. As the number of holes in the physical array elements increases, the corresponding number of holes in the virtual array also increases. When the number of holes is large, the signal recovery performance advantage of the proposed method is more obvious, and the RMSE result of the proposed method is better than that of the traditional non-network ANM method.
[0208] The present invention provides a sparse array direction-of-arrival estimation method based on local and global self-supervised networks. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.
Claims
1. A method for estimating the direction of arrival of a sparse array based on a local and global self-supervised network, characterized in that, Including the following steps: Step 1, perform signal processing: First, construct the received signals of an arbitrary sparse array, and perform difference set virtual array expansion to obtain the measurement data of the difference set virtual array; Subsequently, construct a neural network training dataset according to the measurement data; Step 2, establish a local and global feature extraction network: Construct a local and global feature extraction network through more than two layers of gated linear units (GLUs) and Transformer networks, input the measurement data of the difference set virtual array, and establish a spatial feature relationship through the local and global feature extraction network and obtain the measurement data missing at the positions of the virtual array holes; Step 3, establish an atomic norm minimization (ANM) self-supervised network; The atomic norm minimization (ANM) self-supervised network converts complex number operations into real-valued matrix operations, and through a complex Toeplitz matrix construction layer, a complex singular value decomposition (SVD) layer, a complex positive semi-definite constraint (PSD) layer, and an atomic norm minimization (ANM) loss function, realizes gridless signal recovery while ensuring the complex structure in the network; Step 4, gridless signal recovery and direction of arrival (DOA) estimation: Input the neural network training dataset into the local and global feature extraction network, utilize the amplitude correlation relationship between array elements to establish a model to recover the missing measurement data; At the same time, embed the atomic norm minimization (ANM) method into the network to realize label-free self-supervised data recovery, and finally perform direction of arrival estimation using the recovered gridless data.
2. The method according to claim 1, wherein Step 1 includes: setting an arbitrary sparse array S, the sparse array S having K h holes, the array aperture being M; K narrowband uncorrelated far-field sources from different directions being θ = [θ1, θ2, …, θ K , where θ K represents the Kth signal source, and the signal x S (t) received by the sparse array S at time t is expressed as: x S (t) = As(t) + n(t) (1), Wherein, s(t) is the signal received at time t, n(t) is independent and identically distributed zero-mean additive Gaussian noise, and A represents the array manifold.
3. The method according to claim 2, wherein In Step 1, the array manifold A is expressed as: A = [a(θ1), a(θ2), … a(θ K )] (2), Among them, the steering vector a(θ k ) of the k-th signal source is specifically expressed as: Among them, the intermediate parameter d = λ / 2, where λ is the wavelength, p1 represents the position of the first array element, and p(M - K h ) represents the position of the (M - K h )-th array element; j represents the imaginary unit, and e represents the natural constant; Covariance R of the signal S is expressed as: Among them, E represents taking the expectation of a matrix, H represents the conjugate transpose of a matrix, I represents the identity matrix, p k represents the power of the k-th signal source, p n represents the noise power.
4. The method according to claim 3, wherein Step 1 further includes: calculating a signal covariance matrix from finite sampling where T s is the number of sampling snapshots.
5. The method according to claim 4, characterized in that Step 1 further includes: vectorizing the covariance matrix to obtain the measurement data corresponding to the difference set virtual array Calculating to obtain the set of position sequence numbers of the difference set virtual array corresponding to the sparse array S where m and n respectively represent the position sequence number of the m-th array element and the position sequence number of the n-th array element in the sparse array S, and duplicate array element positions are deleted from the difference set virtual array position set to obtain the final difference set virtual array position set S diff ; Remove the difference set virtual array measurement data corresponding to the repeated array element positions to obtain the difference set virtual array position set S diff The corresponding virtual array signal where, b(θ k ) represents the steering vector corresponding to the difference set virtual array; i represents the result after removing duplicate elements from the vectorized unit matrix I; Virtual array signal Interpolate 0 at the positions of the holes in the virtual array element positions, and the input data set of the local and global feature extraction networks consists of the real part of the interpolated measurement data and the imaginary part ; Denote the set of real numbers; Collect D training samples with different signal directions, different numbers and positions of missing array elements. The d-th training sample Input data set Y = {Y (1) , …, Y (D)}; d ranges from 1 to D.
6. The method according to claim 5, wherein Step 1 further includes: using the original training data to generate data in the loss function, i.e., and G, where is the vector obtained by vectorizing and interpolating the sampling covariance in formula (5), is the Toeplitz matrix constructed, and G represents a binary matrix; then, the self-supervised data representation corresponding to the d-th training sample is expressed as Samples are collected under different signal directions and different sparse array configurations, and the input self-supervised data set is expressed as P = {P (1) , …, P (D)}; The training data consists of data pairs (Y (d) , P (d) ). The training data set is specifically represented as Data = {(Y (1) , P (1) ), …, (Y (D) , P (D) )}.
7. The method according to claim 6, characterized in that, Step 2 includes: The local and global feature extraction network includes a local feature modeling layer and a global feature modeling layer; The local feature modeling layer performs the following operations: The local feature modeling layer includes a fully connected feature expansion layer, and the fully connected feature expansion layer includes a complex linear connection layer, a connection layer, a convolutional layer, and an activation layer. The training data set first passes through the fully connected feature expansion layer to obtain data First, the data features of adjacent 3 array elements are grouped by a sliding window, and the grouped data is represented as where L = 2M - 1 is the total number of the virtual array after inserting 0 at the virtual array holes; B represents the training batch size; d f represents the expanded feature dimension; R represents the real number field; Perform local feature aggregation through convolution, with convolution kernel k e being 3*1, and the output X of the convolutional layer s is expressed as: Among them, d g represents the feature dimension after passing through the convolutional layer; Then, the output data of the convolutional layer is passed through more than two layers of gated linear units (GLUs). For the z-th layer of GLU, the gating threshold δ z is obtained from the linear connection layer and the sigmoid activation layer: δ z = Sigmoid(X s W + b) (8), Wherein, W is the weight of the linear connection layer, and b is the bias of the linear connection layer; The output X of the z-th layer gated linear unit (GLU) lz is expressed as: X lz = δ z ·(X s W + b)(9), Set the gated linear unit GLU to have Z layers, and obtain the output data of Z layers, denoted as X l1 ,..., X lZ , and stack X l1 ,..., X lZ to obtain the output X of the local feature modeling layer l : Among them means directly stacking and paralleling vectors; The global feature modeling layer performs the following operations: establishing a Transformer network to achieve parallel processing of X through the Transformer network d and global feature modeling to capture the long-range dependencies of the virtual array measurement data features. The Transformer network includes position encoding, a multi-head attention layer, and a feed-forward layer; The positional encoding models the relative or absolute positions of the features of the array elements, and the feature X at the l-th array element d,l The feature encoding X p,l is expressed as: Wherein, i is the index of the input feature, sine coding is used for even indices, and cosine coding is used for odd indices; Model the signal feature relationship of different array element positions through the multi-head attention layer. For the h-th head attention mechanism head h Specifically expressed as: head h = Attention(Q h , K h , V h ) (12), Wherein, the attention mechanism Attention is expressed as: Among them, Q h , K h , V h respectively represent the query, key, and value of the h-th attention mechanism, and are obtained by linearly projecting the data X p through different linear transformation layers. d h is the dimension of Q h and K h ; Connect the attention results of a H to obtain the multi-head attention MultiHead feature X a : Among them, head1 represents the attention result of the first head, represents the attention result of the a Hth head, is the multi-head connection matrix, d v represents the dimension of V, d model represents the dimension expected to be obtained after multi-head concatenation; Concat represents the concatenation operation; Q, K, and V respectively represent the query, key, and value of the multi-head attention after concatenating multiple attention results; The feed-forward layer includes two linear layers and a rectified linear unit (Relu) activation function, and the output X of the feed-forward layer f is expressed as: X f = Relu((X a W1)W2) (15), Among them, W1 and W2 represent the network weights of two linear layers; X a represents the MultiHead feature of the multi-head attention; After residual sum and layer normalization processing, the output of the Transformer network is denoted as X g , then, the learned local features are concatenated with the global features to obtain the output X of the local and global feature extraction network gl .
8. The method according to claim 7, wherein Step 3 includes: The ANM self-supervised network includes a complex Toeplitz reconstruction layer, a complex singular value decomposition (SVD) layer, a complex positive semi-definite constraint (PSD) layer, and an ANM loss function, which is used to realize gridless data recovery; The complex Toeplitz reconstruction layer performs the following operations: For the output of the local and global feature extraction networks Take X gl The first row of X gl,1 As the real part data, the second row of X gl,2 Data as the imaginary part data, according to the selected data, the complex Toeplitz matrix X T The real part of T,real And the imaginary part of X T,imag Are constructed as: Among them, <X gl,1 > M and <X gl,2 > M respectively represent the M-th value of the first row of data and the M-th value of the second row of data; gl The plural singular value decomposition (SVD) layer performs the following operations: perform SVD decomposition on the Toeplitz matrix. First, use formula (16) to obtain the real part X T of the complex Toeplitz matrix X T,real and the imaginary part X T,imag to construct a real-valued combined matrix C: Then, perform singular value decomposition (SVD) on the matrix C: [U C , Σ C , V C = SVD(C) (19), Among them, U C and V C are eigenvectors, and Σ C is an eigenvalue; The real parts U real and V real of the eigenvectors of the complex matrix, and the imaginary parts U imag and V imag are respectively: U real = U C (1:M, 1:2:2M) (20) U imag = U C (M + 1:2M, 1:2:2M) (21), V real = V C (1:M, 1:2:2M) (22), V imag = V C (M + 1:2M, 1:2:2M) (23), X T The eigenvalue Σ of Σ = Σ C (1:2:2M, 1:2:2M) (24), The complex positive semi - definite constrained positive semi - definite (PSD) layer performs the following operations: For the (i, i) - th eigenvalue, define the eigenvalue threshold activation function f max (Σ): where ∈ is a positive threshold; Σ (i,i) represents the value of the i-th row and i-th column in the eigenvalue Σ; Perform complex semi - positive definite constraint PSD algorithm operations to obtain intermediate parameter X p : The outputs of the complex positive semi - definite constrained PSD layers are the real and imaginary parts of X respectively p ; Virtual array signals along the positive axis of the difference set virtual array Interpolation is performed, and the noise-free signal y of the positive-axis virtual array after interpolation is expressed as: Among them, z(θ k ) represents the steering vector corresponding to the positive axis of the virtual array; Based on the atomic norm theory, the atom set of the virtual array signal is defined as: Atomic norm minimization problem of y Constructed as: Among them, represents taking the infimum; The atomic norm minimization problem is equivalent to a semi-definite programming form: where min T(y) represents minimization, s.t. represents subject to, represents the PSD constraint: Tr represents the trace; represents the calculation of and the Frobenius norm of the matrix after subtraction; Toeplitz matrix constructed from vector y Specifically expressed as: Matrix It is represented as: Among them, represents taking the L-th value in the item vector, and * represents taking the conjugate of the matrix; In formula (30), for the PSD constraint is implemented through a complex SVD layer and a PSD constraint layer, and the convex optimization problem equivalent to atomic norm minimization is expressed as the loss function of the self-supervised network Specifically expressed as:
9. An electronic device, characterized in that, Including a processor and a memory, the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that, Stored with a computer program or instruction, when the computer program or instruction runs on a computer, it executes the steps of the method according to any one of claims 1 to 8.
Citation Information
Cited By
Two-dimensional underwater direction of arrival estimation method based on equilateral triangle array
CN121208745A
A two-dimensional underwater direction of arrival estimation method based on equilateral triangle array
CN121208745B
DOA estimation method and system based on global dynamic convolutional neural network
CN121457531A