DOA estimation method based on triple uniform linear array and multi-scale attention network

By constructing a differential co-array using a triple uniform linear array and a multi-scale attention network, the problems of virtual array continuity and degree of freedom in sparse array design and signal processing are solved, achieving high-precision DOA estimation in complex scenarios.

CN121805938APending Publication Date: 2026-04-07JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-28
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing sparse array designs have room for optimization in terms of virtual array continuity, maximization of degrees of freedom, and structural versatility. Traditional high-resolution algorithms suffer severe performance degradation in low signal-to-noise ratio and low snapshot scenarios, while deep learning methods fail to fully exploit the array's geometric advantages and signal spatial characteristics.

Method used

A differential co-array is constructed using a triple uniform linear array to generate a hole-free virtual uniform linear array. This is combined with a multi-scale attention network for signal processing. The accuracy of DOA estimation is improved through multi-scale feature extraction, feature fusion, and attention mechanisms.

Benefits of technology

It achieves high-precision and robust DOA estimation in scenarios with low signal-to-noise ratio and few snapshots, improves angle resolution and estimation accuracy, adapts to different array configurations and has good applicability and flexibility, and has high training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121805938A_ABST
    Figure CN121805938A_ABST
Patent Text Reader

Abstract

The invention relates to a DOA (Direction of Arrival) estimation method based on a triple uniform linear array and a multi-scale attention network, and belongs to a DOA estimation method in the field of array signal processing. Comprising the steps of constructing a triple uniform linear array, performing original data set and preprocessing, constructing and training a multi-scale attention network model, testing the multi-scale attention network model on a test set and performing actual data estimation. The method has the advantages that the regularized and high-degree-of-freedom sparse array design is adopted, the deep coupling multi-scale attention mechanism neural network is combined, the estimation performance of the signal direction is jointly improved, and through verification of analog data and actual data, under the complex conditions of a low signal-to-noise ratio, few snapshots and the like, the method has the advantage that the estimation performance of the signal direction is improved. And high-precision and high-robustness DOA estimation can still be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to direction-of-arrival (DOA) estimation methods in the field of array signal processing. It integrates sparse array hardware design with deep learning signal processing technology, and in particular, it relates to a DOA estimation method based on a triple uniform linear array and a multi-scale attention network. Background Technology

[0002] Direction of arrival (DOA) estimation is a core task in array signal processing, and its performance directly depends on the array's aperture and structure. Traditional uniform linear arrays are simple in structure, but their degrees of freedom are limited by the number of physical array elements, and improving resolution requires high hardware costs. To overcome this limitation, sparse arrays such as nested arrays and coprime arrays have been proposed. These arrays optimize element positions and utilize the concept of difference comatrix to obtain more virtual array elements than physical array elements, thereby increasing the degrees of freedom. However, existing sparse array designs still have room for optimization in terms of virtual array continuity, maximizing degrees of freedom, and structural versatility. For example, some arrays have "holes" in their difference comatrix, resulting in discontinuous virtual arrays and limiting their performance upper limit; other arrays have complex construction rules, making it difficult to flexibly adapt to scenarios with different numbers of array elements. At the signal processing algorithm level, traditional high-resolution algorithms based on subspace decomposition (such as MUSIC and ESPRIT) rely on ideal signal statistical characteristics, and their performance degrades severely in scenarios with low signal-to-noise ratio, few snapshots, or coherent signals. Although techniques such as spatial smoothing can resolve coherence, they further reduce the effective array aperture. In recent years, deep learning-based DOA estimation methods have shown potential, but their performance is limited by two key factors: first, the quality of the input features. If the original array data or a simple covariance matrix is ​​used directly, the geometric advantages of the array cannot be fully exploited; second, the design of the network structure. Most methods use general convolutional networks, which are not capable of modeling the multi-scale and channel correlation of signal spatial features.

[0003] Therefore, existing technologies lack a collaborative optimization scheme from the array physical layer to the signal processing algorithm layer. There is an urgent need for a novel sparse array structure that can maximize the generation of continuous virtual apertures under a given physical array element. At the same time, a matching intelligent signal processing framework is needed that can efficiently utilize the extended virtual aperture features to achieve ultra-high accuracy DOA estimation in complex environments. Summary of the Invention

[0004] This invention provides a DOA estimation method based on a triple uniform linear array and a multi-scale attention network. The aim is to achieve high-precision and robust DOA estimation in complex scenarios with low signal-to-noise ratio and few snapshots by using a regularized, high-degree-of-freedom sparse array design and a deeply coupled dedicated neural network processing flow.

[0005] The technical solution adopted by this invention includes the following steps:

[0006] Step 1: Construct a triple uniform linear array; Based on the set total number of physical array elements and according to the preset regularized mathematical relationship, determine the array element arrangement of the triple uniform linear array composed of three uniform linear subarrays.

[0007] Step 2: Construct the original dataset; Based on the mathematical model of the triple uniform linear array established in Step 1, generate the original dataset required for the multi-scale attention network; The dataset includes: samples of the received signal matrix of the triple uniform linear array. And the actual DOA angle corresponding to each array signal sample. ;

[0008] Step 3: Preprocessing the raw dataset; This involves processing the array received signal matrix obtained in Step 2. The process involves sequentially performing covariance matrix calculation, vectorization, redundancy removal based on array structure, spatial smoothing, and separation of real and imaginary parts. This ultimately constructs a dual-channel real-valued feature map characterizing the spatial properties of the extended virtual aperture. Thus, the preprocessed dataset includes the input Output ;Finish After preprocessing the original dataset of each sample, the fully preprocessed original dataset is divided into training set, validation set and test set in a ratio of 7:2:1.

[0009] Step 4: Construction of a multi-scale attention network model; Construct a deep network that includes multi-scale feature extraction, feature fusion, multi-scale attention mechanism, and output layer;

[0010] Step 5: Training the multi-scale attention network model;

[0011] Step 6: Test the multi-scale attention network model on the test set;

[0012] Step 7: Estimation of actual data.

[0013] The first step of this invention, which involves constructing a triple uniform linear array, is as follows:

[0014] Consider adopting A triple uniform linear array with individual elements is used for signal reception, wherein... , , These represent the number of element intervals in each subarray. , , These represent the element spacing within each subarray, and are variables. , , representing the in a triple uniform linear array The position of each array element, when The element spacing and element spacing of each subarray can be determined by the following formula. To simplify the formula, the element spacing of the first subarray is set to 1. The specific formula is as follows:

[0015] when When the number is even, the number of element intervals within each subarray , and the element spacing of each subarray , ,as follows:

[0016] (1)

[0017] when When the number is odd, the number of element intervals within each subarray , and the element spacing of each subarray , ,as follows:

[0018] (2)

[0019] The position of each element in the three uniform linear array is determined by formulas (1) and (2). The difference between the positions of the array elements is obtained by calculating the differential covariance matrix (DCA), denoted as... This generates a virtual array, which is a continuous uniform linear array without holes. The number of useful virtual elements in the virtual array, i.e., the total number of effective virtual elements in the differential array, is used... The formula is as follows:

[0020] (3)

[0021] Based on the actual array element configuration, an additional scaling factor is introduced. The scaling factor This represents the element spacing of the virtual uniform linear array after differential co-array processing, and the actual element positions of the triple uniform linear array. It can be expressed by the following formula:

[0022] (4)

[0023] In this way, the spacing between the expanded virtual array elements can be adjusted according to different needs while maintaining the proportional relationship of the element spacing. Thus, the triple uniform linear array model is completed.

[0024] The process of constructing the original dataset in step two of this invention is as follows:

[0025] Consider the DOA problem in a real-world system: A far-field narrowband signal was The formula for the received signal in a triple uniform linear array is as follows:

[0026] (5)

[0027] in, express The received signal vector at time t, This represents the total number of snapshots.

[0028] Direction matrix It contains the steering vector corresponding to each direction of arrival. ,in, , indicating the first The incident angle of each signal source carrier frequency The corresponding wavelength, Wave speed;

[0029] Let be the signal source vector, where , indicating the first A signal source in The signal input at a given time; This represents additive white Gaussian noise;

[0030] all The signal captured by each snapshot is represented by the following formula, where... For the first The received signal vector of a snapshot, :

[0031] (6)

[0032] The above steps construct a sample dataset. To improve the model's generalization ability, random information sources are used to input angles. Random input signal source Random signal-to-noise ratio, where randomness is represented by a uniform distribution, is constructed from this. A dataset of samples, each sample is processed by formula (6) to obtain all... The signal matrix collected by a quick snapshot The output is the set of directions of arrival for all corresponding signal sources. This establishes the array received signal matrix for each sample. With the direction of the port A one-to-one correspondence between them.

[0033] Step three of the present invention, raw data preprocessing, specifically includes:

[0034] (1) The covariance matrix of the observed data is expressed as:

[0035] (7)

[0036] in, , Indicates the first The power of each signal source, Indicates noise power. for 3D identity matrix;

[0037] (2) In order to expand the triple uniform linear array and obtain a virtual array, the covariance matrix is... Vectorization, as shown below:

[0038] (8)

[0039] in, This represents matrix vectorization operations, which stack the matrices column-wise into a column vector.

[0040] , This represents the Kronecker product operation. express The complex conjugate matrix;

[0041] , ;

[0042] (3) Remove redundancy, vector Each element corresponds to a virtual sensor. Since these positional differences are redundant, redundant elements need to be removed and rearranged to obtain... This vector is used to model the virtual uniform linear array signal reception vector based on a single snapshot;

[0043] (4) Covariance matrix of virtual uniform linear array signal receiving vector based on a single snapshot As shown below:

[0044] (9)

[0045] (5) For the generated Spatial smoothing is performed to reduce the correlation between virtual array elements and to improve the quality of the new matrix by averaging the covariance matrices of different subarrays or snapshots. The rank of the space smoothing process can be expressed by the following formula:

[0046] (10)

[0047] in, Indicates the first The covariance matrix of each subarray The triple uniform linear array of array elements can be expanded into A uniform linear array of elements, the element spacing of which is denoted as . ;

[0048] (6) Normalize the matrix using the following formula:

[0049] (11)

[0050] in Representation matrix The Euclidean norm;

[0051] (7) Transform the matrix The real and imaginary parts are separated into two channels, denoted as . ;

[0052] R c A dual-channel real-valued feature map characterizing the spatial properties of the received signal after spread. , For the number of channels, real part, and imaginary part, and For spatial dimensions, Finally, the complete The original dataset of each sample was preprocessed and divided into training set, validation set and test set in a ratio of 7:2:1.

[0053] The construction of the multi-scale attention network in step four of this invention specifically includes:

[0054] (1) Multi-scale feature extraction: Multi-scale convolution operation through Implemented with parallel convolutional branches, each branch Output The calculation is as follows, where :

[0055] (12)

[0056] in, Indicates the first Two-dimensional convolution operation with branches, For the convolution kernel weight tensor of this branch, Number of output channels The kernel size is the convolution kernel size. For the corresponding bias vector, This indicates the batch normalization operation performed on the output of this branch. The nonlinear activation function is represented by the SiLU function, which is defined as follows: ;

[0057] all The outputs of each branch are concatenated along the channel dimension to form a fused multi-scale feature map. :

[0058] (13)

[0059] Total number of output channels ;

[0060] (2) Feature fusion: After concatenating all branch outputs along the channel dimension, the feature fusion is performed... Convolutional feature fusion yields:

[0061] (14)

[0062] in This represents the total number of output channels for the multi-scale feature extraction block. This indicates that the number of input and output channels is both of convolution;

[0063] (3) Multi-scale attention mechanism: The output of the multi-scale attention block is obtained through grouping, weight calculation, intra-scale feature enhancement, weighting, cross-scale information exchange and residual connection. ;

[0064] (4) Output layer: The output of the multi-scale attention block Mapping to DOA estimates, first through... Number of channels in convolution compression:

[0065] (25)

[0066] Will Flattened into a one-dimensional vector, the formula is as follows:

[0067] (26)

[0068] The above process yields a highly compressed one-dimensional vector, which is then passed through a linear regression layer to obtain the DOA estimate, as shown in the following formula:

[0069] (27)

[0070] in, The set of directions of arrival (DOAs) of the signal sources estimated by the model; The bias vector of the linear layer. This is the weight matrix of the linear layer.

[0071] The multi-scale attention mechanism described in (3) of this invention specifically includes:

[0072] 3.1) Grouping: Evenly divided in the channel dimension Group, corresponding Obtained at different scales The formula is as follows:

[0073] (15)

[0074] Each group ;

[0075] 3.2) Weight Calculation

[0076] For input features Perform global average pooling to obtain channel statistics:

[0077] (16)

[0078] Pass through the first layer Convolution reduces the channel dimension, as shown in the following formula:

[0079] (17)

[0080] Through the second layer Convolution generates scale attention weights:

[0081] (18)

[0082] in express Attention weights at each scale, where This represents the attention weight for the g-th group; This represents the Softmax function, another commonly used non-linear activation function, which is suitable for the output layer representation of multi-classification tasks in neural networks, thereby obtaining the attention weights for each scale.

[0083] 3.3) Intra-scale feature enhancement: For each scale group Perform independent feature enhancement processing (depthiable separable convolution):

[0084] (19)

[0085] in This represents a depthwise convolution with a kernel size of . The number of groups is , This represents point convolution with a kernel size of . ;

[0086] 3.4) Weighting: Weighting attention points Expand to the same spatial dimension as the feature map and weight the enhanced features:

[0087] (20)

[0088] Concatenate all weighted scale group features along the channel dimension:

[0089] (twenty one)

[0090] 3.5) Cross-scale information exchange: This involves the exchange of spliced ​​features. To conduct cross-scale information exchange, the first step is through a layer Convolution dimensionality reduction:

[0091] (twenty two)

[0092] Then through another layer Number of convolution recovery channels:

[0093] (twenty three)

[0094] 3.6) Residual Connectivity: To prevent gradient vanishing or exploding, residual connectivity is used here, as shown in the following formula:

[0095] (twenty four).

[0096] The training of the multi-scale attention network in step five of this invention includes:

[0097] The training phase configuration is as follows: the loss function is the mean squared error loss (MSELoss), used to measure the deviation between the model's predicted values ​​and the true values; the optimizer is the Adam optimizer; the initial learning rate is set to 0.001, and the learning rate scheduler uses the ReduceLROnPlateau strategy, with the validation loss as the monitoring metric for dynamic parameter tuning. When the validation loss stagnates for four consecutive rounds, the learning rate is halved; simultaneously, torch.amp.GradScaler is enabled to achieve mixed-precision training, supporting FP16 precision calculation, effectively improving training speed and reducing GPU memory usage while ensuring training accuracy. The model that performs best on the validation set is saved.

[0098] In step six of this invention, the root mean square error (RMSE) is used as a model performance evaluation index during the testing phase to quantify the accuracy of the model's prediction results. The formula is as follows:

[0099] (28)

[0100] in Indicates the number of samples in the test set. The number of signal sources; Indicates the first Model prediction values ​​for the angle of each test sample; Indicates the first The true value of the angle of each test sample.

[0101] Step seven of this invention includes building a corresponding physical array hardware platform, collecting array received signals in the actual environment, performing data preprocessing, and inputting the data into a trained network model to obtain the direction of arrival (DOA) estimation result of the actual signal.

[0102] The present invention has the following advantages:

[0103] 1. Virtual Aperture Expansion: By constructing a difference array on the triple uniform linear array and performing deduplication, a virtual uniform linear array with a larger aperture and no holes is generated, which effectively improves the angular resolution and estimation accuracy without increasing the number of physical array elements.

[0104] 2. Powerful feature extraction capability: The designed multi-scale attention network can fully extract multi-scale features from the spatial covariance matrix of the signal, and adaptively focus on key information by means of scale attention and channel attention mechanisms, which significantly enhances the feature representation capability and robustness of the network.

[0105] 3. Excellent robustness to low signal-to-noise ratio and few snapshots: Combining the superior statistical properties brought by virtual aperture expansion with the strong nonlinear fitting ability of deep learning, this method significantly outperforms traditional high-resolution algorithms and existing simple deep learning models in low signal-to-noise ratio and few snapshot scenarios.

[0106] 4. Good applicability and flexibility: By providing a general mathematical expression for triple uniform linear arrays, this method can be flexibly adapted to different array configurations. In practical applications, only the dataset needs to be adjusted, without modifying the multi-scale attention network and preprocessing procedures.

[0107] 5. Highly efficient training and inference performance: The model has only about 0.03M parameters and a lightweight architecture. Using a parallel training and mixed-precision training strategy, training on a dataset of 10,000 samples takes only about 4 seconds per epoch based on two NVIDIA GeForce RTX 5060 GPUs. Single-sample inference time is approximately 6 milliseconds, providing a feasible foundation for real-time DOA estimation.

[0108] The invention has a wide range of applications, including but not limited to massive MIMO user positioning and beam management in fifth-generation and future mobile communications (5G / 6G), IoT terminal positioning, and microphone array voice enhancement. Attached Figure Description

[0109] Figure 1 This is a schematic diagram of the array position of the triple uniform linear array of the present invention;

[0110] Figure 2 This is a schematic diagram of the triple uniform linear array signal reception of the present invention;

[0111] Figure 3 This is a flowchart of the method of the present invention;

[0112] Figure 4 This is a diagram showing the experimental results of the triple uniform linear array in Experiment Example 1 of this invention;

[0113] Figure 5 This is a diagram showing the experimental results of uniform linear array and nested array in Experiment Example 1 of this invention;

[0114] Figure 6 This is a diagram showing the experimental results of Experiment Example 2 of the present invention;

[0115] Figure 7 This is a graph showing the ablation experiment results of Experiment Example 3 of this invention;

[0116] Figure 8 It is the physical verification platform in Example 4 of this invention;

[0117] Figure 9 This is a test result diagram of the physical verification platform in Example 4 of the present invention. Detailed Implementation

[0118] Includes the following steps:

[0119] Step 1: Construct a triple uniform linear array; based on the set total number of physical array elements and according to the preset regularized mathematical relationships, determine the element arrangement of the triple uniform linear array composed of three uniform linear subarrays. Consider using... A triple uniform linear array with 100 elements is used for signal reception, and the array structure is shown in Figure 1. , , These represent the number of element intervals in each subarray. , , These represent the element spacing within each subarray, and are variables. ( ) represents the th element in a triple uniform linear array. The position of each array element. When The element spacing and element spacing of each subarray can be determined using the following formula. To simplify the formula, the element spacing of the first subarray is first set to 1. The specific formula is as follows:

[0120] when When the number is even, the number of element intervals within each subarray ( ) and the element spacing of each subarray ( ) as follows:

[0121] (1)

[0122] when When the number is odd, the number of element intervals within each subarray ( ) and the element spacing of each subarray ( ) as follows:

[0123] (2)

[0124] The position of each element in a three uniform linear array can be determined using formulas (1) and (2). Calculating the differential common matrix (DCA) yields a set of element position differences, denoted as... This generates a virtual array, which is a continuous uniform linear array without holes. The number of useful virtual elements in the virtual array (i.e., the total number of effective virtual elements in the differential array) is used... The formula is as follows:

[0125] (3)

[0126] Based on the actual array element configuration, an additional scaling factor *r* is introduced. This scaling factor *r* represents the element spacing of the virtual uniform linear array after differential co-array processing of the triple uniform linear array. The actual element positions of the triple uniform linear array are... This can be expressed by the following formula:

[0127] (4)

[0128] In this way, the spacing between the expanded virtual array elements can be adjusted according to different needs while maintaining the proportional relationship of the element spacing. Thus, the triple uniform linear array model is completed.

[0129] Step 2: Construct the original dataset; Based on the mathematical model of the triple uniform linear array established in Step 1, generate the original dataset required for the multi-scale attention network; The dataset includes: samples of the received signal matrix of the triple uniform linear array. And the actual DOA angle corresponding to each array signal sample. This step is used to supervise network training and evaluate model performance; it only generates the basic dataset, as the dataset still requires further preprocessing in the next step.

[0130] Consider the DOA problem in a real-world system: A far-field narrowband signal was The triple uniform linear array receiver of the array elements, such as Figure 2 As shown. The formula for receiving the signal is as follows:

[0131] (5)

[0132] in, express The received signal vector at time t, This represents the total number of snapshots.

[0133] Direction matrix It contains the steering vector corresponding to each direction of arrival. ,in, Indicates the first The incident angle of each signal source carrier frequency The corresponding wavelength, Wave speed;

[0134] Let be the signal source vector, where Indicates the first A signal source in The signal input at a given time; This represents additive white Gaussian noise;

[0135] all The signal captured by a snapshot can be expressed by the following formula, where... For the first The received signal vector of a snapshot ( );

[0136] (6)

[0137] The above steps construct a sample dataset. To improve the model's generalization ability, random information sources are used to input angles. Random input signal source Random signal-to-noise ratio, where randomness is represented by a uniform distribution, is constructed from this. A dataset of samples, each sample is processed by formula (6) to obtain all... The signal matrix collected by a quick snapshot The output is the set of directions of arrival for all corresponding signal sources. This establishes the array received signal matrix for each sample. With the direction of the port The one-to-one correspondence between them. To achieve the virtual aperture expansion effect of the triple uniform linear array, the received signal matrix also needs to be... Perform the preprocessing of the original dataset in step three.

[0138] Step 3: Preprocessing the raw dataset; This involves processing the array received signal matrix obtained in Step 2. The process involves sequentially performing covariance matrix calculation, vectorization, redundancy removal based on array structure, spatial smoothing, and separation of real and imaginary parts. This ultimately constructs a dual-channel real-valued feature map characterizing the spatial properties of the extended virtual aperture. Thus, the preprocessed dataset includes the input Output ;Finish After preprocessing the original dataset of each sample, the fully preprocessed original dataset is divided into training, validation, and test sets in a ratio of 7:2:1; specifically including:

[0139] (1) The covariance matrix of the observed data is expressed as:

[0140] (7)

[0141] in, , Indicates the first The power of each signal source, Indicates noise power. for 3D identity matrix;

[0142] (2) In order to expand the triple uniform linear array and obtain a virtual array, the covariance matrix is... Vectorization, as shown below:

[0143] (8)

[0144] in, This represents matrix vectorization operations, which stack the matrices column-wise into a column vector.

[0145] , This represents the Kronecker product operation. express The complex conjugate matrix;

[0146] , ;

[0147] (3) Remove redundancy, vector Each element corresponds to a virtual sensor. Since these positional differences are redundant, redundant elements need to be removed and rearranged to obtain... This vector is used to model the virtual uniform linear array signal reception vector based on a single snapshot;

[0148] (4) Covariance matrix of virtual uniform linear array signal receiving vector based on a single snapshot As shown below:

[0149] (9)

[0150] (5) For the generated Spatial smoothing is performed to reduce the correlation between virtual array elements and to improve the quality of the new matrix by averaging the covariance matrices of different subarrays or snapshots. The rank of the space smoothing process can be expressed by the following formula:

[0151] (10)

[0152] in, Indicates the first The covariance matrix of each subarray The triple uniform linear array of array elements can be expanded into A uniform linear array of elements, the element spacing of which is denoted as . ;

[0153] (6) Normalize the matrix using the following formula:

[0154] (11)

[0155] in Representation matrix The Euclidean norm;

[0156] (7) Transform the matrix The real and imaginary parts are separated into two channels, denoted as . ;

[0157] R c A dual-channel real-valued feature map characterizing the spatial properties of the received signal after spread. , For the number of channels, real part, and imaginary part, and For spatial dimensions, Finally, the complete The original dataset of each sample was preprocessed and divided into training set, validation set and test set in a ratio of 7:2:1;

[0158] Step 4: Construction of a multi-scale attention network; Construct a deep network that includes multi-scale feature extraction, feature fusion, multi-scale attention mechanism, and feature compression and mapping modules;

[0159] (1) Multi-scale feature extraction: Multi-scale convolution operation through Implemented with parallel convolutional branches, each branch Output The calculation is as follows, where :

[0160] (12)

[0161] in, Indicates the first Two-dimensional convolution operation with branches, For the convolution kernel weight tensor of this branch, Number of output channels The kernel size is the convolution kernel size. For the corresponding bias vector, This indicates the batch normalization operation performed on the output of this branch. The nonlinear activation function is represented by the SiLU function, which is defined as follows: ;

[0162] all The outputs of each branch are concatenated along the channel dimension to form a fused multi-scale feature map. :

[0163] (13)

[0164] Total number of output channels ;

[0165] (2) Feature fusion: After concatenating all branch outputs along the channel dimension, the feature fusion is performed... Convolutional feature fusion yields:

[0166] (14)

[0167] in This represents the total number of output channels for the multi-scale feature extraction block. This indicates that the number of input and output channels is both of convolution;

[0168] (3) Multiscale attention mechanism:

[0169] 3.1) Grouping: Evenly divided in the channel dimension Group, corresponding Obtained at different scales The formula is as follows:

[0170] (15)

[0171] Each group ;

[0172] 3.2) Weight Calculation

[0173] For input features Perform global average pooling to obtain channel statistics:

[0174] (16)

[0175] Pass through the first layer Convolution reduces the channel dimension, as shown in the following formula:

[0176] (17)

[0177] Through the second layer Convolution generates scale attention weights:

[0178] (18)

[0179] in express Attention weights at each scale, where This represents the attention weight for the g-th group; This represents the Softmax function, another commonly used non-linear activation function, which is suitable for the output layer representation of multi-classification tasks in neural networks, thereby obtaining the attention weights for each scale.

[0180] 3.3) Intra-scale feature enhancement: For each scale group Perform independent feature enhancement processing (depthiable separable convolution):

[0181] (19)

[0182] in This represents a depthwise convolution with a kernel size of . The number of groups is , This represents point convolution with a kernel size of . ;

[0183] 3.4) Weighting: Weighting attention points Expand to the same spatial dimension as the feature map and weight the enhanced features:

[0184] (20)

[0185] Concatenate all weighted scale group features along the channel dimension:

[0186] (twenty one)

[0187] 3.5) Cross-scale information exchange: This involves the exchange of spliced ​​features. To conduct cross-scale information exchange, the first step is through a layer Convolution dimensionality reduction:

[0188] (twenty two)

[0189] Then through another layer Number of convolution recovery channels:

[0190] (twenty three)

[0191] 3.6) Residual Connectivity: To prevent gradient vanishing or exploding, residual connectivity is used here, as shown in the following formula:

[0192] (twenty four)

[0193] (4) Output layer: The output of the multi-scale attention block Mapping to DOA estimates, first through... Number of channels in convolution compression:

[0194] (25)

[0195] Will Flattened into a one-dimensional vector, the formula is as follows:

[0196] (26)

[0197] The above process yields a highly compressed one-dimensional vector, which is then passed through a linear regression layer to obtain the DOA estimate, as shown in the following formula:

[0198] (27)

[0199] in, The set of directions of arrival (DOAs) of the signal sources estimated by the model; The bias vector of the linear layer. This is the weight matrix of the linear layer.

[0200] Step 5: Training the multi-scale attention network model; the training phase configuration is as follows: the loss function uses mean squared error loss (MSELoss) to measure the deviation between the model's predicted values ​​and the true values; the optimizer is the Adam optimizer; the initial learning rate is set to 0.001, and the learning rate scheduler uses the ReduceLROnPlateau strategy, using the validation loss as the monitoring metric for dynamic parameter tuning. When the validation loss stagnates for four consecutive rounds, the learning rate is halved; simultaneously, torch.amp.GradScaler is enabled to achieve mixed-precision training, supporting FP16 precision calculation, effectively improving training speed and reducing memory usage while ensuring training accuracy. The model that performs best on the validation set is saved.

[0201] Step 6: Test the multi-scale attention network model on the test set;

[0202] During the testing phase, the root mean square error (RMSE) is used as the model performance evaluation metric to quantify the accuracy of the model's predictions. The formula is as follows:

[0203] (28)

[0204] in Indicates the number of samples in the test set. The number of signal sources; Indicates the first Model prediction values ​​for the angle of each test sample; Indicates the first The true value of the angle of each test sample;

[0205] Step 7: Estimation of actual data;

[0206] A corresponding physical array hardware platform is built to collect array received signals in a real-world environment. Since the collected signals are real-valued, they need to be subjected to a Hilbert transform to obtain analytic signals. Then, the data is processed according to the preprocessing procedure that is completely consistent with step three of the model training phase. Subsequently, the processed signals are input into the trained network model to obtain the direction of arrival (DOA) estimation results of the actual signals.

[0207] This invention integrates sparse array hardware design with deep learning signal processing techniques, and particularly relates to a virtual aperture expansion method based on a Triple Uniform Linear Array (TULA), and a multi-scale attention-based deep neural network specifically designed to handle such expanded features. This technical solution aims to fundamentally improve the system's angular resolution, estimation accuracy, and robustness to complex scenarios such as low signal-to-noise ratio and limited snapshots when physical array elements are constrained.

[0208] The effects of the present invention will be further illustrated by the following experimental examples.

[0209] Experimental Example 1:

[0210] To verify the advantage of triple uniform linear array in terms of the number of signal sources, a triple uniform linear array consisting of 6 array elements was built according to the instructions. The received signal was preprocessed using the same method as before training the multi-scale network. Experiments were conducted based on this array and combined with the MUSIC algorithm. The specific experimental parameters are shown in Table 1.

[0211] Table 1 Experimental parameters for Example 1

[0212] Experimental results are as follows Figure 4 As shown, the 6-element triple uniform linear array successfully estimated 12 signal sources, effectively demonstrating its performance in improving the signal source estimation capacity. Each spectral peak of the MUSIC algorithm corresponds to the DOA estimation value of a signal source. The sharper the spectral peak, the better the estimation effect of the algorithm.

[0213] In comparison, it can be seen that neither a uniform linear array with 6 elements nor a nested array with 6 elements can estimate 12 signal sources, failing to achieve a similar number of signal sources. The experimental results are as follows: Figure 5 As shown.

[0214] Experimental Example 2:

[0215] To systematically evaluate the robustness of the proposed method, experiments were conducted to further analyze the performance of the multi-scale convolutional model under different signal-to-noise ratio conditions, and RMSE was used as the main evaluation metric for comparison. Specific parameters for model training settings are shown in Table 2. During the training phase, parallel training was performed using two NVIDIA GeForce RTX 5060 GPUs.

[0216] Table 2 Experimental parameters for Example 2

[0217] Figure 6 The RMSE variation curves of each method under different signal-to-noise ratios during the testing phase are presented. The results show that the RMSE of all methods decreases as the signal-to-noise ratio increases; under the same signal-to-noise ratio conditions, the multi-scale convolutional model of our proposed method consistently achieves the lowest RMSE, and maintains good estimation accuracy, especially in the low signal-to-noise ratio range, verifying its good noise resistance and stability.

[0218] As can be seen from the comparison, the triple uniform linear array outperforms the nested array in the configuration of sparse array. In addition, the method of this invention not only surpasses many classic deep learning algorithms, but also significantly outperforms the traditional MUSIC algorithm in terms of RMSE, demonstrating its superior estimation accuracy and stability.

[0219] Experimental Example 3:

[0220] To evaluate the effectiveness of the multi-scale attention mechanism module and the multi-scale convolutional feature extraction module, this experiment designed the following three sets of ablation experiments for comparison: the complete multi-scale attention network method, the network method without the multi-scale attention mechanism module, and the benchmark CNN method. RMSE was used as the main performance indicator, and the signal-to-noise ratio was set from -5dB to 15dB.

[0221] This experimental example focuses on the ablation study of the proposed multi-scale convolutional attention network. By systematically comparing the performance of different module combinations, the effectiveness of the multi-scale feature extraction module and the multi-scale attention mechanism in the network is verified, and the contribution of each module to the performance improvement of this method is demonstrated.

[0222] The results are as follows Figure 7 As shown, the complete multi-scale network achieves the lowest RMSE under all signal-to-noise ratio conditions, significantly outperforming its ablation variant and benchmark CNN methods. This result demonstrates that the multi-scale feature extraction structure makes a crucial contribution to performance, and the multi-scale attention mechanism module further enhances the model's feature discrimination and noise resistance.

[0223] Experiment Example 4:

[0224] This experiment constructed a physical acoustic microphone array platform based on the array element parameters listed in Table 2, such as... Figure 8 As shown, a 6-element triple uniform linear array was constructed using a microphone array, with a single signal source being a 5000Hz sine wave. Calibration was performed using methods such as lasers and levels. The rotating platform adopted a PLC 200SMART to ensure angular accuracy. Estimation and verification were also performed for single signal source scenarios.

[0225] The experimental procedure followed the received signal preprocessing method defined in step two of the model test, followed by an RMSE evaluation of the system performance. The physical test was conducted from -40° to 40° in 10° increments, with nine experiments performed. The difference between the actual angle and the model-estimated angle was measured in each experiment.

[0226] Experimental results are as follows Figure 9 As shown, the measured RMSE curve is consistent with the trend of the simulation results, which effectively verifies the feasibility and reliability of the proposed method in the actual experimental environment.

[0227] In summary, this invention employs a regularized and highly flexible sparse array design, combined with a deeply coupled multi-scale attention mechanism neural network, to jointly improve the estimation performance of signal direction. Verification with simulated and real-world data shows that this method can still achieve high-precision and robust DOA estimation under complex conditions such as low signal-to-noise ratio and limited snapshots.

Claims

1. A DOA estimation method based on a triple uniform linear array and a multi-scale attention network, characterized in that, Includes the following steps: Step 1: Construct a triple uniform linear array; Based on the set total number of physical array elements and according to the preset regularized mathematical relationship, determine the array element arrangement of the triple uniform linear array composed of three uniform linear subarrays. Step 2: Construct the original dataset; Based on the mathematical model of the triple uniform linear array established in step one, the original dataset required for the multi-scale attention network is generated; the dataset includes: samples of the received signal matrix of the triple uniform linear array. And the actual DOA angle corresponding to each array signal sample. ; Step 3: Preprocessing the raw dataset; This involves processing the array received signal matrix obtained in Step 2. The process involves sequentially performing covariance matrix calculation, vectorization, redundancy removal based on array structure, spatial smoothing, and separation of real and imaginary parts. This ultimately constructs a dual-channel real-valued feature map characterizing the spatial properties of the extended virtual aperture. Thus, the preprocessed dataset includes the input Output ;Finish After preprocessing the original dataset of each sample, the fully preprocessed original dataset is divided into training set, validation set and test set in a ratio of 7:2:

1. Step 4: Construction of a multi-scale attention network model; Construct a deep network that includes multi-scale feature extraction, feature fusion, multi-scale attention mechanism, and output layer; Step 5: Training the multi-scale attention network model; Step 6: Test the multi-scale attention network model on the test set; Step 7: Estimation of actual data.

2. The DOA estimation method based on a triple uniform linear array and a multi-scale attention network according to claim 1, characterized in that, The first step of constructing a triple uniform linear array is as follows: Consider adopting A triple uniform linear array with individual elements is used for signal reception, wherein... , , These represent the number of element intervals in each subarray. , , These represent the element spacing within each subarray, and are variables. , , representing the in a triple uniform linear array The position of each array element, when The element spacing and element spacing of each subarray can be determined by the following formula. To simplify the formula, the element spacing of the first subarray is set to 1. The specific formula is as follows: when When the number is even, the number of element intervals within each subarray , and the element spacing of each subarray , ,as follows: (1) when When the number is odd, the number of element intervals within each subarray , and the element spacing of each subarray , ,as follows: (2) The position of each element in the three uniform linear array is determined by formulas (1) and (2). The difference between the positions of the array elements is obtained by calculating the differential covariance matrix (DCA), denoted as... This generates a virtual array, which is a continuous uniform linear array without holes. The number of useful virtual elements in the virtual array, i.e., the total number of effective virtual elements in the differential array, is used... The formula is as follows: (3) Based on the actual array element configuration, an additional scaling factor is introduced. The scaling factor This represents the element spacing of the virtual uniform linear array after differential co-array processing, and the actual element positions of the triple uniform linear array. It can be expressed by the following formula: (4) In this way, the spacing between the expanded virtual array elements can be adjusted according to different needs while maintaining the proportional relationship of the element spacing. Thus, the triple uniform linear array model is completed.

3. The DOA estimation method based on a triple uniform linear array and a multi-scale attention network according to claim 1, characterized in that, The process of constructing the original dataset in step two is as follows: Consider the DOA problem in a real-world system: A far-field narrowband signal was The formula for the received signal in a triple uniform linear array is as follows: (5) in, express The received signal vector at time t, This represents the total number of snapshots. Direction matrix It contains the steering vector corresponding to each direction of arrival. ,in, , indicating the first The incident angle of each signal source carrier frequency The corresponding wavelength, Wave speed; Let be the signal source vector, where , indicating the first A signal source in The signal input at a given time; This represents additive white Gaussian noise; all The signal captured by each snapshot is represented by the following formula, where... For the first The received signal vector of a snapshot, : (6) The above steps construct a sample dataset. To improve the model's generalization ability, random information sources are used to input angles. Random input signal source Random signal-to-noise ratio, where randomness is represented by a uniform distribution, is constructed from this. A dataset of samples, each sample is processed by formula (6) to obtain all... The signal matrix collected by a quick snapshot The output is the set of directions of arrival for all corresponding signal sources. This establishes the array received signal matrix for each sample. With the direction of the port A one-to-one correspondence between them.

4. The DOA estimation method based on a triple uniform linear array and a multi-scale attention network according to claim 1, characterized in that, Step three, raw data preprocessing, specifically includes: (1) The covariance matrix of the observed data is expressed as: (7) in, , Indicates the first The power of each signal source, Indicates noise power. for 3D identity matrix; (2) In order to expand the triple uniform linear array and obtain a virtual array, the covariance matrix is... Vectorization, as shown below: (8) in, This represents matrix vectorization operations, which stack the matrices column-wise into a column vector. , This represents the Kronecker product operation. express The complex conjugate matrix; , ; (3) Remove redundancy, vector Each element corresponds to a virtual sensor. Since these positional differences are redundant, redundant elements need to be removed and rearranged to obtain... This vector is used to model the virtual uniform linear array signal reception vector based on a single snapshot; (4) Covariance matrix of virtual uniform linear array signal receiving vector based on a single snapshot As shown below: (9) (5) For the generated Spatial smoothing is performed to reduce the correlation between virtual array elements and to improve the quality of the new matrix by averaging the covariance matrices of different subarrays or snapshots. The rank of the space smoothing process can be expressed by the following formula: (10) in, Indicates the first The covariance matrix of each subarray The triple uniform linear array of array elements can be expanded into A uniform linear array of elements, the element spacing of which is denoted as . ; (6) Normalize the matrix using the following formula: (11) in Representation matrix The Euclidean norm; (7) Transform the matrix The real and imaginary parts are separated into two channels, denoted as . ; R c A dual-channel real-valued feature map characterizing the spatial properties of the received signal after spread. , For the number of channels, real part, and imaginary part, and For spatial dimensions, Finally, the complete The original dataset of each sample was preprocessed and divided into training set, validation set and test set in a ratio of 7:2:

1.

5. The DOA estimation method based on a triple uniform linear array and a multi-scale attention network according to claim 1, characterized in that, The construction of the multi-scale attention network in step four specifically includes: (1) Multi-scale feature extraction: Multi-scale convolution operation through Implemented with parallel convolutional branches, each branch Output The calculation is as follows, where : (12) in, Indicates the first Two-dimensional convolution operation with branches, For the convolution kernel weight tensor of this branch, Number of output channels The kernel size is the convolution kernel size. For the corresponding bias vector, This indicates the batch normalization operation performed on the output of this branch. The nonlinear activation function is represented by the SiLU function, which is defined as follows: ; all The outputs of each branch are concatenated along the channel dimension to form a fused multi-scale feature map. : (13) Total number of output channels ; (2) Feature fusion: After concatenating all branch outputs along the channel dimension, the feature fusion is performed... Convolutional feature fusion yields: (14) in This represents the total number of output channels for the multi-scale feature extraction block. This indicates that the number of input and output channels is both of convolution; (3) Multi-scale attention mechanism: The output of the multi-scale attention block is obtained through grouping, weight calculation, intra-scale feature enhancement, weighting, cross-scale information exchange and residual connection. ; (4) Output layer: The output of the multi-scale attention block Mapping to DOA estimates, first through... Number of channels in convolution compression: (25) Will Flattened into a one-dimensional vector, the formula is as follows: (26) The above process yields a highly compressed one-dimensional vector, which, after passing through a linear regression layer, provides the DOA estimate, as shown in the following formula: (27) in, The set of directions of arrival (DOAs) of the signal sources estimated by the model; The bias vector of the linear layer. This is the weight matrix of the linear layer.

6. The DOA estimation method based on a triple uniform linear array and a multi-scale attention network according to claim 5, characterized in that, The multi-scale attention mechanism mentioned in (3) specifically includes: 3.1) Grouping: Evenly divided in the channel dimension Group, corresponding Obtained at different scales The formula is as follows: (15) Each group ; 3.2) Weight Calculation For input features Perform global average pooling to obtain channel statistics: (16) Pass through the first layer Convolution reduces the channel dimension, as shown in the following formula: (17) Through the second layer Convolution generates scale attention weights: (18) in express Attention weights at each scale, where This represents the attention weight for the g-th group; This represents the Softmax function, another commonly used non-linear activation function, which is suitable for the output layer representation of multi-classification tasks in neural networks, thereby obtaining the attention weights for each scale. 3.3) Intra-scale feature enhancement: For each scale group Perform independent feature enhancement processing (depthiable separable convolution): (19) in This represents a depthwise convolution with a kernel size of . The number of groups is , This represents point convolution with a kernel size of . ; 3.4) Weighting: Weighting attention points Expand to the same spatial dimension as the feature map and weight the enhanced features: (20) Concatenate all weighted scale group features along the channel dimension: (21) 3.5) Cross-scale information exchange: This involves the exchange of spliced ​​features. To conduct cross-scale information exchange, the first step is through a layer Convolution dimensionality reduction: (22) Then through another layer Number of convolution recovery channels: (23) 3.6) Residual Connectivity: To prevent gradient vanishing or exploding, residual connectivity is used here, as shown in the following formula: (24)。 7. The DOA estimation method based on a triple uniform linear array and a multi-scale attention network according to claim 1, characterized in that, Step five, training the multi-scale attention network, includes: The training phase configuration is as follows: the loss function uses mean squared error loss (MSELoss) to measure the deviation between the model's predicted values ​​and the true values; the optimizer is the Adam optimizer; the initial learning rate is set to 0.001, and the learning rate scheduler uses the ReduceLROnPlateau strategy, with the validation loss as the monitoring metric for dynamic parameter tuning. When the validation loss stagnates for four consecutive rounds, the learning rate is halved; simultaneously, torch.amp.GradScaler is enabled to achieve mixed-precision training, supporting FP16 precision calculation, effectively improving training speed and reducing GPU memory usage while ensuring training accuracy, and saving the model that performs best on the validation set.

8. The DOA estimation method based on a triple uniform linear array and a multi-scale attention network according to claim 1, characterized in that, In step six, the root mean square error (RMSE) is used as a model performance evaluation metric during the testing phase to quantify the accuracy of the model's prediction results. The formula is as follows: (28) in Indicates the number of samples in the test set. The number of signal sources; Indicates the first Model predictions for the angle of each test sample; Indicates the first The true value of the angle of each test sample.

9. The DOA estimation method based on a triple uniform linear array and a multi-scale attention network according to claim 1, characterized in that, Step seven includes building a corresponding physical array hardware platform, collecting array received signals in the actual environment, performing data preprocessing, and inputting the data into the trained network model to obtain the direction of arrival (DOA) estimation results of the actual signal.