Dynamic direction of arrival estimation method based on fusion neural network and beamforming technology

By integrating neural networks and beamforming technology, and combining multimodal feature extraction and physical prior constraints, the challenge of angle of arrival estimation in dynamic source scenarios was solved, achieving high-precision and robust signal positioning.

CN121456780BActive Publication Date: 2026-07-24QINGDAO UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO UNIV OF SCI & TECH
Filing Date
2025-09-15
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing angle-of-arrival estimation methods are inadequate in dynamic source scenarios, low signal-to-noise ratio, and multipath effects. They are particularly affected by marine environmental noise, animal calls, and multipath effects, and lack dynamic output capabilities.

Method used

A dynamic angle of arrival (DOA) estimation method based on fusion neural network and beamforming technology is adopted. Through multimodal feature extraction, gated cross-attention mechanism and physical prior constraints, combined with Transformer decoder and Bellhop model, dynamic DOA estimation is achieved.

Benefits of technology

High-accuracy signal positioning was achieved in complex marine environments, breaking through the performance bottleneck of traditional methods. It can dynamically estimate the signal angle of arrival, improving estimation accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456780B_ABST
    Figure CN121456780B_ABST
Patent Text Reader

Abstract

The application belongs to the field of signal processing, and relates to a dynamic direction of arrival (DOA) estimation method based on a fusion neural network and a beamforming technology. The method comprises the following steps: S1, receiving an observation signal, using a multi-modal feature extraction method to extract complementary features from three dimensions of a time-frequency domain, a spatial domain and a perception domain, and constituting a basis of multi-modal feature representation; S2, based on the multi-modal features extracted in S1, using a gated cross attention mechanism to adaptively fuse the features; S3, based on the adaptively fused features in S2, using a multi-level optimization framework that fuses physical priori and data driving, and through the synergistic effect of neural network parameter learning and beamforming theory constraints, realizing DOA estimation in a complex scene. The application can efficiently and accurately dynamically estimate the DOA of a signal, improve the signal receiving quality, and provide a new idea with theoretical rigor and engineering practicability for underwater acoustic monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of signal processing and relates to a dynamic angle of arrival estimation method based on fused neural networks and beamforming technology. Background Technology

[0002] Angle of arrival (DOA) estimation is a crucial task in underwater acoustic array signal processing, as it forms the basis for beamforming and source localization. DOA estimation techniques include beamforming methods such as Capon beamformers and maximum likelihood (ML) methods, as well as subspace-based methods. The superior performance of these algorithms hinges on an understanding of array characteristics such as gain, phase, and position. For DOA estimation of uniform linear arrays, numerous high-resolution algorithms have been proposed, including the Multiple Signal Classification (MUSIC) algorithm, the Rotation Invariant Techniques for Signal Parameter Estimation (ESPRIT) algorithm, the Propagator Method (PM), and the DOA matrix algorithm. The traditional DOA matrix method is a two-dimensional (2-D) DOA estimation method for multiple narrowband signals from two parallel linear arrays. It utilizes autocorrelation and cross-correlation information to construct a new matrix called the DOA matrix. Furthermore, two-dimensional DOA estimation can also be achieved by performing eigenvalue decomposition (EVD) of the DOA matrix. In particular, the traditional DOA matrix method avoids spectral peak search and offers higher computational efficiency compared to the MUSIC and ESPRIT algorithms. However, traditional direction-of-arrival (DOA) matrix methods partially discard autocorrelation and cross-correlation information, which inherently reduces estimation accuracy. Based on the above research, current DOA estimation algorithms are observed to be severely affected by marine environmental noise. Animal calls may appear suddenly or overlap, and multipath effects and reverberation affect signal quality. Existing algorithms suffer from drawbacks such as relying on prior information about the number of sound sources, while neural network methods are mostly designed for cases with a fixed number of sound sources and lack dynamic output capabilities. Summary of the Invention

[0003] The purpose of this invention is to propose a dynamic angle of arrival estimation method based on the fusion of neural networks and beamforming technology to overcome the shortcomings of existing technologies.

[0004] To achieve the above-mentioned objectives, the present invention employs the following technical solution: The dynamic angle-of-arrival estimation method based on the fusion of neural networks and beamforming technology includes the following steps: 1. A dynamic angle-of-arrival estimation method based on a fusion neural network and beamforming technology, characterized by comprising the following steps: S1: Receive the observed signal and use the multimodal feature extraction method to extract complementary features from the three dimensions of time-frequency domain, spatial domain and perception domain, which constitute the basis of multimodal feature representation; S2: Based on the multimodal features extracted by S1, a gated cross-attention mechanism is used for adaptive fusion of features; S3: Based on the adaptive fusion features of S2, a multi-level optimization framework that integrates physical priors and data-driven approaches is adopted. Through the synergistic effect of neural network parameter learning and beamforming theoretical constraints, DOA estimation in complex scenarios is achieved.

[0005] Preferably, in step S1, the time-frequency feature extraction specifically includes: The time-frequency characteristics of the signal are analyzed using STFT transformation; The multi-channel STFT spectral feature branch is input and processed in parallel across multiple channels to construct a dimension of... The time-frequency feature tensor, where For the number of receivers, For frequency, For time.

[0006] Preferably, in step S1, spatial feature extraction specifically includes: Constructing spatial features based on the spatial covariance matrix of array signals: ; In the formula For multi-channel frequency domain vectors, the received signal In the frequency domain, it is represented as: ; in, For frequency, For time frame index ( ), It is the transpose symbol; For the first Each receiving element in The frequency domain signal at time t.

[0007] Calculate the covariance matrix of the signal at a single frequency point: ; in, This indicates the conjugate transpose. This represents the total number of time frames. Calculate the actual symmetric part for each frequency point. and virtual antisymmetry : ; ; and These respectively reflect the energy coupling and encoded phase difference information between array elements; Final generation 3D feature matrix.

[0008] Preferably, in step S1, the acoustic feature extraction specifically includes: firstly, optimizing the Mel filter bank based on the vocal range of marine mammals, adjusting the center frequency of the filter, with the lower limit of the filter's center frequency covering the lowest vocal frequency of the species; the upper limit is based on the sampling rate. Set as 80%; the optimal scale for marine mammals is: ; The frequency resolution is enhanced by logarithmic transformation using the above formula to make it suitable for the harmonic structure of whale songs, and the bandwidth is compressed using power law transformation to adapt it to the broadband characteristics of dolphin ultrasonic signals. Let the number of filters be Then the center frequency Determine using the following steps: Map the target frequency band to the optimization scale: ; ; Uniformly divide the perception scale: ; Inverse mapping back to linear frequency: ; Frequency response function of each filter Defined as: ; in That is, the center frequency of the previous filter. , that is, the center frequency of the next filter; The received signal is passed through a filter to obtain the received signal energy. (Mel filter bank, number 1) The output energy of each channel is: ; in, For the spectrum at frequency The range, It is the first The frequency response of each filter; For output energy The specific steps for performing nonlinear compression are as follows: ; in, To avoid taking the logarithm of zero or negative values, the parameter This is the compressive strength control coefficient. The value is adjusted according to the characteristics of the target signal; The time difference for calculating static MFCC coefficients includes two enhancement levels: first-order difference and second-order difference. The final eigenvector is obtained based on the static MFCC, first-order difference, and second-order difference of the signal: ; in, , , These are the time differences, first-order differences, and second-order differences of the static MFCC coefficients, respectively. Final composition The feature matrix of dimension, where This represents the total number of time frames. The cepstral coefficient dimension.

[0009] Preferably, in S2, the gated cross-attention mechanism specifically includes: (1) Dimensional normalization is performed on each modal feature, and a query vector is generated through linear transformation of the modality setting: ; in, To obtain learnable parameters, each modality is projected onto a unified [parameter]. Dimensional attention space; (2) Cross-modal attention calculation: Construct a trimodal bidirectional attention network; Calculate the cross-attention paths to form a complete feature interaction topology; (3) Dynamic gating fusion: Design an independent gating unit for each attention path: ; Here, || represents the concatenation operation. It is the sigmoid function; The final fusion output is a weighted sum of the gating attention results for each path: ; Represents the set of all bidirectional attention paths. Element-wise multiplication; is the gating coefficient.

[0010] Preferably, in step S3, the multi-level optimization framework integrating physical priors and data-driven approaches specifically includes: (1) Physically guided deep learnable beamformer: ; in, For beamformers in The weight vector of the direction, dimension , The number of array elements; For parameterized neural networks, For the input multi-channel time-frequency signal, dimension ; The target azimuth angle, It is the set of learnable parameters for neural networks; the rationality of the theory is guaranteed by the dual constraints of Cramer-Rao bound and popular consistency. (2) Dynamic output architecture training strategy: A sequence-to-sequence output mechanism based on Transformer; Solving the problem of ambiguity between prediction and actual permutation using the Hungarian loss function; ; in, To arrange from all possible permutations The permutation function selected. It is the first The prediction angles are arranged... The value after that, It is the first A true perspective It is the number of dynamically changing information sources; Using contrastive learning regularization terms to enhance inter-class discrimination: ; in, For positive sample similarity, For negative sample similarity, This is the temperature coefficient, with a value ranging from 0.05 to 0.2, and is usually set to 0.1. The number of negative samples; (3) Channel-adaptive reinforcement learning: Constructing a multipath enhancement dataset based on the Bellhop sound field model: ; in, It is the raw signal data. It is the number of original, pure signal samples. The channel impulse response is randomly generated. It is the number of randomly generated multipath channel impulse responses.

[0011] Preferably, the generalization ability is improved through the following mechanisms: (1) Add path delay compensation in cross-modal attention: ; Among them, here The query vector has a dimension of , The key vector has a dimension of . , The delay compensation matrix has dimensions of . , here For feature dimensions; (2) Dynamically adjust the sound velocity profile during training: ; For depth The speed of sound at a certain point, in m / s. It is the reference speed of sound. It is the amplitude of the change in sound speed, in m / s. It is the period of sound speed change.

[0012] Preferably, a self-supervised pre-training strategy is adopted, and a two-level pre-training task is designed, specifically including: (1) Mask spectrum reconstruction Randomly mask 50% of the spectrum and recover it using an autoencoder: ; in, It is the original spectrum, dimension F × T , here It is a random binary mask matrix, with dimensions... F × T , It is an encoder network. It is a decoder network; (2) Spatial Consistency Learning Constructing positive and negative samples using inter-array geometric constraints: ; in, and These are the estimated angles for different arrays. and These are the array position coordinates, with dimensions 3×1. It's the speed of sound. It's the tolerance error. Preferably, the joint optimization objective is to achieve a final training objective that is a weighted sum of multi-task losses: ; It is the main loss function. These are adaptive weighting coefficients. .

[0013] The advantages and technical effects of this invention are as follows: Addressing the performance bottlenecks of traditional methods in dynamic source scenarios, low signal-to-noise ratio (SNR), and multipath effects, this invention first utilizes a multimodal feature fusion network. It extracts the time-frequency, spatial, and acoustic features of the signal through parallel CNNs, and combines this with an attention mechanism to achieve adaptive weighted feature fusion. Then, based on a dynamic output architecture, a Transformer decoder generates variable-length DOA sequences. Combined with Hungarian loss and contrastive learning regularization terms, this overcomes the traditional limitation of a fixed number of sources, achieving dynamic estimation. Next, a learnable beamformer is constructed using physical-guided deep learning to replace the traditional MVDR. Cramer-Rao bound constraints are introduced to improve theoretical optimality, and multipath data is synthesized based on the Bellhop model to enhance generalization ability. Finally, based on a self-supervised pre-training strategy, masked spectrum reconstruction and spatial consistency learning are used to strengthen the model's DOA estimation performance under low SNR and multipath effects. This invention can dynamically, accurately, and efficiently estimate the signal's angle of arrival (OA), ultimately achieving high-accuracy signal localization. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the dynamic angle of arrival estimation method based on fused neural network and beamforming technology of the present invention; Figure 2 This is a network array receiving model diagram according to an embodiment of the present invention; Figure 3 This is an overall network structure diagram of one embodiment of the present invention; Figure 4 This is a schematic diagram of cross-modal attention fusion according to an embodiment of the present invention; Figure 5 This is a multimodal joint optimization strategy diagram of an embodiment of the present invention. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0016] Example 1: Addressing the shortcomings of existing algorithms, such as reliance on prior information about the number of sources and the limitations of neural network methods (primarily applicable to scenarios with a fixed number of sources and lacking dynamic output capabilities), this invention proposes a dynamic angle-of-arrival (DOA) estimation method based on a fusion of neural networks and beamforming technology for dynamic underwater source DOA estimation. The main steps are as follows: First, feature extraction is performed on the array-received signal, including three modal features: time-frequency features, spatial features, and acoustic features. Then, cross-modal attention fusion is performed on these three modal features. Next, a physically guided beamformer is used, and finally, Hungarian matching is applied to obtain the final output angle result. However, neural network methods are primarily applicable to scenarios with a fixed number of sources and lack dynamic output capabilities. How to achieve dynamic source number estimation in complex marine environments such as multipath propagation and low SNR is the technical problem this embodiment aims to solve.

[0017] The dynamic angle-of-arrival estimation method based on the fusion of neural networks and beamforming technology proposed in this embodiment, such as Figure 1 and Figure 3 As shown, the specific steps include: S1: Establish a signal reception model and, for the multimodal feature extraction method of array signal processing, the system extracts complementary features from three dimensions: time-frequency domain, spatial domain, and sensing domain. The specific steps are as follows: In this invention, a uniform linear array as shown in Figure 2 is used. The uniform linear array contains... Each receiving array element corresponds to a separate channel. Element spacing. Less than half the signal wavelength. The first element is used as a reference element. Uniform linear array receiver. There are array elements, with the first element used as a reference to another array element. Assume there are... A far-field narrowband signal is incident on the array. The frequency of each signal is... The speed of signal propagation underwater is The signal received by a single array element is represented as: ; in, Indicates the first The receiving gain of each receiving array element Indicates the time, also known as the time frame index. Indicates the first One transmitted signal, express The noise received by each receiving element Indicates the first The delay of each receiving array element relative to the reference array element The angle of arrival is expressed as follows: .

[0018] S1-1: Extraction of time-frequency features (STFT spectral features): The time-frequency characteristics of the signal are analyzed using windowed short-time Fourier transform (STFT). Given an array signal... Its STFT transformation is defined as: ; In the formula, As the center of time, For frequency, This represents the window function, which varies with time... move, Window width Adaptive modulation is applied based on signal characteristics. The window function typically has a large value near the center and rapidly decays to 0 further away, ensuring that only the signal within the window contributes to the current analysis. This invention uses a Gaussian window, which effectively balances time and frequency resolution, making it particularly suitable for analyzing time-varying signals. The complex spectrum is decomposed into its real part (…). ) and imaginary part ( The signals are input separately, thus preserving the signal energy distribution and phase information.

[0019] After performing the STFT transformation, the multi-channel STFT spectral feature branches are input and processed in parallel through multiple channels to construct a dimension of The time-frequency feature tensor, where For the number of receivers, For frequency, For time.

[0020] S1-2: Spatial Feature (Covariance Matrix) Extraction: Constructing spatial features based on the spatial covariance matrix of array signals: ; In the formula, For multi-channel frequency domain vectors, the received signal In the frequency domain, it is represented as: ; in, For frequency, For time ( ), It is the transpose symbol; For the first Each receiving element in The frequency domain signal at time t.

[0021] Calculate the single-frequency covariance matrix of the signal using the following formula: in, This indicates the conjugate transpose. For time.

[0022] The actual symmetric part is calculated for each frequency point using the following two formulas. and virtual antisymmetry : ; ; and These respectively reflect the energy coupling and encoded phase difference information between array elements.

[0023] Final generation A 3D feature matrix can resist directional interference.

[0024] S1-3: Acoustic Feature Extraction (Mel Frequency Cepstral Coefficients MFCC): To address the vocal characteristics of mammals, the traditional MFCC feature extraction process is improved. First, the center frequency of the filter is adjusted according to the vocal range of marine mammals. The signal is then passed through the filter to obtain the output energy. Next, the output energy is nonlinearly compressed. Finally, the dynamic features of the MFCC are obtained by calculating the time difference of the static MFCC coefficients. The specific steps are as follows: (1) First, optimize the Mel filter bank by adjusting the center frequency of the filter according to the vocal range of marine mammals.

[0025] lower limit : The lowest vocal frequency of the covered species (e.g., 20Hz).

[0026] upper limit Based on sampling rate Set as 80% (to avoid high-frequency noise), if ,but 。

[0027] The optimal scale for marine mammals is: ; The frequency resolution is enhanced by logarithmic transformation using the above formulas to make it suitable for the harmonic structure of whale songs, and the bandwidth is compressed using power-law transformation to adapt it to the broadband characteristics of dolphin ultrasonic signals.

[0028] Let the number of filters be Then the center frequency Determine using the following steps: (a) Mapping the target frequency band to the optimization scale: ; .

[0029] (b) Uniformly dividing the perceptual scale: ; (c) Inverse mapping back to linear frequency: ; (d) Frequency response function of each filter Defined as: ; in, , that is, the center frequency of the previous filter; That is, the center frequency of the next filter.

[0030] (e) The received signal is passed through a filter to obtain the received signal energy. The Mel filter bank... The output energy of each filter is: ; in, For the spectrum at frequency The range, It is the first The frequency response of each filter.

[0031] (2) Regarding energy Perform nonlinear compression: ; Among them, the increment operation ( Avoid taking the logarithm of zero or negative values, parameters This is the compressive strength control coefficient. The value is adjusted according to the characteristics of the target signal: .

[0032] By nonlinearly compressing the energy of the received signal, the background noise of the signal can be reduced, the details of low-energy signals can be preserved, and the signal feature discrimination can be improved, making it easier to separate frequency bands with small energy differences in the feature space through logarithmic transformation.

[0033] (3) Finally, the dynamic features of MFCC are obtained by calculating the time difference of the static MFCC coefficients, which includes two enhancement levels: first-order difference and second-order difference.

[0034] The first-order difference is calculated as follows: ; in, This represents the radius of the difference window, typically taken as 2-3 frames, corresponding to a time span of 20-30ms. This is the index for the current time frame. The first-order difference reflects the slope change of the spectral envelope.

[0035] The second-order difference is calculated as follows: ; Second-order difference can detect transient features such as bulges / dips in the spectrum.

[0036] The final eigenvector is obtained based on the static MFCC, first-order difference, and second-order difference of the signal: ; Final composition The feature matrix of dimension, where This represents the total number of time frames. Using first-order and second-order difference features for the cepstral coefficient dimension can improve the model’s time-domain dynamic sensitivity.

[0037] S2: Based on the multimodal features extracted in S1, a gated cross-attention mechanism is proposed to achieve adaptive fusion between features. This mechanism achieves refined cross-modal information interaction through triple gating control, such as... Figure 4 As shown, the specific steps are as follows: S2-1: Feature Preprocessing and Query Generation Dimensional normalization is performed on each modal feature: STFT features ( Flattened along the frequency-time dimension The matrix, the covariance matrix ( Maintain the spatial structure and flatten it out. Vector, MFCC features ( The temporal structure is preserved. Then, a query vector is generated through a mode-specific linear transformation: ; in, For learnable parameters, For bias, This indicates that the features are flattened; superscript 1 indicates the parameters corresponding to the STFT features, superscript 2 indicates the parameters corresponding to the feature covariance matrix, and superscript 3 indicates the parameters corresponding to the MFCC features, projecting each mode onto a unified plane. Dimensional attention space.

[0038] S2-2: Cross-modal attention calculation: Construct a trimodal bidirectional attention network to... For example, the path: ; ; in, It is the query vector corresponding to the STFT feature. It is an activation function. , and All are learnable parameters. This represents the MFCC feature.

[0039] Similarly, calculate the remaining 4 sets of cross-attention paths ( ): ; ; ; ; The above six sets of cross-attention paths form a complete feature interaction topology.

[0040] S2-3: Dynamic Gated Fusion Design an independent gating unit for each attention path: ; Here, || represents the concatenation operation. For the sigmoid function, This indicates the output of the independent gating unit. Learnable parameters It is a bias.

[0041] The final fusion output is a weighted sum of the gating attention results for each path: ; Represents the set of all bidirectional attention paths. For element-wise multiplication, Indicates an independent gating unit. Indicates features, It is a cross-attention path. (Using gating coefficients) It achieves soft selection of feature contributions and uses residual connections to preserve key information of the original modes, while also improving the utilization rate of cross-modal complementarity.

[0042] S3: Based on the adaptively fused features of S2, a multi-level optimization framework integrating physical priors and data-driven approaches is proposed. Through the synergistic effect of neural network parameter learning and beamforming theoretical constraints, DOA estimation of the system in complex scenarios such as low signal-to-noise ratio and multipath interference is achieved. Figure 5 As shown, the specific steps are as follows: S3-1: Physically Guided Deep Beamformer (Forces the Network to Obey Acoustic Physics Laws): Traditional MVDR beamformers experience a sharp performance degradation when array mismatch occurs. This invention constructs a differentiable beamforming layer to replace traditional computation: ; in, For the beamformer at the target azimuth angle Weight vector of direction (dimensional) , (Number of array elements) The covariance matrix representing the noise. Indicates the direction corresponding to the target. The guide vector, This indicates the conjugate transpose. For parameterized neural networks, For the input multi-channel time-frequency signal (dimension) ), This is the set of learnable parameters for the neural network. The theoretical validity is ensured by the dual constraints of Cramer-Rao bound and popular consistency. The Cramer-Rao bound constraint involves adding a CRB regularization term to the loss function. ; To estimate the covariance matrix (dimensions) of the angle for the model M × M , M (Number of sources) The theoretical value of the lower bound of Cramer-Rao (dimension) M × M ), It is the Frobenius norm.

[0043] Manifold consistency constraints can force the network output to satisfy the orthogonality of the array manifold: ; For array manifold matrix (dimension) ), For all beam weight matrices output by the network (dimensions) ), identity matrix (dimension) ).

[0044] S3-2: Dynamic Output Architecture Training Strategy (Handling Sudden Changes in Whale Population): To adapt to changes in the number of information sources, a sequence-to-sequence output mechanism based on Transformer is designed: Solving the problem of ambiguity between prediction and true permutation using the Hungarian loss function: ; in, To arrange from all possible permutations The permutation function selected in; It is the first The prediction angles are arranged... The value after that, It is the first A true perspective It is the number of information sources that changes dynamically.

[0045] Enhance inter-class discrimination by using contrastive learning regularization terms: ; in, For positive sample similarity (such as angle prediction from the same information source). For negative sample similarity (such as prediction from the perspective of different information sources). This is the temperature coefficient, with a value ranging from 0.05 to 0.2, and is usually set to 0.1. This represents the number of negative samples.

[0046] S3-3: Channel Adaptive Reinforcement Learning (Automatically Adapting to Echo, Noise, and Other Interference): Constructing a multipath enhancement dataset based on the Bellhop sound field model: ; in, It is the raw signal data. It is the number of original, pure signal samples. The channel impulse response is randomly generated. It is the first A true perspective This represents the number of randomly generated multipath channel impulse responses. Generalization capability is improved through the following mechanisms: (1) Multipath attention mask: Add path delay compensation in cross-modal attention ; Among them, here The query vector has a dimension of , The key vector has a dimension of . , Indicates transpose. The delay compensation matrix has dimensions of . , here For feature dimensions.

[0047] (2) Time-varying channel simulation: dynamically adjusting the sound velocity profile during training. ; For depth The speed of sound at a certain point, in m / s. It is the reference speed of sound, usually taken as 1500 m / s. This is the amplitude of the change in sound speed, measured in m / s, with a range of 20-100 m / s. It is the period of sound speed change, such as 100 m.

[0048] S3-4: Self-supervised pre-training paradigm: Design a two-level pre-training task: (1) Masked spectrum reconstruction: Randomly mask 50% of the spectrum region and recover it through an autoencoder. ; in, It is the original spectrum, with dimensions F×T, here. It is a random binary mask matrix, with dimensions... , It is an encoder network. It is a decoder network.

[0049] (2) Spatial Consistency Learning: Constructing positive and negative samples using geometric constraints between arrays ; in, and These are the estimated angles for different arrays. and These are the array position coordinates, with dimensions 3×1. It's the speed of sound. It is the tolerance error, with a value range of 0.05-0.1 rad.

[0050] S3-5: Joint Optimization Objective: The ultimate training objective is a weighted sum of multi-task losses: ; It is the main loss function (such as mean squared error). ( ) is the adaptive weighting coefficient.

[0051] To evaluate the performance of the method of the present invention, the present invention provides the following evaluation experiments.

[0052] 1. Algorithm Evaluation Metrics 1) Detection performance The formula for calculating accuracy is: ; in, (True Positive) is the predictive angle. With a certain real angle The number of samples when the error is less than the threshold of 5°. (False Positive) is the number of predicted angles that do not match any real sound source.

[0053] The detection coverage of real sound sources is: (False Negative) indicates that no predicted angle was matched (true angle). With all predictive angles The number of true sound source samples with a minimum error > 5°. The F1-score is used as a comprehensive indicator to evaluate both detection accuracy and recall, as shown in the formula: ; 2) Estimation accuracy Root Mean Square Error (RMSE): ; This represents the number of samples.

[0054] Cramiro Boundary Achievement Rate (CRB Ratio): .

[0055] 3) Robustness Multipath Tolerance (MPT) is used to evaluate the robustness of the algorithm. MPT reflects the algorithm's ability to resist multipath attacks, as shown in the formula: ; in. These represent performance metrics (F1-score, RMSE, etc.) under pure signal (no multipath) conditions. This represents the performance index value under multipath interference signals (with the same SNR conditions). This invention uses the robustness of MPT based on F1-score and MPT based on RMSE to evaluate the performance of these algorithms.

[0056] 2. Performance Comparison of Different Algorithms The algorithm proposed in this invention was compared with several existing ODA algorithms, including RAB-DOA, DANM-2D, FBCK, TRDML, and CNN-OG, and the experimental results are shown in Table 1.

[0057] Table 1 Average Experimental Results The proposed algorithm achieves an F1 score 7.06% higher than TRDML (the second highest), thanks to its physical constraint architecture. The proposed algorithm has a mean root mean square error of 2.1, 0.8 lower than TRDML. A CRB ratio of 1.2 indicates its theoretical optimality, superior to TRDML's 1.6. The proposed algorithm also boasts the highest MPT among all algorithms, ensuring performance even with multipath effects. Unfortunately, the proposed algorithm does not have the shortest runtime; the shortest runtime is achieved by the CNN-OG algorithm. Considering all performance metrics, the proposed algorithm performs best.

[0058] 3. Comparison of algorithm performance in different simulation environments Referring to Table 2, the algorithm proposed in this invention maintains a maximum path loss (MPT) of over 85% in all scenarios. This is because the gated cross-attention mechanism in the algorithm effectively suppresses false targets caused by reflections. In shallow water areas, due to strong sea surface reflections and dense multipath interference, all algorithms perform poorly. The algorithm proposed in this invention performs best, followed by the TRDML algorithm, which improves the accuracy of DOA estimation by using a deep mutual learning (DML) model. The RAB-DOA algorithm performs third best because it uses a linearly constrained minimum variance (LCMV) beamforming algorithm and phase adjustment to optimize signal directivity and suppress interference. The DANM-2D, FBCK, and CNN-OG algorithms perform the worst.

[0059] Table 2 MPT Comparison Aside from pristine environments, all algorithms performed best in deep water. However, due to weak seabed reflections and low-frequency signal attenuation, the performance of traditional algorithms was only moderate, decreasing by approximately 10-25%. In the thermocline, due to acoustic channel effects and signal distortion, the performance of all algorithms decreased by approximately 14-36% compared to pristine environments.

[0060] 4. Ablation test To verify the role of key components in the algorithm of this invention, firstly, physical constraint components such as CRB constraints and manifold consistency loss were removed, while other components in the network remained unchanged, and the average F1 score was calculated after 100 experiments. Then, the attention fusion mechanism was removed, and the gated cross-attention mechanism was replaced with feature concatenation, while other components in the network remained unchanged, and the average F1 score was calculated after 100 experiments. Finally, the network output was fixed at 4 DOAs, while other components in the network remained unchanged, and the average F1 score was calculated after 100 experiments. Table 3 shows that when physical constraint components such as CRB constraints and manifold consistency loss were removed, the F1 score decreased by 8.8%. When the attention fusion mechanism was removed and the gated cross-attention mechanism was replaced with feature concatenation, the F1 score decreased by 13.2%, indicating that angle confusion errors will increase in multi-source scenarios. When the Transformer decoder and Hungarian loss were removed, and the output was fixed at 4 DOAs, the F1 score decreased by 20.9%, and due to redundant computation, the processing time increased from 92ms to 115ms. Experimental results verify the effectiveness of the "physical guidance + data-driven" joint design concept of the algorithm of this invention.

[0061] Table 3 Results of the ablation experiment This invention can solve the problem of dynamic DOA estimation of mammalian vocalizations in complex marine environments, and overcome the performance bottleneck of traditional methods in dynamic source scenarios, low signal-to-noise ratio and multipath effects.

[0062] The above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Therefore, any changes made in accordance with the shape and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A dynamic angle-of-arrival estimation method based on a fusion of neural networks and beamforming technology, characterized in that, Includes the following steps: S1. Receive the observed signal and use the multimodal feature extraction method to extract complementary features from the three dimensions of time-frequency domain, spatial domain and perception domain, which constitute the basis of multimodal feature representation; S2. Based on the multimodal features extracted in S1, a gated cross-attention mechanism is used for adaptive fusion of features; specifically: (1) The dimensions of each modal feature are normalized, and a query vector is generated through a linear transformation of the modality setting: ; in, To obtain learnable parameters, each modality is projected onto a unified [parameter]. Dimensional attention space; Indicates MFCC features; F stft Indicates STFT characteristics; F cov Represents the characteristics of the covariance matrix; For bias, superscript 1 indicates the parameter corresponding to the STFT feature, superscript 2 indicates the parameter corresponding to the feature covariance matrix, and superscript 3 indicates the parameter corresponding to the MFCC feature. (2) Cross-modal attention calculation: Construct a trimodal bidirectional attention network; Calculate the cross-attention paths to form a complete feature interaction topology; (3) Dynamic gating fusion: Design an independent gating unit for each attention path: ; Here, || represents the concatenation operation. It is the sigmoid function; Learnable parameters It is a bias; The final fusion output is a weighted sum of the gating attention results for each path: ; Represents the set of all bidirectional attention paths. Element-wise multiplication; The gating factor; It is a cross-attention path; S3. Based on the adaptively fused features of S2, a multi-level optimization framework integrating physical priors and data-driven approaches is adopted. Through the synergistic effect of neural network parameter learning and beamforming theoretical constraints, DOA estimation in complex scenarios is achieved. The multi-level optimization framework integrating physical priors and data-driven approaches is as follows: (1) Physically guided deep learnable beamformer: ; in, For beamformers in The weight vector of the direction, dimension , The number of array elements; For parameterized neural networks, For the input multi-channel time-frequency signal, dimension ; The target azimuth angle, It is the set of learnable parameters for neural networks; the rationality of the theory is guaranteed by the dual constraints of Cramer-Rao bound and popular consistency. The covariance matrix representing the noise; α(θ) Indicates the direction corresponding to the target. θ The guide vector; (2) Dynamic output architecture training strategy: A sequence-to-sequence output mechanism based on Transformer; Solving the problem of ambiguity between prediction and actual permutation using the Hungarian loss function; ; in, To arrange from all possible permutations The permutation function selected in the selection; It is the first The prediction angles are arranged The value after that, It is the first A true perspective It is the number of dynamically changing information sources; Enhance inter-class discrimination by using contrastive learning regularization terms: ; in, For positive sample similarity, For negative sample similarity, This is the temperature coefficient, with a value ranging from 0.05 to 0.

2. The number of negative samples; (3) Channel-adaptive reinforcement learning: Constructing a multipath enhancement dataset based on the Bellhop sound field model: ; in, It is the raw signal data. It is the number of original, pure signal samples. The channel impulse response is randomly generated. It is the number of randomly generated multipath channel impulse responses.

2. The dynamic angle of arrival estimation method based on fused neural network and beamforming technology according to claim 1, characterized in that, In step S1, the time-frequency feature extraction specifically includes: The time-frequency characteristics of the signal are analyzed using STFT transform; the multi-channel STFT spectral feature branches are input and processed in parallel through multiple channels to construct a dimension... The time-frequency feature tensor, where For the number of receivers, For frequency, For time.

3. The dynamic angle of arrival estimation method based on fused neural network and beamforming technology according to claim 1, characterized in that, In S1, spatial feature extraction specifically includes: Constructing spatial features based on the spatial covariance matrix of array signals: ; In the formula For multi-channel frequency domain vectors, the received signal In the frequency domain, it is represented as: ; in, For frequency, time( ), It is the transpose symbol; For the first Each receiving element in Frequency domain signal at time; Calculate the covariance matrix of the signal at a single frequency point: ; in, This indicates the conjugate transpose. For time; Calculate the actual symmetric part for each frequency point. and virtual antisymmetry : ; ; and These respectively reflect the energy coupling and encoded phase difference information between array elements; Final generation 3D feature matrix.

4. The dynamic angle of arrival estimation method based on fused neural network and beamforming technology according to claim 1, characterized in that, In step S1, acoustic feature extraction specifically includes: First, the Mel filter bank was optimized based on the vocal range of marine mammals, adjusting the filter's center frequency so that its lower limit covers the lowest vocal frequency of the species; the upper limit is determined by the sampling rate. Set as 80%; the optimal scale for marine mammals is: ; The frequency resolution is enhanced by logarithmic transformation using the above formula to make it suitable for the harmonic structure of whale songs, and the bandwidth is compressed using power law transformation to adapt it to the broadband characteristics of dolphin ultrasonic signals. Let the number of filters be Then the center frequency Determine using the following steps: Map the target frequency band to the optimization scale: ; ; Uniformly divide the perception scale: ; Inverse mapping back to linear frequency: ; Frequency response function of each filter Defined as: ; in That is, the center frequency of the previous filter. ; The received signal is passed through a filter to obtain the received signal energy. (Mel filter bank, number 1) The output energy of each channel is: ; in, For the spectrum at frequency The range, It is the first The frequency response of each filter; For output energy The specific steps for performing nonlinear compression are as follows: ; in, To avoid taking the logarithm of zero or negative values, the parameter This is the compressive strength control coefficient. The value is adjusted according to the characteristics of the target signal; The time difference for calculating static MFCC coefficients includes two enhancement levels: first-order difference and second-order difference. The final eigenvector is obtained based on the static MFCC, first-order difference, and second-order difference of the signal: ; in, , , These are the time differences, first-order differences, and second-order differences of the static MFCC coefficients, respectively. Final composition The feature matrix of dimension, where For time, The cepstral coefficient dimension.

5. The dynamic angle of arrival estimation method based on fused neural network and beamforming technology according to claim 1, characterized in that, Generalization ability can be improved through the following mechanisms: (1) Add path delay compensation in cross-modal attention: ; Among them, here The query vector has a dimension of , The key vector has a dimension of . , The delay compensation matrix has dimensions of . , here For feature dimensions; (2) Dynamically adjust the sound velocity profile during training: ; For depth The speed of sound at a certain point, in m / s. It is the reference speed of sound. It is the amplitude of the change in sound speed, in m / s. It is the period of sound speed change.

6. The dynamic angle of arrival estimation method based on fused neural network and beamforming technology according to claim 1, characterized in that, A self-supervised pre-training strategy is adopted, and a two-level pre-training task is designed, specifically including: (1) Mask spectrum reconstruction Randomly mask 50% of the spectrum and recover it using an autoencoder: ; in, It is the original spectrum, dimension , here It is a random binary mask matrix, with dimensions... , It is an encoder network. It is a decoder network; (2) Spatial Consistency Learning Constructing positive and negative samples using inter-array geometric constraints: ; in, and These are the estimated angles for different arrays. and These are the array position coordinates, with dimensions 3×1. It's the speed of sound. It is the tolerance error, with a value range of 0.05-0.1 rad.

7. The dynamic angle of arrival estimation method based on fused neural network and beamforming technology according to claim 1, characterized in that, The joint optimization objective is to ultimately train a weighted sum of multi-task losses. ; It is the main loss function. These are adaptive weighting coefficients. ; Represents the loss of the Clameros boundary; This represents the loss of array manifold consistency.

Citation Information

Patent Citations

  • Wireless signal and thermal imaging multi-mode cooperative detection and directional interference method and system

    CN119395688A

  • Method for jointly estimating gain-phase error and direction of arrival (DOA) based on unmanned aerial vehicle (UAV) array

    US20230160991A1