A dual-mode sound source direction of arrival estimation method

By employing a dual-modal direction-of-arrival (DOA) estimation method for acoustic targets, combining electromagnetic and acoustic modes for joint sensing and feature fusion, the stability and accuracy issues of DOA estimation in complex environments are resolved, achieving higher estimation stability and robustness.

CN121679473BActive Publication Date: 2026-05-15HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2026-02-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing direction-of-arrival (DOA) estimation methods suffer from insufficient stability and accuracy in complex environments. In particular, in environments with low signal-to-noise ratios or strong reflections, a single mode is difficult to provide a reliable direction criterion, and multi-mode fusion methods fail to fully utilize modal differences, resulting in unstable direction estimation.

Method used

A dual-modal target direction-of-arrival estimation method is adopted, which combines electromagnetic and acoustic modes for joint sensing. Modal importance weights are generated through modal adaptive gating and cross-modal attention mechanisms, and feature fusion is performed. Adjustable metasurfaces are used to reduce hardware complexity and improve estimation stability and accuracy.

Benefits of technology

The stability and robustness of direction-of-arrival estimation are improved in complex environments, the system's adaptability to asynchronous and multi-source information scenarios is enhanced, and the discriminative performance and generalization ability of direction estimation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121679473B_ABST
    Figure CN121679473B_ABST
Patent Text Reader

Abstract

The application discloses a dual-mode sound emission target direction-of-arrival estimation method, which comprises the following steps: firstly, synchronously collecting electromagnetic echo signals and acoustic signals of a target; obtaining an electromagnetic mode phase matrix and an acoustic mode phase matrix respectively; then, extracting an electromagnetic characteristic vector and an acoustic characteristic vector; taking the electromagnetic characteristic vector and the acoustic characteristic vector as inputs, and outputting electromagnetic target mode characteristics and acoustic target mode characteristics through a mode adaptive gating module respectively; taking the electromagnetic target mode characteristics as source mode characteristics of the acoustic target mode, and taking the acoustic target mode characteristics as source mode characteristics of the electromagnetic target mode; obtaining a dual-mode fusion characteristic vector; finally, inputting the dual-mode fusion characteristic vector into a classifier to output a target direction-of-arrival probability distribution, and selecting a direction with the maximum probability as a final direction-of-arrival estimation result. The method applies electromagnetic mode and acoustic mode joint perception for estimation, and can fully utilize the complementary characteristics between different physical modes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound source back estimation technology, specifically a method for estimating the direction of arrival of a dual-mode sound-emitting target. Background Technology

[0002] Direction-to-arrival (DOA) estimation is a key technical problem in the fields of acoustic signal processing and electromagnetic sensing. Its estimation accuracy directly affects the overall performance of application systems such as speech enhancement, robot navigation, radar systems, wireless communication, and intelligent monitoring.

[0003] Existing direction-of-arrival (DOA) estimation methods typically rely on a single physical mode or a single type of sensor to obtain spatial orientation information, but they still have significant limitations under complex environmental conditions. In ODA methods based solely on acoustic modes, acoustic signals are easily affected by environmental noise, reverberation, multipath propagation, and obstruction, leading to phase information distortion and thus orientation estimation errors. Especially in low signal-to-noise ratio or strong reflection environments, a single acoustic mode is insufficient to provide a stable and reliable orientation criterion. In ODA methods based solely on electromagnetic modes, although electromagnetic signals have good anti-interference capabilities in some scenarios, the echo signal strength decreases and the orientation information becomes unstable when the target is a non-metallic material, has a low scattering cross section, or has weak electromagnetic reflection characteristics, making it difficult to meet the requirements for high-precision ODA on its own.

[0004] To improve the robustness of direction-of-arrival (DOA) estimation, some studies have attempted to introduce multi-sensor or multi-modal fusion strategies. However, existing multi-modal methods often employ simple data concatenation or fixed-weighting methods for fusion, failing to fully consider the differences in sampling mechanisms, time scales, and physical characteristics among different modalities. When the information from different modalities is asynchronous, misaligned, or of varying quality, the fusion result is easily affected by low-quality modal information, thus reducing the stability and accuracy of the DOA estimation.

[0005] Therefore, how to fully utilize the complementary characteristics between different physical modes under complex acoustic and electromagnetic environments, and introduce effective modal reliability assessment and fusion mechanisms to improve the stability, robustness and accuracy of direction of arrival estimation in complex environments, remains a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and propose a dual-modal sound-emitting target direction-of-arrival estimation method. This method uses joint sensing of electromagnetic and acoustic modes for estimation, which can make full use of the complementary characteristics between different physical modes and introduces an effective modal reliability assessment and fusion mechanism to improve the stability, robustness and accuracy of direction-of-arrival estimation in complex environments.

[0007] To achieve the above objectives, the technical solution specifically adopted by the present invention is as follows:

[0008] A method for estimating the direction of arrival (DOA) of a dual-modal acoustic target includes the following steps:

[0009] Step S1: Construct a dual-mode signal acquisition hardware platform to simultaneously acquire the electromagnetic echo signal and acoustic signal of the target;

[0010] Step S2: Perform orthogonal demodulation on the acquired electromagnetic echo signal to obtain the electromagnetic mode phase matrix; perform time-frequency transformation on the acquired multi-channel acoustic signal, extract the phase spectrum, and obtain the acoustic mode phase matrix.

[0011] Step S3: Input the electromagnetic mode phase matrix and the acoustic mode phase matrix into the feature extraction module respectively to obtain the electromagnetic feature vector and the acoustic feature vector;

[0012] Step S4: Input the electromagnetic feature vector and acoustic feature vector into the modal adaptive gating module, generate modal importance weights based on the reliability of each mode, and use the weights to adjust the feature vectors. At the same time, fuse the position encoding information and output the target modal features, which are electromagnetic target modal features and acoustic target modal features, respectively.

[0013] Step S5: Use the electromagnetic target modal features as the source modal features of the acoustic target mode, and use the acoustic target modal features as the source modal features of the electromagnetic target mode. Inject the respective source modal features into the corresponding target modal features using a cross-modal attention mechanism to obtain the respective target cross-modal features in the dual-modal domain. Then, perform self-attention extraction on the target cross-modal features to obtain their respective fused modal features. Finally, perform a weighted concatenation of the respective fused modal features of the dual modes based on the modal importance weights generated by the reliability of each mode to obtain the dual-modal fused feature vector.

[0014] Step S6: Input the dual-modal fusion feature vector into the classifier, output the direction-of-arrival probability distribution of the target, and select the direction with the highest probability as the final direction-of-arrival estimation result.

[0015] Preferably, in step S1, the dual-mode signal acquisition hardware platform includes an electromagnetic mode acquisition module and an acoustic mode acquisition module;

[0016] The electromagnetic mode acquisition module includes an adjustable metasurface, a transmitting and receiving antenna, and a modulation and demodulation processor, used to acquire and demodulate the phase information of the target echo;

[0017] The acoustic modal acquisition module includes an array of multiple linearly arranged microphones for synchronously acquiring multi-channel acoustic signals.

[0018] Preferably, step S2 specifically includes:

[0019] The electromagnetic echo signal is quadrature demodulated, the phase sequence of each channel is extracted, and the electromagnetic mode phase matrix is ​​constructed.

[0020] The multi-channel acoustic signals are framed, windowed, and subjected to short-time Fourier transform to extract phase information at each frequency point and construct the acoustic mode phase matrix.

[0021] Preferably, the electromagnetic mode phase matrix and acoustic mode phase matrix are paired with the corresponding true direction-of-arrival angle labels to construct a dataset, which is then divided into a training set and a test set in a 7:3 ratio.

[0022] Preferably, the modal adaptive gating module in step S4 performs the following operations:

[0023] The input feature vector is subjected to linear transformation and nonlinear activation to generate intermediate feature representations.

[0024] The intermediate feature representation is then linearly transformed again and activated by the Sigmoid function to generate initial modal gating weights;

[0025] A preset confidence threshold is introduced to suppress modal information corresponding to initial gating weights that are below the threshold.

[0026] The retained initial gating weights are Softmax normalized to obtain the final modal importance weights;

[0027] The modality importance weights are used to perform weighted fusion of the feature vectors of the corresponding modalities and the positional codes to output the adjusted target modal features.

[0028] Preferably, step S5 includes the following sub-steps:

[0029] Step S5-1: Calculate the query vector, key vector, and value vector of the source modality and the target modality respectively through the cross-modal attention mechanism, and calculate the cross-modal attention output based on the query vector, key vector, and value vector to realize the injection of source modality features into target modality features;

[0030] Step S5-2: The cross-modal attention output is weighted and adjusted using the modal importance weights generated by the modal adaptive gating module, and the weighted cross-modal attention output is merged with the target modal features through residual connection to obtain the cross-modal enhanced target modal features;

[0031] Step S5-3: Perform self-attention calculation on the cross-modal enhanced target modal features, extract the structural information inside the modality, and obtain the fused modal feature representation;

[0032] Step S5-4: Reweight the fusion features obtained in step S5-3 according to the importance weights of the source mode and the target mode, and concatenate the weighted features of the source mode and the target mode to form the final dual-modal fusion feature vector.

[0033] Preferably, the model training process also includes a loss function construction step:

[0034] Calculate the cross-entropy loss for a classification task;

[0035] Calculate the cross-example embedding regularization loss, which is obtained by imposing intra-class compactness and inter-class separability constraints on the feature embedding space;

[0036] The cross-entropy loss and the cross-example embedding regularization loss are added together with preset weights to form the total loss function, which is used to optimize the model parameters.

[0037] Preferably, the cross-example embedding regularization loss is calculated as follows:

[0038] Normalize the feature embeddings of a batch of samples;

[0039] Calculate the feature similarity matrix between each pair of samples;

[0040] Based on the similarity matrix and sample labels, a loss term is constructed to make the features of samples of the same class similar and the features of samples of different classes different.

[0041] Preferably, the feature extraction module in step S3 uses a one-dimensional convolutional neural network to perform convolutional encoding on the electromagnetic mode phase matrix and the acoustic mode phase matrix respectively, and maps the encoded features to a unified dimension.

[0042] Preferably, the method includes constructing a dual-modal sound source direction estimation model based on the Touchformer framework, and implementing steps S3-S6 through the dual-modal sound source direction estimation model.

[0043] Preferably, the dual-modal sound source direction estimation model is trained using a training set and tested using a test set. During the training phase, an adaptive moment estimation algorithm is used to dynamically update the model parameters, and a cosine annealing strategy is used to adjust the learning rate.

[0044] Preferably, the classifier in step S6 is a fully connected neural network layer, whose number of output nodes is the same as the number of directions of arrival categories to be estimated, and finally outputs the probability of each direction category through the Softmax function.

[0045] This invention has the following characteristics and beneficial effects:

[0046] 1. By introducing tunable electromagnetic metasurfaces as electromagnetic sensing units, the ability to acquire directional information is guaranteed while reducing the dependence of traditional multi-element electromagnetic sensing systems on complex hardware structures, thereby effectively reducing the overall hardware complexity and deployment cost of the system.

[0047] 2. By jointly sensing electromagnetic and acoustic modes, the complementary characteristics of the two physical modes in complex environments are fully utilized to compensate for the instability of direction information under noise interference, reverberation or echo attenuation conditions of a single mode, thereby improving the stability and reliability of the direction of arrival estimation results.

[0048] 3. By dynamically evaluating the quality and reliability of input information from different modalities and suppressing or weakening low-reliability modal information, the impact of noise and abnormal information on the direction estimation results is reduced at the input level, thereby improving the robustness of the overall direction estimation process.

[0049] 4. By introducing a feature fusion method that combines cross-modal correlation fusion with intramodal temporal modeling, effective fusion of multimodal directional features can be achieved without strict time alignment of signals from different modalities, thereby enhancing the system's adaptability to asynchronous, multi-source information input scenarios.

[0050] 5. By applying structured constraints to the directional feature representation, the consistency of samples in the same direction in the feature space is enhanced, while the ability to distinguish between samples in different directions is improved, thereby improving the discrimination performance and generalization ability of the direction of arrival estimation under complex environmental conditions. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the overall method framework for a dual-modal acoustic target direction-of-arrival estimation method according to an embodiment of the present invention.

[0052] Figure 2 This is a schematic diagram of a dual-modal acoustic target direction-of-arrival estimation method according to an embodiment of the present invention. Detailed Implementation

[0053] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0054] A method for estimating the direction of arrival (DOA) of a dual-modal acoustic target, such as Figure 1 As shown, it includes the following steps:

[0055] Step S1: Construct a dual-mode signal acquisition hardware platform to simultaneously acquire the electromagnetic echo signal and acoustic signal of the target.

[0056] Specifically, the dual-mode signal acquisition hardware platform includes an electromagnetic mode acquisition module and an acoustic mode acquisition module; the electromagnetic mode acquisition module includes an adjustable metasurface, a transmitting and receiving antenna, and a modulation and demodulation processor, used to acquire and demodulate the phase information of the target echo; the acoustic mode acquisition module includes an array of multiple linearly arranged microphones, used to synchronously acquire multi-channel acoustic signals.

[0057] In this embodiment, a dual-modal parallel signal acquisition hardware platform is constructed using an adjustable metasurface and a linear microphone array.

[0058] An electromagnetic wave signal acquisition device is composed of a two-dimensional patch periodically tunable metasurface with period p, a transmitting and receiving antenna front-end, and a modulation / demodulation processor. An eight-linear microphone array is integrated in front of the array for audio signal acquisition.

[0059] Step S2: Perform orthogonal demodulation on the acquired electromagnetic echo signal to obtain the electromagnetic mode phase matrix; perform time-frequency transformation on the acquired multi-channel acoustic signal, extract the phase spectrum, and obtain the acoustic mode phase matrix.

[0060] Specifically, the electromagnetic echo signal is orthogonally demodulated, the phase sequence of each channel is extracted, and an electromagnetic mode phase matrix is ​​constructed.

[0061] The multi-channel acoustic signals are framed, windowed, and subjected to short-time Fourier transform to extract phase information at each frequency point and construct the acoustic mode phase matrix.

[0062] In this embodiment, a 10 Mbps Walsh-Hadamard sequence is selected as the quadrature modulation and demodulation code. A 10 GHz electromagnetic wave is transmitted forward using a transmitting antenna. Upon contact with the target, the electromagnetic wave generates an echo, and... The incident angle irradiates the periodic A metasurface consisting of M=8 rows. The transmitted wave spectrum, the received signal spectrum before demodulation, and the received signal spectra of each channel after demodulation can be obtained. After demodulating the received signal spectra of each channel, the phase of each channel can be obtained. Construct 1 based on the phase diagram of the signal. The electromagnetic mode phase matrix of M is used as the electromagnetic mode input feature of the model.

[0063] The array signal from 8 microphones is framed and subjected to Short-Time Fourier Transform (STFT). Each frame is set to 32ms in length, with a 50% overlap between frames (i.e., a frame shift of 16ms). A Discrete Fourier Transform (DFT) with a point size of 512 is performed on each frame. The STFT coefficients of the nth frame from the mth microphone are then obtained as follows:

[0064]

[0065] in This represents the amplitude of the coefficient. This represents the phase component of the coefficient, where the value of k represents the k-th frequency point, with a maximum value of 257. According to... Build size m The acoustic modal phase matrix of k is called the phase map and is used as an acoustic modal input feature.

[0066] Furthermore, in this embodiment, a signal acquisition hardware platform is placed in an open area, and the sound-emitting test target is randomly placed at different angles in front of the hardware platform. The metasurface subsystem and microphone array system of the hardware platform start working simultaneously. Data is recorded, including the demodulated 8-channel phase sequence metasurface modal data at time t, the acoustic modal data of the phase map of the 8-channel microphone array, and the angle of the test target in front of the hardware platform at this time (36 categories with a resolution of 5°). A total of 10,000 cleaned time frames containing sound samples are collected, with a sample ratio of 7:3 between the training set and the test set.

[0067] Step S3: Input the electromagnetic mode phase matrix and the acoustic mode phase matrix into the feature extraction module respectively to obtain the electromagnetic feature vector and the acoustic feature vector.

[0068] Specifically, a one-layer one-dimensional convolutional neural network is used to perform convolutional encoding on the electromagnetic mode phase matrix and the acoustic mode phase matrix respectively, and the encoded features are mapped to a unified dimension.

[0069] In this embodiment, the phase matrix of the electromagnetic mode and the phase matrix of the acoustic mode received by the metasurface are initially extracted using two one-dimensional convolutional feature extraction modules. To facilitate cross-attention fusion of the two modes later, the metasurface feature matrix is ​​padded with zeros to achieve the same feature dimension as the microphone, i.e., 257. The one-dimensional convolutional module of the metasurface uses a 1*257 convolutional kernel, with a stride of 1 and padding of 128. The microphone one-dimensional convolutional module is identical.

[0070] Step S4: Input the electromagnetic feature vector and acoustic feature vector into the modal adaptive gating module, generate modal importance weights based on the reliability of each mode, and use the weights to adjust the feature vectors. At the same time, fuse position encoding information and output the target modal features.

[0071] Specifically, for the characterization features of the extracted bimodal time-frequency graph, reliability weights are assigned through a modal adaptive gating module;

[0072] The modality adaptive gating module employs a two-layer feedforward neural network. First, it calculates intermediate feature representations through linear transformation of the first layer and the nonlinear activation function ReLU. :

[0073]

[0074] Another linear transformation and a sigmoid activation function are applied to the intermediate feature representation through the second layer. To generate mode-specific gating weights :

[0075]

[0076] Further preset confidence thresholds. The gating weight is below this threshold ( Modal information from the electromagnetic metasurface signal (E) and microphone array signal (V) is considered insufficient or noisy and is therefore discarded to prevent contamination of the fused multimodal representation. Finally, to clarify the relative contribution of each mode (electromagnetic metasurface signal E, microphone array signal V) during quantization fusion, the gating weights are... Apply the softmax normalization function to obtain the final modality importance weights. :

[0077]

[0078] The target modal features adjusted by the adaptive module are calculated as follows:

[0079]

[0080] in Indicates positional embedding, This represents element-wise multiplication. Position embedding uses sine and cosine coding; assuming the position coding matrix is... The element in the pos-th row and 2i-th column represents the encoding of position pos and dimension 2i. The specific formulas for position pos and dimension i (starting from 0) are as follows:

[0081]

[0082] .

[0083] Step S5: Inject the target modal features of another modality into the target modal features of the current modality as the source modal features of the current modality through cross-modal attention to obtain the dual-modal fusion feature vector Z.

[0084] Understandably, in this embodiment, the electromagnetic target modal features are used as the source modal features of the acoustic target modes, and thus the acoustic target modal features are used as the source modal features of the electromagnetic target modes.

[0085] Specifically, it models both intra-modal and inter-modal interactions simultaneously to capture intra-modal temporal structures and cross-modal semantic dependencies. It mainly comprises a two-layer attention network, with the first layer being an inter-modal attention layer. In this embodiment, the target modality is used. The Q-matrix of the domain and the source mode Calculating cross-modal attention using K and V matrices within the domain .

[0086] Understandably, in a bimodal model where one mode is the target mode, the other is the source mode. The Q, K, and V matrices are calculated using the same method as in the general Transformer model:

[0087]

[0088] Simultaneously, this attention output is subsequently modulated by a modality importance weight, which is the computational source modality of the modality adaptive gating module in step S4. Modal importance weight results ,therefore = . Source mode The modal importance weights reflect the source modes. Higher weights indicate greater contribution to the target mode, reflecting its reliability.

[0089] By using residual connections, the cross-modal output is merged with the target modal features to obtain the target cross-modal features. :

[0090]

[0091] Subsequently, the target modal features are self-attentionally extracted through the second-layer intra-modal attention layer to obtain the fused modal features. :

[0092]

[0093] Finally, the two modes were used as target modes respectively to obtain their corresponding fused modal features, and the fused modal features of the two modes were then combined. By its importance coefficient The features are reweighted and concatenated to form a bimodal fusion feature vector Z:

[0094] .

[0095] Furthermore, a total loss function is constructed using cross-example embedding regularization loss and cross-entropy loss function.

[0096] Based on the principle of contrastive learning, structural constraints are imposed on the embedding space from a cross-example perspective. Global supervision signals are used to promote intra-class compactness and inter-class separability, thereby enhancing the clarity and separability of the overall representation space. For a batch of N samples, its Standardized embedding is , and its corresponding tags are We construct a similarity matrix Cross-example embedding regularization loss Defined as:

[0097]

[0098] in It is a temperature over-parameter and needs to be manually adjusted to a value of 0.1. It is an indicator function, only when... The value is 1 when the condition is met, and 0 otherwise.

[0099] By jointly minimizing the classification loss and cross-example embedding regularization loss As a loss function, Using the standard cross-entropy loss function:

[0100]

[0101] The final loss function is a weighted sum of the two types of losses:

[0102]

[0103] in, Indicates the predicted label, Indicates the true label, It is a weighting factor used to control Impact on the overall objective. Minimize. It can enhance the robustness of multi-peak representation and the model's generalization ability in classification tasks.

[0104] Step S6: Input the dual-modal fusion feature vector into the classifier, output the direction-of-arrival probability distribution of the target through a fully connected layer, and select the direction with the highest probability as the final direction-of-arrival estimation result.

[0105] It should be further noted that in this embodiment, steps S3-S6 are achieved by constructing and training a dual-modal sound source direction estimation model based on the Touchformer architecture to obtain the final direction of arrival estimation result.

[0106] Specifically, such as Figure 1 and Figure 2As shown, the dual-modal sound source direction estimation model includes a one-dimensional convolutional feature extraction module, a modality adaptive gating module, a cross-modal and same-modal attention fusion module, a cross-example embedding regularization module, and a classifier.

[0107] First, a one-dimensional convolutional feature extraction module is used to extract the representation features of the two modalities. Then, a modality adaptive gating module is used to assign weights according to the reliability of the modal features. Cross-modal attention is used to inject the source modal features into the target modal features to obtain the target cross-modal features in the two modal domains. Then, self-attention is performed on the target cross-modal features to obtain their respective fused modal features. After the fused modal features of the two modalities are fused, a cross-example embedding regularization module is used to apply structural constraints to the embedding space from the perspective of cross-examples. Finally, a classifier is used to perform classification output.

[0108] Furthermore, the training method for the dual-modal sound source direction estimation model is as follows:

[0109] The preprocessed training set samples are input into a dual-modal sound source direction estimation network to train the network. The network weights are updated using the Adamy optimizer, which is as follows:

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116] in Represents the loss function Regarding model parameters The gradient at time step t. This represents the gradient operator. Represents initialization to 0 The first moment estimate. Represents initialization to 0 The second moment estimate. The exponential decay rate is estimated by the first moment and is set to 0.9. The exponential decay rate estimated by the second moment is 0.9. This is the first-order moment estimate after bias correction. This is the second-order moment estimate after bias correction. The numerical stability constant is set to 10e-8. The learning rate is represented by a cosine annealing strategy to dynamically adjust it. The learning rate changes according to a cosine function curve, slowly decreasing from its initial value to 0. Its formula can be simplified to:

[0117]

[0118] in The initial learning rate is 0.1. =0, This represents the current iteration number. This is the total number of iterations (50 epochs), and the batch size is set to 32.

[0119] Finally, the preprocessed test set data is sequentially input into the trained dual-modal sound target direction of arrival estimation model to obtain the recognition result of each sample in the test set.

[0120] Specifically, the fused features Z are input into a fully connected layer for orientation classification. The number of nodes in the output layer corresponds to the number of orientations in the orientation estimation task (36 possibilities). Finally, the probability distribution of each orientation is output through a softmax function, and the orientation with the highest probability is selected as the final prediction result.

[0121] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for estimating the direction of arrival (DOA) of a dual-mode acoustic target, characterized in that, Includes the following steps: Step S1: Construct a dual-mode signal acquisition hardware platform to simultaneously acquire the electromagnetic echo signal and acoustic signal of the target; Step S2: Perform orthogonal demodulation on the acquired electromagnetic echo signal to obtain the electromagnetic mode phase matrix; perform time-frequency transformation on the acquired multi-channel acoustic signal, extract the phase spectrum, and obtain the acoustic mode phase matrix. Step S3: Input the electromagnetic mode phase matrix and the acoustic mode phase matrix into the feature extraction module respectively to obtain the electromagnetic feature vector and the acoustic feature vector; Step S4: Input the electromagnetic feature vector and acoustic feature vector into the modal adaptive gating module, generate modal importance weights based on the reliability of each mode, and use the weights to adjust the feature vectors. At the same time, fuse position encoding information and output the target modal features. Step S5: Using a cross-modal attention mechanism, the target modal features of another modality are injected into the target modal features of the current modality as the source modal features of the current modality to obtain a dual-modal fusion feature vector; Step S6: Input the dual-modal fusion feature vector into the classifier, output the direction-of-arrival probability distribution of the target, and select the direction with the highest probability as the final direction-of-arrival estimation result.

2. The method according to claim 1, characterized in that, In step S1, the dual-mode signal acquisition hardware platform includes an electromagnetic mode acquisition module and an acoustic mode acquisition module; The electromagnetic mode acquisition module includes an adjustable metasurface, a transmitting and receiving antenna, and a modulation and demodulation processor, used to acquire and demodulate the phase information of the target echo; The acoustic modal acquisition module includes an array of multiple linearly arranged microphones for synchronously acquiring multi-channel acoustic signals.

3. The method according to claim 1, characterized in that, Step S2 specifically includes: The electromagnetic echo signal is quadrature demodulated, the phase sequence of each channel is extracted, and the electromagnetic mode phase matrix is ​​constructed. The multi-channel acoustic signals are framed, windowed, and subjected to short-time Fourier transform to extract phase information at each frequency point and construct the acoustic mode phase matrix.

4. The method according to claim 3, characterized in that, The electromagnetic mode phase matrix and acoustic mode phase matrix are paired with the corresponding true direction-of-arrival angle labels to construct a dataset, which is then divided into a training set and a test set in a 7:3 ratio.

5. The method according to claim 1, characterized in that, The modal adaptive gating module in step S4 performs the following operations: The input feature vector is subjected to linear transformation and nonlinear activation to generate intermediate feature representations. The intermediate feature representation is then linearly transformed again and activated by the Sigmoid function to generate initial modal gating weights; A preset confidence threshold is introduced to suppress modal information corresponding to initial gating weights that are below the threshold. The retained initial gating weights are Softmax normalized to obtain the final modal importance weights; The modality importance weights are used to perform weighted fusion of the feature vectors and positional codes of the corresponding modalities, and the adjusted target modality features are output.

6. The method according to claim 1, characterized in that, Step S5 includes the following sub-steps: Step S5-1: Calculate the query vector, key vector, and value vector of the source modality and the target modality respectively through the cross-modal attention mechanism, and calculate the cross-modal attention output based on the query vector, key vector, and value vector to realize the injection of source modality features into target modality features; Step S5-2: The cross-modal attention output is weighted and adjusted using the modal importance weights generated by the modal adaptive gating module, and the weighted cross-modal attention output is merged with the target modal features through residual connection to obtain the cross-modal enhanced target modal features; Step S5-3: Perform self-attention calculation on the cross-modal enhanced target modal features, extract the structural information inside the modality, and obtain the fused modal feature representation; Step S5-4: Reweight the fusion features obtained in step S5-3 according to the importance weights of the source mode and the target mode, and concatenate the weighted features of the source mode and the target mode to form the final dual-modal fusion feature vector.

7. The method according to claim 1, characterized in that, The model training process also includes the step of constructing the loss function: Calculate the cross-entropy loss for a classification task; Calculate the cross-example embedding regularization loss, which is obtained by imposing intra-class compactness and inter-class separability constraints on the feature embedding space; The cross-entropy loss and the cross-example embedding regularization loss are added together with preset weights to form the total loss function, which is used to optimize the model parameters.

8. The method according to claim 7, characterized in that, The cross-example embedding regularization loss is calculated as follows: Normalize the feature embeddings of a batch of samples; Calculate the feature similarity matrix between each pair of samples; Based on the similarity matrix and sample labels, a loss term is constructed to make the features of samples of the same class similar and the features of samples of different classes different.

9. The method according to claim 1, characterized in that, The feature extraction module in step S3 uses a one-dimensional convolutional neural network to perform convolutional encoding on the electromagnetic mode phase matrix and the acoustic mode phase matrix respectively, and maps the encoded features to a unified dimension.

10. The method according to claim 1, characterized in that, The method includes constructing a dual-modal sound source direction estimation model based on the Touchformer architecture, and implementing steps S3-S6 through the dual-modal sound source direction estimation model.

11. The method according to claim 10, characterized in that, The dual-modal sound source direction estimation model is trained using a training set and tested using a test set. During the training phase, an adaptive moment estimation algorithm is used to dynamically update the model parameters, and a cosine annealing strategy is used to adjust the learning rate.

12. The method according to claim 1, characterized in that, The classifier in step S6 is a fully connected neural network layer, whose output node count is the same as the number of directions of arrival categories to be estimated, and finally outputs the probability of each direction category through the Softmax function.