Sound source localization method based on NV-SRP neural network model

By employing an NV-SRP neural network-based sound source localization method, which utilizes acoustic vector sensors and deep learning technology, and combines sound pressure, sound velocity, and array geometry information, the robustness and accuracy issues of traditional sound source localization methods in complex acoustic environments are resolved, achieving higher localization accuracy and stability.

CN121995309APending Publication Date: 2026-05-08CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF GEOSCIENCES (WUHAN)
Filing Date
2025-12-03
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Traditional sound source localization methods are not robust enough in low signal-to-noise ratio and strong reverberation environments, resulting in poor localization accuracy and stability.

Method used

A sound source localization method based on the NV-SRP neural network model is adopted. Training and test datasets are constructed using an acoustic vector sensor signal receiving model and the SRP-PHAT algorithm. High-level representations are extracted from the time delay features of sensor pairs through a deep spatiotemporal network. Combined with sound pressure, sound velocity and array geometry information, a flexible multi-stage fusion framework is designed for sound source localization.

Benefits of technology

It significantly improves the accuracy and robustness of sound source localization in environments with strong reverberation and high noise, providing a more effective acoustic sensing solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121995309A_ABST
    Figure CN121995309A_ABST
Patent Text Reader

Abstract

The invention discloses a sound source localization method based on an NV-SRP neural network model, and relates to the technical field of sound source localization, and the sound source localization method based on the NV-SRP neural network model mainly comprises the steps: constructing a training and testing data set according to an acoustic vector sensor signal receiving model and an SRP-PHAT algorithm, and training the sound source localization network to obtain a trained sound source localization network, and predicting target sound source localization to obtain a prediction result. By implementing the sound source localization method based on the NV-SRP neural network model provided by the invention, the sound source localization precision and stability can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound source localization technology, and more specifically, to a sound source localization method based on the NV-SRP neural network model. Background Technology

[0002] Traditional sound source localization methods, such as the Generalized Cross-Correlation (GCC) algorithm based on Time Difference of Arrival (TDOA), the SRP-PHAT method, and the high-resolution algorithm based on subspace (MUSIC, ESPRIT), are theoretically mature and perform well in ideal acoustic environments. However, these methods exhibit a sharp decline in performance and lack robustness when faced with harsh acoustic environments such as low signal-to-noise ratio (SNR) and strong reverberation.

[0003] While deep learning techniques have been introduced to improve the robustness of sound source localization and some progress has been made by combining neural networks with traditional frameworks (such as Neural-SRP), most of these methods still rely on traditional microphone arrays that can only acquire scalar sound pressure information. This physical limitation fundamentally restricts the dimensions of the sound field information that can be acquired. When the sound pressure signal is severely contaminated by reverberation and noise, the potential for improving algorithm performance remains limited.

[0004] Improving positioning accuracy and stability in environments with strong reverberation and high background noise is a pressing technical problem that needs to be solved. Summary of the Invention

[0005] The purpose of this invention is to provide a sound source localization method based on the NV-SRP neural network model, which can provide sound source localization accuracy and stability.

[0006] This invention provides 1. a sound source localization method based on an NV-SRP neural network model, characterized by comprising the following steps: S1: Construct training and testing datasets based on the acoustic vector sensor signal reception model and the SRP-PHAT algorithm; S2: Construct a sound source localization network, and train the sound source localization network using the training and testing datasets to obtain a trained sound source localization network; S3: Use the trained sound source localization network to predict the localization of the target sound source and obtain the prediction result.

[0007] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the sound source localization method based on the NV-SRP neural network model described above.

[0008] The sound source localization method based on the NV-SRP neural network model provided by this invention has the following beneficial effects: This invention addresses the shortcomings of existing technologies in terms of positioning accuracy and robustness in complex acoustic environments. Based on an acoustic vector sensor signal receiving model and the SRP-PHAT algorithm, training and testing datasets are constructed to train a sound source localization network. The network then predicts the location of the target sound source. This invention utilizes acoustic vector sensors for robust direction of arrival (DOA) estimation in complex acoustic environments. The sound source localization network model first extracts high-level representations from the temporal delay features of sensor pairs through a deep spatiotemporal network. This includes a flexible multi-stage fusion framework that not only explicitly integrates the physical geometry information of the sensor array into the feature learning process but also adaptively combines the original acoustic velocity vector with deep audio features. By synergistically utilizing sound pressure, sound velocity, and array geometry, the model significantly improves positioning accuracy and robustness in challenging scenarios such as strong reverberation and high noise, providing a more effective and complete solution for acoustic perception based on acoustic vector sensors. This invention can be applied to various fields requiring precise spatial positioning of sound sources, such as robot hearing, speech recognition, intelligent conferencing systems, acoustic monitoring, immersive communication, and UAV airborne perception. Attached Figure Description

[0009] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of the sound source localization method based on the NV-SRP neural network model provided by the present invention; Figure 2 This is a flowchart of the NV-SRP algorithm provided by the present invention; Figure 3 This is a diagram of the NV-SRP network architecture provided by the present invention; Figure 4 This is a flowchart of the vector fusion module structure provided by the present invention; Figure 5 This is a diagram of the NV-SRP module architecture provided by the present invention; Figure 6 This is a schematic diagram illustrating the effect of signal-to-noise ratio on the positioning error of different SRP variants under a reverberation time RT60=0.4s, as provided by the present invention. Figure 7 This is a schematic diagram illustrating the impact of reverberation time and channel configuration on the positioning error of the SRP-based method under a signal-to-noise ratio (SNR) of 20dB, as provided by the present invention. Figure 8 This is a sound source location coordinate diagram provided by the present invention; Figure 9 These are real experimental scene diagrams provided by this invention; Figure 10 This is the SRP-PHAT three-dimensional thermal map provided by the present invention. Detailed Implementation

[0010] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0011] Figure 1 A schematic diagram of the sound source localization method based on the NV-SRP neural network model in this embodiment is shown. In this embodiment, the sound source localization method based on the NV-SRP neural network model includes the following steps: S1: Construct training and testing datasets based on the acoustic vector sensor signal reception model and the SRP-PHAT algorithm.

[0012] In one exemplary embodiment, the acoustic vector sensor signal receiving model is as follows:

[0013]

[0014]

[0015] in, t For time indexing; It is the 4D output vector of the acoustic vector sensor, containing a 1D sound pressure signal. and 3D particle velocity signal The acoustic vector sensor includes an omnidirectional sound pressure sensor and three orthogonal dipole sensors; I is the number of sound sources in the acoustic environment; It is the source signal of the i-th sound source. Let be the 4D room impulse response of the source signal of the i-th sound source; It is additive noise; and It is the short-time Fourier transform domain representation of the received signal and the i-th source signal, where k is the time frame index and n is the frequency bin index; This indicates that the acoustic transfer function (ATF) between the i-th source and each sensor in the AVS is modeled. Indicates modeling error; It is the short-time Fourier transform domain representation of the noise signal; This represents the direct sound path response of the i-th sound source under the far-field assumption; Where K is the discrete angular frequency, and K is the FFT transform length; This indicates the sampling delay of the direct path impulse response relative to the first sample of frame 0; It is determined by the azimuth angle and pitch angle The direction vector of the i-th sound source is determined.

[0016] In one exemplary embodiment, the training and test datasets are generated using a simulation strategy.

[0017] S2: Construct a sound source localization network, and train the sound source localization network using the training and testing datasets to obtain a trained sound source localization network.

[0018] In one exemplary embodiment, the sound source localization network includes paired networks and a global decoder, as shown in the formula:

[0019] in, This represents the output of the sound source localization network; Indicates a global decoder; This represents a pairwise network; m and l represent the m-th and l-th elements in the microphone array, respectively; M represents the number of elements in the microphone array. Indicates microphone pair ( m , l The generalized cross-correlation characteristics of ) For microphone pairs ( m , l The corresponding sensor position coordinates.

[0020] In one exemplary embodiment, the pairwise networks employ a parameter-sharing convolutional recurrent neural network architecture; the pairwise networks are configured as follows: The generalized cross-correlation features of each microphone pair are processed using a three-layer two-dimensional convolutional block. The convolutional kernel of the three-layer two-dimensional convolutional block has a fixed temporal dimension size of 1 and does not perform temporal pooling. The convolution stride is a unit value. The processed features are flattened and then temporal features are extracted using a bidirectional gated cyclic unit. The microphone pairs ( m , l The corresponding sensor position coordinates are fused with the temporal features and converted into a unified-dimensional spatial likelihood code using a multilayer perceptron.

[0021] In one exemplary embodiment, the global decoder is configured as follows: The spatial likelihood codes of all microphone pairs are weighted and summed to obtain the weighted sum result; Based on the original multi-channel signals and the weighted summation results, the fusion features are obtained using the acoustic vector fusion module; Based on the fusion features, the activity probability and three-dimensional Cartesian coordinates of each sound source are obtained using the activity detection branch and the position regression branch.

[0022] In one exemplary embodiment, the weighted summation is as follows:

[0023] in, This is the weighted summation result; These are learnable weights; For microphone pairs ( m , l Spatial likelihood encoding.

[0024] In one exemplary embodiment, the activity detection branch outputs the activity probability of all sound sources, and the position regression branch outputs the three-dimensional Cartesian coordinates of all sound sources.

[0025] In one exemplary embodiment, the acoustic vector fusion module is configured as follows: Separating three-dimensional velocity components from the original multi-channel signal Calculate the three-dimensional velocity components Four-dimensional statistical characteristics are obtained from the spatial average value and velocity norm in the sensor dimension. ; Based on the aforementioned 4-dimensional statistical characteristics Deep velocity representations are extracted using cascaded 1D convolutional layers, aligned to the time dimension T via linear interpolation, and projected onto the feature space to obtain velocity features. ; Using a two-stream gating mechanism, first process the audio feature tensor... and the speed characteristics After performing layer normalization and then concatenating the layers, the concatenated features are obtained. ; Based on the spliced ​​features Calculate the gating coefficient; The final fusion feature is obtained based on the gating coefficient.

[0026] In one exemplary embodiment, the formula for the gating coefficient is:

[0027] in, The gating coefficient; Use the Sigmoid activation function; Features after splicing; This is the weight matrix for the linear transformation; This is a bias term.

[0028] In one exemplary embodiment, the calculation formula for the fusion feature is:

[0029] in, Features of fusion; For audio feature tensors; This indicates element-wise multiplication; These are three-dimensional velocity components.

[0030] It should be noted that the audio feature tensor is the result of the weighted summation described above.

[0031] In one exemplary embodiment, the loss function of the sound source localization network is:

[0032] in, The loss function; Euclidean positioning error coefficients weighted by the activity state of the sound source; The activity status of the target sound source; The direction of arrival matrix of the target sound source; Output matrix for the network; The binary cross-entropy activity detection loss coefficient; Represents the binary cross-entropy; This indicates the activity status of the output sound source.

[0033] S3: Use the trained sound source localization network to predict the localization of the target sound source and obtain the prediction result.

[0034] In one exemplary embodiment, the sound source localization method based on the NV-SRP neural network model further includes: evaluating the prediction results using spatial deviation.

[0035] In one exemplary embodiment, the formula for calculating the spatial deviation is:

[0036] in, Spatial deviation; , , and These represent the actual azimuth and elevation angles, and the predicted azimuth and elevation angles, respectively.

[0037] In some embodiments, the sound source localization method based on the NV-SRP neural network model described above can also be implemented in the following ways.

[0038] In this embodiment, the sound source localization network first extracts high-level representations from the temporal delay features of the sensor pair using a deep spatiotemporal network. Its core innovation lies in a flexible, multi-stage fusion framework: it not only explicitly integrates the physical geometry information of the sensor array into the feature learning process but also designs a dedicated dynamic fusion module to adaptively combine the original acoustic velocity vector with deep audio features. By synergistically utilizing sound pressure, sound velocity, and array geometry, this method aims to significantly improve the model's localization accuracy and robustness in challenging scenarios such as strong reverberation and high noise, providing a more effective and complete solution for acoustic perception based on acoustic vector sensors.

[0039] In this embodiment, the sound source localization method based on the NV-SRP neural network model includes: (1) First, establish a signal receiving model for the acoustic vector sensor to provide a theoretical basis for subsequent processing; then, construct training and testing datasets containing simulation and measured data in parallel. (2) Design a deep learning-based NV-SRP neural network model and perform simulation verification and performance evaluation of the above algorithm; (3) Based on the simulation results, conduct experimental design and build an airborne array hardware system for UAVs; use the platform to conduct experimental tests, collect data and perform processing and performance analysis.

[0040] The algorithm's innovations and flowchart are as follows: Figure 2 As shown.

[0041] 1. Acoustic Vector Sensor Signal Model Assuming there are I sound sources in the acoustic environment, the received signal model of an acoustic vector sensor (AVS) consisting of an omnidirectional sound pressure sensor and three orthogonal dipole sensors can be modeled as follows: (1) in, t For time indexing; It is the 4-dimensional output vector of AVS, which contains a 1-dimensional sound pressure signal. and 3D particle velocity signal ; It is the source signal of the i-th sound source. Its corresponding 4-dimensional room impulse response (RIR); This is additive noise; to perform frequency domain analysis, this embodiment converts the signal to the time-frequency domain using a short-time Fourier transform (STFT): (2) in and It is the STFT field representation of the received signal and the i-th source signal, where k is the time frame index and n is the frequency bin index; This indicates that the acoustic transfer function (ATF) between the i-th source and each sensor in the AVS is modeled. Indicates modeling error. This is the STFT domain representation of the noise signal; in many localization algorithms, the direct sound path component plays a decisive role; therefore, under the far-field assumption, the direct sound path response of the i-th sound source can be approximated as: (3) In the formula: Where K is the discrete angular frequency, and K is the FFT transform length; This indicates the sampling delay of the direct path impulse response relative to the first sample of frame 0; It is determined by the azimuth angle and pitch angle The determined direction vector of the sound source.

[0042] 2.SRP-PHAT (Steered Response Power with Phase Transform) algorithm SRP-PHAT is a classic method in the field of sound source localization. It first calculates the generalized cross-correlation (GCC) function of the signals received by any two microphones m and l in the array. To enhance robustness against noise and reverberation, phase transform (PHAT) weighting is typically used, and its GCC-PHAT spectrum can be expressed as: (4) in, It is the frequency domain representation of the m-th microphone signal; This indicates taking the complex conjugate. This is the time delay between the two microphones; subsequently, SRP-PHAT constructs the spatial spectrum P(q) by scanning all possible spatial directions q and aggregating the cross-correlation energy of all microphone pairs: (5) Ultimately, the direction corresponding to the peak of the spatial spectrum is the estimated location of the sound source: (6) 3. NV-SRP Model The sound source localization network proposed in this embodiment adopts a Neural Vector SRP (Neural Vector Steered Response Power, NV-SRP) architecture, such as... Figure 3 As shown, it consists of pairwise networks and global decoder composition: (7) in Indicates microphone pair ( m , l The generalized cross-correlation characteristics of ) For the corresponding sensor position coordinates; pairwise network Employing a parameter-sharing convolutional recurrent neural network (CRNN) architecture: GCC features for each microphone pair First, the data is processed using three layers of two-dimensional convolutional blocks. To maintain temporal causality, the convolutional kernels have a fixed size of 1 in the temporal dimension and no temporal pooling is performed. The convolutional stride is a single value. After flattening, the obtained features are extracted using a bidirectional gated recurrent unit (GRU) to extract temporal features, while simultaneously incorporating microphone coordinate information. (Including normalized position difference vector and spacing scalar) is fused with GRU output features—in the later fusion mode, it is concatenated with GRU output features, and finally converted into a unified-dimensional spatial likelihood code through MLP (Multilayer Perceptron). .

[0043] Global Decoder First, a weighted sum is calculated for the spatial likelihood codes of all microphone pairs: (8) The weights are then processed by a dedicated vector fusion module to extract the velocity components from the original signal. Calculate the average vibration velocity sum norm Statistics are fused with spatial likelihood and velocity features through an adaptive gating mechanism; the final result is generated by a dual-branch output structure: the activity detection branch output. I The activity probability of each sound source, output of the position regression branch. I The three-dimensional Cartesian coordinates of each sound source are implicitly related—when the activity detection indicates that a sound source does not exist, the output at its corresponding position automatically becomes invalid; in the implementation, dynamic dimension processing is used to adapt to different array sizes, the sensor reuses effective features to avoid zero-filling when the number of sensors is insufficient, and training stability is ensured through layer normalization, gradient clipping, and layered dropout.

[0044] This embodiment proposes an innovative Vector Fusion Module (VFM) to address the multi-channel characteristics of acoustic vector sensors. This module, through a gated residual mechanism, achieves adaptive fusion of sound pressure features and three-dimensional velocity vectors, significantly enhancing the model's perception capability in complex sound field environments. Figure 4 As shown, this module uses audio feature tensors The original multi-channel signal is used as input and processed in four stages: First, the three-dimensional velocity components are separated from the four-channel signal. Calculate its spatial average value and velocity norm in the sensor dimension to form a 4-dimensional statistical feature. ; Secondly, deep velocity representations are extracted through cascaded 1D convolutional layers, aligned to the time dimension T via linear interpolation, and projected onto the feature space to obtain... The crucial fusion process employs a dual-stream gating mechanism, first processing audio features... and velocity characteristics After performing layer normalization and then splicing, we get Then calculate the gating coefficient: (9) It is the weight matrix of the linear transformation. It is a bias term; The Sigmoid activation function compresses the output to the range [0,1], representing the gating strength; the final fused features are constructed using a residual structure. (10) This design preserves the core audio information while introducing adaptively adjusted speed features; the feature enhancement layer expands the fusion result to a higher dimension, significantly improving the expressive power of the features.

[0045] In the feature processing flow, this module implements deep encoding and adaptive fusion of the three-dimensional velocity field: when velocity features are highly correlated with audio features (such as in direct sound-dominated scenarios), the gating coefficient... This allows velocity characteristics to fully participate in the fusion process; when the velocity characteristic noise is high (such as in a strong reverberation environment), the gating coefficient... Automatic suppression of interference signals; residual connections ensure that the integrity of the audio backbone is preserved even in cases of feature conflict.

[0046] 4. Simulation verification of NV-SRP model This embodiment employs an end-to-end joint optimization strategy to train the pairwise network. With global decoder ;definition The target DOA matrix (each column is a unit vector representing the direction of the sound source). The network output matrix, and The target and the output are respectively A 3D activity state vector; the loss function design considers the physical characteristics of the sound source localization task and the training stability requirements: in the case of a single sound source, the loss function is defined as: (11) The first item is the sound source activity status. The weighted Euclidean positioning error, with the second term being the binary cross-entropy activity detection loss; consistent with the latest research findings, the Euclidean distance loss converges faster and improves positioning accuracy than the spherical coordinate angle error, which stems from the more stable gradient propagation characteristic in the Cartesian coordinate system.

[0047] This embodiment trains a network with reverberation time and SNR uniformly distributed from 0.2 s to 1 s and from 5 dB to 30 dB, respectively. However, this embodiment found that training converges faster as the signal-to-noise ratio increases. Therefore, this embodiment follows a course learning strategy, using a simulation dataset with an SNR of 30 dB and a duration of four hours for the first 20 iterations. In the subsequent iterations, this embodiment uses a simulation dataset with a full range of SNRs and reduces the learning rate from 1e-4 to 1e-5.

[0048] This study aims to comprehensively evaluate the performance of the proposed Neural Vector SRP (NV-SRP) model under different acoustic conditions through a series of simulation experiments. To ensure the rigor of the evaluation, this study compares NV-SRP with three representative baseline methods, including the traditional SRP-PHAT algorithm and two advanced deep learning localization methods—Neural-SRP and Cross3D. All experiments are conducted under diverse signal-to-noise ratio (SNR) and reverberation time (RT60) conditions to systematically test the robustness of each method. To ensure fairness, the architectural parameters and training procedures of all baseline models are reproduced in accordance with the descriptions in their respective original papers.

[0049] Figure 5 The complete Neural Vector SRP network architecture is demonstrated. This architecture begins by receiving multi-sensor acoustic feature input from the left side, followed by three levels of core convolution processing (the first level uses a 64-channel 3×3 convolution kernel, the second level expands to 128 channels, and the third level boosts to 256 channels). Each level is equipped with 2×2 max pooling and PReLU activation functions for feature compression and nonlinear enhancement. The 256-dimensional features then enter a bidirectional GRU module (configured with 128 hidden units) to capture temporal dependencies. Its output is flattened and fused with location metadata, then processed by a feedforward neural network (FNN) for the 128-dimensional features. The fused 128-dimensional features are input into an innovative vector fusion module, which achieves adaptive sound pressure-velocity fusion through spatial statistics extraction and gated residual mechanisms, outputting 256-dimensional enhanced features. Finally, a multilayer perceptron (MLP) simultaneously generates the 3D Cartesian coordinate system localization result and the sound source existence probability. All networks are implemented using the PyTorch library; the Adam optimizer is used for backpropagation.

[0050] The entire network model is implemented using the PyTorch deep learning framework and trained end-to-end using the Adam optimizer. For the Cross3D baseline that requires a spatial spectral map as input, this embodiment adopts the same configuration as the original paper: first, a 64×32 SRP-PHAT rectangular grid (corresponding to the azimuth and elevation angles respectively) is generated, and this spectral map is used as the input to its network.

[0051] 4.1 Training Dataset This embodiment employs a simulation strategy to generate training data to address the scarcity of labeled data for real moving sound sources. The training data is constructed based on the LibriSpeech corpus, which contains 960 hours of clear speech signals with a sampling rate of 16 kHz. These signals are sampled from audiobooks in the LibriVox project, covering various accents and reading styles. It is a large, publicly available dataset widely used in speech recognition research. To avoid background noise interfering with model learning, a WebRTC speech activity detector (VAD) is used to identify silent segments and completely remove noise components, ensuring the realism of the acoustic scene.

[0052] The acoustic environment parameters were sampled randomly according to a uniform distribution. The sampling rate used for simulation was 16 kHz, and the total duration of the simulation signal used for training was 4 hours. The specific parameter ranges are shown in Table 1. Table 1: Parameter range of the simulated dataset

[0053] This embodiment uses the gpuRIR Python library to achieve efficient sound field simulation. This library is based on the Image Source Method (ISM) principle and computes the Room Impulse Response (RIR) in parallel on the GPU. Compared with the traditional CPU implementation, gpuRIR improves the computational efficiency by two orders of magnitude through the GPU parallel architecture, reducing the simulation time of a single trajectory from minutes to milliseconds, thus meeting the data generation requirements during training. The core of this library lies in the efficient modeling of moving sound sources: by dynamically updating the room impulse response between the sound source and the microphone through GPU parallel computation, it avoids the time discretization error of traditional methods.

[0054] At the signal processing level, gpuRIR uses the overlapping addition method to achieve fast frequency domain convolution. By utilizing the parallel FFT computing power of the GPU, the speed of 1024-point convolution is increased by 120 times, which greatly optimizes the processing efficiency of long sequence signals. The acoustic parameter modeling follows physical laws: the sound absorption coefficient of the wall independently controls the reflection characteristics of the six walls, and the reverberation time is accurately mapped to ensure the physical accuracy of the acoustic simulation.

[0055] The simulation system has the limitation of lacking directional noise, and directional noise sources in real-world scenarios (such as fans) can cause the model to mislocalize. To address this, a two-stage processing mechanism is introduced: first, a WebRTC speech activity detector is used to identify silent frames and force the corresponding feature maps to zero; second, isotropic Gaussian noise is added, and the noise power is dynamically controlled according to the signal-to-noise ratio requirements. This scheme improves the simulation efficiency by orders of magnitude.

[0056] 4.2 Evaluation Indicators The primary metric used in this experiment is the root mean square angle error (RMSAE), which uses spherical geometry principles to accurately determine the spatial deviation between the predicted direction and the true orientation; it is defined as a pair of positions... Each position has an azimuth angle and an elevation angle. and This metric avoids the dimensional bias of traditional Euclidean distance, quantifying the chord distance between two points on a unit sphere from a three-dimensional spatial perspective, as shown below: (12) (18) is the average value of all frames in the dataset; the algorithm performs acoustic energy weighted averaging on the azimuth deviation of all valid frames (areas with sound source activity): the positioning error of high acoustic energy frames is given higher weight, which accurately reflects the correlation between sound source discernibility and signal strength in actual auditory perception.

[0057] 4.3 Experimental Results Experiment 1: LibriSpeech Simulation Data In this experiment, this embodiment evaluates the performance of traditional SRP, Cross 3D, Neural-SRP, and Neural Vector SRP with independent white Gaussian noise (WGN) added to each sensor. The experiment was trained and tested on a pseudo-spherical array of 12 microphones; the specific microphone coordinates can be found in the benchmark2 array of task 1 in the LOCATA dataset. The neural model uses 10... -4 The learning rate was used to train for 80 epochs; this embodiment used a frame size of 256 ms and a hop size of 192 ms; all three networks were trained using the SNR range and reverberation time range defined in Table 1 and tested using a simulated dataset of unseen source signals from the Librispeech test set.

[0058] Figure 6 The figure shows the angular errors of four sound source localization methods under different signal-to-noise ratios (SNR) and reverberation times (RT60); the figure clearly reveals the impact of noise and reverberation on the performance of various algorithms.

[0059] The subplot above, with a fixed reverberation time of 0.4 seconds, shows the trend of localization error as the signal-to-noise ratio (SNR) changes from 5 dB to 30 dB. The graph shows that the performance of all methods improves with increasing SNR. However, different methods exhibit significant differences in their sensitivity to noise. The traditional SRP algorithm (green dashed line) performs the worst under low SNR conditions, with errors far exceeding those of other methods, indicating its high sensitivity to additive noise. In contrast, the three neural network-based methods demonstrate stronger noise robustness. Cross3D (orange dashed line) and NSRP (blue dotted line) show very similar performance curves, with significantly lower errors than SRP. Notably, the NVSRP method proposed in this embodiment (red solid line) exhibits the lowest localization error across all tested SNR ranges. Especially under the harsh conditions of a low SNR of 5 dB, NVSRP's advantages are particularly evident, demonstrating its superior performance and stability in noisy environments.

[0060] The subplot below, with the signal-to-noise ratio fixed at 20 dB, evaluates the change in localization error as the reverberation time increases from 0.2 seconds to 1.0 seconds. The results show that reverberation poses a challenge to all algorithms, with longer reverberation times leading to higher localization errors. The performance of the SRP algorithm deteriorates sharply with increasing reverberation time, confirming its inherent weakness in strong reverberation environments. Cross3D and NSRP outperform SRP, demonstrating some reverberation resistance, but their errors also steadily increase with increasing RT60. In contrast, the NVSRP method again exhibits the best performance; its error curve is the flattest, and even under strong reverberation conditions with RT60 reaching 1.0 seconds, its error growth is less than all other methods. This fully demonstrates that NVSRP has the strongest robustness to reverberation distortion.

[0061] Experiment 2: Comparison of AVS and traditional microphones.

[0062] This experiment delves into the fundamental differences in sound source localization performance between acoustic vector sensors (AVS) and traditional microphones. A unified tetrahedral array architecture (radius 8 cm) was designed and rigorously compared under independent white Gaussian noise (WGN) conditions. Three system configurations were set up: traditional microphone arrays with 4 channels (basic) and 16 channels (high density), and acoustic vector sensor arrays with 4 AVS units (equivalent to 16 channels). All test systems were operated under the same acoustic conditions and trained using the SNR range and reverberation time range defined in Table 1. The test signals were clean speech samples from LibriSpeech that were not used in the training. After being recorded by each array, they were input into three processing flows: traditional SRP, Neural-SRP, and the Neural Vector SRP proposed in this embodiment.

[0063] The experimental procedure strictly controlled variables: all neural models used 10 -4 The learning rate is trained for 80 cycles. In terms of array configuration, traditional microphone arrays adopt a uniform layout scheme, with each microphone independently receiving WGN interference of the same power; while AVS arrays give full play to the advantages of four-dimensional signals, with each unit synchronously acquiring sound pressure and three-dimensional vibration velocity components. Figure 7 The experimental results intuitively reveal the impact of different sensor types and channel numbers on positioning accuracy.

[0064] The subplot above compares the performance of each system as the signal-to-noise ratio changes, with a fixed reverberation time of 0.4 seconds. First, a general trend is that increasing the number of sensors can effectively improve positioning accuracy. For both traditional SRP and Neural-SRP, the 16-channel (dark green / dark blue dashed lines) outperforms the corresponding 4-channel version (light green / light blue dotted lines). Second, the neural network-based method significantly outperforms the traditional SRP algorithm, especially in the low signal-to-noise ratio region, where Neural-SRP shows stronger robustness compared to Classic SRP.

[0065] The most crucial finding is that the NV-SRP method proposed in this embodiment (solid red line), although using only 4 acoustic vector sensors (4-AVS), outperforms the Neural-SRP (16-ch) system using 16 traditional microphones under all signal-to-noise ratio conditions. This means that although both have the same total number of channels input to the neural network of 16, the sound field vector information (sound pressure and three-dimensional vibration velocity) provided by AVS is more effective in suppressing noise interference than simply increasing the number of scalar microphones.

[0066] Figure 7 The subplot below fixes the signal-to-noise ratio at 20dB to evaluate the system's performance at different reverberation times. The results show that the error of all systems increases with the reverberation time (RT60). The traditional SRP method is particularly sensitive to reverberation, and the error grows rapidly. Increasing the number of microphones (from 4 channels to 16 channels) can alleviate the performance degradation caused by reverberation to some extent, but the effect is limited.

[0067] Experiment 3: Microphone Array Sound Source Localization Experiment The experimental setup consists of three parts: a microphone array, a sound source transmitter, and a signal acquisition device. The microphone array receives signals from the spatial sound field, achieving spatial sound field sampling through a reasonable layout and dense arrangement. The sound source transmitter generates stable and controllable sound source signals, ensuring a stable excitation signal during the experiment. The signal acquisition device collects the signals received by the microphones in real time and performs subsequent processing and analysis on the signals using data analysis software to obtain the sound field information required for the experiment. A relatively quiet and open outdoor site was selected to reduce reflection and reverberation. The coordinate system is with east as the x-axis and north as the y-axis. The UAV's center coordinates are at the origin (0, 0), at a height of 5m. The relationship between the UAV and the array is fixed, and subsequent calculations require the UAV's own orientation and altitude information. The sound source location is a 5×5 orthogonal grid system composed of the X-axis (horizontal) and Y-axis (vertical).

[0068] Detailed layout diagram as follows: Figure 8 The experiment used a loudspeaker as the sound source; it played audio signals in Chinese, English, and German, with a duration of 76 seconds. The audio was divided into three segments: the first segment was kept at 70dB, the second segment was increased by 10dB, the third segment was decreased by 10dB, and a 10-second white noise segment was added at the end. These signals were saved to different USB drives and then played through the loudspeaker. Before the experiment, a sound level meter was used to measure the speech at a distance of 1m from the loudspeaker. The volume of the loudspeaker was adjusted, and the setting of the sound level meter at 70dB was recorded and kept constant.

[0069] The experimental procedure is as follows: Before the experiment, the drone was launched with its speakers turned off. A 10-second ambient sound recording was made using a microphone array for reference. Then, a sound level meter was used to measure the volume at a distance of 1 meter from the speaker, which was adjusted to 70 dB. The speaker level was recorded to complete the following experiment: hovering drone; First set of tests: Connect the power supply to the signal acquisition device and set the file save location on the computer; place the sound source (speaker) at the specified coordinates, and insert the audio USB flash drive into the speaker. This experiment uses only one speaker. The drone takes off to a height of 5m and starts the signal acquisition device to collect data. Adjust the speaker to the pre-marked setting, set it to emit a voice signal, and save the data when the audio ends; The drone landed; The acquired files are converted into audio and verified using a pre-prepared algorithm to check the accuracy of the positioning. If the result is significantly off, the speaker volume is increased, the file is re-acquired and verified. If the result is accurate, proceed to the next step. If the deviation is still significant, carefully check other experimental setups for problems.

[0070] Second group of tests: Set the file save location on your computer; place the sound source (speaker) at the specified coordinates, and insert the audio USB drive into the speaker; The drone takes off to a height of 5m and starts the signal acquisition device to collect data. Adjust the speaker to the pre-marked setting, set it to emit a voice signal, and save the data when the audio ends; at the same time, prepare for the next location in advance, including the sound source and the shelf.

[0071] Subsequent group testing: Following the process of the second set of tests, the data collection of 20 coordinates was completed in a loop; the actual experimental image is shown below. Figure 9 As shown; The collected real data is then synthesized into data received by a four-channel acoustic vector sensor, and the results are verified using the algorithm of this invention. Figure 10 As shown; This experiment, based on the SRP-PHAT method, utilizes a microphone array for spatial localization of the sound source. The experimental results show that in the two-dimensional spectrum composed of azimuth and elevation angles and its corresponding three-dimensional heatmap, the energy distribution exhibits a distinct ridge structure with local energy peaks. The color change from dark blue to bright red visually reflects the distribution of spatial spectrum values ​​from low to high. The detection system successfully located the sound source, with an azimuth (Theta) of 38.00° and an elevation (Phi) of -18.77°, and a calculated angle estimation error of 2.9393 degrees. This result demonstrates that the present invention has good localization accuracy under the current experimental configuration and can effectively estimate the direction of the sound source in three-dimensional space.

[0072] This embodiment provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the sound source localization method based on the NV-SRP neural network model described above.

[0073] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A sound source localization method based on an NV-SRP neural network model, characterized in that, Includes the following steps: S1: Construct training and testing datasets based on the acoustic vector sensor signal reception model and the SRP-PHAT algorithm; S2: Construct a sound source localization network, and train the sound source localization network using the training and testing datasets to obtain a trained sound source localization network; S3: Use the trained sound source localization network to predict the localization of the target sound source and obtain the prediction result.

2. The sound source localization method based on the NV-SRP neural network model according to claim 1, characterized in that, The acoustic vector sensor signal receiving model is as follows: in, t For time indexing; It is the 4D output vector of the acoustic vector sensor, containing a 1D sound pressure signal. and 3D particle velocity signal The acoustic vector sensor includes an omnidirectional sound pressure sensor and three orthogonal dipole sensors; I is the number of sound sources in the acoustic environment; It is the source signal of the i-th sound source. Let be the 4D room impulse response of the source signal of the i-th sound source; It is additive noise; and It is the short-time Fourier transform domain representation of the received signal and the i-th source signal, where k is the time frame index and n is the frequency bin index; This indicates that the acoustic transfer function (ATF) between the i-th source and each sensor in the AVS is modeled. Indicates modeling error; It is the short-time Fourier transform domain representation of the noise signal; This represents the direct sound path response of the i-th sound source under the far-field assumption; Where K is the discrete angular frequency, and K is the FFT transform length; This indicates the sampling delay of the direct path impulse response relative to the first sample of frame 0; It is determined by azimuth angle and pitch angle The direction vector of the i-th sound source is determined.

3. The sound source localization method based on the NV-SRP neural network model according to claim 1, characterized in that, The sound source localization network includes paired networks and a global decoder, as shown in the formula: , in, This represents the output of the sound source localization network; Indicates a global decoder; This represents a pairwise network; m and l represent the m-th and l-th elements in the microphone array, respectively; M represents the number of elements in the microphone array. Indicates microphone pair ( m , l The generalized cross-correlation characteristics of ) For microphone pairs ( m , l The corresponding sensor position coordinates.

4. The sound source localization method based on the NV-SRP neural network model according to claim 3, characterized in that, The pairwise network is configured as follows: The generalized cross-correlation features of each microphone pair are processed using a three-layer two-dimensional convolutional block. The convolutional kernel of the three-layer two-dimensional convolutional block has a fixed temporal dimension size of 1 and does not perform temporal pooling. The convolution stride is a unit value. The processed features are flattened and then temporal features are extracted using a bidirectional gated cyclic unit. The microphone pairs ( m , l The corresponding sensor position coordinates are fused with the temporal features and converted into a unified-dimensional spatial likelihood code using a multilayer perceptron.

5. The sound source localization method based on the NV-SRP neural network model according to claim 3, characterized in that, The global decoder is configured as follows: The spatial likelihood codes of all microphone pairs are weighted and summed to obtain the weighted sum result; Based on the original multi-channel signals and the weighted summation results, the fusion features are obtained using the acoustic vector fusion module; Based on the fusion features, the activity probability and three-dimensional Cartesian coordinates of each sound source are obtained using the activity detection branch and the position regression branch.

6. The sound source localization method based on the NV-SRP neural network model according to claim 5, characterized in that, The acoustic vector fusion module is configured as follows: Separating three-dimensional velocity components from the original multi-channel signal Calculate the three-dimensional velocity components Four-dimensional statistical characteristics are obtained from the spatial average value and velocity norm in the sensor dimension. ; Based on the aforementioned 4-dimensional statistical characteristics Deep velocity representations are extracted using cascaded 1D convolutional layers, aligned to the time dimension T via linear interpolation, and projected onto the feature space to obtain velocity features. ; Using a two-stream gating mechanism, first process the audio feature tensor... and the speed characteristics After performing layer normalization and then concatenating the layers, the concatenated features are obtained. ; Based on the spliced ​​features Calculate the gating coefficient; The final fusion feature is obtained based on the gating coefficient.

7. The sound source localization method based on the NV-SRP neural network model according to claim 1, characterized in that, The loss function of the sound source localization network is: , in, The loss function; Euclidean positioning error coefficients weighted by the activity state of the sound source; The activity status of the target sound source; The direction of arrival matrix of the target sound source; Output matrix for the network; The binary cross-entropy activity detection loss coefficient; Represents the binary cross-entropy; This indicates the activity status of the output sound source.

8. The sound source localization method based on the NV-SRP neural network model according to claim 1, characterized in that, The sound source localization method based on the NV-SRP neural network model further includes: evaluating the prediction results using spatial deviation.

9. The sound source localization method based on the NV-SRP neural network model according to claim 8, characterized in that, The formula for calculating the spatial deviation is: in, Spatial deviation; , , and These represent the actual azimuth and elevation angles, and the predicted azimuth and elevation angles, respectively.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the sound source localization method based on the NV-SRP neural network model as described in any one of claims 1-9.