Audio-visual binaural sound source localization method based on pulse neural network
By constructing a large-scale sound source positioning data set and adopting multimodal feature extraction and fusion technology, combining pulse neural networks and attention mechanisms, the noise interference and front-to-back confusion problems of sound source positioning in complex sound fields are solved, and high-precision sound source positioning is achieved.
Patent Information
- Application Number
- CN202510198055.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing sound source positioning technology faces problems of noise, reverberation interference and multi-sound source environment in complex sound fields, resulting in large errors in binaural clues extracted by the network, affecting the effect of sound source positioning. Moreover, it is difficult to determine whether the sound source comes from the front half plane or the back half plane, which limits its application range.
A method of audio-visual binaural sound source positioning based on pulsed neural network is proposed. By constructing a large-scale sound source positioning data set containing the first viewing angle depth map and binaural audio, multimodal feature extraction and fusion technology are adopted, including feature extraction in the frequency domain, time domain and spatial domain, and feature fusion is performed in combination with attention mechanism, and finally the azimuth angle and distance estimation of the sound source is realized through regression prediction.
This method significantly improves the accuracy and robustness of sound source positioning, effectively suppresses noise and reverberation interference in complex sound fields, can accurately locate the sound source in dynamic and multi-sound source environments, and solves the problem of front and back confusion, achieving high-precision prediction of the sound source position.
Smart Images

Figure CN120124463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer auditory technology, and particularly to an audiovisual binaural sound source localization method based on a spiking neural network. Background Art
[0002] Sound source localization refers to the technology of determining the position of a sound source by using the received audio signal. The position of the sound source is generally represented by three coordinates: azimuth, elevation, and distance. In recent years, with the development of intelligent products, the research on sound source localization technology, which can locate a certain sound source in a noisy environment like humans, has become a research hotspot in the fields of acoustics and signal processing. This technology has important practical value in application scenarios such as robot navigation, human-computer interaction, security monitoring, and autonomous driving.
[0003] Most traditional sound source localization technologies are based on microphone arrays. Sound signals are collected by multiple microphones, and methods such as beamforming, time difference estimation, and sound intensity difference are used to estimate the position of the sound source. For example, the widely used Generalized Cross-Correlation Phase Transform (GCC-PHAT) algorithm can achieve high positioning accuracy in a complex sound field by estimating the Time Difference of Arrival (TDOA) of the signal. In addition, beamforming technologies such as MVDR and GSC can enhance the signal in a specific direction while suppressing the interference sound in other directions, improving the reliability of sound source localization. Although these technologies have high positioning accuracy, they often rely on complex array designs and precise microphone calibrations, and the additional microphone arrays also bring additional hardware costs. Secondly, these methods have poor robustness to interference such as environmental noise and reverberation, which limits their application in dynamic and complex environments.
[0004] In contrast, the binaural sound source localization technology simulates the working mode of the human auditory system and relies only on two receivers (i.e., binaural microphones). By analyzing features such as the Interaural Time Difference (ITD) and the Interaural Level Difference (ILD), it realizes sound source localization. The binaural sound source localization technology was first applied to the research of psychoacoustics and auditory modeling, and its algorithm model has gradually developed from the early classical delay estimation to the complex network structure based on deep learning. For example, the weighted model based on ITD and ILD has been widely used in static scenarios, while the Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) proposed in recent years show stronger adaptability in dynamic scenarios. However, binaural sound source localization faces challenges in complex sound fields, such as noise, reverberation interference, and multi-source environments, which makes the binaural cues extracted by the network have large errors and seriously affects the effect of sound source localization. Secondly, since there are only two microphones in binaural sound source localization, there is a problem of front-back confusion in binaural cues. It is very difficult for current methods to determine whether the sound source comes from the front half-plane or the back half-plane. Therefore, current methods usually can only perform sound source localization in the front half-plane.
[0005] With the development of multi-modal perception technology, researchers have gradually found that fusing visual information with audio information can significantly improve the robustness and accuracy of sound source localization. Visual information can provide intuitive spatial cues, such as depth information in the scene, the distribution of obstacles in the scene, and the material conditions, etc. In recent years, the sound source localization method based on audio-visual fusion has been widely studied. Its basic idea is to achieve accurate estimation of the sound source position through feature extraction and joint modeling of audio and visual signals. However, the current multi-modal sound source localization methods still have problems such as insufficient feature extraction and fusion mechanisms, failure to effectively solve the problem of front-back confusion, and failure to fully utilize the spatial cues hidden in visual information, which seriously affects the robustness of the sound source localization method and its application in dynamic and complex environments.
[0006] The quality and diversity of the sound source localization dataset are crucial for the sound source localization method based on audio-visual fusion. However, the current sound source localization datasets are mostly collected in laboratory environments or based on simulations, with relatively small scales, limited covered scenarios and sound source types, inaccurate data annotation, and mostly only contain the auditory modality, lacking surrounding visual information.
[0007] In response to the above problems, the present invention collects a large-scale sound source localization dataset containing first-person depth maps and binaural audio, and proposes an audio-visual binaural localization network (AVBLNet) based on a pulse neural network to estimate the azimuth and distance of the sound source. This method uses binaural audio signals and first-person depth information, adopts a front-to-back information collection mechanism, combines multimodal feature extraction technology in frequency domain, time domain and space domain, performs feature fusion through an attention mechanism, and finally realizes the azimuth and distance estimation of the sound source through regression prediction. Summary of the invention
[0008] Based on the investigation and analysis of the existing sound source prediction technology, the present invention innovatively uses multi-domain feature extraction and fusion to realize the audio-visual sound source localization network AVBLNet, which includes a frequency domain feature extraction module, a time domain feature extraction module, a spatial clue extraction module, a multi-source feature adaptive fusion module and a positioning module. The frequency domain feature extraction module is used to extract the frequency distribution information in binaural audio; the time domain feature extraction module uses a pulse neural network to extract time domain features from binaural audio and mine the temporal dynamic characteristics of the sound; the spatial clue extraction module is used to extract spatial geometric features, such as the geometric structure of the scene, the distribution of obstacles and material properties; the multi-source feature adaptive fusion module strengthens the expression ability of key features by adaptively fusing multimodal features, and effectively suppresses the interference of redundant information; the positioning module combines the fused multimodal features for regression analysis to determine the azimuth and distance of the sound source.
[0009] The technical solution of the present invention is as follows: a method for audio-visual binaural sound source localization based on a pulse neural network, comprising the following steps:
[0010] Step 1: Build an audio-visual sound source localization dataset;
[0011] The sound source position, receiving position and receiving direction are randomly selected from the simulation environment to form an audio-visual pair, which is an audio-visual sound source localization dataset; each audio-visual pair provides a depth image of the front and rear perspectives of the sound receiving position, binaural audio signals of the front and rear perspectives, and the azimuth and distance of the sound source relative to the current position;
[0012] Step 2: Construct the audio-visual sound source localization network AVBLNet;
[0013] The audio-visual sound source localization network includes a frequency domain feature extraction module, a time domain feature extraction module, a spatial clue extraction module, a multi-source feature adaptive fusion module and a localization module;
[0014] The binaural audio signals are respectively subjected to frequency-domain feature extraction module and time-domain feature extraction module to extract frequency-domain features and time-domain features; the depth image is subjected to spatial cue extraction module to extract spatial geometric features;
[0015] The multi-source feature adaptive fusion module effectively fuses the frequency-domain features, time-domain features and spatial geometric features to form the fusion feature F fused ;
[0016] The positioning module is based on the fusion feature F fused , and predicts the azimuth angle and distance
[0017] Step 3: Use the data in the front and back directions respectively to input the audiovisual sound source localization network for sound source prediction training. After the total loss function meets the conditions, the trained audiovisual sound source localization network is obtained for audiovisual binaural sound source localization.
[0018] The frequency-domain feature extraction module is used to extract frequency-domain features from binaural audio signals and capture the frequency distribution information of sounds; input the binaural audio signals in the audiovisual sound source localization dataset, and convert them from the time domain to the frequency domain through Fourier transform to generate the amplitude spectrogram and phase spectrogram of the binaural audio signals; use the pre-trained ResNet network as a feature extractor to extract high-dimensional frequency-domain features F freq .
[0019] The time-domain feature extraction module is used to extract time dynamic characteristics and time correlation features from binaural audio signals to provide timing information for sound source localization; input the binaural audio signals, which are the left ear signal A L (t) and the right ear signal A R (t); the time-domain feature extraction module directly performs pulse coding on the binaural audio signals through short-time window segmentation in the time domain, and then inputs them into the spiking neural network SNN;
[0020] The left ear signal A L (t) and the right ear signal A R (t) are pulse-coded through an encoder to convert the continuous signal into pulse sequences S L (t) and S R (t), representing the pulse excitation patterns of the binaural audio signals in the time dimension; the pulse sequences are input into the spiking neural network, and feature extraction is performed through a series of spiking neuron layers; the output of each layer of neurons depends on the intensity and time of the input pulses, and the specific formula is as follows:
[0021]
[0022] Among them, V i(t) is the membrane potential of neuron i, and W ij represents the connection weight, and φ i is the threshold of the neuron; when the membrane potential V i (t) exceeds the threshold, neuron i fires a pulse S i (t); after being processed by the multi-layer spiking neural network SNN, the time-domain feature extraction module finally outputs the time-domain feature representation F time , including the time dynamic pattern of the binaural audio signal and the binaural time difference information.
[0023] The spatial cue extraction module is used to extract key spatial geometric features from the depth image to provide scene information for audiovisual sound source localization; the input of the spatial cue extraction module is a depth image D∈R H×W , and through the depth image information, it captures the geometric structure, obstacle distribution, and material properties of the scene; the spatial cue extraction module uses a convolutional neural network as the feature extraction backbone;
[0024] First, through a series of convolution and pooling operations, the multi-scale features of the depth image are gradually extracted to capture the geometric information at different levels in the scene; the specific formula for the convolution operation is as follows:
[0025] F k =σ(W k *D + b k )
[0026] where W k is the convolution kernel, b k is the bias term, * represents the convolution operation, σ is the activation function, and F k represents the feature map of the k-th layer;
[0027] The feature map passes through the spatial attention mechanism to generate the spatial weight map A space , which is used to strengthen the important feature regions and suppress redundant information at the same time; the specific formula for the attention weight is as follows:
[0028] A space =Softmax(Conv(f))
[0029] where F is a set of multiple feature maps, Conv represents the convolution operation, and Softmax is used to normalize the weights;
[0030] The final spatial feature is represented by weighted fusion as:
[0031] F space =A space ⊙F
[0032] where ⊙ represents the per-pixel weighted operation, and F space is the spatial geometric feature.
[0033] The multi-source feature adaptive fusion module adaptively allocates weights of different features through an attention mechanism;
[0034] The features input to the multi-source feature adaptive fusion module include frequency domain features Time domain characteristics and spatial geometric features Its three features are aligned and weighted;
[0035] First, the features are transformed into the same dimension through a fully connected layer:
[0036] F aligned,l =α l (F l )
[0037] Among them, l∈freq,time,space,α l represents the alignment operation of the lth type of features, F aligned,l Represents the lth class feature after alignment;
[0038] Then, the attention mechanism is used to adaptively assign fusion weights to each feature; the calculation formula of the fusion weight is as follows:
[0039] Z=Softmax(W a Concat(F aligned,freq ,F aligned,time ,F aligned,space ))
[0040] Where W a is the learnable parameter matrix generated for the attention weights, Concat represents the feature concatenation operation, Z = [z freq ,z time ,z space ] represents the weights of various features; finally, according to the calculated weight matrix Z, the aligned features are weighted and summed to obtain the fused features as follows:
[0041]
[0042] where z l Represents the weight of the lth class feature, fusion feature F fused Integrate multimodal information in frequency, time and space domains.
[0043] The positioning module is a lightweight regression network; through a global average pooling operation, the fusion feature is mapped to a global feature vector f global ∈R C , to compress the spatial dimension and preserve the global information;
[0044]
[0045] The global feature vector is input into two fully connected layers. Each layer uses an activation function to introduce nonlinear capabilities, and Dropout is used to prevent overfitting. Finally, two independent regression branches are designed to predict the azimuth angle. and distance Each regression branch consists of a fully connected layer.
[0046] The total loss function is
[0047]
[0048] Among them, θ,d represents the true azimuth angle and the true distance of the sound source. represents the azimuth and range predicted using the forward data input, indicates the azimuth and range predicted using the backward data input; and It is the weight coefficient used to balance the losses of each part.
[0049] Beneficial effects of the present invention:
[0050] (1) Innovativeness of the method
[0051] The audio-visual sound source localization method (AVBLNet) proposed in the present invention integrates multimodal features in frequency domain, time domain and space domain, and innovatively introduces pulse neural network (SNN) to model time domain features. The network accurately captures the temporal correlation and instantaneous changes of binaural audio signals through the pulse mechanism of bionic neurons, effectively solves the influence of noise and reverberation interference on the sound source localization results in complex sound field environments, and adopts the attention mechanism to realize the adaptive fusion of multi-source features. Specifically, the frequency domain feature extraction module uses Fourier transform combined with deep neural network to extract the spectral features of audio; the time domain feature extraction module uses pulse neural network to capture the temporal dynamic information of sound signals; the spatial clue extraction module extracts geometric spatial features in the depth map through convolutional neural network to provide key visual spatial information. Through the feature fusion of attention mechanism, the correlation between multimodal features is strengthened, the interference of redundant information is suppressed, and the high-precision prediction of the azimuth and distance of the sound source is realized through the positioning module. This method not only improves the accuracy and robustness of sound source localization in complex scenes, but also demonstrates significant technical advantages in multiple links such as feature extraction, fusion and regression analysis, opening up a new path for the application of multimodal fusion technology in sound source localization.
[0052] (2) Information collection mechanism before and after use
[0053] The present invention effectively solves the front-back confusion problem in binaural sound source localization by designing a front-back information collection mechanism. The mechanism uses the forward and backward audio-visual inputs to independently predict the azimuth and distance, and finally optimizes them in combination with the balanced loss function. It can explicitly learn the difference between the front and back sound sources, so that the model can learn a more complex nonlinear relationship between audio-visual features and the sound source location. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 Data distribution diagram of azimuth angle in the audio-visual sound source localization dataset constructed by the present invention; (a) is the training set data distribution, (b) is the validation set data distribution, and (c) is the test set data distribution.
[0055] Figure 2 This is the network structure of AVBLNet of the present invention. DETAILED DESCRIPTION
[0056] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.
[0057] A method for audio-visual sound source localization based on multi-domain feature extraction and fusion, the steps are as follows:
[0058] Step 1: Build an audio-visual sound source localization dataset
[0059] The dataset randomly selects the sound source position, receiving position and receiving direction from the simulation environment to form an audio-visual pair, which contains 33,166 audio-visual pairs in total. Each audio-visual pair provides the depth image of the front and rear perspectives of the sound receiving position, the binaural audio signals of the front and rear perspectives, and the azimuth and distance of the sound source relative to the current position; the data is collected from 9 different room scenes, covering a variety of environmental changes in daily life, including room size, obstacle distribution, and material properties. The dataset is randomly divided into training and test sets for model training and evaluation.
[0060] The present invention constructs an audio-visual sound source localization dataset covering a rich range of scenarios for the sound source localization task. The dataset contains multimodal data from a variety of real and simulated scenes, including first-person depth maps and binaural audio signals. The scene diversity covers different room sizes, sound source locations, obstacle distributions, and material properties. The scale and complexity of the dataset significantly improve the generalization ability of the training model, laying the foundation for achieving highly robust and high-precision audio-visual sound source localization. In addition, the design of this dataset provides extensive support for complex situations such as multi-sound source environments, dynamic scenes, and noise interference in existing sound source localization tasks.
[0061] Step 2: Build the audio-visual sound source localization network AVBLNet
[0062] The audio-visual sound source localization network AVBLNet is roughly composed of a frequency domain feature extraction module, a time domain feature extraction module, a spatial clue extraction module, a multi-source feature adaptive fusion module and a positioning module.
[0063] The frequency domain feature extraction module is designed to extract frequency domain features from binaural audio and capture the frequency distribution information of the sound. The input is the binaural audio in the dataset. First, it is converted from the time domain to the frequency domain through Fourier transform to generate binaural amplitude spectrogram and phase spectrogram to more intuitively express the frequency characteristics of the sound. Then, the pre-trained ResNet network is used as a feature extractor to extract high-dimensional frequency domain features from the amplitude spectrogram and phase spectrogram. The ResNet network can effectively capture multi-scale frequency domain features through the residual connection structure, while avoiding the gradient vanishing problem, improving training efficiency and feature expression capabilities.
[0064] The time domain feature extraction module is mainly responsible for extracting time dynamic characteristics and time correlation characteristics from binaural audio signals to provide timing information for sound source localization. The input is binaural audio signals, which are left ear signal A and L (t) and right ear signal A R (t). Different from the frequency domain feature extraction module, in order to effectively extract the temporal dynamic features of binaural audio, the time domain feature extraction module directly pulse encodes the input signal in the time domain by segmenting it in a short time window, and then inputs it into the spiking neural network (SNN).
[0065] SNN can capture the time correlation and instantaneous changes of input signals by simulating the pulse emission mechanism of biological neurons, and has strong anti-interference ability. The input signal is processed by the encoder, and the continuous signal is converted into a pulse sequence S L (t) and S R (t), represents the pulse excitation pattern of binaural signals in the time dimension. The pulse sequence is then input into the spiking neural network and features are extracted through a series of spiking neuron layers. The output of each layer of neurons depends on the intensity and time of the input pulse, and the specific formula is as follows:
[0066]
[0067] Where V i (t) is the membrane potential of the neuron, W ij represents the connection weight, φ i is the threshold of the neuron. When the membrane potential V i (t) When the threshold is exceeded, the neuron fires a pulse S i (t). After being processed by multiple layers of SNN, the module finally outputs the time domain feature representation F time, which contains the temporal dynamic pattern of binaural audio signals and binaural time difference information (such as ITD, etc.). The use of the time domain feature extraction module can capture richer temporal characteristics and provide accurate time domain information support for subsequent multimodal feature fusion.
[0068] The spatial clue extraction module is mainly used to extract key spatial geometric features from the depth image (Depth image) of the first perspective to provide scene information for audio-visual sound source positioning. The input of the spatial clue extraction module is a depth image D∈R H×W Through depth information, the spatial cue extraction module can capture the geometric structure of the scene, the distribution of obstacles, and the material properties, which are important for improving the spatial perception of sound source localization. In order to fully extract the spatial features in the depth image, the spatial cue extraction module uses a convolutional neural network (CNN) as the feature extraction backbone.
[0069] First, through a series of convolution and pooling operations, the multi-scale features of the depth image are gradually extracted to capture geometric information at different levels in the scene. The specific formula of the convolution operation is as follows:
[0070] F k =σ(W k *D+b k )
[0071] Where W k is the convolution kernel, b k is the bias term, * represents the convolution operation, σ is the activation function, F k Represents the feature map of the kth layer.
[0072] In order to highlight the focus on key spatial areas, the spatial clue extraction module introduces a spatial attention mechanism to generate a spatial weight map A space , which is used to strengthen important feature areas while suppressing redundant information. The specific calculation formula of attention weight is as follows:
[0073] A space =Softmax(Conv(F))
[0074] Where F is a collection of multiple feature maps, Conv represents the convolution operation, and Softmax is used to normalize the weights. The final spatial features are expressed by weighted fusion as:
[0075] F space =A space ⊙F
[0076] Where ⊙ represents the pixel-by-pixel weighted operation, F space It is the spatial geometric feature.
[0077] Finally, the feature representation F output by the spatial cue extraction module is space It contains multi-level information of the geometric structure of the scene, providing accurate spatial clues for subsequent multimodal feature fusion.
[0078] The multi-source feature adaptive fusion module is one of the core components of AVBLNet, which aims to effectively fuse frequency domain features, time domain features and spatial features to form a unified multimodal feature representation. This module adaptively allocates weights of different features through the attention mechanism, thereby enhancing the expressiveness of key features, while suppressing the interference of redundant information and improving the robustness and accuracy of sound source localization.
[0079] The features input to the multi-source feature adaptive fusion module include frequency domain features Time domain characteristics and spatial characteristics Since these features are distributed in different domains and have different dimensions and semantics, the fusion module needs to align and weight them. First, the features are transformed into the same dimension through the fully connected layer:
[0080] F aligned,l =α l (F l )
[0081] Among them, l∈freq,time,space,α l represents the alignment operation of the lth type of features, F aligned,l Represents the lth class feature after alignment;
[0082] Then, the attention mechanism is used to adaptively assign weights to each feature. The calculation formula for the fusion weight is as follows:
[0083] Z=Softmax(W a Concat(F aligned,freq ,F aligned,time ,F aligned,space ))W a is the parameter matrix generated by the attention weight, Concat represents the feature concatenation operation, Z = [z freq ,z time ,z space ] represents the weights of various features; finally, according to the calculated weight matrix Z, the aligned features are weighted and summed to obtain the fused features as follows:
[0084]
[0085] The final output fusion feature F fusedIt integrates multimodal information in frequency domain, time domain and space domain, has stronger expression ability and can provide reliable input for the positioning module.
[0086] The main function of the positioning module is to integrate the fusion features F output by the multi-source feature adaptive fusion module. fused , predict the azimuth of the sound source and distance To achieve this goal, the module designs a lightweight regression network that can efficiently and accurately complete the positioning task. First, a global average pooling (GAP) operation is performed to map the fused features into a global feature vector f global ∈R C , to compress the spatial dimension and preserve the global information.
[0087]
[0088] Next, the global feature vector is input into two fully connected layers (FC). Each layer uses an activation function to introduce nonlinear capabilities, and Dropout is used to prevent overfitting. Finally, two independent regression branches are designed to predict the azimuth angle. and distance Each branch consists of a fully connected layer.
[0089] Step 3: Training process
[0090] During training, the training set data of the data set is first sent to the multi-domain feature extraction module of the network. The extracted features are adaptively fused through the multi-source feature adaptive fusion module, and then sent to the positioning module to output the azimuth and distance of the sound source. In order to improve the training effect and solve the problem of front and back confusion in sound source localization, this method proposes a front and back information collection mechanism, that is, using the data input from the front and back directions to predict the sound source. Therefore, the loss function of this method is as follows:
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097] in, represents the azimuth and distance predicted using the forward input, represents the azimuth and range predicted using the backward input; and It is the weight coefficient used to balance the losses of each part.
[0098] The present invention uses a self-made audio-visual sound source localization dataset, which contains 33,166 samples and is divided into a training set, a validation set, and a test set in a ratio of 8:1:1. To create the dataset, a scene is randomly selected in the simulation environment, and then the coordinates of the sound source are randomly generated. Coordinates of where the sound is received Towards rotation; then render the visual observation V before and after the current orientation d1 、V d2 , using the original audio and the room response function RIR to render binaural audio A through Fourier convolution L ,A R ; Based on the sound source coordinates and the coordinates of where the sound is received As well as the direction, the azimuth and distance of the sound source relative to the current position can be calculated.
[0099] In AVBLNet, the input binaural audio is 1 second long, and the binaural audio and depth image are respectively passed through three different feature extractors. The extracted features are adaptively fused through the multi-source feature adaptive fusion module, and then sent to the positioning module to output the azimuth and distance of the sound source. The implementation of the method is based on PyTorch, and the loss function used combines multiple mean square error loss functions MSE. The mainstream optimizer Adam in this field is used, the learning rate is set to 0.001, the weight decay rate is set to 0.00001, and the batch size is set to 64; the operating system of the present invention is Ubuntu16.04, the CPU model is IntelXeon CPU E5-2650 v4@2.20GHz, and the GPU uses NVIDIA GeForce TITAN V, with a video memory of 12G.
Claims
1. A method for audio-visual binaural sound source localization based on a pulse neural network, characterized in that: The steps include: Step 1: Build an audio-visual sound source localization dataset; The sound source position, receiving position and receiving direction are randomly selected from the simulation environment to form an audio-visual pair, which is an audio-visual sound source localization dataset; each audio-visual pair provides a depth image of the front and rear perspectives of the sound receiving position, binaural audio signals of the front and rear perspectives, and the azimuth and distance of the sound source relative to the current position; Step 2: Construct the audio-visual sound source localization network AVBLNet; The audio-visual sound source localization network includes a frequency domain feature extraction module, a time domain feature extraction module, a spatial clue extraction module, a multi-source feature adaptive fusion module and a localization module; The binaural audio signals are respectively subjected to the frequency domain feature extraction module and the time domain feature extraction module to extract the frequency domain features and the time domain features; the depth image is subjected to the spatial clue extraction module to extract the spatial geometric features; The multi-source feature adaptive fusion module effectively fuses the frequency domain features, time domain features and spatial geometric features to form the fusion feature F fused ; The positioning module is based on the fusion feature F fused , predict the azimuth of the sound source and distance Step 3: Use the data from the front and back directions to input the audio-visual sound source localization network for sound source prediction training. When the total loss function meets the conditions, the trained audio-visual sound source localization network is obtained for audio-visual binaural sound source localization.
2. The method for audio-visual binaural sound source localization based on pulse neural network according to claim 1, characterized in that: The frequency domain feature extraction module is used to extract frequency domain features from binaural audio signals and capture the frequency distribution information of the sound; the binaural audio signals in the audio-visual sound source localization dataset are input, and the binaural audio signals are converted from the time domain to the frequency domain through Fourier transform to generate the amplitude spectrogram and phase spectrogram of the binaural audio signals; the high-dimensional frequency domain features F are extracted from the amplitude spectrogram and the phase spectrogram by using the pre-trained ResNet network as a feature extractor. freq .
3. The method for audio-visual binaural sound source localization based on pulse neural network according to claim 1, characterized in that: The time domain feature extraction module is used to extract time dynamic characteristics and time correlation characteristics from binaural audio signals to provide timing information for sound source localization; the binaural audio signals are input, respectively, the left ear signal A L (t) and right ear signal A R (t); The time domain feature extraction module directly pulse encodes the binaural audio signal in the time domain by segmenting it in a short time window, and then inputs it into the pulse neural network SNN; Left ear signal A L (t) and right ear signal A R (t) The continuous signal is converted into a pulse sequence S by pulse encoding through the encoder. L (t) and S R (t) represents the pulse excitation pattern of the binaural audio signal in the time dimension; the pulse sequence is input into the spiking neural network, and feature extraction is performed through a series of spiking neuron layers; the output of each layer of neurons depends on the intensity and time of the input pulse, and the specific formula is as follows: Among them, V i (t) is the membrane potential of neuron i, W ij represents the connection weight, φ i is the threshold of the neuron; when the membrane potential V i (t) When the threshold is exceeded, neuron i emits a pulse S i (t); After being processed by the multi-layer spiking neural network SNN, the time domain feature extraction module finally outputs the time domain feature representation F time , including the temporal dynamic patterns of binaural audio signals and interaural time difference information.
4. The method for audio-visual binaural sound source localization based on pulse neural network according to claim 1, characterized in that: The spatial clue extraction module is used to extract key spatial geometric features from the depth image to provide scene information for audio-visual sound source positioning; the input of the spatial clue extraction module is a depth image D∈R H×W , through the depth image information, the geometric structure, obstacle distribution and material properties of the scene are captured; the spatial clue extraction module uses a convolutional neural network as the feature extraction backbone; First, through a series of convolution and pooling operations, the multi-scale features of the depth image are gradually extracted to capture geometric information at different levels in the scene; the specific formula of the convolution operation is as follows: F k =σ(W k *D+b k ) Among them, W k is the convolution kernel, b k is the bias term, * represents the convolution operation, σ is the activation function, F k Represents the feature map of the kth layer; The feature map is processed through the spatial attention mechanism to generate a spatial weight map A. space , which is used to strengthen important feature areas while suppressing redundant information. The specific calculation formula of the attention weight is as follows: A space =Softmax(Conv(F)) Where F is a collection of multiple feature maps, Conv represents the convolution operation, and Softmax is used to normalize the weights; The final spatial feature is expressed by weighted fusion as: F space =A space ⊙F Where ⊙ represents the pixel-by-pixel weighted operation, F space It is the spatial geometric feature.
5. The method for audio-visual binaural sound source localization based on pulse neural network according to claim 1, characterized in that: The multi-source feature adaptive fusion module adaptively allocates weights of different features through an attention mechanism; The features input to the multi-source feature adaptive fusion module include frequency domain features Time domain characteristics and spatial geometric features Its three features are aligned and weighted; First, the features are transformed into the same dimension through a fully connected layer: F aligned,l =a l (F l ) Among them, l∈freq,time,space,α l represents the alignment operation of the lth type of features, F aligned,l Represents the lth class feature after alignment; Then, the attention mechanism is used to adaptively assign fusion weights to each feature; the calculation formula of the fusion weight is as follows: Z=Softmax(W a ·Concat(F aligned,freq ,F aligned,time ,F aligned,space )) Where W a is the learnable parameter matrix generated for the attention weights, Concat represents the feature concatenation operation, Z = [z freq ,z time ,z space ], indicating the weight of each type of feature; Finally, according to the calculated weight matrix Z, the aligned features are weighted and summed to obtain the fused features as follows: where z l Represents the weight of the lth class feature, fusion feature F fused Integrate multimodal information in frequency, time and space domains.
6. The method for audio-visual binaural sound source localization based on pulse neural network according to claim 1, characterized in that: The positioning module is a lightweight regression network; through a global average pooling operation, the fused features are mapped to a global feature vector f global ∈R C , to compress the spatial dimension and preserve the global information; The global feature vector is input into two fully connected layers. Each layer uses an activation function to introduce nonlinear capabilities, and Dropout is used to prevent overfitting. Finally, two independent regression branches are designed to predict the azimuth angle. and distance Each regression branch consists of a fully connected layer.
7. The method for audio-visual binaural sound source localization based on pulse neural network according to claim 1, characterized in that: The total loss function is Among them, θ,d represents the true azimuth and distance of the sound source. represents the azimuth and range predicted using the forward data input, indicates the azimuth and range predicted using the backward data input; and It is the weight coefficient used to balance the losses of each part.
Citation Information
Cited By
DAS-based audio time difference prediction method and system, electronic equipment and storage medium
CN120748449A
DAS-based audio time difference prediction method and system, electronic device and storage medium
CN120748449B