Sound source direction positioning method and system based on deep learning feature mapping
By combining deep learning feature mapping and state space model, the accuracy and real-time performance issues of traditional sound source localization in complex environments are solved, achieving high-precision 3D sound source localization and immersive sound field reconstruction, and improving the continuity and accuracy of sound source tracking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANG DONG ULTRA PICTURES CULTURE COMM CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-04-24
AI Technical Summary
In indoor environments with strong reverberation or multipath effects, existing technologies are prone to reduced positioning accuracy due to interference from reflected waves caused by traditional sound source localization algorithms. Furthermore, deep learning solutions have high computational complexity on edge devices, making it difficult to achieve real-time 3D sound source localization and immersive sound field reconstruction.
By employing a deep learning-based feature mapping approach, and utilizing the speech sensor array and field-programmable gate array at the edge computing front end, a high signal-to-noise ratio acoustic feature tensor is generated through neural generalized cross-correlation phase transformation and state-space modeling. Spatiotemporal modeling is then performed to calculate the three-dimensional spatial coordinates, driving the holographic sound rendering unit to reconstruct an immersive sound field.
High-precision three-dimensional sound source localization was achieved in complex environments, eliminating false peaks, reducing computational complexity and data transmission requirements, ensuring real-time synchronization between the virtual and actual sound source movements, and reconstructing an immersive sound field that is consistent with audiovisual perception.
Smart Images

Figure CN121918062A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound source localization technology, specifically to a sound source direction localization method and system based on deep learning feature mapping. Background Technology
[0002] Sound source localization technology is the core of intelligent interaction and environmental perception. With the development of multimedia interaction technology, how to use distributed voice sensors to accurately perceive the three-dimensional spatial location of real sound sources and construct an immersive sound field with high audiovisual integration has become a key research direction for improving user interaction experience.
[0003] Existing technologies typically employ generalized cross-correlation or multi-signal classification algorithms to estimate direction of arrival (DOA). However, in environments with strong indoor reverberation or multipath effects, traditional algorithms are susceptible to spurious peaks caused by reflected waves, leading to a sharp drop in positioning accuracy. While deep learning solutions based on convolutional neural networks or Transformers have improved robustness, the computational complexity of the Transformer architecture's self-attention mechanism increases quadratically with sequence length, facing significant computational bottlenecks and high latency issues when processing long-duration speech streams, making it difficult to run in real-time on edge devices. Furthermore, existing solutions are mostly limited to two-dimensional orientation estimation, lacking accurate calculation of the distance to the sound source.
[0004] In summary, overcoming computational resource limitations in strong reverberation environments to achieve three-dimensional sound source localization and solving the problem that the localization results cannot effectively drive real-time reconstruction of the immersive sound field are currently urgent technical challenges to be addressed.
[0005] To address this, a sound source direction localization method and system based on deep learning feature mapping are proposed. Summary of the Invention
[0006] The purpose of this invention is to provide a sound source direction localization method and system based on deep learning feature mapping. By combining edge neural filtering and state space model processing, it can achieve high-precision three-dimensional localization and real-time reconstruction of holographic sound field in strong reverberation environment.
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] Sound source direction localization methods based on deep learning feature mapping include:
[0009] The system utilizes an edge computing front-end that integrates a voice sensor array and a field-programmable gate array to acquire multi-channel raw audio signals; controls the field-programmable gate array to perform neural generalized cross-correlation phase transform processing on the multi-channel raw audio signals; performs temporal filtering on the multi-channel raw audio signals through a built-in learnable filtering layer; and generates an acoustic feature tensor with displacement equivariance.
[0010] The acoustic feature tensor is input into a deep sound source localization network based on a state-space model. The linear computational complexity of the state-space model is used to perform spatiotemporal modeling on a long sequence of audio streams, and the azimuth, pitch, and distance data of the target sound source are calculated to form three-dimensional spatial coordinates.
[0011] The three-dimensional spatial coordinates are output to the holographic sound rendering unit, which drives the holographic sound rendering unit to run the wave field synthesis algorithm. Based on the three-dimensional spatial coordinates, the excitation amplitude parameters and phase delay parameters of each speaker unit in the speaker array are calculated in real time. The focusing position of the virtual sound source is controlled to keep synchronized with the actual motion trajectory of the target sound source in real time, and an immersive holographic sound field is reconstructed in the physical space.
[0012] Preferably, acquiring the multi-channel raw audio signal includes: synchronously receiving multiple pulse density modulation signals output by the voice sensor array through a parallel digital audio interface logic unit configured in the field-programmable gate array; performing low-pass filtering on the multiple pulse density modulation signals using cascaded integrator comb filter resources instantiated within the field-programmable gate array; and downsampling the filtered multiple pulse density modulation signals to convert them into multi-channel pulse code modulation digital signals, which serve as the multi-channel raw audio signal.
[0013] Preferably, the specific process for generating the acoustic feature tensor includes: calling the Fast Fourier Transform (FFT) hard core inside the field-programmable gate array (FPGA) to convert the original multi-channel audio signal in the time domain into a frequency domain signal; calculating the cross-power spectrum of the frequency domain signal for each pair of sensors in the speech sensor array, and performing multiplication weighting on the cross-power spectrum using the frequency domain weighting coefficients generated by the learnable filter layer; performing an inverse Fast Fourier Transform on the weighted cross-power spectrum to obtain a time-domain cross-correlation function containing time delay information; and stacking the time-domain cross-correlation functions according to a preset channel order to construct the acoustic feature tensor.
[0014] Preferably, the specific steps for forming the three-dimensional spatial coordinates are as follows: inputting the acoustic feature tensor into the state space model of the deep sound source localization network; performing a linear recursive scan operation on the acoustic feature tensor using the state space model to extract deep spatiotemporal latent features containing sound source trajectory information; inputting the deep spatiotemporal latent features into a parallel spatial orientation prediction branch and a sound source distance regression branch; controlling the spatial orientation prediction branch to output a discretized spatial likelihood probability map, and obtaining the azimuth angle data and the pitch angle data by performing peak search and interpolation calculation on the spatial likelihood probability map; controlling the sound source distance regression branch to directly output the radial distance value of the target sound source as the distance data; combining and transforming the azimuth angle data, the pitch angle data, and the distance data to form the three-dimensional spatial coordinates.
[0015] Preferably, the specific workflow of the spatial orientation prediction branch includes: pre-setting a discretized angle grid covering the entire space, using a classification fully connected layer to map the deep spatiotemporal latent features into feature vectors with dimensions matching the discretized angle grid, and applying a normalized exponential function to convert the feature vectors into a spatial likelihood probability map representing the probability distribution of the sound source's existence; the specific workflow of the sound source distance regression branch includes: using a regression fully connected layer to extract the sound energy attenuation mode and direct reverberation ratio features implied in the deep spatiotemporal latent features, and directly regressing and predicting the geometric distance value of the target sound source relative to the speech sensor array through nonlinear mapping, as the distance data.
[0016] Preferably, the specific workflow of the holographic sound rendering unit includes: mapping the received three-dimensional spatial coordinates to virtual sound source location points in a three-dimensional virtual sound field model; calculating the sound wave propagation path length from the virtual sound source location point to each speaker unit in the speaker array in real time, based on the Huygens principle followed by the wave field synthesis algorithm and combined with the geometric distribution parameters of the speaker array; determining the phase delay parameter of each speaker unit according to the sound wave propagation path length, and determining the excitation amplitude parameter of each speaker unit according to the relative distance between the virtual sound source location point and each speaker unit; using the excitation amplitude parameter and the phase delay parameter to perform multi-channel modulation processing on the target audio signal, driving the speaker array to radiate multiple wavelets, and synthesizing an immersive holographic sound field with wavefront curvature consistent with the virtual sound source location point in physical space through the coherent superposition of the wavelets.
[0017] A sound source direction localization system based on deep learning feature mapping includes:
[0018] The edge feature extraction module uses an edge computing front-end that integrates a voice sensor array and a field-programmable gate array to acquire multi-channel raw audio signals; controls the field-programmable gate array to perform neural generalized cross-correlation phase transformation processing on the multi-channel raw audio signals; performs temporal filtering on the multi-channel raw audio signals through a built-in learnable filtering layer; and generates an acoustic feature tensor with displacement equivariance.
[0019] The depth localization calculation module inputs the acoustic feature tensor into a depth sound source localization network based on a state-space model. It uses the linear computational complexity of the state-space model to perform spatiotemporal modeling on a long sequence of audio streams, and calculates the azimuth, pitch, and distance data of the target sound source to form three-dimensional spatial coordinates.
[0020] The holographic sound field rendering module outputs the three-dimensional spatial coordinates to the holographic sound rendering unit, drives the holographic sound rendering unit to run the wave field synthesis algorithm, calculates the excitation amplitude parameters and phase delay parameters of each speaker unit in the speaker array in real time based on the three-dimensional spatial coordinates, controls the focusing position of the virtual sound source to keep it synchronized with the actual motion trajectory of the target sound source in real time, and reconstructs an immersive holographic sound field in physical space.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0022] 1. This invention utilizes a built-in learnable filtering layer to replace traditional fixed-parameter filters, enabling it to adaptively learn and suppress strong reverberation and multipath interference in the environment, effectively eliminating spurious peaks in the cross-correlation function. Furthermore, by using hardware-accelerated acoustic feature tensors with displacement equivariance, it maximizes the preservation of the original sound field's spatial structure information, solving the problem of feature distortion caused by low signal-to-noise ratio in complex, enclosed spaces in traditional algorithms. This provides high-quality, clean feature input for subsequent high-precision positioning while significantly reducing data transmission bandwidth requirements.
[0023] 2. This invention employs a deep sound source localization network based on a state-space model, leveraging its unique linear computational complexity advantage to overcome the drawback of traditional Transformer architectures where computational complexity increases quadratically with sequence length. This enables the system to efficiently perform spatiotemporal modeling of long-duration continuous audio streams on low-power edge devices, capturing the dynamic trajectory of moving sound sources in real time. Simultaneously, through parallel spatial orientation prediction and sound source distance regression branch design, this invention not only achieves high-precision azimuth and pitch angle estimation but also successfully solves the radial distance calculation problem that is difficult to obtain using traditional techniques, thereby constructing complete and accurate three-dimensional spatial coordinates and significantly improving the continuity and accuracy of moving sound source tracking.
[0024] 3. This invention uses the calculated high-precision three-dimensional spatial coordinates directly as the driving source to control the holographic sound rendering unit to run the wave field synthesis algorithm. By calculating and adjusting the excitation amplitude and phase delay parameters of each unit in the speaker array in real time, the system can dynamically synthesize a virtual sound source according to the actual position of the target sound source, ensuring that the focus position of the virtual sound is synchronized with the target's movement trajectory at the millisecond level. This physical sound field reconstruction based on Huygens' principle overcomes the limitations of traditional stereo technology, such as narrow listening range and lack of depth, creating an immersive holographic sound field with a high degree of audiovisual consistency. Attached Figure Description
[0025] Figure 1 This is a flowchart of the sound source direction localization method based on deep learning feature mapping proposed in this invention;
[0026] Figure 2 This is a flowchart illustrating the sound source direction localization method based on deep learning feature mapping proposed in this invention.
[0027] Figure 3 This is a system architecture diagram of the sound source direction localization system based on deep learning feature mapping proposed in this invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Example 1
[0030] Please see Figures 1 to 3 This invention provides a sound source direction localization method based on deep learning feature mapping, the technical solution of which is as follows:
[0031] Sound source direction localization methods based on deep learning feature mapping, such as Figures 1-2 As shown, it includes:
[0032] The system utilizes an edge computing front-end that integrates a voice sensor array and a field-programmable gate array to acquire multi-channel raw audio signals; controls the field-programmable gate array to perform neural generalized cross-correlation phase transform processing on the multi-channel raw audio signals; performs temporal filtering on the multi-channel raw audio signals through a built-in learnable filtering layer; and generates an acoustic feature tensor with displacement equivariance.
[0033] The acoustic feature tensor is input into a deep sound source localization network based on a state-space model. The linear computational complexity of the state-space model is used to perform spatiotemporal modeling on a long sequence of audio streams, and the azimuth, pitch, and distance data of the target sound source are calculated to form three-dimensional spatial coordinates.
[0034] The three-dimensional spatial coordinates are output to the holographic sound rendering unit, which drives the holographic sound rendering unit to run the wave field synthesis algorithm. Based on the three-dimensional spatial coordinates, the excitation amplitude parameters and phase delay parameters of each speaker unit in the speaker array are calculated in real time. The focusing position of the virtual sound source is controlled to keep synchronized with the actual motion trajectory of the target sound source in real time, and an immersive holographic sound field is reconstructed in the physical space.
[0035] Further, acquiring the multi-channel raw audio signal includes: synchronously receiving multiple pulse density modulation signals output by the voice sensor array through a parallel digital audio interface logic unit configured in the field-programmable gate array; performing low-pass filtering on the multiple pulse density modulation signals using cascaded integrator comb filter resources instantiated within the field-programmable gate array; and downsampling the filtered multiple pulse density modulation signals to convert them into multi-channel pulse code modulation digital signals, which serve as the multi-channel raw audio signal.
[0036] Specifically, the parallel digital audio interface logic unit is configured to send a synchronization clock signal to the voice sensor array, and based on the time-division multiplexing principle, acquire pulse density modulation data of the first channel on the rising edge of the synchronization clock signal and acquire pulse density modulation data of the second channel on the falling edge of the synchronization clock signal, thereby separating the mixed signal transmitted on a single data line into independent multi-channel signals; the use of the cascaded integrator-comb filter resources instantiated inside the field-programmable gate array to perform low-pass filtering and downsampling processing on the multi-channel pulse density modulation signals refers to the cascaded operation of the multi-stage integrator and comb filter to map the high-sampling-rate 1-bit-width pulse density modulation stream to a low-sampling-rate multi-bit-width data stream, thereby completing the conversion from pulse density modulation format to pulse code modulation format.
[0037] This invention utilizes dedicated hardware logic within a field-programmable gate array (FPGA) to achieve simultaneous dual-edge acquisition and cascaded integral comb filtering downsampling of pulse density modulation (PDM) signals. Time-division multiplexing (TDM) reduces physical interface overhead and ensures strict synchronization of multi-channel signals. This design employs a fully hardware pipeline to map high-sampling-rate 1-bit pulse streams into high-bit-width pulse-code modulation (PCM) signals in real time, effectively improving the signal-to-noise ratio (SNR) and completing format conversion. This significantly reduces the preprocessing load on the backend deep neural network, substantially enhancing the system's real-time response and data throughput efficiency.
[0038] Further, the specific process for generating the acoustic feature tensor includes: calling the Fast Fourier Transform (FFT) hard core inside the field-programmable gate array (FPGA) to convert the original multi-channel audio signal in the time domain into a frequency domain signal; calculating the cross-power spectrum of the frequency domain signal for each pair of sensors in the speech sensor array, and performing multiplication weighting operations on the cross-power spectrum using the frequency domain weighting coefficients generated by the learnable filter layer; performing an inverse Fast Fourier Transform on the weighted cross-power spectrum to obtain a time-domain cross-correlation function containing time delay information; and stacking the time-domain cross-correlation functions according to a preset channel order to construct the acoustic feature tensor.
[0039] Specifically, the learnable filter layer refers to a set of frequency domain weight parameters that are pre-trained and converged through deep learning and stored in the on-chip block memory of the field-programmable gate array; the multiplication and weighting operation of the cross power spectrum using the frequency domain weighting coefficients generated by the learnable filter layer refers to reading the frequency domain weight parameters from the on-chip block memory, completely replacing the fixed phase transformation whitening factor in the traditional generalized cross-correlation algorithm, and performing adaptive reweighting of the amplitude at each frequency point of the cross power spectrum through a complex multiplier to directly suppress the reverberation and noise-dominated frequency bands in the frequency domain; the acoustic feature tensor refers to a two-dimensional feature map matrix constructed by concatenating the time-domain cross-correlation function vectors obtained by combining all sensors in the first dimension according to the spatial index order of the sensor combination and expanding them in the second dimension according to the time delay index.
[0040] The frequency domain weight parameters of the learnable filter layer are obtained during the offline deep learning training phase before system deployment. Specific steps include: using acoustic simulation software to construct a system containing multiple room sizes (ranging from 10...). Up to 100 A large-scale virtual acoustic impulse response dataset with various reverberation times (0.2s to 1.0s) and signal-to-noise ratios (0dB to 30dB) was constructed. A lightweight frequency-domain weighted prediction network (e.g., an encoder structure consisting of three fully connected layers or one-dimensional convolutional layers) was built, with the cross-power spectrum of the noisy signal as input and the corresponding frequency-domain filter mask as output. The loss function was defined as the generalized cross-correlation peak localization error, i.e., minimizing the Euclidean distance between the peak position of the weighted cross-correlation function and the time delay of the real sound source. The stochastic gradient descent algorithm was used for iterative updates. Training was stopped when the localization accuracy on the validation set no longer improved, and the weight vector of the final output layer of the network was extracted as the frequency-domain weight parameters.
[0041] To accommodate the parallel computing characteristics of field-programmable gate arrays (FPGAs) and reduce storage overhead, the frequency domain weight parameters are designed as globally shared one-dimensional vectors. Assuming the Fast Fourier Transform has N points (e.g., 512 points), the number of effective frequency points is... (i.e., 257 frequency points). The dimension of the frequency domain weighting parameter is set to strictly correspond to the number of effective frequency points, i.e., the length is... A real-valued vector. For all sensor pairs combined in a speech sensor array (e.g., 64 microphones). (Multiple combinations), reusing the same set of frequency domain weighting parameters. This is based on the acoustic principle that the distribution characteristics of reverberation and noise in the frequency domain depend primarily on the environment and the frequency itself, and are weakly correlated with the specific microphone pair's position. This sharing mechanism compresses parameter storage by thousands of times.
[0042] When generating the acoustic feature tensor, an auditory saliency detection mechanism is introduced. The field-programmable gate array (FPGA) calculates the energy proportion of each frequency channel and compares it with a preset auditory masking threshold curve. Frequency channels with energy below the auditory masking threshold are identified as redundant channels, and their corresponding cross-correlation function values are forcibly set to zero, retaining only the data of saliency frequency channels. Subsequently, the sparsified data is non-uniformly compressed and packaged according to frequency order and transmitted to the backend network. The original dimension is reconstructed at the network input end using an index table. Through the auditory saliency detection mechanism, the amount of data transmission from the edge to the central processing end is significantly reduced, lowering the requirements for bus bandwidth.
[0043] This invention achieves frequency-domain adaptive reweighting of the cross-power spectrum by integrating a learnable filter layer within a field-programmable gate array (FPGA) and using on-chip block memory to store pre-trained weights, replacing the traditional fixed whitening factor. This design efficiently suppresses reverberation and noise-dominated frequency bands during the signal feature extraction stage using hardware-level complex multiplication, significantly improving the purity and shift variability of the feature tensor. This provides a high-quality input with strong anti-interference capabilities for the backend deep network, thereby greatly improving the localization robustness in complex acoustic environments.
[0044] Further, the specific steps for forming the three-dimensional spatial coordinates are as follows: inputting the acoustic feature tensor into the state space model of the deep sound source localization network; performing a linear recursive scan operation on the acoustic feature tensor using the state space model to extract deep spatiotemporal latent features containing sound source trajectory information; inputting the deep spatiotemporal latent features into a parallel spatial orientation prediction branch and a sound source distance regression branch; controlling the spatial orientation prediction branch to output a discretized spatial likelihood probability map, and obtaining the azimuth angle data and the pitch angle data by performing peak search and interpolation calculation on the spatial likelihood probability map; controlling the sound source distance regression branch to directly output the radial distance value of the target sound source as the distance data; combining and transforming the azimuth angle data, the pitch angle data, and the distance data to form the three-dimensional spatial coordinates.
[0045] Specifically, the step of performing a linear recursive scan operation on the acoustic feature tensor using the state-space model refers to discretizing the continuous-time parameters of the state-space model based on the zero-order preservation principle, expanding the acoustic feature tensor into a sequence input along the time dimension, and performing a linear recursive update step using the discretized parameter matrix to output a hidden state sequence containing historical context information in parallel with linear time complexity. The step of combining and transforming the azimuth, elevation, and distance data refers to using the transformation relationship between spherical and Cartesian coordinate systems, with the geometric center of the speech sensor array as the origin, to map the azimuth, elevation, and distance data into Cartesian three-dimensional spatial coordinates containing X-axis, Y-axis, and Z-axis values to adapt to the spatial input requirements of the wave field synthesis algorithm.
[0046] The state-space model is mathematically defined as a model derived from the input function. To the output function The mapping, whose original continuous form is governed by a linear ordinary differential equation:
[0047] ;
[0048] ;
[0049] Specifically, setting The HiPPO-LegS matrix is mathematically proven to compress and project historical inputs onto an orthogonal polynomial basis, thus theoretically guaranteeing the model's ability to remember infinitely long historical information. For hidden state dimensions, a setting of 16 to 64 is preferred. and These are learnable linear projection parameters, which are updated during deep learning training using the backpropagation algorithm.
[0050] To process sampled discrete audio signals at the digital edge computing front end, a time step parameter must be introduced to discretize the aforementioned continuous equation. The input signal at the current moment is then processed through a linear layer. Mapped to =Softplus(Linear( This ensures that the step size is always positive. This design allows the model to dynamically adjust the temporal resolution of interest based on rapid changes in the signal content (such as sudden transient sounds) or slow changes (such as background noise).
[0051] Based on the dynamically adjusted time step parameters, the model parameters are discretized using the zero-order hold principle. Specifically, the discrete state transition matrix is obtained by calculating the matrix exponent of the product of the continuous state matrix and the time step; simultaneously, the continuous input matrix is integrally transformed into a discrete input matrix that reflects the cumulative effect of the signal within the sampling interval.
[0052] Based on this, the specific logic for performing the linear recursive scan operation is as follows: the deep spatiotemporal latent feature vector at the current time step is obtained by linearly transforming the latent feature vector at the previous time step through the discrete state transition matrix, and then superimposing the current input signal through the discrete input matrix. To ensure the stability of the recursive process and eliminate initial bias, the initial latent state vector of the recursive scan is set to an all-zero vector. Since the above recursive process involves only fixed matrix-vector multiplication operations at each time step, its overall computational complexity is strictly linearly proportional to the length of the audio sequence, rather than the quadratic relationship of traditional attention mechanisms, thus ensuring low latency and high throughput when processing long audio sequences on edge computing devices.
[0053] The training phase of the deep sound source localization network employs a multi-task physical constraint mechanism: a synthetic dataset containing clean direct sound labels and reverberant interference labels is constructed; an auxiliary branch is added to the network specifically for predicting the ratio of direct sound to reverberant energy in the input audio stream; this energy ratio is used as a weighting factor to dynamically adjust the attention given to different feature dimensions by the distance regression branch: when the proportion of reverberant energy is high, the network is forced to rely more on temporal envelope features rather than phase difference features to calculate the distance, thereby achieving adaptive decoupling of features.
[0054] This invention achieves parameter discretization of the state-space model through the zero-order preservation principle, and efficiently captures the deep spatiotemporal features of long audio sequences while maintaining linear time complexity using a linear recursive update mechanism, effectively breaking through the computational bottleneck of traditional architectures in processing long-duration streaming media. Simultaneously, by establishing a precise mapping from spherical coordinates to Cartesian coordinates, it seamlessly adapts to the spatial input standard of wave field synthesis algorithms, ensuring the closed-loop data flow and geometric consistency from sound source perception to holographic sound field rendering.
[0055] Furthermore, the specific workflow of the spatial orientation prediction branch includes: pre-setting a discretized angle grid covering the entire space, using a classification fully connected layer to map the deep spatiotemporal latent features into feature vectors with dimensions matching the discretized angle grid, and applying a normalized exponential function to convert the feature vectors into a spatial likelihood probability map representing the probability distribution of the sound source's existence; the specific workflow of the sound source distance regression branch includes: using a regression fully connected layer to extract the sound energy attenuation mode and direct reverberation ratio features implied in the deep spatiotemporal latent features, and directly regressing and predicting the geometric distance value of the target sound source relative to the speech sensor array through nonlinear mapping, as the distance data.
[0056] Specifically, the discretized angle grid refers to a set of two-dimensional grid points divided in a spherical coordinate system according to a preset azimuth and pitch resolution; the peak search and interpolation calculation of the spatial likelihood probability map refers to first locating the index of the grid point with the largest probability value in the probability map as a rough estimated position, then extracting the probability distribution data of the grid point and its neighborhood, and using a quadratic curve fitting algorithm to estimate the sub-grid level coordinates of the true probability peak, thereby correcting the quantization error introduced by grid discretization; the direct regression prediction of the geometric distance value of the target sound source relative to the speech sensor array through nonlinear mapping refers to using a non-negative activation function to force the regression output result to be a positive real number, so as to conform to the non-negative characteristics of distance in physical space.
[0057] This invention achieves multi-dimensional joint estimation through a parallel branching architecture. By utilizing a two-dimensional grid in the orientation branch combined with quadratic curve fitting interpolation technology, it effectively corrects quantization errors, achieving sub-grid-level high-precision positioning that surpasses grid resolution. Simultaneously, it extracts acoustic energy attenuation features through a distance branch and applies non-negative physical constraints to accurately regress the target's radial distance. This design solves the problem of traditional algorithms lacking depth information, ensuring the geometric accuracy and physical rationality of the three-dimensional spatial coordinates.
[0058] Furthermore, the specific workflow of the holographic sound rendering unit includes: mapping the received three-dimensional spatial coordinates to virtual sound source location points in a three-dimensional virtual sound field model; calculating the sound wave propagation path length from the virtual sound source location point to each speaker unit in the speaker array in real time, based on the Huygens principle followed by the wave field synthesis algorithm and combined with the geometric distribution parameters of the speaker array; determining the phase delay parameter of each speaker unit based on the sound wave propagation path length, and determining the excitation amplitude parameter of each speaker unit based on the relative distance between the virtual sound source location point and each speaker unit; using the excitation amplitude parameter and the phase delay parameter to perform multi-channel modulation processing on the target audio signal, driving the speaker array to radiate multiple wavelets, and synthesizing an immersive holographic sound field with wavefront curvature consistent with the virtual sound source location point in physical space through the coherent superposition of the wavelets.
[0059] Specifically, the multi-channel modulation processing of the target audio signal using the excitation amplitude parameter and the phase delay parameter includes three cascaded signal processing sub-steps: First, source pre-equalization filtering is performed on the target audio signal, and a high-pass filter with a slope of positive 3dB / octave is applied to compensate for the low-frequency boost effect introduced by the coherent superposition of the linear array, ensuring the spectral flatness of the synthesized sound field; second, the calculated phase delay parameter is converted into a time delay value, and a fractional delay filter is used to perform a sub-sampling level precision time delay operation on the pre-equalized signal to eliminate the influence of digital sampling quantization error on the wavefront reconstruction accuracy; finally, the time-delayed signal is multiplied by the excitation amplitude parameter to complete the simulation of geometric diffusion attenuation, thereby generating the final analog voltage signal used to drive each speaker unit.
[0060] This invention employs a wavefield synthesis algorithm combined with source pre-equalization filtering and fractional delay filtering techniques to effectively compensate for low-frequency coloration caused by coherent superposition of linear arrays and eliminate phase errors introduced by digital sampling quantization. By accurately simulating geometric diffusion attenuation and subsampling-level delay, this design ensures a high degree of consistency between the synthesized wavefront curvature and the virtual sound source, reconstructing a spectrum-flat, precisely positioned, and realistically deep immersive holographic sound field in physical space, significantly enhancing the naturalness of the auditory experience.
[0061] This invention first utilizes a learnable filtering layer integrated by a field-programmable gate array to adaptively suppress reverberation and environmental noise at the signal acquisition source, generating high signal-to-noise ratio acoustic features. Second, leveraging the linear computational advantages of the state-space model, it overcomes the computational bottleneck of real-time processing of long audio stream sequences, achieving millisecond-level high-precision three-dimensional positioning including the radial distance dimension. Finally, based on a precise coordinate-driven wave field synthesis algorithm, it ensures that the virtual sound source's focusing position and the target's actual motion trajectory remain synchronized in real time, reconstructing an immersive holographic sound field with consistent audiovisual perception in physical space.
[0062] Example 2
[0063] See Figure 3 This embodiment provides a sound source direction localization system based on deep learning feature mapping, applied to immersive holographic remote teaching interactive scenarios. The technical solution is as follows:
[0064] A sound source direction localization system based on deep learning feature mapping includes:
[0065] The edge feature extraction module uses an edge computing front-end that integrates a voice sensor array and a field-programmable gate array to acquire multi-channel raw audio signals; controls the field-programmable gate array to perform neural generalized cross-correlation phase transformation processing on the multi-channel raw audio signals; performs temporal filtering on the multi-channel raw audio signals through a built-in learnable filtering layer; and generates an acoustic feature tensor with displacement equivariance.
[0066] The depth localization calculation module inputs the acoustic feature tensor into a depth sound source localization network based on a state-space model. It uses the linear computational complexity of the state-space model to perform spatiotemporal modeling on a long sequence of audio streams, and calculates the azimuth, pitch, and distance data of the target sound source to form three-dimensional spatial coordinates.
[0067] The holographic sound field rendering module outputs the three-dimensional spatial coordinates to the holographic sound rendering unit, drives the holographic sound rendering unit to run the wave field synthesis algorithm, calculates the excitation amplitude parameters and phase delay parameters of each speaker unit in the speaker array in real time based on the three-dimensional spatial coordinates, controls the focusing position of the virtual sound source to keep it synchronized with the actual motion trajectory of the target sound source in real time, and reconstructs an immersive holographic sound field in physical space.
[0068] The system aims to accurately capture the movement trajectory of the lecturer in the classroom and render a holographic sound field that follows the teacher's movement to the remote listening terminal or the classroom sound reinforcement system in real time, so that the sound perception is always consistent with the teacher's physical position.
[0069] The system mainly consists of a hardware layer and an algorithm logic layer, specifically including:
[0070] The edge feature extraction module is deployed at the edge computing front end in the center of the classroom ceiling. It physically integrates a voice sensor array consisting of 64 high signal-to-noise ratio MEMS microphones, as well as a high-performance field-programmable gate array development board.
[0071] The voice sensor array picks up multiple pulse density modulation signals generated by the teacher's lecture in real time. The parallel digital audio interface logic unit configured inside the field-programmable gate array synchronously receives these signals through a dual-edge triggering mechanism, and uses the internally instantiated cascaded integrator comb filter resources to directly complete the conversion from pulse density modulation format to pulse code modulation format and downsampling processing at the hardware level, outputting a clean multi-channel original audio signal.
[0072] Subsequently, the field-programmable gate array (FPGA) invokes its internal Fast Fourier Transform (FFT) hard core to convert the audio to the frequency domain and reads the pre-trained learnable filter layer parameters stored in the on-chip block memory. A complex multiplier is used to perform frequency-domain weighting of the cross-power spectrum, adaptively suppressing strong reverberation reflections within the classroom. After the inverse FFT, the module outputs a two-dimensional acoustic feature tensor with displacement equivariance, which is transmitted to the next stage via a high-speed bus.
[0073] The depth localization solution module runs in an embedded AI accelerator at the edge computing front end, and is loaded with a depth sound source localization network based on a state space model.
[0074] This module receives acoustic feature tensors and utilizes the linear computational complexity of the state-space model to efficiently scan long audio streams generated by teachers' continuous lectures over extended periods. It extracts deep spatiotemporal latent features containing movement trajectories, avoiding computational delays caused by sequence growth in traditional attention mechanism models.
[0075] The deep spatiotemporal latent features are fed into two parallel branches. The spatial orientation prediction branch outputs high-precision azimuth and elevation angles through quadratic curve fitting and interpolation; the sound source distance regression branch, combined with non-negative physical constraints, directly regresses the radial distance of the teacher relative to the sensor array. Finally, the module converts these three data points into Cartesian three-dimensional spatial coordinates with the classroom center as the origin.
[0076] The holographic sound field rendering module is connected to a horizontal linear speaker array (e.g., consisting of 128 independent drive units) deployed on the front wall of the classroom, and has a built-in holographic sound rendering unit.
[0077] The module receives real-time updated three-dimensional spatial coordinates and maps them to a "virtual teacher position" in the virtual sound field. Based on Huygens' principle followed by the wave field synthesis algorithm, it calculates in real time the sound wave propagation path length from the virtual position to each speaker unit on the wall.
[0078] The system calculates the required phase delay parameter for each speaker (which is then converted to a subsampled time delay using a fractional delay filter) and the excitation amplitude parameter (simulating geometric diffusion attenuation). These two parameters are used to modulate the audio signal, and source pre-equalization filtering is performed to compensate for the low-frequency boost effect of the linear array.
[0079] The processed signal drives the speaker array to radiate wavelets, which coherently superimpose in physical space to reconstruct an immersive holographic sound field with a wavefront curvature perfectly matching the actual teacher's position. Whether the teacher moves to the left or right of the podium, students can feel the sound precisely emanating from the teacher's direction, achieving a "sound follows the student" teaching experience.
[0080] This invention integrates a learnable filtering layer at the edge computing front end to adaptively suppress reverberation and noise at the signal source, generating high signal-to-noise ratio acoustic features. By leveraging the linear computational complexity advantage of the state-space model, it overcomes the computational bottleneck of real-time processing of long audio stream sequences, achieving millisecond-level high-precision three-dimensional positioning including the radial distance dimension. Finally, based on a precise coordinate-driven wave field synthesis algorithm, it ensures real-time synchronization between the virtual sound source and the target trajectory, reconstructing an immersive holographic sound field with high audiovisual consistency.
[0081] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A sound source direction localization method based on deep learning feature mapping, characterized in that, include: The system utilizes an edge computing front-end that integrates a voice sensor array and a field-programmable gate array to acquire multi-channel raw audio signals; controls the field-programmable gate array to perform neural generalized cross-correlation phase transform processing on the multi-channel raw audio signals; performs temporal filtering on the multi-channel raw audio signals through a built-in learnable filtering layer; and generates an acoustic feature tensor with displacement equivariance. The acoustic feature tensor is input into a deep sound source localization network based on a state-space model. The linear computational complexity of the state-space model is used to perform spatiotemporal modeling on a long sequence of audio streams, and the azimuth, pitch, and distance data of the target sound source are calculated to form three-dimensional spatial coordinates. The three-dimensional spatial coordinates are output to the holographic sound rendering unit, which drives the holographic sound rendering unit to run the wave field synthesis algorithm. Based on the three-dimensional spatial coordinates, the excitation amplitude parameters and phase delay parameters of each speaker unit in the speaker array are calculated in real time. The focusing position of the virtual sound source is controlled to keep synchronized with the actual motion trajectory of the target sound source in real time, and an immersive holographic sound field is reconstructed in the physical space.
2. The sound source direction localization method based on deep learning feature mapping according to claim 1, characterized in that, Acquiring the multi-channel raw audio signal includes: synchronously receiving multiple pulse density modulation signals output by the voice sensor array through a parallel digital audio interface logic unit configured in the field-programmable gate array; performing low-pass filtering on the multiple pulse density modulation signals using cascaded integrator comb filter resources instantiated within the field-programmable gate array; and downsampling the filtered multiple pulse density modulation signals to convert them into multi-channel pulse code modulation digital signals, which serve as the multi-channel raw audio signal.
3. The sound source direction localization method based on deep learning feature mapping according to claim 1, characterized in that, The specific process for generating the acoustic feature tensor includes: calling the Fast Fourier Transform (FFT) hard core inside the field-programmable gate array (FPGA) to convert the original multi-channel audio signal in the time domain into a frequency domain signal; calculating the cross-power spectrum of the frequency domain signal for each pair of sensors in the speech sensor array, and performing multiplication weighting on the cross-power spectrum using the frequency domain weighting coefficients generated by the learnable filter layer; performing an inverse Fast Fourier Transform on the weighted cross-power spectrum to obtain a time-domain cross-correlation function containing time delay information; and stacking the time-domain cross-correlation functions according to a preset channel order to construct the acoustic feature tensor.
4. The sound source direction localization method based on deep learning feature mapping according to claim 1, characterized in that, The specific steps for forming the three-dimensional spatial coordinates are as follows: inputting the acoustic feature tensor into the state space model of the deep sound source localization network; performing a linear recursive scan operation on the acoustic feature tensor using the state space model to extract deep spatiotemporal latent features containing sound source trajectory information; inputting the deep spatiotemporal latent features into a parallel spatial orientation prediction branch and a sound source distance regression branch; controlling the spatial orientation prediction branch to output a discretized spatial likelihood probability map, and obtaining the azimuth angle data and the pitch angle data by performing peak search and interpolation calculation on the spatial likelihood probability map; controlling the sound source distance regression branch to directly output the radial distance value of the target sound source as the distance data. The azimuth data, elevation data, and distance data are combined and transformed using a coordinate system to form the three-dimensional spatial coordinates.
5. The sound source direction localization method based on deep learning feature mapping according to claim 4, characterized in that, The specific workflow of the spatial orientation prediction branch includes: pre-setting a discretized angle grid covering the entire space; using a classification fully connected layer to map the deep spatiotemporal latent features into feature vectors with dimensions matching the discretized angle grid; and applying a normalized exponential function to convert the feature vectors into a spatial likelihood probability map representing the probability distribution of the sound source's existence. The specific workflow of the sound source distance regression branch includes: using a regression fully connected layer to extract the sound energy attenuation mode and direct reverberation ratio features implied in the deep spatiotemporal latent features; and directly regressing and predicting the geometric distance value of the target sound source relative to the speech sensor array through nonlinear mapping, as the distance data.
6. The sound source direction localization method based on deep learning feature mapping according to claim 1, characterized in that, The specific workflow of the holographic sound rendering unit includes: mapping the received three-dimensional spatial coordinates to virtual sound source location points in a three-dimensional virtual sound field model; calculating the sound wave propagation path length from the virtual sound source location point to each speaker unit in the speaker array in real time, based on the Huygens principle followed by the wave field synthesis algorithm and combined with the geometric distribution parameters of the speaker array; determining the phase delay parameter of each speaker unit based on the sound wave propagation path length, and determining the excitation amplitude parameter of each speaker unit based on the relative distance between the virtual sound source location point and each speaker unit; using the excitation amplitude parameter and the phase delay parameter to perform multi-channel modulation processing on the target audio signal, driving the speaker array to radiate multiple wavelets, and synthesizing an immersive holographic sound field in physical space through the coherent superposition of the wavelets, with the wavefront curvature consistent with the virtual sound source location point.
7. A sound source direction localization system based on deep learning feature mapping, characterized in that, include: The edge feature extraction module uses an edge computing front-end that integrates a voice sensor array and a field-programmable gate array to acquire multi-channel raw audio signals; controls the field-programmable gate array to perform neural generalized cross-correlation phase transformation processing on the multi-channel raw audio signals; performs temporal filtering on the multi-channel raw audio signals through a built-in learnable filtering layer; and generates an acoustic feature tensor with displacement equivariance. The depth localization calculation module inputs the acoustic feature tensor into a depth sound source localization network based on a state-space model. It uses the linear computational complexity of the state-space model to perform spatiotemporal modeling on a long sequence of audio streams, and calculates the azimuth, pitch, and distance data of the target sound source to form three-dimensional spatial coordinates. The holographic sound field rendering module outputs the three-dimensional spatial coordinates to the holographic sound rendering unit, drives the holographic sound rendering unit to run the wave field synthesis algorithm, calculates the excitation amplitude parameters and phase delay parameters of each speaker unit in the speaker array in real time based on the three-dimensional spatial coordinates, controls the focusing position of the virtual sound source to keep it synchronized with the actual motion trajectory of the target sound source in real time, and reconstructs an immersive holographic sound field in physical space.