Sound source localization algorithm implementation method and system accelerated by FPGA (Field Programmable Gate Array)

By combining FPGA parallel processing and large-capacity external memory, the problem of data access latency in high-channel sound source localization algorithms is solved, realizing low-latency, high-precision real-time sound source localization, which is suitable for sound source detection and security monitoring applications in multi-channel, large-data-volume environments.

CN121899753APending Publication Date: 2026-04-21ZHENGZHOU UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHENGZHOU UNIV
Filing Date
2026-01-27
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing sound source localization algorithms suffer from data access delays in high-channel, high-data-volume scenarios, making it difficult to guarantee real-time performance. The physical capacity limitations and storage architecture characteristics of BRAM result in excessive computational burden, making it unable to dynamically adapt to complex environments.

Method used

An FPGA-accelerated sound source localization algorithm is adopted. The audio signal is transformed into a frequency domain signal and stored in parallel to an external large-capacity memory through a protocol with parallel transmission capability. Combined with the parallel calculation of the cross power spectral density and response intensity of the microphone array, the parallel processing capability of the FPGA is used for parallel calculation and storage, thereby optimizing data storage and transmission.

Benefits of technology

It achieves low-latency, high-precision real-time sound source localization, is suitable for multi-channel, high-frequency, and large-data-volume environments, improves the efficiency of sound source localization algorithms, and is applicable to fields such as sound source detection, smart speakers, and security monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121899753A_ABST
    Figure CN121899753A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of sound source localization, and particularly relates to an FPGA accelerated sound source localization algorithm implementation method and system. The method comprises the following steps: acquiring an audio signal through a microphone array with more than two channels to obtain a time domain digital signal corresponding to the audio signal; after the digital signal is converted into a frequency domain signal through the FPGA, the frequency domain signal is stored in a memory outside the FPGA in parallel through a parallel interface of a protocol with parallel transmission capability; the capacity of the memory is greater than that of storage resources in the FPGA chip; acquiring the stored frequency domain signal from the memory by using the protocol through the FPGA, calculating the cross-power spectral density of each microphone pair in the microphone array in parallel according to the frequency domain signal, and storing the calculated cross-power spectral density in the memory in parallel through a parallel interface of the protocol; through the FPGA, the cross-power spectral density is obtained from the memory, and in combination with the obtained pre-stored TDOA value, the response intensity of each microphone pair in the microphone array is calculated in parallel for realizing sound source localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of sound source localization, specifically relating to an FPGA-accelerated sound source localization algorithm implementation method and system. Background Technology

[0002] Sound source localization typically relies on multiple microphone arrays to capture sound signals. Common algorithms include techniques based on TDOA (Time Difference of Origin) and correlation, such as GCC-PHAT (Generalized Cross-Correlation Phase Transform) and SRP-PHAT (Spatial Response Mapping Phase Transform). These algorithms infer the location of the sound source by analyzing the signal differences between different microphone arrays. Although these algorithms can provide high localization accuracy, they usually involve a large amount of computation, especially when facing high-channel, large-data, and complex noise environments. The algorithms need to process multi-channel data in real time, resulting in a huge computational burden. Existing patents mainly optimize the SRP-PHAT algorithm through conjugate symmetry and subarray planar scanning, but it has shortcomings in TDOA pre-computation and data transmission. Moreover, for high-channel, large-data scenarios, existing computations often use BRAM multiplexing, which prolongs data access time and cannot guarantee the real-time performance of sound source localization.

[0003] While BRAM, as an on-chip memory resource for FPGAs, exhibits extremely low single-access latency and low power consumption advantages in low-channel-count scenarios, its physical capacity limitations and storage architecture characteristics result in significant disadvantages in high-channel-count and large-data-volume scenarios. For example, when the channel size exceeds 64 channels, the limited capacity of BRAM (typically only tens of MB) forces the system to process data in fragments. This fragmentation operation not only requires frequent switching of storage areas and rearrangement of data formats but also introduces additional control logic overhead. For instance, in a 128-channel sound source localization scenario, the BRAM solution needs to split the data into 8 groups for processing, and loading and switching each group consumes tens of nanoseconds of implicit time cost, resulting in significant fragmentation latency.

[0004] Furthermore, contention for BRAM ports among multiple computing units can lead to access conflicts, causing arbitration wait times to increase exponentially, making it difficult to realize the theoretical low-latency advantage in practical systems. In addition, the fixed partitioning characteristic of BRAM prevents the system from dynamically adapting to complex environments; for example, when pre-stored noise models or high-precision parameters, the rigid partitioning of storage space severely restricts the algorithm's degrees of freedom. Therefore, achieving low latency and real-time processing while ensuring positioning accuracy has become a key challenge in the development of sound source localization technology. Summary of the Invention

[0005] The purpose of this invention is to provide an FPGA-accelerated sound source localization algorithm implementation method and system, which solves the problem that existing sound source localization algorithm implementation schemes suffer from prolonged data access time and difficulty in guaranteeing real-time sound source localization.

[0006] To achieve the above objectives, this invention provides an FPGA-accelerated sound source localization algorithm implementation method, comprising: Audio signals are acquired through a microphone array with two or more channels to obtain the corresponding time-domain digital signal; the time-domain digital signal corresponding to the audio signal is transformed into a frequency-domain signal through an FPGA, and then stored in parallel to an external memory of the FPGA through a parallel interface with a protocol that has parallel transmission capability; the capacity of the memory is greater than the capacity of the on-chip storage resources of the FPGA. The FPGA uses the protocol to retrieve the stored frequency domain signal from the memory, and calculates the cross power spectral density of each microphone pair in the microphone array in parallel based on the frequency domain signal. The calculated cross power spectral density is then stored in parallel to the memory through the parallel interface of the protocol. The cross-power spectral density is obtained from the memory using the FPGA, and combined with the obtained pre-stored TDOA value, the response intensity of each microphone pair in the microphone array is calculated in parallel to achieve sound source localization.

[0007] Furthermore, the srp-phat algorithm is used to calculate the response intensity of each microphone pair in the microphone array via FPGA; In the SRP-PHAT algorithm, when performing the inverse transformation of the frequency domain signal into a digital signal using an FPGA, the process of inversely transforming the frequency domain signal of each frame into a digital signal is carried out in parallel with the process of transforming the time domain digital signal corresponding to the audio signal of the next frame into the frequency domain signal of the next frame using the FPGA.

[0008] Furthermore, the pre-stored TDOA values ​​are obtained based on the pre-stored TDOA values ​​of each microphone pair in the microphone array calculated by the FPGA in all hypothetical sound source directions; The process of calculating the TDOA value of each microphone in all hypothetical sound source directions using an FPGA is performed in parallel with the process of calculating the cross power spectral density between the individual microphones.

[0009] Furthermore, methods for parallel computing the response intensity of each microphone pair in the microphone array using an FPGA include: The cross-power spectral density is obtained by performing an inverse Fourier transform on the FPGA to obtain the cross-correlation value in the time domain. The pre-stored TDOA value corresponding to each microphone pair in the microphone array is converted into the corresponding sample delay by the FPGA in parallel. Then, based on the corresponding sample delay, the absolute value of the cross-correlation value is accumulated at the corresponding position in the SRP matrix by the FPGA in parallel to obtain the response intensity of each microphone pair in the microphone array.

[0010] Furthermore, the pre-stored TDOA values ​​are stored in the FPGA's Block RAM in the form of a three-dimensional matrix of the microphone array and the direction of the imaginary sound source; so that the FPGA can obtain the pre-stored TDOA values ​​of the corresponding microphone pair in a certain imaginary sound source direction through the microphone index and the direction index.

[0011] Furthermore, while performing calculations for the current frame that require data retrieved from the memory, the process of retrieving the data required for the next frame calculation from the memory using the protocol is performed in parallel.

[0012] Furthermore, during the process of storing data in parallel to an external memory via a parallel interface with a protocol capable of parallel transmission through the FPGA, the calculation result data and intermediate data are transferred to different storage areas of the memory; the calculation result data includes the calculated cross-power spectral density, the transformed frequency domain signal, and the calculated response intensity; the intermediate data includes the data that needs to be stored generated during the calculation of the cross-power spectral density, the transformed frequency domain signal, and the calculated response intensity.

[0013] Furthermore, protocols with parallel transmission capabilities also have incremental and surround burst modes; The methods for storing data to be stored in parallel to external memory of the FPGA through the parallel interface of the protocol include: On each parallel interface of the protocol, the FPGA uses the burst mode of the protocol to transmit the data to be stored, and divides the data to be stored into multiple parallel transmission units by setting different transmission burst lengths and burst types; wherein, when the data to be stored is audio signal data, the burst type set by the FPGA is incremental type and the burst length is a first set length; when the data to be stored is the calculation result data, the burst type set by the FPGA is surround type and the burst length is a second set length; the second set length is less than the first set length.

[0014] Furthermore, the method of storing the data to be stored in parallel to the external memory of the FPGA through the parallel interface of the protocol also includes: During the process of storing the data to be stored in parallel to the external memory of the FPGA through the parallel interface of the protocol, the FPGA sends a handshake sequence corresponding to the transmission of different data to the memory according to different priorities. The memory performs a handshake with the FPGA according to the handshake sequence. After the handshake is completed, it receives and stores the corresponding data so that the data with higher priority is stored in the external memory of the FPGA with higher priority. The methods for transforming the time-domain digital signal corresponding to the audio signal into the frequency-domain signal using an FPGA include: using an FPGA and an FFT IP core to perform Fourier transform on the data, thereby transforming the time-domain digital signal corresponding to the audio signal into the frequency-domain signal.

[0015] The above-described technical solution of this invention provides a novel FPGA-accelerated sound source localization algorithm implementation method, the beneficial effects of which include: Based on the implementation of a sound source localization algorithm using FPGA, a protocol with parallel transmission capabilities and a large-capacity external memory are employed. The FPGA's calculation results are stored in the external memory and then retrieved in parallel when needed, optimizing the data storage and transmission of intermediate and final localization results. This design not only solves the traditional memory access bottleneck problem but also provides reliable hardware support for efficient data processing of large-scale microphone arrays. Utilizing the parallel processing capabilities of the FPGA, large amounts of microphone array data requiring extensive computation can be processed more quickly. Combined with a protocol for parallel and rapid storage to external memory, the efficiency of the sound source localization algorithm is improved, resulting in lower latency and higher computational accuracy.

[0016] This invention also provides an FPGA-accelerated sound source localization algorithm implementation system, including a processor storing executable program instructions. The processor is an FPGA processor, and the executable program instructions are executed to implement the following FPGA-accelerated sound source localization algorithm implementation method, specifically including: Audio signals are acquired through a microphone array with two or more channels to obtain the corresponding time-domain digital signal; the time-domain digital signal corresponding to the audio signal is transformed into a frequency-domain signal through an FPGA, and then stored in parallel to an external memory of the FPGA through a parallel interface with a protocol that has parallel transmission capability; the capacity of the memory is greater than the capacity of the on-chip storage resources of the FPGA. The FPGA uses the protocol to retrieve the stored frequency domain signal from the memory, and calculates the cross power spectral density of each microphone pair in the microphone array in parallel based on the frequency domain signal. The calculated cross power spectral density is then stored in parallel to the memory through the parallel interface of the protocol. The cross-power spectral density is obtained from the memory using the FPGA, and combined with the obtained pre-stored TDOA value, the response intensity of each microphone pair in the microphone array is calculated in parallel to achieve sound source localization.

[0017] Furthermore, the srp-phat algorithm is used to calculate the response intensity of each microphone pair in the microphone array via FPGA; In the SRP-PHAT algorithm, when performing the inverse transformation of the frequency domain signal into a digital signal using an FPGA, the process of inversely transforming the frequency domain signal of each frame into a digital signal is carried out in parallel with the process of transforming the time domain digital signal corresponding to the audio signal of the next frame into the frequency domain signal of the next frame using the FPGA.

[0018] Furthermore, the pre-stored TDOA values ​​are obtained based on the pre-stored TDOA values ​​of each microphone pair in the microphone array calculated by the FPGA in all hypothetical sound source directions; The process of calculating the TDOA value of each microphone in all hypothetical sound source directions using an FPGA is performed in parallel with the process of calculating the cross power spectral density between the individual microphones.

[0019] Furthermore, methods for parallel computing the response intensity of each microphone pair in the microphone array using an FPGA include: The cross-power spectral density is obtained by performing an inverse Fourier transform on the FPGA to obtain the cross-correlation value in the time domain. The pre-stored TDOA value corresponding to each microphone pair in the microphone array is converted into the corresponding sample delay by the FPGA in parallel. Then, based on the corresponding sample delay, the absolute value of the cross-correlation value is accumulated at the corresponding position in the SRP matrix by the FPGA in parallel to obtain the response intensity of each microphone pair in the microphone array.

[0020] Furthermore, the pre-stored TDOA values ​​are stored in the FPGA's Block RAM in the form of a three-dimensional matrix of the microphone array and the direction of the imaginary sound source; so that the FPGA can obtain the pre-stored TDOA values ​​of the corresponding microphone pair in a certain imaginary sound source direction through the microphone index and the direction index.

[0021] Furthermore, while performing calculations for the current frame that require data retrieved from the memory, the process of retrieving the data required for the next frame calculation from the memory using the protocol is performed in parallel.

[0022] Furthermore, during the process of storing data in parallel to an external memory via a parallel interface with a protocol capable of parallel transmission through the FPGA, the calculation result data and intermediate data are transferred to different storage areas of the memory; the calculation result data includes the calculated cross-power spectral density, the transformed frequency domain signal, and the calculated response intensity; the intermediate data includes the data that needs to be stored generated during the calculation of the cross-power spectral density, the transformed frequency domain signal, and the calculated response intensity.

[0023] Furthermore, protocols with parallel transmission capabilities also have incremental and surround burst modes; The methods for storing data to be stored in parallel to external memory of the FPGA through the parallel interface of the protocol include: On each parallel interface of the protocol, the FPGA uses the burst mode of the protocol to transmit the data to be stored, and divides the data to be stored into multiple parallel transmission units by setting different transmission burst lengths and burst types; wherein, when the data to be stored is audio signal data, the burst type set by the FPGA is incremental type and the burst length is a first set length; when the data to be stored is the calculation result data, the burst type set by the FPGA is surround type and the burst length is a second set length; the second set length is less than the first set length.

[0024] Furthermore, the method of storing the data to be stored in parallel to the external memory of the FPGA through the parallel interface of the protocol also includes: During the process of storing the data to be stored in parallel to the external memory of the FPGA through the parallel interface of the protocol, the FPGA sends a handshake sequence corresponding to the transmission of different data to the memory according to different priorities. The memory performs a handshake with the FPGA according to the handshake sequence. After the handshake is completed, it receives and stores the corresponding data so that the data with higher priority is stored in the external memory of the FPGA with higher priority. The methods for transforming the time-domain digital signal corresponding to the audio signal into the frequency-domain signal using an FPGA include: using an FPGA and an FFT IP core to perform Fourier transform on the data, thereby transforming the time-domain digital signal corresponding to the audio signal into the frequency-domain signal.

[0025] The technical solution of the FPGA-accelerated sound source localization algorithm implementation system described above can achieve the same beneficial effects as the FPGA-accelerated sound source localization algorithm implementation method described above. Attached Figure Description

[0026] Figure 1This is a block diagram illustrating the principle of the FPGA-accelerated sound source localization algorithm implementation method in the embodiment of the present invention. Figure 2 This is a block diagram illustrating the principle of the FPGA-accelerated sound source localization algorithm in the implementation method of the FPGA-accelerated sound source localization algorithm of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0028] FPGA-accelerated sound source localization algorithm implementation method This embodiment presents a technical solution for implementing an FPGA-accelerated sound source localization algorithm. The computational operations required for each microphone pair in the sound source localization algorithm are performed in parallel by the FPGA to accelerate the sound source localization process. Furthermore, the computational results are stored in parallel in an external, larger-capacity memory to support efficient storage and transmission of data computed in parallel by the FPGA for sound source localization using multi-channel microphone data. This enables real-time and efficient execution of the sound source localization algorithm in multi-channel, high-frequency, and large-data-volume environments, exhibiting low latency and high computational accuracy. It is suitable for high-precision applications requiring real-time sound source localization, such as sound source detection, smart speakers, and security monitoring.

[0029] Reference Figure 1 The method includes: Audio signals are acquired by a microphone array with two or more channels to obtain the corresponding time-domain digital signal of the audio signal; the time-domain digital signal of the audio signal is transformed into a frequency-domain signal by the FPGA, and then stored in parallel to the external memory of the FPGA (referred to as external memory) through a parallel interface with a protocol that has parallel transmission capability; the capacity of the external memory is greater than the capacity of the on-chip memory resources of the FPGA. The frequency domain signal stored in the external memory is obtained by the FPGA using a protocol with parallel transmission capability. Based on the frequency domain signal, the cross power spectral density of each microphone pair in the microphone array is calculated in parallel. The calculated cross power spectral density is then stored in parallel to the external memory through the parallel interface of the protocol with parallel transmission capability. The cross power spectral density is obtained from the external memory using an FPGA, and combined with the obtained pre-stored TDOA value, the response intensity of each microphone pair in the microphone array is calculated in parallel to achieve sound source localization.

[0030] Therefore, based on the implementation of the sound source localization algorithm using FPGA, a protocol with parallel transmission capabilities and a large-capacity external memory are employed. The FPGA's calculation results are stored in the external memory and then retrieved in parallel when needed, optimizing the data storage and transmission of intermediate and final localization results. This design not only solves the traditional memory access bottleneck problem but also provides reliable hardware support for efficient data processing of large-scale microphone arrays. Utilizing the parallel processing capabilities of FPGA, it is possible to process large amounts of microphone array data requiring extensive computation more quickly. Combined with the protocol for parallel and rapid storage to external memory, the efficiency of the sound source localization algorithm is improved, exhibiting lower latency and higher computational accuracy. It is suitable for high-precision applications requiring real-time sound source localization, such as sound source detection, smart speakers, and security monitoring.

[0031] Specifically, the method of converting digital signals into frequency domain signals using FPGA includes: using the FPGA and the FFT IP core to perform Fourier transform on the data, thereby converting the digital signals into frequency domain signals.

[0032] The Fast Fourier Transform (FFT) is an efficient algorithm for computing the Discrete Fourier Transform (DFT). The Fourier Transform is a mathematical tool for transforming signals from the time domain (or spatial domain) to the frequency domain. By reducing computational complexity, the FFT makes frequency domain analysis of signals faster and more efficient, making it more suitable for the rapid processing of large-scale audio signals in sound source localization algorithms.

[0033] In one specific embodiment, the transform length of the FFT IP core is set to 1024 and the architecture choice is set to architecture choice. Finally, the data input and output formats are set. Using the FPGA, the FFT IP core performs a Fourier transform on the data, ultimately obtaining a result with M frequency domain points, i.e., a frequency domain signal X(k) of length M.

[0034] Subsequently, using an FPGA, the data is stored in parallel to an external memory via a parallel interface of a protocol with parallel transmission capabilities. In a preferred embodiment, the protocol with parallel transmission capabilities is the AXI4 protocol, and the external memory of the FPGA is DDR memory. The AXI4+DDR solution achieves a systemic breakthrough in high-channel, big data scenarios by "trading space for time".

[0035] The calculation of the pre-stored TDOA value is shown in the following example: First, the azimuth parameter (elevation) and pitch parameter (azimuth) of the hypothetical sound source point in the positioning space are converted into Cartesian coordinates (x,y,z) by calculating cosine and sine. Then, using the microphone coordinates and the coordinates of the hypothetical sound source, the distance from the microphone to the sound source is calculated using the Euclidean distance formula; For each microphone pair, based on the difference in distance between the two microphones and the sound source, since the division of the hypothetical sound source points in the positioning space is represented by the azimuth parameter elevation and the pitch parameter azimuth, with the azimuth ranging from 0 to 360 degrees and the pitch ranging from 0 to 90 degrees, and the division step size defined in this embodiment being 5 degrees, all elevation and azimuth values ​​are stored in the elevations and azimuths arrays respectively, thus obtaining a positioning space grid of size 73×19; based on the positioning space grid, a total of 73×19 hypothetical sound source points can be obtained, that is, the number of hypothetical sound source directions is 73×19. Because calculating the tdoa value requires calculating the combination of each microphone pair in all hypothetical sound source directions, and the microphones used in this embodiment are 128 channels, the tdoa matrix is ​​a three-dimensional matrix of size 128 (number of microphones) × 128 (number of microphones) × 1,387 (number of directions).

[0036] The cross-power spectral density R(xy) describes the correlation between two signals in the frequency domain. In this embodiment, when calculating the cross-power spectral density of each microphone pair in the microphone array in parallel, the GCC-PHAT algorithm is used for the calculation of the cross-power spectral density. A specific example is as follows: First, the FPGA executes the computation modules corresponding to each gcc-phat algorithm in parallel, retrieving the previously stored FFT calculation results from the DDR memory in parallel. That is, it obtains the Fourier transforms X(f) and Y(f) of the time-domain signals x(t) and y(t) of the two microphones in the microphone pair, and then obtains their product in the frequency domain (taking the conjugate complex number of one of the signals) as the cross-power spectral density R. xy (f); The function of the conjugate complex number is to align the phase of the signal in the complex plane.

[0037] Then, the PHAT filter operation is applied. The key idea of ​​the PHAT filter is to preserve the phase information of the signal while removing the amplitude information by standardizing the cross-power spectral density. This reduces the impact of amplitude variations on time delay estimation. PHAT achieves this by standardizing the cross-power spectral density R... xy (f) Divide by its magnitude |R xy (f)| is used to achieve this. The significance of this operation is to normalize the amplitude part of the cross-power spectrum to 1, retaining only the phase information. Thus, the gcc-phat operation is completed.

[0038] Using an FPGA, the cross-power spectrum is calculated and PHAT filtering is performed on each microphone pair in parallel using the gcc-phat operation described above. Finally, the cross-power spectrum density after the gcc-phat operation is stored in a ddr via the axi4 protocol, waiting for the FPGA to read it when calculating the response intensity of each microphone pair in the microphone array.

[0039] In this embodiment, the srp-phat algorithm is used to calculate the response intensity of each microphone pair in the microphone array via FPGA; In the SRP-PHAT algorithm, when performing the inverse transformation of the frequency domain signal into a digital signal using an FPGA, the process of inversely transforming the frequency domain signal of each frame into a digital signal is carried out in parallel with the process of transforming the time domain digital signal corresponding to the audio signal of the next frame into the frequency domain signal of the next frame using the FPGA.

[0040] In practical design, the operation of the SRP-PHAT algorithm implemented on the FPGA can be regarded as a single operation module, which can be called the SRP-PHAT module. Since the SRP-PHAT algorithm itself requires an inverse Fourier transform operation when calculating the cross-correlation value—that is, a process step of inversely transforming the frequency domain signal into a digital signal—and considering that after the audio signal is acquired through the microphone array, a Fourier transform operation needs to be performed on the FPGA to transform the time-domain digital signal corresponding to the audio signal into a frequency-domain signal, both of these algorithms can be implemented using an FFT IP core (or, in other implementations, using a regular DFT algorithm submodule, etc.). Furthermore, the DSP and Bram resources consumed by the FPGA using the FFT IP core are limited. Therefore, a pipelined approach is used to rationally allocate the resources consumed by FFT to different processing operations at specific times (i.e., the inverse Fourier transform operation of the current frame and the Fourier transform operation of the next frame), thereby ensuring maximum resource utilization.

[0041] Furthermore, pipeline thinking can also be applied to reducing external memory access time. Specifically, when performing calculations on the current frame that require data retrieved from external memory, the process of retrieving the data needed for the next frame calculation from external memory in parallel using a protocol with parallel transmission capabilities can be performed.

[0042] In other words, when the FPGA is processing the Nth frame of data, it is simultaneously accessing external memory to retrieve the data required for the N+1th frame (for example, when calculating the cross-power spectral density of each microphone pair in a microphone array, the data required for the N+1th frame's cross-power spectral density is the N+1th frame's frequency domain signal). This setup ensures that the initial access latency to external memory (e.g., the initial access latency of DDR memory is typically 30ns-50ns) is completely hidden by the parallel operation of the computation unit that inversely transforms the frequency domain signal into a digital signal. While the FPGA is processing the Nth frame of data, the DDR memory preloads the data required for the N+1th frame into its cache in parallel and transmits it to the FPGA, forming a zero-wait pipeline with overlapping computation and transmission. This architecture makes the end-to-end processing latency depend only on the computation itself, rather than the memory access time and transmission time, reducing the latency in the sound source localization algorithm implementation.

[0043] Based on the SRP-PHAT algorithm for calculating the response intensity of each microphone pair in a microphone array using an FPGA, the methods for parallel calculation of the response intensity of each microphone pair in a microphone array using an FPGA include: The cross-power spectral density is obtained by performing an inverse Fourier transform on the FPGA to obtain the cross-correlation value in the time domain. The pre-stored TDOA value corresponding to each microphone pair in the microphone array is converted into the corresponding sample delay by FPGA parallel processing. Then, based on the corresponding sample delay, the absolute value of the cross-correlation value in the time domain is accumulated at the corresponding position in the SRP matrix by FPGA parallel processing to obtain the response intensity of each microphone pair in the microphone array.

[0044] In one specific embodiment, the external memory is DDR memory; the large capacity of DDR memory (usually more than 4GB) allows for the complete storage of full-resolution TDOA pre-calculation results, multiple versions of noise templates, and continuous multiple frames of raw frequency domain signal data, completely avoiding the accuracy loss caused by data compression or approximate calculations in the BRAM (FPGA on-chip memory resource) scheme, which is usually only tens of MB in capacity.

[0045] In this embodiment, an example of how the response intensity of each microphone pair in the microphone array is calculated in parallel using an FPGA is as follows: Extract the cross-power spectral density data from the DDR and cross-correlate it with the phase transform G in the frequency domain. xy Performing an inverse Fourier transform yields the time-domain cross-correlation function R. xy The cross-correlation function represents the actual time delay information between signals.

[0046] Furthermore, the pre-stored TDOA value of the corresponding microphone pair is retrieved, and the time delay difference between the hypothetical sound source and this microphone pair is converted into a sample delay, sample_delay.

[0047] In the above formula, f s It's the sampling frequency, TDOA ij This represents the pre-stored TDOA value of a microphone pair consisting of the i-th microphone in the row direction and the j-th microphone in the column direction of the microphone array.

[0048] Finally, the absolute value of the cross-correlation value Rxy[sample_delay] is accumulated at the corresponding position in the SRP matrix to represent the response intensity in that direction, thus realizing the calculation of the response intensity of the current microphone pair. The above operation of calculating the response intensity of each microphone pair in the microphone array is performed in parallel on all microphone pairs on the FPGA.

[0049] Furthermore, to shorten the latency of the entire sound source localization algorithm process, the pre-stored TDOA value is obtained based on the pre-stored TDOA value of each microphone pair in the microphone array calculated by the FPGA in all hypothetical sound source directions. The process of calculating the TDOA value of each microphone in all hypothetical sound source directions using an FPGA is performed in parallel with the process of calculating the cross power spectral density between the individual microphones.

[0050] Reference Figure 2 This setup takes into account that calculating the response intensity of each microphone pair in the microphone array (e.g., the result of the SRP-PHA algorithm) requires the results of the cross power spectral density between each microphone (e.g., the result of GCC-PHA in the SRP-PHA algorithm) and the TDOA value as input. However, the resources required for calculating the response intensity of the microphone pair and the TDOA value do not conflict. Therefore, these two parts are calculated in parallel in the top-level logic, thereby saving computation time and improving the efficiency of the algorithm.

[0051] Based on this, in order to further reduce the computation delay, the pre-stored TDOA value is stored in the FPGA's Block RAM in the form of a three-dimensional matrix of the microphone array and the direction of the imaginary sound source; so that the FPGA can obtain the pre-stored TDOA value of the corresponding microphone pair in a certain imaginary sound source direction through the microphone index and the direction index.

[0052] In this way, the calculated TDOA matrix is ​​pre-stored in the Block RAM of the FPGA. When calculating the response intensity of each microphone pair in the microphone array, if the TDOA value of the corresponding microphone pair in a certain hypothetical sound source direction is needed, the corresponding data can be directly retrieved from the Block RAM according to the microphone index and direction index for calculation. There is no need to retrieve data from external memory. This can further reduce latency, shorten data lookup time, and improve response speed and positioning accuracy.

[0053] Furthermore, in this embodiment, during the parallel storage of data from the FPGA to an external memory via a parallel interface with a parallel transmission protocol, the calculation result data and intermediate data are transferred to different storage areas of this memory. The calculation result data includes the calculated cross-power spectral density, the transformed frequency domain signal, and the calculated response intensity. The intermediate data includes data that needs to be stored generated during the calculation of the cross-power spectral density, the transformed frequency domain signal, and the calculated response intensity. Therefore, partitioned storage can isolate important calculation result data from relatively unimportant intermediate data, facilitating indexing when retrieving the calculation result data.

[0054] Furthermore, protocols with parallel transmission capabilities also have both incremental and surround burst modes; Methods for storing data to be stored in parallel to external memory of the FPGA via a parallel interface of a protocol with parallel transmission capabilities include: On each parallel interface of the protocol, the FPGA uses the burst mode of the protocol to transmit the data to be stored, and divides the data to be stored into multiple parallel transmission units by setting different transmission burst lengths and burst types; wherein, when the data to be stored is audio signal data, the burst type set by the FPGA is incremental type and the burst length is a first set length; when the data to be stored is calculation result data, the burst type set by the FPGA is surround type and the burst length is a second set length; the second set length is less than the first set length.

[0055] In one specific embodiment, taking the AXI4 protocol with parallel transmission capability as an example and DDR memory outside the FPGA as an example, the method of storing the data to be stored in parallel to the memory outside the FPGA through the parallel interface of the protocol with parallel transmission capability is as follows: The FPGA accesses multiple memory channels in parallel through the various parallel interfaces of the AXI4 protocol, optimizing the storage efficiency of DDR memory. In implementation, the FPGA utilizes the burst transfer mode of the AXI4 protocol to transfer data in blocks, ensuring efficient data storage in DDR. To achieve this parallel storage optimization, Xilinx's AXI4Master IP core can be used, combined with multiple AXI4 interfaces for parallel data transfer. The FPGA configures multiple AXI4 Master interfaces (i.e., the various parallel interfaces of AXI4) to transfer computation results and intermediate data in parallel to different memory areas (i.e., the storage areas for computation results and intermediate data are different), with each memory channel handling a different data stream. This parallel transfer method significantly reduces potential bandwidth bottlenecks during storage and improves data storage throughput.

[0056] On each AXI4 Master interface, the FPGA divides the data stream into multiple parallel transmission units by setting different burst lengths and burst types (such as INCR type). These units can simultaneously write data on multiple memory paths. Xilinx's AXI4 Burst Transaction function supports flexible settings for burst transmission length and rate, allowing the FPGA to dynamically adjust the amount of data transmitted for each data block, thereby ensuring high efficiency in the storage process and access efficiency. In this way, the FPGA can perform data transmission and storage operations in parallel across multiple memory channels, avoiding the latency and processing bottlenecks caused by traditional serial storage. For example, for audio signal data, which is usually large in volume and sequential, the burst type is configured as INCR (incrementing) type, and the burst length is set to 64 or 128 bytes (first set length), enabling efficient continuous writing to DDR memory. For smaller data blocks such as sound source localization calculation results, the burst type usually adopts the WRAP (surround) type burst mode, setting a shorter burst length (i.e., a second set length less than the first set length), such as 16 bytes, so that the data can be stored quickly and read around, reducing the waiting time in the storage path.

[0057] Furthermore, methods for storing data to be stored in parallel to external FPGA memory via a parallel interface using a protocol with parallel transmission capabilities also include: During the process of storing the data to be stored in parallel to the memory outside the FPGA through the parallel interface of the protocol with parallel transmission capability, the FPGA sends the handshake sequence corresponding to the transmission of different data to the memory according to the set different priorities. The memory performs a handshake with the FPGA according to the handshake sequence. After the handshake is completed, it receives and stores the corresponding data so that the data with higher priority is stored in the memory first. In one specific embodiment, the FPGA manages data storage through the AXI4 protocol's Write Address (AW) and Write Data (WD) channels. The FPGA sends data to external memory (DDR memory) via Xilinx's AXI4 Master IP core. During data transmission, the data storage order can be dynamically adjusted based on data priority by changing the handshake sequence. For example, the result data of sound source localization will be stored first, prioritizing other intermediate data to ensure that critical data is not delayed. Each transmitted data block is acknowledged via the AXI4 Write Response (B) channel to ensure successful data transmission.

[0058] FPGA-accelerated sound source localization algorithm implementation system implementation method This embodiment provides a technical solution for an FPGA-accelerated sound source localization algorithm implementation system, including a processor containing executable program instructions. The processor is an FPGA processor, and the executable program instructions are executed to implement the FPGA-accelerated sound source localization algorithm implementation method described below. Specifically, the method includes: Audio signals are acquired by a microphone array with two or more channels to obtain the corresponding time-domain digital signal of the audio signal; the time-domain digital signal of the audio signal is transformed into a frequency-domain signal by the FPGA, and then stored in parallel to the external memory of the FPGA (referred to as external memory) through a parallel interface with a protocol that has parallel transmission capability; the capacity of the external memory is greater than the capacity of the on-chip memory resources of the FPGA. The frequency domain signal stored in the external memory is obtained by the FPGA using a protocol with parallel transmission capability. Based on the frequency domain signal, the cross power spectral density of each microphone pair in the microphone array is calculated in parallel. The calculated cross power spectral density is then stored in parallel to the external memory through the parallel interface of the protocol with parallel transmission capability. The cross power spectral density is obtained from the external memory using an FPGA, and combined with the obtained pre-stored TDOA value, the response intensity of each microphone pair in the microphone array is calculated in parallel to achieve sound source localization.

[0059] Therefore, based on the implementation of the sound source localization algorithm using FPGA, a protocol with parallel transmission capabilities and a large-capacity external memory are employed. The FPGA's calculation results are stored in the external memory and then retrieved in parallel when needed, optimizing the data storage and transmission of intermediate and final localization results. This design not only solves the traditional memory access bottleneck problem but also provides reliable hardware support for efficient data processing of large-scale microphone arrays. Utilizing the parallel processing capabilities of FPGA, it is possible to process large amounts of microphone array data requiring extensive computation more quickly. Combined with the protocol for parallel and rapid storage to external memory, the efficiency of the sound source localization algorithm is improved, exhibiting lower latency and higher computational accuracy. It is suitable for high-precision applications requiring real-time sound source localization, such as sound source detection, smart speakers, and security monitoring.

[0060] Specifically, the method of converting digital signals into frequency domain signals using FPGA includes: using the FPGA and the FFT IP core to perform Fourier transform on the data, thereby converting the digital signals into frequency domain signals.

[0061] The Fast Fourier Transform (FFT) is an efficient algorithm for computing the Discrete Fourier Transform (DFT). The Fourier Transform is a mathematical tool for transforming signals from the time domain (or spatial domain) to the frequency domain. By reducing computational complexity, the FFT makes frequency domain analysis of signals faster and more efficient, making it more suitable for the rapid processing of large-scale audio signals in sound source localization algorithms.

[0062] In one specific embodiment, the transform length of the FFT IP core is set to 1024 and the architecture choice is set to architecture choice. Finally, the data input and output formats are set. Using the FPGA, the FFT IP core performs a Fourier transform on the data, ultimately obtaining a result with M frequency domain points, i.e., a frequency domain signal X(k) of length M.

[0063] Subsequently, using an FPGA, the data is stored in parallel to an external memory via a parallel interface of a protocol with parallel transmission capabilities. In a preferred embodiment, the protocol with parallel transmission capabilities is the AXI4 protocol, and the external memory of the FPGA is DDR memory. The AXI4+DDR solution achieves a systemic breakthrough in high-channel, big data scenarios by "trading space for time".

[0064] The calculation of the pre-stored TDOA value is shown in the following example: First, the azimuth parameter (elevation) and pitch parameter (azimuth) of the hypothetical sound source point in the positioning space are converted into Cartesian coordinates (x,y,z) by calculating cosine and sine. Then, using the microphone coordinates and the coordinates of the hypothetical sound source, the distance from the microphone to the sound source is calculated using the Euclidean distance formula; For each microphone pair, based on the difference in distance between the two microphones and the sound source, since the division of the hypothetical sound source points in the positioning space is represented by the azimuth parameter elevation and the pitch parameter azimuth, with the azimuth ranging from 0 to 360 degrees and the pitch ranging from 0 to 90 degrees, and the division step size defined in this embodiment being 5 degrees, all elevation and azimuth values ​​are stored in the elevations and azimuths arrays respectively, thus obtaining a positioning space grid of size 73×19; based on the positioning space grid, a total of 73×19 hypothetical sound source points can be obtained, that is, the number of hypothetical sound source directions is 73×19. Because calculating the tdoa value requires calculating the combination of each microphone pair in all hypothetical sound source directions, and the microphones used in this embodiment are 128 channels, the tdoa matrix is ​​a three-dimensional matrix of size 128 (number of microphones) × 128 (number of microphones) × 1,387 (number of directions).

[0065] The cross-power spectral density R(xy) describes the correlation between two signals in the frequency domain. In this embodiment, when calculating the cross-power spectral density of each microphone pair in the microphone array in parallel, the GCC-PHAT algorithm is used for the calculation of the cross-power spectral density. A specific example is as follows: First, the FPGA executes the computation modules corresponding to each gcc-phat algorithm in parallel, retrieving the previously stored FFT calculation results from the DDR memory in parallel. That is, it obtains the Fourier transforms X(f) and Y(f) of the time-domain signals x(t) and y(t) of the two microphones in the microphone pair, and then obtains their product in the frequency domain (taking the conjugate complex number of one of the signals) as the cross-power spectral density R. xy (f); The function of the conjugate complex number is to align the phase of the signal in the complex plane.

[0066] Then, the PHAT filter operation is applied. The key idea of ​​the PHAT filter is to preserve the phase information of the signal while removing the amplitude information by standardizing the cross-power spectral density. This reduces the impact of amplitude variations on time delay estimation. PHAT achieves this by standardizing the cross-power spectral density R... xy (f) Divide by its magnitude |R xy (f)| is used to achieve this. The significance of this operation is to normalize the amplitude part of the cross-power spectrum to 1, retaining only the phase information. Thus, the gcc-phat operation is completed.

[0067] Using an FPGA, the cross-power spectrum is calculated and PHAT filtering is performed on each microphone pair in parallel using the gcc-phat operation described above. Finally, the cross-power spectrum density after the gcc-phat operation is stored in a ddr via the axi4 protocol, waiting for the FPGA to read it when calculating the response intensity of each microphone pair in the microphone array.

[0068] In this embodiment, the srp-phat algorithm is used to calculate the response intensity of each microphone pair in the microphone array via FPGA; In the SRP-PHAT algorithm, when performing the inverse transformation of the frequency domain signal into a digital signal using an FPGA, the process of inversely transforming the frequency domain signal of each frame into a digital signal is carried out in parallel with the process of transforming the time domain digital signal corresponding to the audio signal of the next frame into the frequency domain signal of the next frame using the FPGA.

[0069] In practical design, the operation of the SRP-PHAT algorithm implemented on the FPGA can be regarded as a single operation module, which can be called the SRP-PHAT module. Since the SRP-PHAT algorithm itself requires an inverse Fourier transform operation when calculating the cross-correlation value—that is, a process step of inversely transforming the frequency domain signal into a digital signal—and considering that after the audio signal is acquired through the microphone array, a Fourier transform operation needs to be performed on the FPGA to transform the time-domain digital signal corresponding to the audio signal into a frequency-domain signal, both of these algorithms can be implemented using an FFT IP core (or, in other implementations, using a regular DFT algorithm submodule, etc.). Furthermore, the DSP and Bram resources consumed by the FPGA using the FFT IP core are limited. Therefore, a pipelined approach is used to rationally allocate the resources consumed by FFT to different processing operations at specific times (i.e., the inverse Fourier transform operation of the current frame and the Fourier transform operation of the next frame), thereby ensuring maximum resource utilization.

[0070] Furthermore, pipeline thinking can also be applied to reducing external memory access time. Specifically, when performing calculations on the current frame that require data retrieved from external memory, the process of retrieving the data needed for the next frame calculation from external memory in parallel using a protocol with parallel transmission capabilities can be performed.

[0071] In other words, when the FPGA is processing the Nth frame of data, it is simultaneously accessing external memory to retrieve the data required for the N+1th frame (for example, when calculating the cross-power spectral density of each microphone pair in a microphone array, the data required for the N+1th frame's cross-power spectral density is the N+1th frame's frequency domain signal). This setup ensures that the initial access latency to external memory (e.g., the initial access latency of DDR memory is typically 30ns-50ns) is completely hidden by the parallel operation of the computation unit that inversely transforms the frequency domain signal into a digital signal. While the FPGA is processing the Nth frame of data, the DDR memory preloads the data required for the N+1th frame into its cache in parallel and transmits it to the FPGA, forming a zero-wait pipeline with overlapping computation and transmission. This architecture makes the end-to-end processing latency depend only on the computation itself, rather than the memory access time and transmission time, reducing the latency in the sound source localization algorithm implementation.

[0072] Based on the SRP-PHAT algorithm for calculating the response intensity of each microphone pair in a microphone array using an FPGA, the methods for parallel calculation of the response intensity of each microphone pair in a microphone array using an FPGA include: The cross-power spectral density is obtained by performing an inverse Fourier transform on the FPGA to obtain the cross-correlation value in the time domain. The pre-stored TDOA value corresponding to each microphone pair in the microphone array is converted into the corresponding sample delay by FPGA parallel processing. Then, based on the corresponding sample delay, the absolute value of the cross-correlation value in the time domain is accumulated at the corresponding position in the SRP matrix by FPGA parallel processing to obtain the response intensity of each microphone pair in the microphone array.

[0073] In one specific embodiment, the external memory is DDR memory; the large capacity of DDR memory (usually more than 4GB) allows for the complete storage of full-resolution TDOA pre-calculation results, multiple versions of noise templates, and continuous multiple frames of raw frequency domain signal data, completely avoiding the accuracy loss caused by data compression or approximate calculations in the BRAM (FPGA on-chip memory resource) scheme, which is usually only tens of MB in capacity.

[0074] In this embodiment, an example of how the response intensity of each microphone pair in the microphone array is calculated in parallel using an FPGA is as follows: Extract the cross-power spectral density data from the DDR and cross-correlate it with the phase transform G in the frequency domain. xy Performing an inverse Fourier transform yields the time-domain cross-correlation function R. xy The cross-correlation function represents the actual time delay information between signals.

[0075] Furthermore, the pre-stored TDOA value of the corresponding microphone pair is retrieved, and the time delay difference between the hypothetical sound source and this microphone pair is converted into a sample delay, sample_delay.

[0076] In the above formula, f s It's the sampling frequency, TDOA ij This represents the pre-stored TDOA value of a microphone pair consisting of the i-th microphone in the row direction and the j-th microphone in the column direction of the microphone array.

[0077] Finally, the absolute value of the cross-correlation value Rxy[sample_delay] is accumulated at the corresponding position in the SRP matrix to represent the response intensity in that direction, thus realizing the calculation of the response intensity of the current microphone pair. The above operation of calculating the response intensity of each microphone pair in the microphone array is performed in parallel on all microphone pairs on the FPGA.

[0078] Furthermore, to shorten the latency of the entire sound source localization algorithm process, the pre-stored TDOA value is obtained based on the pre-stored TDOA value of each microphone pair in the microphone array calculated by the FPGA in all hypothetical sound source directions. The process of calculating the TDOA value of each microphone in all hypothetical sound source directions using an FPGA is performed in parallel with the process of calculating the cross power spectral density between the individual microphones.

[0079] This setup takes into account that calculating the response intensity of each microphone pair in the microphone array (e.g., the result of the SRP-PHA algorithm) requires the results of the cross power spectral density between each microphone (e.g., the result of GCC-PHA in the SRP-PHA algorithm) and the TDOA value as input. However, the resources required for calculating the response intensity of the microphone pair and the TDOA value do not conflict. Therefore, these two parts are calculated in parallel in the top-level logic, thereby saving computation time and improving the efficiency of the algorithm.

[0080] Based on this, in order to further reduce the computation delay, the pre-stored TDOA value is stored in the FPGA's Block RAM in the form of a three-dimensional matrix of the microphone array and the direction of the imaginary sound source; so that the FPGA can obtain the pre-stored TDOA value of the corresponding microphone pair in a certain imaginary sound source direction through the microphone index and the direction index.

[0081] In this way, the calculated TDOA matrix is ​​pre-stored in the Block RAM of the FPGA. When calculating the response intensity of each microphone pair in the microphone array, if the TDOA value of the corresponding microphone pair in a certain hypothetical sound source direction is needed, the corresponding data can be directly retrieved from the Block RAM according to the microphone index and direction index for calculation. There is no need to retrieve data from external memory. This can further reduce latency, shorten data lookup time, and improve response speed and positioning accuracy.

[0082] Furthermore, in this embodiment, during the parallel storage of data from the FPGA to an external memory via a parallel interface with a parallel transmission protocol, the calculation result data and intermediate data are transferred to different storage areas of this memory. The calculation result data includes the calculated cross-power spectral density, the transformed frequency domain signal, and the calculated response intensity. The intermediate data includes data that needs to be stored generated during the calculation of the cross-power spectral density, the transformed frequency domain signal, and the calculated response intensity. Therefore, partitioned storage can isolate important calculation result data from relatively unimportant intermediate data, facilitating indexing when retrieving the calculation result data.

[0083] Furthermore, protocols with parallel transmission capabilities also have both incremental and surround burst modes; Methods for storing data to be stored in parallel to external memory of the FPGA via a parallel interface of a protocol with parallel transmission capabilities include: On each parallel interface of the protocol, the FPGA uses the burst mode of the protocol to transmit the data to be stored, and divides the data to be stored into multiple parallel transmission units by setting different transmission burst lengths and burst types; wherein, when the data to be stored is audio signal data, the burst type set by the FPGA is incremental type and the burst length is a first set length; when the data to be stored is calculation result data, the burst type set by the FPGA is surround type and the burst length is a second set length; the second set length is less than the first set length.

[0084] In one specific embodiment, taking the AXI4 protocol with parallel transmission capability as an example and DDR memory outside the FPGA as an example, the method of storing the data to be stored in parallel to the memory outside the FPGA through the parallel interface of the protocol with parallel transmission capability is as follows: The FPGA accesses multiple memory channels in parallel through the various parallel interfaces of the AXI4 protocol, optimizing the storage efficiency of DDR memory. In implementation, the FPGA utilizes the burst transfer mode of the AXI4 protocol to transfer data in blocks, ensuring efficient data storage in DDR. To achieve this parallel storage optimization, Xilinx's AXI4Master IP core can be used, combined with multiple AXI4 interfaces for parallel data transfer. The FPGA configures multiple AXI4 Master interfaces (i.e., the various parallel interfaces of AXI4) to transfer computation results and intermediate data in parallel to different memory areas (i.e., the storage areas for computation results and intermediate data are different), with each memory channel handling a different data stream. This parallel transfer method significantly reduces potential bandwidth bottlenecks during storage and improves data storage throughput.

[0085] On each AXI4 Master interface, the FPGA divides the data stream into multiple parallel transmission units by setting different burst lengths and burst types (such as INCR type). These units can simultaneously write data on multiple memory paths. Xilinx's AXI4 Burst Transaction function supports flexible settings for burst transmission length and rate, allowing the FPGA to dynamically adjust the amount of data transmitted for each data block, thereby ensuring high efficiency in the storage process and access efficiency. In this way, the FPGA can perform data transmission and storage operations in parallel across multiple memory channels, avoiding the latency and processing bottlenecks caused by traditional serial storage. For example, for audio signal data, which is usually large in volume and sequential, the burst type is configured as INCR (incrementing) type, and the burst length is set to 64 or 128 bytes (first set length), enabling efficient continuous writing to DDR memory. For smaller data blocks such as sound source localization calculation results, the burst type usually adopts the WRAP (surround) type burst mode, setting a shorter burst length (i.e., a second set length less than the first set length), such as 16 bytes, so that the data can be stored quickly and read around, reducing the waiting time in the storage path.

[0086] Furthermore, methods for storing data to be stored in parallel to external FPGA memory via a parallel interface using a protocol with parallel transmission capabilities also include: During the process of storing the data to be stored in parallel to the memory outside the FPGA through the parallel interface of the protocol with parallel transmission capability, the FPGA sends the handshake sequence corresponding to the transmission of different data to the memory according to the set different priorities. The memory performs a handshake with the FPGA according to the handshake sequence. After the handshake is completed, it receives and stores the corresponding data so that the data with higher priority is stored in the memory first. In one specific embodiment, the FPGA manages data storage through the AXI4 protocol's Write Address (AW) and Write Data (WD) channels. The FPGA sends data to external memory (DDR memory) via Xilinx's AXI4 Master IP core. During data transmission, the data storage order can be dynamically adjusted based on data priority by changing the handshake sequence. For example, the result data of sound source localization will be stored first, prioritizing other intermediate data to ensure that critical data is not delayed. Each transmitted data block is acknowledged via the AXI4 Write Response (B) channel to ensure successful data transmission.

[0087] It should be understood that the above-described specific embodiments of the present invention are merely illustrative or explanatory of the principles of the present invention, and do not constitute a limitation thereof.

Claims

1. A method for implementing an FPGA-accelerated sound source localization algorithm, characterized in that, include: Audio signals are acquired through a microphone array with two or more channels to obtain the corresponding time-domain digital signal; the time-domain digital signal corresponding to the audio signal is transformed into a frequency-domain signal through an FPGA, and then stored in parallel to an external memory of the FPGA through a parallel interface with a protocol that has parallel transmission capability; the capacity of the memory is greater than the capacity of the on-chip storage resources of the FPGA. The FPGA uses the protocol to retrieve the stored frequency domain signal from the memory, and calculates the cross power spectral density of each microphone pair in the microphone array in parallel based on the frequency domain signal. The calculated cross power spectral density is then stored in parallel to the memory through the parallel interface of the protocol. The cross-power spectral density is obtained from the memory using the FPGA, and combined with the obtained pre-stored TDOA value, the response intensity of each microphone pair in the microphone array is calculated in parallel to achieve sound source localization.

2. The FPGA-accelerated sound source localization algorithm implementation method according to claim 1, characterized in that, The SRP-PHAT algorithm is used to calculate the response intensity of each microphone pair in the microphone array via FPGA. In the SRP-PHAT algorithm, when performing the inverse transformation of the frequency domain signal into a digital signal using an FPGA, the process of inversely transforming the frequency domain signal of each frame into a digital signal is carried out in parallel with the process of transforming the time domain digital signal corresponding to the audio signal of the next frame into the frequency domain signal of the next frame using the FPGA.

3. The FPGA-accelerated sound source localization algorithm implementation method according to claim 1 or 2, characterized in that, The pre-stored TDOA value is obtained based on the pre-stored TDOA value of each microphone pair in the microphone array calculated by the FPGA in all hypothetical sound source directions; The process of calculating the TDOA value of each microphone in all hypothetical sound source directions using an FPGA is performed in parallel with the process of calculating the cross power spectral density between the individual microphones.

4. The FPGA-accelerated sound source localization algorithm implementation method according to claim 2, characterized in that, Methods for parallel computation of the response intensity of each microphone pair in a microphone array using an FPGA include: The cross-power spectral density is obtained by performing an inverse Fourier transform on the FPGA to obtain the cross-correlation value in the time domain. The pre-stored TDOA value corresponding to each microphone pair in the microphone array is converted into the corresponding sample delay by the FPGA in parallel. Then, based on the corresponding sample delay, the absolute value of the cross-correlation value is accumulated at the corresponding position in the SRP matrix by the FPGA in parallel to obtain the response intensity of each microphone pair in the microphone array.

5. The FPGA-accelerated sound source localization algorithm implementation method according to claim 3, characterized in that, The pre-stored TDOA values ​​are stored in the FPGA's Block RAM in the form of a three-dimensional matrix of the microphone array and the direction of the imaginary sound source; This allows the FPGA to obtain the pre-stored TDOA value of the corresponding microphone in the direction of a hypothetical sound source by using the microphone index and direction index.

6. The FPGA-accelerated sound source localization algorithm implementation method according to claim 1 or 2, characterized in that, While performing calculations for the current frame that require data retrieved from the memory, the process of retrieving the data needed for the next frame calculation from the memory using the protocol is performed in parallel.

7. The FPGA-accelerated sound source localization algorithm implementation method according to claim 1 or 2, characterized in that, During the process of storing data in parallel to an external memory via a parallel interface of a protocol with parallel transmission capability through an FPGA, the calculation result data and intermediate data are transferred to different storage areas of the memory. The calculation result data includes the calculated cross-power spectral density, the transformed frequency domain signal, and the calculated response intensity. The intermediate data includes the data that needs to be stored generated during the calculation of the cross-power spectral density, the transformed frequency domain signal, and the calculation of the response intensity.

8. The FPGA-accelerated sound source localization algorithm implementation method according to claim 7, characterized in that, Protocols capable of parallel transmission also have incremental and surround burst modes; The methods for storing data to be stored in parallel to external memory of the FPGA through the parallel interface of the protocol include: On each parallel interface of the protocol, the FPGA uses the burst mode of the protocol to transmit the data to be stored, and divides the data to be stored into multiple parallel transmission units by setting different transmission burst lengths and burst types; wherein, when the data to be stored is audio signal data, the burst type set by the FPGA is incremental type and the burst length is a first set length; when the data to be stored is the calculation result data, the burst type set by the FPGA is surround type and the burst length is a second set length; the second set length is less than the first set length.

9. The FPGA-accelerated sound source localization algorithm implementation method according to claim 1 or 2, characterized in that, The method of storing the data to be stored in parallel to the external memory of the FPGA through the parallel interface of the protocol also includes: During the process of storing the data to be stored in parallel to the external memory of the FPGA through the parallel interface of the protocol, the FPGA sends a handshake sequence corresponding to the transmission of different data to the memory according to different priorities. The memory performs a handshake with the FPGA according to the handshake sequence. After the handshake is completed, it receives and stores the corresponding data so that the data with higher priority is stored in the external memory of the FPGA with higher priority. The methods for transforming the time-domain digital signal corresponding to the audio signal into the frequency-domain signal using an FPGA include: using an FPGA and an FFT IP core to perform Fourier transform on the data, thereby transforming the time-domain digital signal corresponding to the audio signal into the frequency-domain signal.

10. A system for implementing an FPGA-accelerated sound source localization algorithm, comprising a processor, wherein the processor stores executable program instructions, characterized in that, The processor is an FPGA processor, and the executable program instructions are executed to implement the FPGA-accelerated sound source localization algorithm implementation method according to any one of claims 1-9.

Citation Information

Cited By

  • A sound source positioning and tracking device based on an STM32H7 microcontroller SAI interface

    CN122488038A