Parallel Acceleration Method for Delay Summation Acoustic Imaging Based on CUDA Technology

Through the parallel acceleration method of delay-sum-refined acoustic imaging based on CUDA technology, the problem of high computational complexity of traditional delay-sum algorithms is solved, and the industrial application of real-time acoustic imaging is realized.

CN116309921BActive Publication Date: 2025-07-22XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310446933.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2025-07-22
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

Traditional delay summing algorithms have high computational complexity and slow calculation speed, making it difficult to meet the real-time imaging needs of industrial production.

Method used

The parallel acceleration method of delay-sum-sum-sum-sum-sum-form-convolution unit based on CUDA technology is adopted to reduce unnecessary calculations through frequency point screening units, convolutional units and grouping convolution units, and parallel computing is performed using convolutional neural networks to improve computing efficiency.

Benefits of technology

Significantly shortens computing time, improves computing efficiency, and realizes real-time acoustic imaging, suitable for industrial applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116309921B_ABST
    Figure CN116309921B_ABST
Patent Text Reader

Abstract

A parallel acceleration method for delay-and-sum acoustic imaging based on CUDA technology belongs to the fields of acoustic signal processing and neural networks and is used to solve problems such as high computational complexity of traditional delay-and-sum algorithms. It includes the following steps: 1) Assume that the steering vector has a spatial arrangement and regard the whole steering vector as a feature map; 2) Perform matrix operations on each bar vector in the feature map and the cross-spectral matrix in turn, and the number of channels remains unchanged after the operation; 3) Multiply the result of the operation with the cross-spectral matrix by the original steering vector to obtain the power of the final bar feature vector. Design the large-dimensional matrices of the cross-spectral matrix and the steering vector into a convolutional network, and use GPU parallel acceleration for the operation. The result can improve the operation time by nearly 100 times compared with the traditional delay-and-sum algorithm, so as to meet the real-time imaging requirements in industrial applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of acoustic imaging and deep neural networks, and in particular relates to a parallel acceleration method for delay and sum acoustic imaging based on CUDA technology. Background Art

[0002] In daily life, the human ear can hear various sounds and identify their locations, which is usually referred to as "locating by sound". Specifically, when someone makes a sound, the human ear can easily know the direction of the sound source; and the human ear can also easily judge the approaching direction of a car passing by, and even roughly know how far the car is. However, the original sound localization function of the human ear is only for solving the problems of life and survival after all, and the localization accuracy is very limited. In order to more accurately receive and locate sound signals, the beam imaging technology based on microphone arrays has begun to attract people's attention, and the acoustic imaging technology that makes "sound visible" has thus been widely used.

[0003] Acoustic imaging, also known as an acoustic camera, refers to using a microphone array to collect multi-channel audio data, using a beamforming algorithm to calculate the sound source distribution information on the scanning plane at a specified frequency, and then fusing the sound source distribution information with a real scene image, so that the spatial position and generation source of the sound source can be quickly determined on the intuitive fused image (Oliver Lylloff, Efrén Fernández-Grande, Finn Agerkvist, Hald, Elisabet Tiana Roig, and Martin S Andersen. 2015. Improving the efficiency of deconvolution algorithms for sound source localization. The journal of the acoustical society of America 138, 1 (2015), 172–180).

[0004] Acoustic imaging technology converts sound information into image signals to visualize the distribution of spatial sound sources. Currently, this technology has been widely applied in various fields such as transportation, noise detection, and industrial anomaly detection (Roberto Merino-Martínez, Pieter Sijtsma, Mirjam Snellen, Thomas Ahlefeldt, Jerome Antoni, Christopher J Bahr, Daniel Blacodon, Daniel Ernst, Arthur Finez, Stefan Funke, et al. 2019. A review of acoustic imaging methods using phased microphone arrays. CEAS Aeronautical Journal 10, 1(2019), 197–230). For example, Boeing and NASA use acoustic imaging technology to detect and locate the noise of the airframe or injector in various acoustic wind tunnel tests and the technology has been quite mature; the BeamformX software system developed by OptiNav company integrates a variety of acoustic imaging algorithms internally and is uniformly named OptiNav BF beamforming; in addition, the company also develops a leak detector - compressed air leak detection and identification system, which can clearly provide images and videos for leak source identification.

[0005] Utilizing the phase shift characteristic of signal delay, the Delay And Sum (DAS) algorithm (Don H Johnson and Dan E Dudgeon. 1992. Array signal processing: concepts and techniques. Simon & Schuster, Inc; Barry D Van Veen and Kevin M Buckley. 1988. Beamforming: A versatile approach to spatial filtering. IEEE assp magazine 5, 2(1988), 4–24) provides robust, fast, and intuitive imaging results through delay compensation and weighted summation. The DAS algorithm describes the core idea of acoustic imaging: using beamforming and time difference to invert the sound source distribution and output an image-like result. However, this algorithm has problems such as large computational complexity and slow calculation speed, resulting in a phenomenon of excessive parameter quantity and redundant computational resources (Austin Reiter and Muyinatu A Lediju Bell. 2017. A machine learning approach to identifying point source locations in photoacoustic data. In Photons Plus Ultrasound: Imaging and Sensing 2017, Vol. 10064. International Society for Optics and Photonics, 100643J), making this algorithm unable to meet the needs of daily industrial production. Summary of the Invention

[0006] The purpose of the present invention is to address the above problems existing in the prior art, provide a parallel acceleration method for delay summation acoustic imaging based on CUDA technology that avoids complex models and long calculation times, is convenient for practical application, solves the problems of high computational complexity of the traditional delay summation algorithm, meets the requirements of actual production applications, and at the same time converts sound information into image signals to visualize the distribution of spatial sound sources, thereby accurately locating the sound.

[0007] The present invention provides a parallel acceleration network for delay summation acoustic imaging based on CUDA technology, including a frequency point screening unit, a convolution unit, and a grouped convolution unit;

[0008] The frequency point screening unit is used to select the maximum value point within the scanning frequency range, so that no points without values need to be introduced in the calculation process, reducing unnecessary calculation time;

[0009] The convolution unit is used to take the steering vector feature map as the input, and perform the convolution operation of the convolutional neural network on the cross-spectral matrix and the large-dimensional matrix of the steering vector to generate a convolution data matrix, thereby improving the operation efficiency.

[0010] When the grouped convolution unit is used for convolution processing, the entire input feature is divided into N groups along the channel direction and convolved respectively to reduce the operation parameters and improve the operation efficiency.

[0011] Furthermore, the frequency point screening unit obtains the sampling values within the scanning frequency range as the input signal of the microphone channel through the number of sampling points and the sampling resolution, and selects several sampling points with the largest values in the input signal to participate in the subsequent matrix operation to improve the calculation efficiency.

[0012] The present invention provides a parallel acceleration method for delay-and-sum acoustic imaging based on CUDA technology, including the following steps:

[0013] 1) Assume that the steering vector has a spatial arrangement, and regard the whole steering vector as a feature map.

[0014] 2) Take each bar vector in the feature map in step 1) as the input of the convolution unit and perform matrix operations with the cross-spectral matrix in sequence, and the number of channels remains unchanged after the operation.

[0015] 3) Perform matrix operations on the result of the matrix operation in step 2) and the original steering vector through the grouped convolution unit to obtain the power of the final bar feature vector.

[0016] In step 1), regarding the whole steering vector as a feature map, the steering vector can be regarded as a feature map of 41×41×64×21, where 41×41 is the number of grid points of the scanned acoustic plane. Assume the size of the scanned plane is [-2, 2 meters] and the scanning resolution is 0.1 meter; 64 is the number of array elements of the microphone array, and 21 is the number of frequency points with larger spectral peaks selected from each microphone channel within the scanning frequency range. The 21 largest frequency points within the swept frequency range are calculated by the frequency point screening unit. This steering vector is used as the input of the feature map to participate in the subsequent convolution operation.

[0017] In step 2), the convolution unit performs a convolution operation on the feature map obtained in step 1) and the cross-spectral matrix. For each pixel on the feature map, the matrix operation steps are the same, which makes this process have the potential for convolution. Since the cross-spectral matrices corresponding to different grid points are the same, only the steering vectors are different, the cross-spectral matrix can be regarded as a feature vector of 1344×64×1×1, that is, the cross-spectral matrix can be regarded as 1344 64-channel 1×1 convolution kernels. After the convolution operation of the steering vector feature map and the cross-spectral matrix, each channel of each convolution output channel is weighted, and finally the convolution layer obtains a feature of 1×1344×41×41.

[0018] In step 3), the feature map obtained in step 2) and the steering vector are successively subjected to convolution operations to obtain the power of the grid point finally. Through the convolution operation of the grouped convolution unit, parallel processing of multi-channel microphone signals is realized, and finally the acoustic power of all grid points on the entire scanning plane is obtained; different from the traditional DAS algorithm that serially processes different microphone signals and different frequency points using a for loop and then adds them, through the convolution operation, parallel processing of multi-channel microphone signals and each frequency point can be realized, greatly shortening the calculation time and improving the calculation efficiency.

[0019] The grouped convolution unit divides the input feature matrix into 21 groups, and the input dimension of each group is At the same time, the convolution kernels are divided into 21 groups, and the dimension of each group is Then each group is convolved separately, and the output is The result of the dimension, and finally the results obtained by each group are cascaded to form the final result; the grouped convolution uses fewer parameters to achieve the same result as the standard convolution, improving the operation speed of the model. At the same time, the implementation of the grouped convolution enables the multi-branch model parallel learning method of the model on multiple GPUs to promote more efficient model training and a better model.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] 1. The present invention combines acoustic signal processing with a convolutional neural network, performs multi-channel parallel processing, extracts effective features of acoustic imaging, improves the processing speed of the model, simplifies the model, and reduces the computational redundancy.

[0022] 2. The present invention can reduce the running time and facilitate the application of the model in actual industrial applications.

[0023] 3. The present invention can accurately locate acoustic signals, and the positioning effect in a poor environment is also more accurate, with good robustness and applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flow chart for simulating the positioning of the time-domain signal of a microphone array.

[0025] Figure 2 This is the convolution operation process diagram of the embodiment of the present invention. Detailed implementation manners

[0026] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0027] The embodiment of the present invention includes the following steps:

[0028] Step 1, sound source modeling: The sound wave generated by a given sound source will present different phases with different propagation distances. Therefore, each microphone in the array will "perceive" different phases, so as to determine the sound source position.

[0029] Most standard beamforming techniques assume that the distance between the source and the observer is large enough and the diameter of the microphone array is small enough, ignoring the directivity of the sound source, so as to be able to approximate the complex sound source well.

[0030] The propagation of sound waves in a continuous medium can be accurately represented by the Navier-Stokes equation. However, in most daily and industrial applications, the propagation of sound waves can be considered a linear isentropic phenomenon. Therefore, the complex Navier-Stokes equation can be significantly simplified to the (Helmholtz) wave equation, which can still accurately reflect the sound wave propagation phenomenon.

[0031]

[0032] Among them, c0 is the speed of sound, is the Laplace operator, q(t) represents a monopole sound source, x0 is the position of the sound source, the Dirac function δ(x - x0) provides the geometric position information of the sound source, x is the position of the microphone array, and p is the sound signal received by the microphone array. In the absence of a sound source, the right side of the above equation is equal to zero.

[0033] The inhomogeneous wave equation has a free-field solution (no reflection or solid boundary):

[0034]

[0035] This illustrates some important aspects of sound wave propagation.

[0036] 1. The amplitude of the sound pressure decays inversely with the observation distance (|x - x0|).

[0037] The sound pressure observation at time t is related to the source characteristics |x - x0| / c0 at the previous moment. That is to say, the sound source information propagates in the medium at a finite speed c0. Therefore, a key concept in the formation of sound wave beams is the delay time (t0), which is formally defined as:

[0038]

[0039] Step 2. Time-domain signal processing: As Figure 1 shown, for an array with M microphone elements, a search grid is set on the plane where the sound source is expected to be located. The scanning plane is divided into N×N grid points, and the distance between the microphone array and the scanning plane is z. The positioning process is as follows:

[0040] 1. Determine the plane where the sound source may be located and divide this plane into (rectangular) grids;

[0041] 2. Scan each grid node. For each node, the time signals measured by each microphone are delayed by their respective delay times (t0 = |x - x0| / c0);

[0042] 3. Add the signals of each microphone and divide by the number of microphones M to obtain the beamformer output map.

[0043] The mathematical expression of the beam output is:

[0044]

[0045] where p m is the signal measured by each microphone, M is the number of microphones, x0 is the sound source position, and x is the microphone position.

[0046] Step 2.1. Obtain the time-domain data of the microphone array: The number of sampling points for each microphone channel is 2400 obtained from the sampling rate and duration, and frame division is performed to obtain the distance from the sound source point to the signal center. The sound pressure level formula is:

[0047]

[0048] where amp is the signal power, which is scaled by the distance from the source to the center of the array to obtain the given SPL at the center, and then a 2400×64-dimensional microphone time-domain signal is obtained through sampling.

[0049] Step 3. Frequency-domain signal processing: The frequency-domain beamforming formula starts from the Fourier transform of the free-field monopole solution:

[0050]

[0051] In the frequency domain, the propagation delay (t0) is represented by a complex number It is expressed that, where P(x, x0, ω) is the frequency-domain representation of the signal measured by each microphone, and Q(ω) is the frequency-domain representation of the sound source signal.

[0052] The Fourier transform of beamforming is:

[0053]

[0054] Among them, it is defined that.g(x, x0, ω) is the steering vector:

[0055]

[0056] The Fourier transform of the signal received by the microphone:

[0057]

[0058] The beamforming Z(ω) can be expressed in matrix form:

[0059]

[0060] The power of the output signal can be calculated by L(x) = |Z| 2 Calculate:

[0061]

[0062] Define the Cross-Spectral Matrix (CSM):

[0063]

[0064] Define the weight vector w of a grid point n :

[0065]

[0066] The expression of conventional beamforming at the x0 grid point is:

[0067]

[0068] Step 3.1: Perform Fourier transform on the collected microphone time-domain signal to obtain the microphone frequency-domain signal. Assume that the dimension of each frame of the collected time-domain signal is 2400×64, and the sampling frequency of the microphone array is fs = 240KHz. Then the frequency resolution after Fourier transform is 100Hz. If the scanning frequency range is from 5KHz to 20KHz, a total of 151 points among the 2400 sampling points are located within the scanning frequency range. Since the spectral peaks of most frequency points among these 151 sampling points are zero, frequency point screening is added here, and only the 21 points with the largest values are selected to participate in the subsequent operations. In this way, the frequency points without values are not calculated, reducing unnecessary calculation time and improving the operation efficiency.

[0069] Step 3.2: Screen out the steering vectors of the specified frequency points according to the sound source signal. The steering vectors in step 3) have a spatial arrangement, and the overall steering vectors are regarded as feature maps; the steering vectors can be regarded as feature maps of 41×41×64×21, where 41×41 is the size of the scanned acoustic plane, 64 is the number of microphones, and 21 is the number of maximum sampled points selected for each microphone channel. This steering vector is used as the input of the feature map to participate in the subsequent convolution operations.

[0070] As Figure 2 shown, each bar vector g H in the feature map is H successively subjected to matrix operations with the cross-spectral matrix pp

[0071] and the number of channels remains unchanged after the operation; for each pixel on the feature map, the steps of its matrix operation are the same, which makes this process have the potential for convolution. Since the cross-spectral matrices corresponding to different grid points are the same, only the steering vectors are different, and the cross-spectral matrix can be regarded as a feature vector of 1344×64×1×1, that is, the cross-spectral matrix can be regarded as 1344 64-channel 1×1 convolution kernels. After the convolution operation of the steering vector feature map and the cross-spectral matrix, each channel of each convolution output channel is weighted, and finally the convolution layer obtains a feature of 1×1344×41×41. H ×PP H ×g; Through continuous convolution operations, parallel processing of multi-channel microphone signals is realized, greatly shortening the calculation time and improving the calculation efficiency.

[0072] Step 3.3: During the convolution operation of the steering vector and the cross-spectral matrix, the method of grouped convolution is adopted. The grouped convolution unit divides the input feature matrix into 21 groups, and the input dimension of each group is At the same time, the convolution kernels are also divided into 21 groups, and the dimension of each group is Then each group is convolved separately, and the output The results of the dimensions are finally cascaded to form the final result. Grouped convolution uses fewer parameters to achieve the same result as standard convolution, improving the network training speed. At the same time, the implementation of grouped convolution enables the multi-branch model parallel learning of the model on multiple GPUs, making the model training more efficient and the model better.

[0073] Table 1 Comparison of GPU and CPU running time results

[0074] The model in question CPU GPU Running time (s / 1000 times) 487.1 4.414

[0075] As can be seen from Table 1, the present invention designs the large-dimensional matrices of the cross-spectrum matrix and the steering vector into a convolutional network and uses GPU parallel acceleration for computing. The result can improve the computing time by nearly 100 times compared with the traditional delay summation algorithm, meeting the real-time imaging requirements in industrial applications.

[0076] In the above embodiments, each step and each unit included are only divided according to functional logic and are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the naming of each unit is only for convenience of distinction and does not limit the protection scope of the present invention.

[0077] The above is only a specific description of the preferred embodiments, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these changes are all included within the scope defined by the claims of this application.

Claims

1. A parallel acceleration method for delay-sum beamforming imaging based on CUDA technology, characterized in that The parallel acceleration network for delay-sum beamforming imaging based on CUDA technology includes a frequency point screening unit, a convolution unit, and a grouped convolution unit; The frequency point screening unit is used to select the maximum value point within the scanning frequency range, so that points with no values do not need to be introduced in the calculation process; The convolution unit takes the steering vector feature map as the input, and performs convolution operation of the convolutional neural network on the cross-spectral matrix and the large-dimensional matrix of the steering vector to generate a convolution data matrix, so as to improve the operation efficiency; When the grouped convolution unit performs convolution processing, the entire input feature is divided into N groups along the channel direction and convolved respectively to reduce the operation parameters and improve the operation efficiency; The parallel acceleration method for delay-sum beamforming imaging based on CUDA technology includes the following steps: 1) Assume that the steering vector has a spatial arrangement, and regard the entire steering vector as a feature map; 2) Take each bar vector in the feature map in step 1) as the input of the convolution unit and perform matrix operations with the cross-spectral matrix in sequence, and the number of channels remains unchanged after the operation; 3) Perform matrix operations on the result of the matrix operation in step 2) and the original steering vector through the grouped convolution unit to obtain the power of the bar vector finally; Perform convolution operations on the feature map obtained in step 2) and the steering vector in sequence, and finally obtain the acoustic power of all grid points on the entire 41×41 scanning plane. Through the convolution operation of the grouped convolution unit, parallel processing of multi-channel microphone signals is realized, and finally the acoustic power of all grid points on the entire scanning plane is obtained; through convolution operations, parallel processing of multi-channel microphone signals and each frequency point is realized, shortening the calculation time and improving the calculation efficiency.

2. The parallel acceleration method for delayed summation acoustic imaging based on CUDA technology according to claim 1, characterized in that The frequency point screening unit obtains the sampling values within the scanning frequency range through the number of sampling points and the sampling resolution as the input signal of the microphone channel, and selects the 21 sampling points with the largest numerical values in the input signal to participate in the subsequent matrix operations to improve the calculation efficiency.

3. The parallel acceleration method for delay summation acoustic imaging based on CUDA technology according to claim 1, characterized in that In step 1), regarding the entire steering vector as a feature map means regarding the steering vector as a feature map of 41×41×64×21, where 41×41 is the number of grid points of the scanned acoustic plane. Assume that the size of the scanning plane is [-2, 2] meters and the scanning resolution is 0.1 meter; 64 is the number of array elements of the microphone array, and 21 is the number of frequency points with larger spectral peaks selected from each microphone channel within the scanning frequency range. The 21 largest frequency points within the swept frequency range are calculated by the frequency point screening unit, and this steering vector is used as the input of the feature map to participate in the subsequent convolution operations.

4. The parallel acceleration method for delayed summation acoustic imaging based on CUDA technology according to claim 1, wherein In step 2), the convolution unit performs a convolution operation on the feature map obtained in step 1) and the cross-spectrum matrix. For each pixel on the feature map, the steps of its matrix operation are the same, which makes this process have the potential for convolution. Since the cross-spectrum matrices corresponding to different grid points are the same, only the steering vectors are different, the cross-spectrum matrix can be regarded as a feature vector of 1344×64×1×1, that is, the cross-spectrum matrix is regarded as 1344 64-channel 1×1 convolution kernels. After the convolution operation of the steering vector feature map and the cross-spectrum matrix, each channel of each convolution output channel is weighted, and finally the convolution layer obtains a feature of 1×1344×41×41.

5. The parallel acceleration method for delayed summation acoustic imaging based on CUDA technology according to claim 1, wherein In step 3), the grouped convolution unit divides the input feature matrix into 21 groups, and the input dimension of each group is Meanwhile, the convolution kernel is divided into 21 groups, and the dimension of each group is Then each group performs convolution separately, and outputs results with a dimension of. Finally, the results obtained from each group are concatenated to form the final result. Grouped convolution uses fewer parameters to achieve the same result as standard convolution, improving the model operation speed. At the same time, grouped convolution enables the multi-branch model of the model on multiple GPUs to learn in parallel, promoting more efficient model training and a better model.

Citation Information

Patent Citations

  • Deconvolution sound source imaging algorithm suitable for coherent and incoherent sound sources

    CN107765221A

  • Automobile noise source acoustic imaging method based on array element random uniform distribution spherical array deconvolution beam forming

    CN112180329A