Implementation method, device and electronic equipment of acoustic camera
By dividing the space into pixels in the acoustic camera, calculating the microphone delay relationship and performing quantile scanning, the problem of high computational load for large microphone arrays is solved, and efficient sound source localization is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI FULLHAN MICROELECTRONICS
- Filing Date
- 2022-12-30
- Publication Date
- 2026-05-19
AI Technical Summary
Existing acoustic cameras require a huge amount of computation when processing a large number of microphones, resulting in low efficiency in locating sound sources.
By pre-dividing the space into pixels, calculating the microphone delay relationship of each pixel, scanning the pixels using the quantile method, filling in the delay, finding the spatial node with the largest output, frequency domain processing is achieved, and the amount of computation for each frame of FFT is reduced.
This reduces the computational load on acoustic cameras, improves the efficiency and accuracy of sound source localization, and reduces the consumption of computing resources.
Smart Images

Figure CN116224230B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio and image signal processing technology, and in particular to a method, apparatus and electronic device for implementing an acoustic camera. Background Technology
[0002] Microphone arrays can deduce the location of a sound source by using the time difference between the sound source's arrival at the microphone. The more microphones there are, the higher the positioning accuracy. When there are a sufficient number of microphones, the location of the sound source can be pinpointed relatively accurately, and a spatial sound field distribution can be generated. Combined with optical image information, the sound source can be imaged, making it easy to intuitively obtain the distribution information of the sound source within the sound field.
[0003] Acoustic cameras, also known as audio-visual cameras, are now widely used in areas such as capturing violations of vehicle horn use. They typically calculate the weighted cross-power spectrum of each microphone signal in the frequency domain, then use an inverse Fourier transform to obtain the cross-correlation function of the signal. The maximum value of the cross-correlation function is used to determine the delay of the microphone signal, thereby determining the signal's location.
[0004] Currently, it is necessary to first perform time-frequency transformation on each frame of signal from each microphone to the frequency domain, calculate the cross-power spectrum and weight it, then perform inverse time-frequency transformation to the time domain to calculate the delay and sound source location information, including approximately Each time-frequency transformation operation is involved. When the number of microphones M is large, the computational load for time-frequency transformation alone is enormous. Summary of the Invention
[0005] To overcome the shortcomings of the existing technology, the present invention aims to provide a method, apparatus and electronic device for implementing an acoustic camera, so as to provide a method for implementing an acoustic camera with low computational requirements.
[0006] To achieve the above objectives, the present invention provides a method for implementing an acoustic camera, comprising the following steps:
[0007] Step S1: Calculate the phase compensation slope tensor based on the microphone position, pickup angle, and pixel division of the acoustic camera.
[0008] Step S2: Frame windowing is performed on the time-domain signal of each microphone. The windowing results are summed according to the number of frames to be processed. Time-frequency analysis is performed on the summed results to obtain the frequency domain signal of each microphone.
[0009] Step S3: Using the quantile method, perform phase compensation on the phase compensation slope tensor for each microphone spectrum, and update the display matrix based on the compensated spectrum;
[0010] Step S4 involves matching the displayed matrix with the actual graphic to accurately locate the sound-emitting object.
[0011] Optionally, step S1 further includes:
[0012] Step S100: Obtain the specific location of each microphone in the acoustic camera microphone array, the angle range to be picked up, and the required resolution;
[0013] Step S101: Project the area to be picked onto a plane, and divide the plane into several pixels according to the resolution.
[0014] Step S102: Calculate the distance from each pixel to each microphone. Define a reference distance for each pixel, calculate the difference between each microphone and the reference distance, and calculate the phase compensation slope based on the difference.
[0015] Step S103: Place the phase compensation slope into the corresponding pixel point to obtain the phase compensation slope tensor.
[0016] Optionally, the phase compensation slope tensor is obtained by the following formula:
[0017]
[0018] in, denoted by , represents the difference between the distance from the i-th and j-th pixels to the i-th microphone and the distance to the center of the array, where fs is the signal sampling rate, c is the speed of sound, and N is the FFT length.
[0019] Optionally, step S2 further includes:
[0020] Step S200: Frame-by-frame windowing is performed on the time-domain signal of each microphone. Then, based on the number of frames to be processed simultaneously, the windowed time-domain signals are summed to obtain the summed signal. ;
[0021] Step S201: Sum the signals from each microphone. Time-frequency analysis was performed to obtain the frequency domain signal of each microphone. .
[0022] Optionally, step S3 further includes:
[0023] Frequency domain signal of each microphone Phase compensation is performed on the corresponding pixels based on the phase compensation slope tensor to obtain the compensation spectrum of each pixel. ;
[0024] The compensation spectra of the M microphones are summed, and the mean of the summed spectrum amplitudes is taken as the display value of the (i,j)th pixel to obtain the display matrix.
[0025] Optionally, step S3 further includes:
[0026] Frequency domain signal of each microphone Phase compensation is performed on the corresponding pixels based on the phase compensation slope tensor to obtain the compensation spectrum of each pixel. ;
[0027] The compensation spectra of the M microphones are summed and differentiated, and the ratio of the summation to the difference is taken as the display value of the (i,j)th pixel to obtain the display matrix.
[0028] Optionally, the frequency domain signal of each microphone Phase compensation is performed on the corresponding pixels based on the phase compensation slope tensor to obtain the compensation spectrum of each pixel. Specifically:
[0029]
[0030] in, This is the compensated spectrum of the m-th microphone, where k is the frequency index. .
[0031] Optionally, the display matrix is obtained as follows:
[0032]
[0033] in, This indicates the selection of a frequency point.
[0034] To achieve the above objectives, the present invention also provides an apparatus for implementing an acoustic camera, comprising:
[0035] The phase compensation slope tensor calculation unit is used to calculate the phase compensation slope tensor based on the microphone position, pickup angle, and pixel division of the acoustic camera.
[0036] The frequency domain signal acquisition unit is used to divide the time domain signal of each microphone into frames and window it. The windowing results are summed according to the number of frames to be processed, and the summation results are analyzed in time and frequency to obtain the frequency domain signal of each microphone.
[0037] The phase compensation and pixel calculation unit is used to perform phase compensation on the corresponding pixel points according to the phase compensation slope tensor of the frequency domain signal of each microphone, to obtain the compensated spectrum of each pixel point, and update the display matrix based on the compensated spectrum.
[0038] The positioning unit is used to match the display matrix with the actual graphic to accurately locate the sound-emitting object.
[0039] To achieve the above objectives, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the above-described method for implementing an acoustic camera.
[0040] Compared with the prior art, the present invention provides an acoustic camera implementation method, device and electronic device. By pre-determining the spatial range, dividing the space into individual pixels, calculating the microphone delay relationship of each pixel, scanning the pixels using the quantile method, filling in the delay, and finding the spatial node with the largest output as the sound source location, the present invention provides an acoustic camera implementation method with low computational load. The present invention processes in the frequency domain, but on average, the number of FFTs used per frame is very small.
[0041] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0042] The above and other objects, features, and advantages of the present invention will become more apparent from the more detailed description of the embodiments of the invention in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same parts or steps.
[0043] Figure 1 This is a flowchart illustrating the implementation method of the acoustic camera provided in Embodiment 1 of the present invention;
[0044] Figure 2 This is a schematic diagram of microphone orientation estimation in this embodiment.
[0045] Figure 3 This is a schematic diagram of the acoustic camera in this embodiment;
[0046] Figure 4 This is a schematic diagram of the pixel division of the sound source incident plane in this embodiment;
[0047] Figure 5 This is a flowchart illustrating the implementation of the acoustic camera in this embodiment;
[0048] Figure 6 This is a simulation diagram of the acoustic camera in an embodiment of the present invention;
[0049] Figure 7 This is a system structure diagram of the acoustic camera implementation device provided in Embodiment 2 of the present invention;
[0050] Figure 8 This is the structure of an electronic device provided in an exemplary embodiment of the present invention. Detailed Implementation
[0051] The following describes the embodiments of the present invention through specific examples and in conjunction with the accompanying drawings. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific examples, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention. Example 1
[0052] Figure 1 This is a flowchart illustrating an implementation method for an acoustic camera provided in an exemplary embodiment of the present invention. This embodiment can be applied to electronic devices, such as... Figure 1 As shown, the method for implementing the acoustic camera includes the following steps:
[0053] Step S1: Calculate the phase compensation slope tensor based on the microphone position, pickup angle, and pixel division of the acoustic camera.
[0054] Specifically, step S1 further includes:
[0055] Step S100: Determine the specific location of each microphone in the acoustic camera microphone array, the required angle range, and the required resolution.
[0056] Step S101: Project the area to be picked onto a plane, and divide the plane into several pixels according to the resolution.
[0057] Step S102: Calculate the distance from each pixel to each microphone. Define a reference distance for each pixel, calculate the difference between each microphone and the reference distance, and calculate the phase compensation slope based on the difference.
[0058] Step S103: Place the phase compensation slope into the corresponding pixel to obtain the phase compensation slope tensor slopeMatrix.
[0059] The above steps only need to be performed once during algorithm initialization, and the computational cost is negligible.
[0060] Figure 2 With two microphones, the incident angle of the far-field sound source The relationship between latency and time delay is determined by Figure 2 It can be seen that the angle of incidence There is a one-to-one correspondence between latency and time delay, that is:
[0061] ,
[0062] in, It's the speed of sound. This refers to the distance between two microphones. For arrays with multiple microphones, there is still a one-to-one correspondence between the incident angle and the phase difference between each microphone. Figure 3 This is a schematic diagram of the acoustic camera in this embodiment. The microphone array is located in a plane, and a planar sound wave is incident in front of the plane at a certain angle to the plane's normal. When the incident direction of the sound wave is fixed, the phase relationship of the microphone array is also fixed, meaning that there is a one-to-one correspondence between the incident direction and the phase difference between the microphones. Just as the viewing angle presented by an optical lens has a certain range, the angle of the sound source received by the microphone array also needs to be defined within a certain range, because as the angle increases, the positioning accuracy will gradually decrease.
[0063] Assuming the maximum viewing angle of the array in all directions is θ, then the range of incident angles of the sound source that the array can display is: If a sound source is incident on a plane perpendicular to the normal in front of the microphone array, and the distance from the incident plane to the array is d, then the length and width range of the incident plane of the sound source are... ,like Figure 4 As shown, the length and width of the incident plane can be divided into... The distance between two adjacent points is tan(θ)*d / dpi, where dpi can be considered as the resolution; the larger the value, the higher the resolution. Each point can be considered as a pixel, and for each pixel of the sound source, there is a unique phase relationship corresponding to it in the microphone array.
[0064] If we place the microphone array at the origin in a Cartesian coordinate system, and take the array's normal as the x-axis, then the spatial coordinates of the incident plane are (d, y, z), where d is the distance from the incident plane to the array, and the values of y and z range from... Calculate the Euclidean distance from each pixel (d, y, z) to each microphone in the microphone array to obtain the tensor distMatrix. The dimension of distMatrix is... , where M represents the number of microphones.
[0065] Choose the distance from each pixel to the origin. As a benchmark, we obtain 3D matrix Find the tensor distMatrix and The difference:
[0066]
[0067] Where dMatrix(i,j,m) represents the difference between the distance from the i-th and j-th pixels to the m-th microphone and the distance to the array center. When the distance d approaches infinity, this difference is equivalent to the distance difference at the incident point of a plane wave. Once the distance difference is obtained, the phase difference slope in the frequency domain can be calculated:
[0068]
[0069] Where fs is the signal sampling rate, c is the speed of sound, and N is the known FFT length. Multiplying slopeMatrix(i,j,m) by the frequency index k represents the phase difference at a specific frequency.
[0070] The phase difference slope tensor slopeMatrix(i,j,m) is thus obtained and used for subsequent compensation of microphone phase difference. It can be seen that once the microphone positions in the array are fixed, the maximum viewing angle θ is determined, and the resolution value is determined, the phase difference slope tensor slopeMatrix is uniquely determined, so it only needs to be calculated once during algorithm initialization.
[0071] Step S2: Frame and window the time-domain signal of each microphone. Sum the windowed results according to the number of frames to be processed. Perform DFT on the summed results to obtain the frequency domain signal of each microphone.
[0072] Specifically, such as Figure 5 As shown, step S2 further includes:
[0073] Step S200: Frame-by-frame windowing is performed on the time-domain signal of each microphone. Then, based on the number of frames to be processed simultaneously, the windowed time-domain signals are summed to obtain the summed signal. .
[0074] Step S201: Sum the signals from each microphone. Time-frequency analysis was performed to obtain the frequency domain signal of each microphone. The subscript m represents the microphone index, m∈[1,M].
[0075] Microphone spectrum at this time The phase difference between the two frames satisfies the same relationship as the phase difference between the spectra of a single frame, which is the basis for the optimization of this invention.
[0076] In this invention, if each microphone signal is calculated separately for each frame, it is easy to produce unstable sound source position results, and the amount of calculation is very large. If each frame signal is first converted to the frequency domain and then multiple frames are considered for calculation at the same time, the phenomenon of unstable sound source position can be improved, but the amount of calculation is still very large.
[0077] Therefore, in order to solve the above problems, this embodiment has made the following treatment:
[0078] First, multiple frames are considered simultaneously. Utilizing the linear property of the DFT, DFT(x) + DFT(y) = DFT(x + y), the summation of multiple frame signals in the frequency domain is transformed into a summation in the time domain, and then back to the frequency domain. Let dm(l) represent the windowed signal of the m-th microphone in the l-th time frame. (l) Perform summation:
[0079]
[0080] This is the summation of the time-domain windowed signals of L frames. A larger L indicates more frames are considered simultaneously. While more stable sound source location information is possible, a value that is too large can lead to insufficient refresh rate. To balance stability and refresh rate, with a frame shift of 10ms, L can be set to 10, allowing for 10 refreshes per second and reducing the number of DFTs to 1 / 10 of the previous value. Perform a DFT operation to obtain the frequency domain signal of the m-th microphone. Then, the phase compensation in equation (4) and the calculation of the display matrix for each pixel in equation (5) are performed.
[0081] Step S3: Using the quantile method, the phase compensation slope tensor is used to perform phase compensation on the spectrum of each microphone. The compensated spectrum is summed and placed into the corresponding pixel position to obtain the display matrix.
[0082] Specifically, the frequency domain signal of each microphone Phase compensation is performed on the corresponding pixels using the slope matrix of the phase compensation slope tensor, resulting in the compensation spectrum for each pixel. In other words, step S2 yields the spectral signal for each microphone. Then, the phase compensation slope tensor is applied to the microphone's spectral signal:
[0083]
[0084] in, It is the compensated spectrum of the m-th microphone, and k is the frequency index, k∈[0,N / 2].
[0085] Then, the compensated spectra of the M microphones are summed, and the average of the summed spectrum amplitudes is used as the display value for the (i,j)th pixel.
[0086]
[0087] Among them, dispMatrix is the final output display matrix. This refers to the selection of frequency points. Because the wavelength of low-frequency signals is much larger than the distance between microphones, the phase difference of each pixel changes very little, resulting in limited positioning effect. High-frequency signals are prone to attenuation during propagation. After reaching the microphone, the signal-to-noise ratio is too low, and the phase relationship is no longer reliable. Therefore, it is necessary to select an appropriate frequency for positioning analysis.
[0088] The frequency can be selected manually in real time, or an algorithm can be designed to adaptively select it. When there is a sound source at the corresponding pixel location, the value of the pixel corresponding to dispMatrix(i,j) will be relatively large, and when displayed after normalization, the color of that location will be relatively dark. By matching the pixels of the dispMatrix matrix with the pixels of the image, the location of the sound source can be accurately displayed, realizing the function of an acoustic camera.
[0089] Step S4 involves matching the displayed matrix with the actual graphic to accurately locate the sound-emitting object.
[0090] Although every pixel needs to be displayed, most pixels are not actually of concern to us. We only need to focus on the pixel where the sound source is located and the pixels near the sound source. Therefore, we can use the quantile method to locate the sound source and calculate the pixels near the sound source step by step. The total number of pixels is... If each pixel is calculated using equation (5), the computational load would be enormous. To reduce the computational load, equation (5) is used once every dpis pixels, which will increase the computational load. For each pixel, find the maximum value (maxValue) and record its location (maxi, maxj). This locates the sound source near the pixel (maxi, maxj). Then, a second search is performed, calculating intervals of dpis2 pixels within the pixel intervals (maxi - dpis1, maxi + dpis1) and (maxj - dpis1, maxj + dpis1). Simultaneously, update the maximum value (maxValue) and its location (maxi, maxj). The number of pixels to be calculated here is... Finally, the pixel range is calculated. All pixels, the number of which needs to be calculated here is .
[0091] After optimization, the number of pixels calculated was reduced from It became If the resolution dpi is 128, dpis1 is 32, and dpis2 is 8, the number of pixels to be calculated decreases from 66049 to 451, reducing the computational load to less than 1 / 100. Combined with the optimization of the DFT computation, even in scenarios with a large number of microphones, CPU deployment is no longer a major issue.
[0092] In addition, to optimize the imaging effect, an optimization scheme is given for formula (5). At the pixel point corresponding to the sound source direction, the phase of each microphone spectrum after phase compensation is consistent, so the sum has a maximum value. On the other hand, the subtraction should also have a minimum value. Therefore, the mean of the sum of the compensation spectra of M microphones divided by the mean of the difference of the compensation spectra of M microphones can also be used as the pixel output, that is, as the display of the (i,j)th pixel point:
[0093]
[0094] The robustness of the algorithm can be improved by using the above formula (6), and there is no need to perform additional normalization operations on the results.
[0095] Figure 6 This invention relates to a circular array with M=8 microphones, an array size of approximately 2.5 dm, and a maximum pickup angle θ. 𝑚 The screenshot showing the sound source localization when the angle is 45 degrees, dpi=128, dpis1=32, and dpis2=8. Example 2
[0096] Figure 7 This is a system structure diagram of an acoustic camera implementation device provided in an exemplary embodiment of the present invention. This embodiment can be applied to electronic devices, such as... Figure 7 As shown, it includes:
[0097] The phase compensation slope tensor calculation unit 701 is used to calculate the phase compensation slope tensor based on the microphone position, pickup angle and pixel division of the acoustic camera.
[0098] Specifically, the phase compensation slope tensor calculation unit 701 further includes:
[0099] The information acquisition unit is used to acquire the specific location of each microphone in the acoustic camera microphone array, the required angle range, and the required resolution.
[0100] The projection unit is used to project the area to be picked up onto a plane, and divides the plane into several pixels according to the resolution.
[0101] The phase compensation slope calculation unit is used to calculate the distance from each pixel to each microphone. Each pixel has a defined reference distance. The difference between each microphone and the reference distance is calculated, and the phase compensation slope is calculated based on the difference.
[0102] The phase compensation slope tensor calculation unit is used to put the phase compensation slope into the corresponding pixel point to obtain the phase compensation slope tensor slopeMatrix.
[0103] The frequency domain signal acquisition unit 702 is used to perform frame-by-frame windowing on the time domain signal of each microphone, sum the windowing results according to the number of frames to be processed, and perform DFT on the summation results to obtain the frequency domain signal of each microphone.
[0104] Specifically, the frequency domain signal acquisition unit 702 further includes:
[0105] The time-domain frame-segmentation, windowing, and summation unit performs frame-segmentation and windowing on the time-domain signal of each microphone. Then, based on the number of frames to be processed simultaneously, it sums the windowed time-domain signals to obtain signal d. 𝑚 ;
[0106] The time-frequency analysis unit is used to sum the signals d from each microphone. 𝑚 Time-frequency analysis was performed to obtain the frequency domain signal 𝐷 𝑚 .
[0107] Phase compensation and pixel calculation unit 703 is used to convert the frequency domain signal of each microphone into a pixel value. 𝑚 Phase compensation is performed on the corresponding pixels based on the slope matrix of the phase compensation tensor, resulting in the compensation spectrum DC of each pixel. 𝑚 The compensated spectrum is summed and placed into the corresponding pixel position to obtain the display matrix.
[0108] Specifically,
[0109] Phase compensation unit, used to convert the frequency domain signal of each microphone to a phase compensation unit. 𝑚 Phase compensation is performed on the corresponding pixels based on the slope matrix of the phase compensation tensor, resulting in the compensation spectrum DC of each pixel. 𝑚 ;
[0110] Pixel calculation unit, used to calculate the compensation spectrum DC of the corresponding pixel. 𝑚 The summation is performed based on the number of microphones to obtain the display matrix `dispMatrix`. Optionally, the summation can be replaced with the ratio of the sum to the difference to increase the algorithm's robustness.
[0111] The positioning unit 704 is used to match the display matrix with the actual graphic to accurately locate the sound-emitting object.
[0112] Exemplary electronic devices
[0113] Figure 8 This is the structure of an electronic device provided in an exemplary embodiment of the present invention. The electronic device may be either or both of a first device and a second device, or a standalone device independent of them, which may communicate with the first device and the second device to receive acquired input signals from them. Figure 8 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Figure 8 As shown, the electronic device includes one or more processors 81 and memory 82.
[0114] The processor 81 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0115] The memory 82 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 81 may execute the program instructions to implement the acoustic camera implementation method of the software program of the various embodiments of this disclosure described above, and / or other desired functions. In one example, the electronic device may also include an input device 83 and an output device 84, these components being interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0116] In addition, the input device 83 may also include, for example, a keyboard, a mouse, etc.
[0117] The output device 84 can output various information to the outside. The output device 84 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0118] Of course, for the sake of simplicity, Figure 8 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.
[0119] Exemplary computer program products and computer-readable storage media
[0120] In addition to the methods and devices described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps in the implementation methods of an acoustic camera according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0121] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0122] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the implementation methods of an acoustic camera according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.
[0123] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0124] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0125] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0126] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0127] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.
[0128] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps are decomposable and / or recombinable. Such decomposition and / or recombination should be considered equivalent to the present disclosure. The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0129] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for implementing an acoustic camera, comprising the following steps: Step S1: Calculate the phase compensation slope tensor based on the microphone position, pickup angle, and pixel division of the acoustic camera. Step S2: Frame and window the time-domain signal of each microphone. Sum the windowed results according to the number of frames to be processed. Perform time-frequency analysis on the summed results to obtain the frequency domain signal of each microphone. Step S3: Using the quantile method, perform phase compensation on the phase compensation slope tensor for each microphone spectrum, and update the display matrix based on the compensated spectrum; Step S4 involves matching the displayed matrix with the actual graphic to accurately locate the sound-emitting object; Step S1 further includes: Step S100: Obtain the specific location of each microphone in the acoustic camera microphone array, the angle range to be picked up, and the required resolution; Step S101: Project the area to be picked onto a plane, and divide the plane into several pixels according to the resolution. Step S102: Calculate the distance from each pixel to each microphone. Define a reference distance for each pixel, calculate the difference between each microphone and the reference distance, and calculate the phase compensation slope based on the difference. Step S103: Place the phase compensation slope into the corresponding pixel point to obtain the phase compensation slope tensor.
2. The method for implementing an acoustic camera as described in claim 1, characterized in that, The phase compensation slope tensor is obtained by the following formula: in, denoted by , represents the difference between the distance from the i-th and j-th pixels to the m-th microphone and the distance to the center of the array, where fs is the signal sampling rate, c is the speed of sound, and N is the FFT length.
3. The method for implementing an acoustic camera as described in claim 2, characterized in that, Step S2 further includes: Step S200: Frame-by-frame windowing is performed on the time-domain signal of each microphone. Then, based on the number of frames to be processed simultaneously, the windowed time-domain signals are summed to obtain the summed signal. ; Step S201: Sum the signals from each microphone. Time-frequency analysis was performed to obtain the frequency domain signal of each microphone. .
4. The method for implementing an acoustic camera as described in claim 3, characterized in that, Step S3 further includes: Frequency domain signal of each microphone Phase compensation is performed on the corresponding pixels based on the phase compensation slope tensor to obtain the compensation spectrum of each pixel. ; The compensation spectra of the M microphones are summed, and the mean of the summed spectrum amplitudes is taken as the display value of the (i,j)th pixel to obtain the display matrix.
5. The method for implementing an acoustic camera as described in claim 3, characterized in that, Step S3 further includes: Frequency domain signal of each microphone Phase compensation is performed on the corresponding pixels based on the phase compensation slope tensor to obtain the compensation spectrum of each pixel. ; The compensation spectra of the M microphones are summed and differentiated, and the ratio of the summation to the difference is taken as the display value of the (i,j)th pixel to obtain the display matrix.
6. The method for implementing the acoustic camera as described in claim 4 or 5, characterized in that, The frequency domain signal of each microphone Phase compensation is performed on the corresponding pixels based on the phase compensation slope tensor to obtain the compensation spectrum of each pixel. : This is the compensated spectrum of the m-th microphone, where k is the frequency index. .
7. The method for implementing an acoustic camera as described in claim 4, characterized in that, The display matrix is obtained as follows: in, This indicates the selection of a frequency point.
8. An apparatus for implementing an acoustic camera, comprising: The phase compensation slope tensor calculation unit is used to calculate the phase compensation slope tensor based on the microphone position, pickup angle, and pixel division of the acoustic camera. The frequency domain signal acquisition unit is used to divide the time domain signal of each microphone into frames and window it. The windowing results are summed according to the number of frames to be processed, and the summation results are analyzed in time and frequency to obtain the frequency domain signal of each microphone. The phase compensation and pixel calculation unit is used to perform phase compensation on the corresponding pixel points according to the phase compensation slope tensor of the frequency domain signal of each microphone, to obtain the compensated spectrum of each pixel point, and update the display matrix based on the compensated spectrum. The positioning unit is used to match the display matrix with the actual graphic to accurately locate the sound-emitting object; The phase compensation slope tensor calculation unit includes: The information acquisition unit is used to acquire the specific location of each microphone in the acoustic camera microphone array, the angle range to be picked up, and the required resolution; The projection unit is used to project the area to be picked up onto a plane, and divide the plane into several pixels according to the resolution. The phase compensation slope calculation unit is used to calculate the distance from each pixel to each microphone. Each pixel is defined with a reference distance. The difference between each microphone and the reference distance is calculated, and the phase compensation slope is calculated based on the difference. The phase compensation slope tensor calculation unit is used to put the phase compensation slope into the corresponding pixel point to obtain the phase compensation slope tensor.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for implementing an acoustic camera as described in any one of claims 1 to 7.