Conference room system and audio processing method
Through the collaborative work of microphone array and processor, using spectrum array data and angular energy sequence calculation, the problem of inaccurate azimuth determination in the prior art is solved, fast and stable sound source positioning is achieved, and the service quality of the video conferencing system is improved.
Patent Information
- Application Number
- CN202210087776.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-21
- Filing Date
- 2022-01-25
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2042-01-25
AI Technical Summary
The existing azimuth estimation method cannot provide fast and stable azimuth determination, resulting in insufficient accuracy in identifying the speaker's position.
Audio data is collected through the microphone array, spectrum array data is calculated and the angle corresponding to the maximum value is the sound source angle. The buffer and processor are used to perform fast Fourier conversion and angle energy sequence calculations, combined with fixed-point operation and lookup table acceleration processor operation, and directly calculate the energy value of the frequency data to determine the sound source angle.
Fast and stable azimuth determination of sound source azimuth is achieved, reducing calculation time and resource consumption, and improving the service quality of video conferencing systems.
Smart Images

Figure CN115379351B_ABST
Abstract
Description
Technical Field
[0001] This case relates to an electronic operating system and method, and more particularly to a conference room system and audio processing method. Background Art
[0002] As society evolves, the use of video conferencing systems is becoming increasingly widespread. Video conferencing systems should be more than just connecting multiple electronic devices to function; they should also incorporate user-friendly design and keep pace with the times. For example, if a video conferencing system can quickly and accurately identify the caller's location, it can provide better service quality.
[0003] However, existing azimuth estimation methods cannot provide fast and stable azimuth angle determination. For those skilled in the art, how to provide more accurate azimuth angle estimation is a technical problem that needs to be solved urgently. Summary of the Invention
[0004] This summary is intended to provide a simplified summary of the present disclosure so that readers can have a basic understanding of the present disclosure. This summary is not a complete overview of the present disclosure and is not intended to identify important / critical elements of the present disclosure or to define the scope of the present disclosure.
[0005] According to one embodiment of the present case, an audio processing method is disclosed, including: capturing audio data through a microphone array to calculate spectrum array data of the audio data; calculating an angle energy sequence using the spectrum array data; and calculating the difference between a maximum value and a minimum value of the angle energy sequence to determine whether the angle corresponding to the maximum value is a source angle relative to the microphone array.
[0006] In one embodiment, an audio processing method includes: storing the audio data in a first buffer according to a sampling rate; reading a certain number of the audio data from the first buffer to perform a fast Fourier transform operation; processing the certain number of the audio data according to a Fourier length and a window shift; and storing the spectrum array data in a second buffer.
[0007] In one embodiment, the spectrum array data stored in the second buffer includes frequency intensity of each frequency of the audio data.
[0008] In one embodiment, the microphone array includes a first microphone and a second microphone, and the first microphone is arranged at a position at a distance relative to the second microphone, wherein the audio processing method further includes: calculating the angle energy sequence based on the frequency intensity of a first spectrum array data corresponding to the first microphone at each frequency and the frequency intensity of a second spectrum array data corresponding to the second microphone at each frequency, wherein the angle energy sequence includes the sound energy at each angle on the plane.
[0009] In one embodiment, the audio processing method further includes calculating a duration between a first audio data from the first microphone and a second audio data from the second microphone.
[0010] In one embodiment, the audio processing method further includes: correcting the time of the first audio data and the second audio data according to the duration to align the waveforms of the first audio data and the second audio data; and obtaining the first spectrum array data and the second spectrum array data using the first audio data and the second audio data with aligned waveforms.
[0011] In one embodiment, the audio processing method further includes: calculating the sum of the squares of the frequency intensity of the first spectrum array data at each frequency and the frequency intensity of the second spectrum array data at each frequency to obtain the angle energy sequence.
[0012] In one embodiment, the audio processing method further includes: determining whether the difference between the maximum value and the minimum value of the angle energy sequence is greater than a threshold value; and when the difference is greater than the threshold value, determining that the angle corresponding to the maximum value is the source angle relative to the microphone array.
[0013] In one embodiment, the audio processing method further includes: when the difference is not greater than the threshold, determining that the audio data corresponding to the maximum value is noise data.
[0014] In one embodiment, the audio processing method further includes: outputting the source angle as an angle of a sound source generating the audio data relative to the microphone array.
[0015] According to another embodiment, a conference room system is disclosed, comprising a microphone array and a processor. The microphone array is configured to capture audio data. The processor, electrically coupled to the microphone array, is configured to: calculate spectrum array data of the audio data; calculate an angle energy sequence using the spectrum array data; and calculate the difference between a maximum value and a minimum value of the angle energy sequence to determine whether the angle corresponding to the maximum value is a source angle relative to the microphone array.
[0016] In one embodiment, the conference room system further includes a first buffer and a second buffer. The first buffer is electrically coupled to the microphone array, wherein the first buffer is configured to store the audio data having a sampling rate. The second buffer is electrically coupled to the first buffer and the processor, wherein the processor is further configured to: read a certain number of audio data from the first buffer to perform a fast Fourier transform operation; calculate spectrum array data for the certain number of audio data based on a Fourier length and a window shift; and store the spectrum array data in the second buffer.
[0017] In one embodiment, the spectrum array data stored in the second buffer includes frequency intensity of each frequency of the audio data.
[0018] In one embodiment, the microphone array includes a first microphone and a second microphone, the first microphone is arranged at a distance relative to the second microphone, wherein the processor is further configured to: calculate the angle energy sequence based on the frequency intensity of each frequency of a first spectrum array data corresponding to the first microphone and the frequency intensity of each frequency of a second spectrum array data corresponding to the second microphone, wherein the angle energy sequence includes the sound energy at each angle on the plane.
[0019] In one embodiment, the processor is further configured to calculate a duration between a first audio data from the first microphone and a second audio data from the second microphone.
[0020] In one embodiment, the processor is further configured to: correct the time of the first audio data and the second audio data according to the duration to align the waveforms of the first audio data and the second audio data; and use the first audio data and the second audio data with aligned waveforms to obtain the first spectrum array data and the second spectrum array data.
[0021] In one embodiment, the processor is further configured to calculate the sum of the squares of the frequency intensity of the first spectrum array data at each frequency and the frequency intensity of the second spectrum array data at each frequency to obtain the angular energy sequence.
[0022] In one embodiment, the processor is further configured to: determine whether the difference between the maximum value and the minimum value of the angle energy sequence is greater than a threshold value; and when the difference is greater than the threshold value, determine that the angle corresponding to the maximum value is the source angle relative to the microphone array.
[0023] In one embodiment, the processor is further configured to: when the difference is not greater than the threshold, determine that the audio data corresponding to the maximum value is noise data.
[0024] In one embodiment, the processor is further configured to: output the source angle as an angle of a sound source generating the audio data relative to the microphone array. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The following detailed description, when read in conjunction with the accompanying drawings, will facilitate a better understanding of the aspects of this disclosure. It should be noted that, due to practical requirements for illustration, the features in the drawings are not necessarily drawn to scale. In fact, the dimensions of the features may be arbitrarily increased or decreased for clarity of discussion.
[0026] Figure 1 A block diagram of a conference room system according to some embodiments of the present invention is shown;
[0027] Figure 2 A flowchart of an audio processing method according to some embodiments of the present invention is shown.
[0028]
Explanation of symbols
[0029] 100: Conference Room System
[0030] 110: Microphone array
[0031] 120: Buffer
[0032] 121: First buffer
[0033] 122: Second buffer
[0034] 140: Processor
[0035] 200: Audio Processing Methods
[0036] S210~S260: Steps DETAILED DESCRIPTION
[0037] The following disclosure provides many different embodiments to implement the different features of the present invention. The following describes embodiments of components and arrangements to simplify the present invention. Of course, these embodiments are merely exemplary and are not intended to be restrictive. For example, the use of terms such as "first" and "second" to describe components in the present invention is merely to distinguish between identical or similar components or operations. These terms are not intended to limit the technical components of the present invention, nor are they intended to limit the order or sequence of operations. In addition, the present invention may repeat component symbols and / or letters in each embodiment, and the same technical terms may use the same and / or corresponding component symbols in each embodiment. This repetition is for the purpose of simplicity and clarity, and does not in itself indicate the relationship between the various embodiments and / or configurations discussed.
[0038] Please refer to Figure 1, which illustrates a block diagram of a conference room system 100 according to some embodiments of the present invention. The conference room system 100 includes a microphone array 110, a buffer 120, and a processor 140. The microphone array 110 is electrically coupled to the buffer 120. The buffer 120 is electrically coupled to the processor 140. In some embodiments, the buffer 120 includes a first buffer 121 (or a ring buffer) and a second buffer 122 (or a moving window buffer). The first buffer 121 is electrically coupled to the second buffer 122. Figure 1 As shown, the first buffer 121 is electrically coupled to the microphone array 110 , and the second buffer 122 is electrically coupled to the processor 140 .
[0039] In some embodiments, microphone array 110 is configured to capture audio data. For example, microphone array 110 includes multiple microphones that are continuously activated to capture any audio data, causing the audio data to be stored in first buffer 121. In some embodiments, the audio data captured by microphone array 110 is stored in first buffer 121 at a sampling rate. For example, the sampling rate may be 48 kHz, which samples the analog audio signal 48,000 times per second, resulting in the audio data being stored in first buffer 121 in a discrete data format.
[0040] In some embodiments, conference room system 100 can instantly detect the angle of a sound's source. For example, microphone array 110 is mounted on a conference table in a conference room. Conference room system 100 can use audio data received by microphone array 110 to determine whether the sound's source is located at an angle or range of angles relative to microphone array 110 within a 360-degree angle. A detailed method for calculating the angle of a sound source is described below.
[0041] In some embodiments, processor 140 calculates spectral array data of audio data. For example, the sampling rate of the audio data stored in first buffer 121 is 48 kHz, which means 48,000 samples per second. To facilitate the calculation of sample data in this embodiment, 1024 samples are used as one frame of data, meaning that one frame lasts approximately 21.3 (1024 / 48,000) milliseconds.
[0042] In some embodiments, the microphone array 110 continuously generates audio data and, after sampling at a sampling rate of 48 kHz, stores multiple frames in the first buffer 121. The size of the first buffer 121 can be a 2-second buffer space, which can be designed or adjusted based on actual needs and is not limited thereto.
[0043] In some embodiments, the processor 140 reads a certain number of audio data (e.g., one frame) from the first buffer 121 as input for a fast Fourier transform (FFT) operation. In some embodiments, when the first buffer 121 does not store any audio data, the processor 140 continuously detects whether the amount of data stored in the first buffer 121 has reached a feasible number of data, i.e., one frame of data. The processor 140 reads each frame of audio data from the first buffer 121 to calculate the fast Fourier transform (FFT) and stores the calculated result in the second buffer 122.
[0044] In some embodiments, the processor 140 calculates spectrum array data for this frame of audio data based on a Fourier length (FFT length) and a window shift (FFT shift). The Fourier length can be 1024 sample data items, and the window shift can be 512 sample data items. It is worth mentioning that the size of the window shift affects the number of frames used to subsequently calculate the degree of arrival (DOA). For example, when the window shift is 512 sample data items, then after 0.75 seconds of audio data is input into the fast Fourier transform operation, approximately 70 frames (0.75 seconds * 48000 / 512) of spectrum array data can be obtained. When the window shift is 1024 sample data items, then after 0.75 seconds of audio data is input into the fast Fourier transform operation, approximately 35 frames (0.75 seconds * 48000 / 1024) of spectrum array data can be obtained. In other words, the size of the window shift affects the accuracy of the subsequent angle-of-arrival calculation. For example, a window shift of 512 allows for more frames of audio data to be used for angle-of-arrival calculation. Therefore, processor 140 can instantly calculate the audio data's spectrum array data based on each newly arrived frame of audio data.
[0045] In some embodiments, processor 140 pre-stores a lookup table that records Fast Fourier Transform (FFT) angles and their corresponding sine function values. During each FFT operation, processor 140 can directly obtain the value from the lookup table without actually performing the FFT operation. This can improve the processing speed of processor 140.
[0046] During each Fast Fourier Transform operation, the processor 140 can directly obtain the sine and cosine values by looking up the pre-established trigonometric function table without having to calculate the trigonometric function values again, thereby speeding up the Fast Fourier Transform operation.
[0047] In some embodiments, the second buffer 122 includes a storage space, such as temporary storage space capable of storing 0.75 seconds of audio data. After the processor 140 calculates the spectrum array data for each frame from the audio data in the first buffer 121, the processor 140 stores the spectrum array data in the second buffer 122. The spectrum array data stored in the second buffer 122 includes the frequency intensity of each frequency of the audio data. For example, the second buffer 122 stores the intensity distribution of each frequency for 0.75 seconds.
[0048] In some embodiments, the processor 122 simply reads 0.75 seconds of audio data from the first buffer 121 in an initial state (e.g., when the second buffer 122 does not store any spectrum array data), and calculates the spectrum array data, so that the second buffer 122 stores 0.75 seconds of spectrum array data. Subsequently, the processor 122 obtains each newly arrived frame of audio data from the first buffer 121 to calculate the spectrum array data, and deletes the oldest frame of data from the 0.75 seconds of data in the second buffer 122, storing the new frame of spectrum array data in the second buffer 122. In other words, when the processor 122 subsequently calculates the energy sequence for each angle from the spectrum array data in the second buffer 122, for example, if the second buffer 122 stores a total of 70 frames of data, of which 69 frames are old data and 1 frame is new data, the energy sequence for each angle can be calculated using the old spectrum array data. Therefore, only the new frame of spectrum array data is needed to calculate the energy sequence for each angle. This reduces the time required to calculate the energy sequence for each angle. The following describes how to calculate the energy sequence for each angle from the spectrum array data.
[0049] In some embodiments, the microphone array 110 includes multiple microphones, each of which captures audio data, allowing the processor 140 to calculate the audio data captured by each microphone to obtain corresponding spectrum array data. Therefore, the processor 140 can calculate the frequency intensity of the audio data of each microphone at each frequency from the audio data of each microphone. In other embodiments, the microphone array 110 includes multiple microphones arranged in a ring, for example, the microphones are arranged in a ring with a radius of 4.17 cm. For ease of explanation, the microphone array 110 is described as an example with two microphones.
[0050] In some embodiments, microphone array 110 includes a first microphone and a second microphone. The first microphone is positioned at a distance from the second microphone. In some embodiments, processor 140 calculates first spectrum array data for the first microphone and second spectrum array data for the second microphone. The spectrum array data calculation process is as described above and will not be repeated here.
[0051] Since the configuration distance between the microphones is a known value and the distance between the microphones is quite small, for the same sound source, the waveforms of the audio data generated by each microphone will be similar, and there will be a time delay between each waveform. In some embodiments, the processor 140 can calculate the source angle of the sound source relative to the microphone array 110 through the time delay or phase angle of the audio data of the first microphone and the audio data of the second microphone. For example, the processor 140 calculates the time delay between the first audio data of the first microphone and the second audio data of the second microphone to correct the time of the first audio data and the second audio data according to the time delay to align the waveforms of the first audio data and the second audio data. Then, the processor 140 uses the first audio data and the second audio data with aligned waveforms to obtain the first spectrum array data and the second spectrum array data. It is worth mentioning that the delay superposition technology can be implemented in the time domain or frequency domain, and the present case is not limited to this.
[0052] In some embodiments, the processor 140 calculates an angular energy sequence based on the frequency intensity of the first spectrum array data of the first microphone at each frequency and the frequency intensity of the second spectrum array data of the second microphone at each frequency. The angular energy sequence includes the sound energy at each angle on the plane. For example, the processor 140 uses the first spectrum array and the second spectrum array to calculate the delayed superposition spectrum from 0° to 360°. The processor 140 calculates the square sum of the frequency intensity of the first spectrum array data at each frequency and the frequency intensity of the second spectrum array data at each frequency to obtain the angular energy sequence. In some embodiments, the processor 140 can calculate the angular energy every 1° angle, or can calculate the energy of an angular range every 10° angle (for example, 0° to 9° angle), but the present invention is not limited to this. In this way, the energy distribution of each angle or angle range from 0° to 360° on the plane can be calculated, for example, the maximum energy is at 40° angle and the minimum energy is at 271° angle.
[0053] It is worth mentioning that in conventional technology, after performing a fast Fourier transform to calculate the frequency data (such as the SRP-PHAT algorithm), it is necessary to perform an inverse Fourier transform (IFFT) operation to convert the frequency data back into time domain data to obtain a time curve. Then, it is necessary to calculate the area of the time curve to obtain the energy value as the angle energy data. However, the area calculated from the frequency curve, that is, the energy value, does not change after the frequency domain is converted to the time domain. Therefore, in this case, after performing a fast Fourier transform (FFT) to calculate the frequency data, it is not necessary to perform an inverse Fourier transform (IFFT) operation. Instead, the frequency data obtained by the fast Fourier transform (FFT) is directly used to calculate the energy value of the angle to obtain the angle energy sequence (the energy value corresponding to each angle or angle range). In this way, the time for performing the inverse Fourier transform (IFFT) operation can be saved, greatly shortening the cost and time of the calculation.
[0054] In some embodiments, the processor 140 determines whether the difference between the maximum value and the minimum value of the angle energy sequence is greater than a threshold value. When the difference is greater than the threshold value, the angle corresponding to the maximum value is determined to be the source angle relative to the microphone array. When the difference is not greater than the threshold value, the audio data corresponding to the maximum value is determined to be noise data. For example, if the difference between the energy of the maximum value (at an angle of 40°) and the energy of the minimum value (at an angle of 271°) is greater than the threshold value, it means that the source of the sound is meaningful, such as someone is talking, and this angle (40° angle) is output to, for example, a display device ( Figure 1 Not shown. On the other hand, if the difference between the maximum energy (at an angle of 40°) and the minimum energy (at an angle of 271°) is not greater than the threshold, it indicates that there is noise or interference in the environment, and the maximum value is simply a loud noise. Therefore, the angle corresponding to the maximum value is not used as the source angle of the sound source.
[0055] In some embodiments, the processor 140 uses fixed-point arithmetic to process the fast Fourier transform operation, and accelerates the processing of audio data by using hardware support for floating-point number conversion to fixed-point number calculation method.
[0056] Please refer to the following instructions Figure 1 and Figure 2 . Figure 2 FIG2 is a flow chart illustrating an audio processing method 200 according to some embodiments of the present invention. The audio processing method 200 may be executed by at least one component in the conference room system 100 .
[0057] In step S210 , audio data is captured through the microphone array 110 to calculate spectrum array data of the audio data.
[0058] In some embodiments, audio data captured by the microphone array 110 is stored in the first buffer 121 at a sampling rate of, for example, 48 kHz. The first buffer 121 is a temporary storage space capable of storing, for example, two seconds of audio signals. As the microphone array 110 continuously captures audio signals, the audio signals are stored in the first buffer 121 in a first-in, first-out order. If a frame of audio data includes 1024 samples, the first buffer 121 can store multiple frames for subsequent Fast Fourier Transform (FFT) calculations.
[0059] In step S220 , the angular energy sequence is calculated using the spectrum array data.
[0060] In some embodiments, the processor 140 reads a certain number of audio data (e.g., one frame) from the first buffer 121 as input for a fast Fourier transform (FFT) operation. In some embodiments, the processor 140 calculates spectrum array data for this frame of audio data based on a Fourier length and a window shift. The Fourier length can be one frame of audio data (e.g., 1024 samples), and the window shift can be 512 samples. The processor 140 performs a FFT operation on each frame of audio data to obtain spectrum array data for each frame. The spectrum array data is stored in the second buffer 122 in a first-in, first-out order. The storage space of the second buffer 122 is, for example, a temporary storage space capable of storing 0.75 seconds of audio data. Therefore, each time the processor 140 calculates a new frame of spectrum array data, it first deletes the oldest frame of data in the second buffer 122, so that the new frame of spectrum array data is stored in the last storage space in the second buffer 122 in a first-in, first-out order.
[0061] In step S230 , the difference between the maximum value and the minimum value of the angle energy sequence is calculated.
[0062] In some embodiments, the microphone array 110 includes multiple microphones. The processor 140 reads the audio data generated by each of these microphones and calculates spectrum array data for each of these audio data. For example, the processor 140 calculates first spectrum array data for the first microphone and second spectrum array data for the second microphone. The process for calculating the spectrum array data is described above and will not be repeated here.
[0063] In some embodiments, the processor 140 can calculate the source angle of the sound source relative to the microphone array 110 through the time delay or phase angle of the audio data of the first microphone and the audio data of the second microphone. In addition, the processor 140 calculates the angle energy sequence based on the frequency intensity of the first spectrum array data of the first microphone at each frequency and the frequency intensity of the second spectrum array data of the second microphone at each frequency. The angle energy sequence includes the sound energy at each angle on the plane. In this way, the sound energy at each angle can be updated each time one frame of spectrum array data is generated. In some embodiments, the processor 140 can obtain the maximum and minimum values of the sound energy from 0° to 360°.
[0064] In step S240, a determination is made as to whether the difference is greater than a threshold value. In some embodiments, when processor 140 determines that the difference between the maximum and minimum values of the angle energy sequence is greater than the threshold value, step S250 is executed. In step S250, if the difference is greater than the threshold value, the angle corresponding to the maximum value is determined to be the source angle relative to the microphone array. If, in step S240, the difference is determined not to be greater than the threshold value, step S260 is executed. In step S260, the audio data corresponding to the maximum value is determined to be noise data.
[0065] In some embodiments, since the audio processing method 200 obtains the source angle of the sound source in real time, the processor 140 further outputs the source angle, for example, to a display device ( Figure 1 Not shown) for relevant personnel to watch, or control another camera according to the source angle to control the camera to rotate to the source angle to shoot the image of the sound source or make relevant close-ups.
[0066] In some embodiments, the processor 140 may be implemented as, but not limited to, a central processing unit (CPU), a system on chip (SoC), an application processor, an audio processor, a digital signal processor (DSP), or a processing chip or controller with specific functions.
[0067] In some embodiments, a non-transitory computer readable recording medium is provided, which can store a plurality of program codes. The program codes are loaded into a computer readable recording medium. Figure 1 After the processor 140 is executed, the processor 140 executes the program code and performs the following Figure 2For example, the processor 140 obtains audio data through the microphone array 110 and calculates spectrum array data of the audio data, and uses the spectrum array data to calculate an angle energy sequence, and calculates the difference between the maximum value and the minimum value of the angle energy sequence to determine whether the angle corresponding to the maximum value is the source angle relative to the microphone array 110.
[0068] In summary, the conference room system and audio processing method of this invention have the following advantages: A lookup table that records angle values and their corresponding sine values eliminates (effectively reduces) the computation time required by processor 140 to calculate each Fourier transform. Furthermore, the provision of a first buffer 121 allows the recording process and angle calculation process to be performed separately. Furthermore, the conference room system includes hardware supporting fixed-point arithmetic, significantly accelerating computation time. Furthermore, after obtaining the spectrum array data, this invention eliminates the need for performing an inverse Fourier transform to convert it into time-domain data. Instead, it directly calculates the frequency data to calculate the energy of the sound source, shortening the calculation time. Furthermore, the second buffer 122 stores a 0.75-second spectrum array. Each time a new frame of data is calculated, the energy value for each angle is updated by simply deleting the oldest frame of spectrum data in the second buffer 122 and adding the new frame. Compared to conventional methods that require two seconds to recalculate each angle, this invention can instantly reflect the current sound source angle.
[0069] Furthermore, the conference room system and audio processing method of this case calculates the difference between the maximum and minimum values each time to determine whether the current maximum value of the sound source is noise, thereby preventing the judgment of the sound source from being interfered with by noise, thereby improving the stability and accuracy of the system.
[0070] The above content summarizes the features of several embodiments so that those skilled in the art can better understand the aspects of this invention. Those skilled in the art should understand that the above content can be easily used as a basis for designing or modifying other variations to achieve the same objectives and / or advantages as the embodiments described herein without departing from the spirit and scope of this invention. The above content should be understood as an example of this invention, and the scope of protection should be determined by the claims.
Claims
1. An audio processing method, characterized in that: include: Capturing audio data through a microphone array to calculate spectrum array data of the audio data, wherein the microphone array includes a first microphone and a second microphone, the first microphone being disposed at a distance relative to the second microphone; directly calculating an angle energy sequence using the spectrum array data, wherein the spectrum array data includes first spectrum array data corresponding to the first microphone and second spectrum array data corresponding to the second microphone, wherein the sum of the frequency intensity of the first spectrum array data at each frequency and the square of the frequency intensity of the second spectrum array data at each frequency are calculated to obtain the angle energy sequence, wherein the angle energy sequence includes sound energy at angles from 0° to 360° on a plane; Calculating a difference between a maximum energy value and a minimum energy value of the angle energy sequence; as well as When the difference is greater than a threshold, it is determined that the angle corresponding to the maximum energy value is a source angle relative to the microphone array.
2. The audio processing method according to claim 1, wherein: Also includes: storing the audio data in a first buffer according to a sampling rate; Reading a certain amount of the audio data from the first buffer to perform a fast Fourier transform operation; The audio data of the data number is processed according to a Fourier length and a window shift; and The spectrum array data is stored in a second buffer.
3. The audio processing method according to claim 2, wherein: The spectrum array data stored in the second buffer includes the frequency intensity of each frequency of the audio data.
4. The audio processing method according to claim 1, wherein: Also includes: A duration between first audio data from the first microphone and second audio data from the second microphone is calculated.
5. The audio processing method according to claim 4, wherein: Also includes: Correcting the time of the first audio data and the second audio data according to the duration to align the waveforms of the first audio data and the second audio data; as well as The first frequency spectrum array data and the second frequency spectrum array data are obtained using the first audio data and the second audio data having aligned waveforms.
6. The audio processing method according to claim 1, wherein: Also includes: When the difference is not greater than the threshold, the audio data corresponding to the maximum energy value is determined to be noise data.
7. The audio processing method according to claim 1, wherein: Also includes: The source angle relative to the microphone array is output as an angle of a sound source generating the audio data relative to the microphone array.
8. A conference room system, characterized in that: include: a microphone array configured to capture audio data, wherein the microphone array includes a first microphone and a second microphone, the first microphone being positioned at a distance relative to the second microphone; and a processor electrically coupled to the microphone array and configured to: Calculating spectrum array data of the audio data; directly calculating an angle energy sequence using the spectrum array data, wherein the spectrum array data includes first spectrum array data corresponding to the first microphone and second spectrum array data corresponding to the second microphone, wherein the sum of the frequency intensity of the first spectrum array data at each frequency and the square of the frequency intensity of the second spectrum array data at each frequency are calculated to obtain the angle energy sequence, wherein the angle energy sequence includes sound energy at angles from 0° to 360° on a plane; Calculating a difference between a maximum energy value and a minimum energy value of the angle energy sequence; as well as When the difference is greater than a threshold, it is determined that the angle corresponding to the maximum energy value is a source angle relative to the microphone array.
9. The conference room system according to claim 8, wherein: Also includes: a first buffer electrically coupled to the microphone array, wherein the first buffer is configured to store the audio data including a sampling rate; and a second buffer electrically coupled to the first buffer and the processor, wherein the processor is further configured to: Reading a certain amount of the audio data from the first buffer to perform a fast Fourier transform operation; Calculate the spectrum array data for the audio data of the data number according to a Fourier length and a window shift; as well as The spectrum array data is stored in the second buffer.
10. The conference room system according to claim 9, wherein: The spectrum array data stored in the second buffer includes the frequency intensity of each frequency of the audio data.
11. The conference room system according to claim 8, wherein: The processor is further configured to: A duration between first audio data from the first microphone and second audio data from the second microphone is calculated.
12. The conference room system according to claim 11, wherein: The processor is further configured to: Correcting the time of the first audio data and the second audio data according to the duration to align the waveforms of the first audio data and the second audio data; as well as The first frequency spectrum array data and the second frequency spectrum array data are obtained using the first audio data and the second audio data having aligned waveforms.
13. The conference room system according to claim 8, wherein: The processor is further configured to: When the difference is not greater than the threshold, the audio data corresponding to the maximum energy value is determined to be noise data.
14. The conference room system according to claim 8, wherein: The processor is further configured to: The source angle relative to the microphone array is output as an angle of a sound source generating the audio data relative to the microphone array.
Citation Information
Patent Citations
Device and method for determining sound source direction
US20070160230A1