Automatic calibration of microphone arrays for telepresence conferencing
By measuring the reverberant sound field generated by the loudspeaker and calculating the power spectral density ratio of the microphone and loudspeaker, a calibration filter is automatically generated, solving the problems of cumbersome and inaccurate microphone and loudspeaker calibration in remote presentation systems, and achieving high-quality audio signal capture and presentation.
Patent Information
- Application Number
- CN202080106843.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-30
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2040-10-30
AI Technical Summary
In existing telepresence conferencing systems, microphone and speaker calibration requires manual intervention, has strict requirements on hardware location, is prone to errors, and traditional methods are cumbersome and inflexible.
By measuring the reverberant sound field generated by the loudspeaker, calculating the power spectral density ratio of the microphone and loudspeaker, and automatically generating a calibration filter, adaptive calibration of the microphone and loudspeaker is achieved, reducing dependence on hardware location and environment.
It enables automatic calibration without human intervention, improves the calibration accuracy and consistency of microphones and speakers, and enhances the audio quality of the telepresence system.
Smart Images

Figure CN116472724B_ABST
Abstract
Description
Technical Field
[0001] This description relates to the calibration of microphones and speakers used in applications such as telepresence conferencing. Background Technology
[0002] A telepresence conferencing system may include a large number of microphones for detecting directional audio signals from users and multiple speakers for providing directional audio signals to users. Summary of the Invention
[0003] In one general aspect, a method may include receiving a reverberant sound field based on an audio signal generated by each speaker in a speaker array via each microphone in a microphone array. The method may further include generating a corresponding power spectral density for each microphone in the microphone array and each speaker in the speaker array, based on the corresponding reverberant sound field generated by the speaker and received at the microphone. The method may further include generating a corresponding calibration filter for each microphone in the microphone array as a ratio of the average power spectral density across the speaker array and microphone array to the average power spectral density across the speaker array. The method may further include recording an acoustic signal from a user via the microphone array using the corresponding calibration filter generated for each of the microphone arrays, each of the microphone arrays recording substantially the same spectrum of the acoustic signal using the corresponding calibration filter.
[0004] In another general aspect, a computer program product includes a non-transitory storage medium, the computer program product including code that, when executed by processing circuitry of a computing device, causes the processing circuitry to perform a method. The method may include generating a corresponding power spectral density for each microphone in a microphone array and each speaker in a speaker array, based on a corresponding reverberant sound field generated by the speaker and received at the microphone. The method may also include generating a corresponding calibration filter for each microphone in the microphone array, as a ratio of the average power spectral density across the speaker array and the microphone array to the average power spectral density across the speaker array. The method may further include recording an acoustic signal from a user through the microphone array using the corresponding calibration filter generated for each of the microphone arrays, each of the microphone arrays recording substantially the same spectrum of the acoustic signal using the corresponding calibration filter.
[0005] In another general aspect, an electronic device includes a memory and control circuitry coupled to the memory. The control circuitry can be configured to receive a reverberant sound field based on an audio signal generated by each speaker in a speaker array via each microphone in a microphone array. The control circuitry can also be configured to generate a corresponding power spectral density for each microphone in the microphone array and each speaker in the speaker array, based on the corresponding reverberant sound field generated by the speaker and received at the microphone. The control circuitry can also be configured to generate a corresponding calibration filter for each microphone in the microphone array, as a ratio of the average power spectral density across the speaker array and microphone array to the average power spectral density across the speaker array. The control circuitry can also be configured to record an acoustic signal from a user via the microphone array using the corresponding calibration filter generated for each of the microphone arrays, each of the microphone arrays recording substantially the same spectrum of the acoustic signal using the corresponding calibration filter.
[0006] In another general aspect, a method may include receiving a reverberant sound field based on an audio signal generated by each speaker in a speaker array via each microphone in a microphone array. The method may also include generating a corresponding power spectral density for each microphone in the microphone array and each speaker in the speaker array, based on the corresponding reverberant sound field generated by the speaker and received at the microphone. The method may further include generating a corresponding calibration filter for each microphone in the microphone array as a ratio of the average power spectral density across the speaker array and microphone array to the average power spectral density across the speaker array. The method may also include generating an acoustic signal by the speaker array using the corresponding calibration filter generated for each speaker in the speaker array, each speaker in the speaker array generating a substantially identical spectrum of acoustic signal in response to the same output stimulus using the corresponding calibration filter.
[0007] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features will be apparent from the description, drawings, and claims. Attached Figure Description
[0008] Figure 1A This is a diagram illustrating an example electronic environment used to implement the technical solutions described herein.
[0009] Figure 1B It is a diagram. Figure 1A The diagram shows an example configuration of microphones and speakers in an electronic environment.
[0010] Figure 1C This is a diagram illustrating an example configuration of microphones and speakers within a remote presentation system.
[0011] Figure 2 The diagram is in Figure 1A The flowchart illustrates an example method for implementing a technical solution in an electronic environment.
[0012] Figure 3 This is a flowchart illustrating an example process for calibrating a microphone array according to the technical solution.
[0013] Figure 4A The diagram is in Figure 1A The diagram shows a graph of the original impulse response function for an example electronic environment ranging from two speakers to four microphones.
[0014] Figure 4B It is a diagram and comes from Figure 4A A graph of an example time-dependent energy metric associated with the original impulse response function.
[0015] Figure 4C The diagram shows the average values across all speakers and microphones, from... Figure 4B Example time-related energy metric curve.
[0016] Figure 4D This is a graph illustrating an example of a decayed normalized impulse response function corresponding to the original impulse response function for four microphones and two speakers.
[0017] Figure 4E It is a diagram. Figure 4D The graph shows an example segment of the decayed normalized impulse response function.
[0018] Figure 4F The diagram is from Figure 4E The graph shows an example multichannel white noise autocorrelation function derived from a segment of the decayed normalized impulse response function.
[0019] Figure 4G The diagram corresponds to Figure 4F The example power spectral density curve of the multichannel white noise autocorrelation function is shown.
[0020] Figure 4H The diagram shows the average value from the speaker. Figure 4G Example power spectral density curve.
[0021] Figure 4I The diagram is from Figure 4H The graph of an example microphone calibration filter derived from the average power spectral density of the loudspeaker.
[0022] Figure 4J The diagram shows the average signal from the microphone. Figure 4G Example power spectral density curve.
[0023] Figure 4K The diagram is from Figure 4J The example speaker calibration filter is plotted using the microphone average power spectral density.
[0024] Figure 5 Examples of computer devices and mobile computer devices that can be used with the circuits described herein are illustrated. Detailed Implementation
[0025] To accurately capture signals from microphones that can be used to generate high-quality, direction-sensitive audio signals, each microphone in the array can be calibrated relative to the other microphones (e.g., microphone gain). Furthermore, to accurately reproduce realistic spatialized output in a telepresence system, each speaker must also be calibrated relative to the other speakers (e.g., speaker gain). Traditional methods of performing such calibrations involve using external hardware, such as sound sources and microphones located at the intended positions of the user / speaker in the telepresence conferencing system.
[0026] However, for such telepresence systems, the technical problems with the conventional methods used to calibrate microphones and speakers are the cumbersome use and storage of the equipment, the need for manual setup and disassembly, and the susceptibility to errors if the hardware is not accurately positioned relative to the actual user's location. Furthermore, the device may also malfunction if the hardware is not properly configured, such as volume knobs or equalizer adjustments.
[0027] Compared to conventional methods for solving the aforementioned technical problems, the technical solution involves generating calibration filters for microphones and / or speakers by deriving the power spectral density at each microphone in response to the signal generated by each speaker. For example, a computer in an improved telepresence system can measure the raw impulse response function corresponding to each channel (i.e., each speaker / microphone pair). In some implementations, the computer normalizes the raw impulse response function based on the contribution of different reverberant reflections to the reverberant sound field energy. The computer then extracts a segment of each impulse response function between the start and end times after the time the speaker generates the signal. The computer then generates a white noise power spectral density for each channel based on the segment. The microphone calibration function is then based on the reciprocal of the power spectral density averaged on the speaker. The speaker calibration function is based on the reciprocal of the power spectral density averaged on the microphone.
[0028] The technical advantages of the above-described solution are that it is insensitive to room configuration and can be executed automatically without human intervention. It is also insensitive to hardware configuration, such as the relative positions of the microphone and speakers. Furthermore, it requires no external hardware beyond what is already present in the telepresence system. Essentially, the user can generate a calibration filter simply by tossing a switch.
[0029] In some implementations, the computer normalizes the raw impulse response function for all channels using the impulse response energy averaged across the speakers and microphones. In some implementations, the start time is based on the distance the echo travels in the reverberant sound field. In some implementations, the end time is based on the background noise associated with the measurement process. In some implementations, the white noise power spectral density of a channel is based on the Fourier transform of the white noise autocorrelation of a sub-segment of that channel. In some implementations, the Fourier transform replaces the windowed version of the white noise autocorrelation function.
[0030] Figure 1A This is a diagram illustrating an example electronic environment 100 in which the aforementioned improved technology can be implemented. For example... Figure 1A As shown, the example electronic environment 100 includes a computer 120.
[0031] Computer 120 includes a network interface 122, one or more processing units 124, and memory 126. The network interface 122 includes, for example, an Ethernet adapter, for converting electronic and / or optical signals received from a network into electronic form for use by computer 120. The group of processing units 124 includes one or more processing chips and / or components. Memory 126 includes volatile memory (e.g., RAM) and non-volatile memory, such as one or more ROMs, disk drives, solid-state drives, etc. The group of processing units 124 and memory 126 together form control circuitry configured and arranged to perform the various methods and functions described herein.
[0032] In some embodiments, one or more components of computer 120 may include a processor (e.g., processing unit 124) configured to process instructions stored in memory 126. Examples of such instructions depicted in FIG1 include a reverberation field manager 130, an impulse response manager 140, a power spectral density manager 150, and a calibration filter manager 160. Moreover, as shown in FIG1, memory 126 is configured to store various types of data, which will be described with respect to the corresponding managers using that data.
[0033] The reverberant sound field manager 130 is configured to generate reverberant sound field data 132, which represents the reverberant sound field generated by the speakers and used to measure the impulse response at the microphone. The sound field is reverberant because once converted into an audio signal at the speakers, the audio signal may be reflected by nearby walls, ceilings, floors, and objects in the room containing the computer 120, speakers, and microphone.
[0034] In some implementations, the reverberant sound field data 132 is generated via a signal of a specific mathematical form emitted from a loudspeaker, which itself helps to measure the frequency response of a microphone with a high signal-to-noise ratio (SNR). The reverberant sound field data 132 is then derived from reflections from, for example, walls, ceilings, floors, and objects in the room including the computer 120, microphone, and loudspeakers (see [link to relevant documentation]). Figure 1B In some embodiments, the mathematical form includes a swept-frequency sinusoidal chirped signal represented by the swept-frequency sinusoidal chirped signal data 134. For example, in some embodiments, the swept-frequency sinusoidal chirped signal x(t) may take the following form:
[0035]
[0036] Where f1 is the starting (first) frequency, f2 is the ending (second) frequency, and T is the chirp duration. The advantage of this swept-frequency sinusoidal chirp signal is that it reduces the impact of frequency response measurements at the microphone on the speaker's nonlinearity.
[0037] Impulse response manager 140 is configured to generate impulse response data 144 representing the impulse response function h(m,s,t) of a speaker-microphone channel, where m represents the microphone index identifying the microphone of the microphone array, s represents the speaker index identifying the speaker of the speaker array, and t represents time. Therefore, each pair (m,s) represents a channel on which impulse response manager 140 measures the original impulse response function h(m,s,t) based on reverberation field data 132. In some embodiments, impulse response manager 140 includes attenuation normalization manager 141 and / or segment manager 142.
[0038] Attenuation normalization manager 141 is configured to provide a consistent normalization factor across multiple channels, represented by normalized data 145. The normalization factor is used to normalize the contribution of different reverberation reflections to the calculation of calibration filters for the microphone and speaker. This normalization factor makes the calibration results insensitive to hardware configuration. In some embodiments, attenuation normalization manager 141 is configured to estimate the impulse response energy e(m, s, t) as a function of time for each channel. In some embodiments, attenuation normalization manager 141 smooths the square h(m, s, t) of the impulse response function over time. 2To perform such an estimation. In some implementations, the attenuation normalization manager 141 performs time smoothing by performing a moving average of the impulse response data over a specified duration (e.g., 1.5 ms), although the duration may be less than or greater than 1.5 ms. In some implementations, the attenuation normalization manager 141 compensates for the delay caused by time smoothing to align the energy estimate with the corresponding impulse response in time. In some implementations, the attenuation normalization manager 141 calculates the normalization factor as the square root of the impulse response energy averaged over the microphone and speaker, i.e.,
[0039]
[0040]
[0041] The segment manager 142 is configured to extract time-based segments of the impulse response function h(m, s, t) between a first time and a second time to produce a segment representing the impulse response h. analysis The sub-segment data 146 is (m, s, t). In some implementations, the first time is based on the distance the echo travels (i.e., the reverberant sound field reflected from walls, ceilings, floors, or objects). Defining the first time in this way can mitigate distance-related energy effects. For example, for a sound speed of 343 m / s, a delay of 45 ms relative to the start of the impulse response corresponds to a distance of approximately 15.4 m for the echo travels.
[0042] The power spectral density manager 150 is configured to generate power spectral density data 154 based on the impulse response data 144. Power spectral data 154 represents the white noise power spectral density S. analysis (m, s, f), where f is the frequency. In some embodiments, the power spectral density manager 150 includes a convolution manager 151 and a transform manager 152.
[0043] Convolution manager 151 is configured to perform convolutions on time-based functions to produce autocorrelation data 155 representing another time-based function. Specifically, convolution manager 151 is configured to generate autocorrelation functions from the impulse response functions of each microphone and speaker. In some implementations, the autocorrelation function r represented by the autocorrelation data 155 is... analysis (m, s, t) is given by the following formula.
[0044]
[0045] Transform manager 152 is configured to perform a frequency-space transformation on a time-based function. Specifically, the time-based function on which transform manager 152 performs the frequency-space transformation is an autocorrelation function r. analysis (m, s, t) to generate a power spectral density Sanalysis The power spectral density data 154 of (m, s, f). In some embodiments, the transform manager 152 is configured to convert the autocorrelation function r analysis (m, s, t) is multiplied by the window function W(t) represented by window data 156. In some implementations, the window function W(t) has a duration of 1 millisecond, although the duration may be greater than or less than 1 millisecond.
[0046] In some implementations, the transform manager applies the Fourier transform to r analysis The product of (m, s, t) and W(t). In some implementations, because the representation of the time-based function may be discrete in time, a Fast Fourier Transform implementation is used to perform the Fourier Transform. In some implementations, the transform manager uses wavelet transform to generate a transform to the frequency space.
[0047] The calibration filter manager 160 is configured to generate microphone calibration data 162 and / or speaker calibration data 164. Specifically, the calibration filter manager 160(i) performs an averaging operation on the power spectral density data 154 on the speaker to generate the speaker average power spectral density S. microphone (m, f), (ii) Perform an averaging operation on the power spectral density data 154 on the microphone to produce the microphone average power spectral density S. loudspeaker (s,f) and (iii) generate the average power spectral density S of the microphone and speaker. avg (f). Specifically...
[0048]
[0049]
[0050]
[0051] The calibration filter manager 160 then calculates the microphone calibration filter as represented by the microphone calibration data 152 as follows.
[0052]
[0053]
[0054] Using these calibration filters, the microphone is calibrated to record the same spectrum in response to the same input stimulus, and the speaker is calibrated to produce the same spectrum in response to the same input stimulus.
[0055] Figure 1BThis is a diagram illustrating an example configuration 170 of microphone 172 and speaker 174, and a computer 120 capable of performing calibration of microphone 172 and speaker 174. Figure 1B The configuration 170 shown has sixteen microphones and two speakers. An example of a microphone 172 that can be used in configuration 170 includes the Invensense ICS-52000TDM microphone. An example of a speaker 174 that can be used in configuration 170 includes the Tymphany TC5FC07-04. Note that any number of microphones and speakers can be considered.
[0056] Figure 1C This diagram illustrates an example configuration of microphones and speakers within a telepresence system 180. The telepresence system 180 can be used by multiple users for, for example, 3D video conferencing communication (e.g., telepresence sessions). Generally speaking, Figure 1C The system 180 shown can be used to capture a user's video and / or images during a 2D or 3D video conference.
[0057] like Figure 1C As shown, the telepresence system 180 is being used by a first user 182 and a second user 182'. For example, users 182 and 182' are using the telepresence system 180 to participate in a 3D telepresence session. In such an example, the telepresence system 180 allows each of users 182 and 182' to see a highly realistic and visually consistent representation of the other, thereby facilitating interaction between users in a manner similar to their physical presence in each other.
[0058] The telepresence system 180 may include one or more 2D or 3D displays. Here, a 3D display 190 is provided for user 182, and a 3D display 192 is provided for user 182'. The 3D displays 190 and 192 may use any of a variety of 3D display technologies to provide an autostereoscopic view for the respective viewer (here, for example, user 102 or user 104). In some embodiments, the 3D displays 190 and 192 may be stand-alone units (e.g., self-supporting or wall-mounted). In some embodiments, the displays 190 and 192 may be 2D displays.
[0059] Typically, displays such as displays 190, 192 can provide images that approximate the 3D optical properties of physical objects in the real world without the use of head-mounted display (HMD) devices. Typically, the displays described herein include flat panel displays, biconvex lenses (e.g., microlens arrays), and / or parallax barriers to redirect images to multiple different viewing areas associated with the display.
[0060] In some example displays, there may be a single location providing a 3D view of the image content (e.g., a user, an object, etc.) presented by such a display. A user can sit in this single location to experience appropriate parallax, minimal distortion, and realistic 3D images. If the user moves to a different physical location (or changes head position or eye gaze position), the image content (e.g., the user, objects worn by the user, and / or other objects) may begin to appear less realistic, 2D, and / or distorted. The systems and techniques described herein can reconfigure the image content projected from the display to ensure that the user can move around but still experience appropriate parallax, low distortion, and realistic 3D images in real time. Therefore, the systems and techniques described herein offer the advantage of maintaining and providing 3D image content and objects for display to the user, regardless of any user movement that occurs while the user is viewing the 3D display.
[0061] As shown in Figure 1, the telepresence system 180 may include one or more networks. Network 198 may be a publicly available network (e.g., the Internet) or a private network, to name just two examples. Network 198 may be wired, wireless, or a combination of both. Network 198 may include or utilize one or more other devices or systems, including but not limited to one or more servers (not shown).
[0062] The telepresence system 180 also includes a microphone array 172 and a speaker array 174 for user 182, and an analog microphone array 172' and a speaker array 174' for user 182'. These are arranged to provide the most realistic audio experience for users 182 and 182'. Speaker arrays 174 and 174' can provide 3D audio signals locally, and microphone arrays 172 and 172' can be used to detect 3D audio signals from users, which can then be encoded and sent to remote users to present a 3D sound field representing the sound in the telepresence system 180 to the remote users.
[0063] Figure 2 This is a flowchart depicting an example method 200 for calibrating a microphone and speaker. Method 200 can be executed by a software architecture described in conjunction with Figure 1, which resides in the memory 126 of the user equipment computer 120 and is run by a set of processing units 124, or can be executed by a software architecture residing in the memory of a computing device different from (e.g., remote from) the user equipment computer 120.
[0064] At 202, the impulse response manager 140 receives a reverberant sound field (e.g., reverberant sound field data 132) via each microphone in the microphone array, which is based on a signal (e.g., swept-frequency sinusoidal chirp signal data 134) generated by the reverberation field manager 130 at each speaker in the speaker array. For example, the reverberation field manager 130 causes a swept-frequency sinusoidal chirp signal to be emitted from the speaker array. In some embodiments, the swept-frequency sinusoidal chirp signal is emitted one after another from each speaker in the speaker array in a time-separated manner. In some embodiments, the swept-frequency sinusoidal chirp signal is emitted simultaneously from multiple speakers in the speaker array. After the swept-frequency sinusoidal chirp signal is emitted from one or more speakers, each microphone in the microphone array records the subsequent reverberant sound field for a duration at least as long as the duration T of the swept-frequency sinusoidal chirp signal according to equation (1). In some embodiments, the microphone records the subsequent reverberant sound field between the start time and the end time. In some embodiments, the difference between the start time and the end time is different from T. In some implementations, the start and / or end times are based on measured characteristics of the reverberant sound field. For example, the start time may be based on the time when the direct sound to the microphone—i.e., sound that has not been reflected and travels directly along the path to the microphone—can be ignored. In such an example, the reverberant sound field manager 130 may increase the amplitude of the reverberant sound field because audio signals reflected from boundaries and / or obstacles may suffer some degradation. Note that by using reverberation instead of a direct sound field, a calibrated microphone may generate high-quality direction-sensitive audio signals.
[0065] At 204, the power spectral density manager 150 generates a corresponding power spectral density (e.g., power spectral density data 154) for each microphone in the microphone array and for each speaker in the speaker array, based on the corresponding reverberant sound field generated by the speaker and received at the microphone. (Reference) Figure 3 Describe in detail the generation of the power spectral density.
[0066] At 206, the calibration filter manager 160 generates a corresponding calibration filter (e.g., microphone calibration data 162) for each microphone in the microphone array based on the ratio of the average power spectral density on the speaker array and the microphone array (Equation (7)) to the average power spectral density on the speaker array (Equation (5)).
[0067] Using these calibration filters, microphones are calibrated to record the same spectrum in response to the same input stimulus, and loudspeakers are calibrated to produce the same spectrum in response to the same input stimulus. Furthermore, because the calibration filters are based on the reverberant sound field rather than the direct sound field, the calibration factors are largely insensitive to the geometry of the environment in which the loudspeakers and microphones are located, as well as any nodes in the direct sound signal, and the calibration filters enable the microphones to produce high-quality, direction-sensitive audio signals.
[0068] At 208, computer 120 uses a corresponding calibration filter generated for each microphone in the microphone array to record acoustic signals from the user through the microphone array, wherein each microphone in the microphone array records substantially the same spectrum of acoustic signals. The signals recorded from the calibration microphones can then be processed to generate spatial audio signals representing sound in the environment of the microphone array (e.g., speech produced by one or more speakers using a telepresence system including the microphone array), and the generated spatial audio signals can be transmitted to a sound presentation system (e.g., a remote telepresence system) for presentation.
[0069] Figure 3 This is a flowchart illustrating an example process 300 for calibrating a microphone array. Process 300 can be executed by a software structure described in conjunction with Figure 1, which resides in the memory 126 of the user equipment computer 120 and is run by a set of processing units 124, or it can be executed by a software structure residing in the memory of a computing device different from (e.g., remote from) the user equipment computer 120.
[0070] At 301, the impulse response manager 140 measures the reverberant (raw) impulse response from each channel, i.e., each microphone / speaker pair. As described above, the impulse response can be derived from a swept-frequency sinusoidal chirp signal generated by the reverberation field manager 130, received at the microphone, or reflected from walls, ceilings, floors, or objects. The actual recording of the signal occurs at a start time that occurs after a sufficiently long period of time has elapsed since the microphone received the direct audio signal. Therefore, the reverberant field measured at the microphone only includes signals reflected from boundaries and obstacles.
[0071] Figure 4A The diagram illustrates example raw impulse response functions for two speakers and four microphones, resulting in eight raw impulse response functions. In some implementations, each raw impulse response function is measured based on the reverberant sound field from a single speaker. In some implementations, each speaker generates a swept-frequency sinusoidal chirp signal at different times. In some implementations, the raw impulse response functions are all measured simultaneously at the microphone array. In some implementations, the raw impulse response functions are measured one at a time at the microphone array.
[0072] At 302, the attenuation normalization manager 141 estimates the impulse response energy as a function of time for each channel. Figure 4B The diagram shows the connection with the source. Figure 4A Example time-dependent energy metrics associated with the original impulse response function; note Figure 4B The vertical axis of the curve represents the square root of the energy.
[0073] At 303, the attenuation normalization manager 141 averages the impulse response energy across the microphone and speaker to produce an average impulse response energy according to equation (3). Figure 4C The diagram illustrates the average output from all speakers and microphones. Figure 4B Example time-related energy measure.
[0074] At 304, the attenuation normalization manager 141 normalizes the raw impulse response of each channel using the average impulse response energy to produce an attenuation normalized impulse response function. Figure 4D The illustration shows an example of a decayed normalized impulse response function corresponding to the original impulse response function for four microphones and two speakers.
[0075] At 305, the segment manager 142 extracts segments of the decayed normalized impulse response function for the time interval (i.e., from the first time to the second time) to generate time-based segments. Figure 4E The diagram shows Figure 4D The example segment of the decayed normalized impulse response function is shown below.
[0076] At 306, the convolution manager 151 generates white noise autocorrelation functions from each segment. Figure 4F The diagram shows from Figure 4E The example multichannel white noise autocorrelation function is derived from a sub-segment of the decayed normalized impulse response function shown.
[0077] At 307, the transform manager 152 performs a Fourier transform on the white noise autocorrelation function over a short time window to generate a power spectral density for each channel according to equation (4). Figure 4G The diagram corresponds to Figure 4F The example power spectral density of the multichannel white noise autocorrelation function is shown.
[0078] At 308, the calibration filter manager 160 generates an average of the power spectral density on the speaker to produce the speaker average power spectral density. Figure 4H The diagram illustrates the average signal from the speaker. Figure 4G Example power spectral density.
[0079] At 309, the calibration filter manager 160 generates a microphone calibration filter as the ratio of the average power spectral density of the microphone and speaker to the average power spectral density of the speaker. Figure 4I The diagram shows from Figure 4H The example microphone calibration filter is derived from the average power spectral density of the loudspeaker.
[0080] Note that 308 and 309 can also be used to generate speaker calibration filters. Figure 4J The diagram illustrates the average signal from the microphone. Figure 4G Example power spectral density. Figure 4K The diagram shows from Figure 4J An example speaker calibration filter derived from the microphone average power spectral density.
[0081] Figure 5 Examples of a general-purpose computer device 500 and a general-purpose mobile computer device 550 are illustrated, which can be used with the technologies described herein.
[0082] like Figure 5 As shown, computing device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 550 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the inventions described and / or claimed in this document.
[0083] Computing device 500 includes a processor 502, a memory 504, a storage device 506, a high-speed interface 508 connected to the memory 504 and a high-speed expansion port 510, and a low-speed interface 412 connected to a low-speed bus 514 and the storage device 506. Each of components 502, 504, 506, 508, 510, and 512 is interconnected using various buses and can be mounted on a common motherboard or otherwise suitably mounted. Processor 502 can process instructions for execution within computing device 500, including instructions stored in memory 504 or on storage device 506 for displaying graphical information for a GUI on an external input / output device (e.g., a display 516 coupled to high-speed interface 508). In other embodiments, multiple processors and / or multiple buses, as well as multiple memories and various memory types, can be suitably used. Furthermore, multiple computing devices 500 can be connected, each providing a portion of the necessary operation (e.g., as a server group, a set of blade servers, or a multiprocessor system).
[0084] Memory 504 stores information within computing device 500. In one embodiment, memory 504 is one or more volatile memory cells. In another embodiment, memory 504 is one or more non-volatile memory cells. Memory 504 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0085] Storage device 506 provides large-capacity storage for computing device 500. In one implementation, storage device 506 may be or contain computer-readable media, such as floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, flash memory, or other similar solid-state storage devices or device arrays, including devices in storage area networks or other configurations. A computer program product may be tangibly embodied in an information carrier. The computer program product may also contain instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 504, storage device 506, or memory on processor 502.
[0086] High-speed controller 508 manages bandwidth-intensive operations of computing device 500, while low-speed controller 512 manages lower bandwidth-intensive operations. This functional allocation is merely exemplary. In one implementation, high-speed controller 508 is coupled to memory 504, display 516 (e.g., via a graphics processor or accelerator), and high-speed expansion port 510, which can accept various expansion cards (not shown). In this embodiment, low-speed controller 512 is coupled to storage device 506 and low-speed expansion port 514. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), may be coupled to one or more input / output devices, such as keyboards, pointing devices, scanners, or network devices (e.g., switches or routers), for example, via a network adapter.
[0087] Computing device 500 can be implemented in a variety of different forms, as shown in the figure. For example, it can be implemented as a standard server 520, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 524. Furthermore, it can be implemented in a personal computer such as a laptop computer 522. Alternatively, components from computing device 500 can be combined with other components (not shown) in a mobile device such as device 550. Each of such devices can contain one or more of computing devices 500, 550, and the entire system can consist of multiple computing devices 500, 550 communicating with each other.
[0088] Computing device 550 includes processor 552, memory 564, input / output devices such as display 554, communication interface 566 and transceiver 568, and other components. Device 550 may also be equipped with storage devices, such as microdrives or other devices, to provide additional storage. Each of components 550, 552, 564, 554, 566, and 568 is interconnected using various buses, and several of the components may be mounted on a common motherboard or otherwise suitably mounted.
[0089] Processor 552 can execute instructions within computing device 450, including instructions stored in memory 564. The processor can be implemented as a chipset including individual and multiple analog and digital processors. The processor can provide, for example, coordination of other components of device 550, such as control of the user interface, applications running on device 550, and wireless communication via device 550.
[0090] Processor 552 can communicate with the user via control interface 558 and display interface 556 coupled to display 554. Display 554 can be, for example, a TFT LCD (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display or other suitable display technology. Display interface 556 may include appropriate circuitry for driving display 554 to present graphics and other information to the user. Control interface 558 can receive commands from the user and translate them for submission to processor 552. Additionally, an external interface 562 can be provided to communicate with processor 552 to enable near-field communication between device 550 and other devices. External interface 562 may provide wired communication in some embodiments, wireless communication in others, and multiple interfaces may be used.
[0091] Memory 564 stores information within computing device 550. Memory 564 can be implemented as one or more computer-readable media, one or more volatile memory cells, or one or more non-volatile memory cells. Extended memory 574 may also be provided and connected to device 550 via an extended interface 572, which may include, for example, a SIMM (Single In-line Memory Module) card interface. Such extended memory 574 can provide additional storage space for device 550, or it may store applications or other information for device 550. Specifically, extended memory 574 may include instructions for performing or supplementing the above processes, and may also include security information. Therefore, for example, extended memory 574 may be provided as a security module of device 550 and can be programmed with instructions that allow secure use of device 550. Furthermore, secure applications and additional information, such as identification information placed on the SIMM card in an unhackable manner, can be provided via a SIMM card.
[0092] The memory may include, for example, flash memory and / or NVRAM memory, as discussed below. In one embodiment, the computer program product is tangibly contained in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 564, extended memory 574, or memory on processor 552, which may be received, for example, via transceiver 568 or external interface 562.
[0093] Device 550 can communicate wirelessly via communication interface 566, which may include digital signal processing circuitry if necessary. Communication interface 566 can provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messages, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS. Such communication can occur, for example, via radio frequency transceiver 568. Furthermore, short-range communication can occur, for example, using Bluetooth, WiFi, or other such transceivers (not shown). Additionally, GPS (Global Positioning System) receiver module 570 can provide device 550 with additional navigation and location-related wireless data, which can be used as appropriate by applications running on device 550.
[0094] Device 550 can also use audio codec 560 for audio communication, which can receive spoken information from a user and convert it into usable digital information. Audio codec 560 can similarly generate audible sounds for the user, for example, through a speaker in the mobile phone of device 550. Such sounds can include sounds from voice phone calls, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications running on device 550.
[0095] The computing device 550 can be implemented in many different forms, as shown in the figure. For example, it can be implemented as a cellular phone 580. It can also be implemented as part of a smartphone 582, a personal digital assistant, or other similar mobile device.
[0096] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system, which includes at least one programmable processor, which may be specialized or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device.
[0097] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" mean any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, which includes a machine-readable medium that receives machine instructions as machine-readable signals. The term "machine-readable signal" means any signal used to provide machine instructions and / or data to a programmable processor.
[0098] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0099] The systems and technologies described herein can be implemented in computing systems that include back-end components (such as data servers), middleware components (such as application servers), front-end components (such as client computers having a graphical user interface or web browser through which users can interact with implementations of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. Components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), and the Internet.
[0100] A computing system can include clients and servers. Clients and servers are typically geographically separated and usually interact through communication networks. The relationship between clients and servers is generated by computer programs running on various computers that have a client-server relationship with each other.
[0101] Referring back to Figure 1, in some embodiments, memory 126 can be any type of memory, such as random access memory, disk drive memory, flash memory, etc. In some embodiments, memory 126 can be implemented as multiple memory components (e.g., multiple RAM components or disk drive memory) associated with components of compression computer 120. In some embodiments, memory 126 can be database memory. In some embodiments, memory 126 can be or may include non-local memory. For example, memory 126 can be or may include memory shared by multiple devices (not shown). In some embodiments, memory 126 can be associated with server devices (not shown) within a network and configured to serve components of compression computer 120.
[0102] Components of the compression computer 120 (e.g., modules, processing unit 124) can be configured to operate on one or more platforms (e.g., one or more similar or different platforms), which may include one or more types of hardware, software, firmware, operating systems, runtime libraries, etc. In some embodiments, components of the compression computer 120 can be configured to operate within a device cluster (e.g., a server farm). In such embodiments, the functionality and processing of the components of the compression computer 120 can be distributed across several devices within the device cluster.
[0103] The components of computer 120 can be or may include any type of hardware and / or software configured to process attributes. In some embodiments, one or more portions of the components shown in the components of computer 120 of FIG1 can be or may include hardware-based modules (e.g., digital signal processors (DSPs), field-programmable gate arrays (FPGAs), memory), firmware modules, and / or software-based modules (e.g., computer code modules, a set of computer-readable instructions executable on a computer). For example, in some embodiments, one or more portions of the components of computer 120 can be or may include software modules configured to be executed by at least one processor (not shown). In some embodiments, the functionality of the components can be included in different modules and / or different components than those shown in FIG1.
[0104] Although not shown, in some embodiments, components (or portions thereof) of computer 120 may be configured to operate in, for example, a data center (e.g., a cloud computing environment), a computer system, one or more server / host devices, etc. In some embodiments, components of computer 120 (or portions thereof) may be configured to operate within a network. Therefore, components of computer 120 (or portions thereof) may be configured to function in various types of network environments that may include one or more devices and / or one or more server devices. For example, the network may be or may include a local area network (LAN), a wide area network (WAN), etc. The network may be or may include a wireless network and / or a wireless network implemented using, for example, gateway devices, bridges, switches, etc. The network may include one or more segments and / or may have portions based on various protocols (e.g., Internet Protocol (IP) and / or proprietary protocols). The network may include at least a portion of the Internet.
[0105] In some embodiments, one or more components of computer 120 may be or may include a processor configured to process instructions stored in memory. For example, depth image manager 130 (and / or a portion thereof), viewpoint manager 140 (and / or a portion thereof), ray casting manager 150 (and / or a portion thereof), SDV manager 160 (and / or a portion thereof), aggregation manager 170 (and / or a portion thereof), rooting manager 180 (and / or a portion thereof), and depth image generation manager 190 (and / or a portion thereof) may be a combination of processor and memory configured to execute process-related instructions to implement one or more functions.
[0106] Several embodiments have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the specification.
[0107] It should also be understood that when an element is referred to as being on, connected to, electrically connected to, coupled to, or electrically coupled to another element, it may be directly on, connected to, or coupled to the other element, or one or more intermediate elements may be present. Conversely, when an element is referred to as being directly on, directly connected to, or directly coupled to another element, no intermediate elements are present. Although the terms "directly on," "directly connected to," or "directly coupled to" may be omitted throughout the detailed description, elements shown as being directly on, directly connected to, or directly coupled to may be referred to in this way. The claims of this application may be modified to describe the exemplary relationships described in the specification or shown in the drawings.
[0108] While certain features of the described embodiments have been shown as described herein, many modifications, substitutions, alterations, and equivalents will now occur to those skilled in the art. Therefore, it should be understood that the appended claims are intended to cover all such modifications and variations falling within the scope of this invention. It should be understood that they are presented by way of example only and not limitation, and various changes in form and detail may be made. Any part of the apparatus and / or methods described herein can be combined in any combination, except for mutually exclusive combinations. The embodiments described herein may include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described.
[0109] Furthermore, the logical flow depicted in the figures does not require the specific or sequential order shown to achieve the desired result. Additionally, other steps may be provided relative to the described flow, or steps may be removed relative to the described flow, and other components may be added to or removed from the described system. Therefore, other embodiments are within the scope of the appended claims.
Claims
1. A method for calibrating a microphone, comprising: The reverberant sound field is received via each microphone in the microphone array based on the audio signal generated by each speaker in the speaker array; For each microphone in the microphone array and each speaker in the speaker array, a corresponding power spectral density is generated for the microphone and the speaker based on the corresponding reverberant sound field generated by the speaker and received at the microphone; Based on the ratio of the average power spectral density on the speaker array and the microphone array to the average power spectral density on the speaker array, a corresponding calibration filter is generated for each microphone in the microphone array. as well as The microphone array records acoustic signals from the user using the corresponding calibration filter generated for each microphone in the microphone array, with each microphone in the microphone array recording the same spectrum of the acoustic signals using the corresponding calibration filter.
2. The method according to claim 1, wherein, Generating the corresponding power spectral density for each microphone in the microphone array and each speaker in the speaker array includes: Based on the corresponding reverberant sound field generated by the loudspeaker and received at the microphone, a corresponding impulse response function is generated for the microphone and the loudspeaker.
3. The method according to claim 2, wherein, Generating the corresponding power spectral density for each microphone in the microphone array and each speaker in the speaker array further includes: Perform autocorrelation on the corresponding impulse response functions of the microphone and speaker to generate autocorrelation impulse response functions; and A frequency space transformation is performed on the autocorrelation impulse response function to generate the power spectral density of the microphone and speaker.
4. The method according to claim 3, wherein, Performing the transformation of the autocorrelation impulse response function to the frequency space includes: Generate a window function that is constant during a specified time interval and equal to zero outside the specified time interval; and Perform a Fourier transform operation on the product of the window function and the autocorrelation impulse response function.
5. The method of claim 2, further comprising, before generating the corresponding impulse response function: A swept sine chirp signal having a frequency between a first frequency and a second frequency is generated at the speaker as the audio signal, and the swept sine chirp signal is received at the microphone.
6. The method according to claim 2, wherein, Generating the corresponding impulse response function includes: For each microphone in the microphone array and each speaker in the speaker array: Measure the raw impulse response function corresponding to the microphone and the speaker; and Generate a time-dependent energy metric associated with the original impulse response function; A normalization factor is generated based on the average of the corresponding time-related energy metrics associated with each microphone in the microphone array and each speaker in the speaker array across the microphone array and the speaker array; and The original impulse response function corresponding to each microphone in the microphone array and each speaker in the speaker array is divided by the normalization factor to produce the attenuated normalized impulse response function corresponding to that microphone and that speaker.
7. The method according to claim 6, wherein, Generating the time-dependent energy metric associated with the corresponding raw impulse response function for each microphone in the microphone array and each speaker in the speaker array includes: Generate a first power corresponding to the absolute value of the respective original impulse response function of the microphone and the speaker; and A smoothing operation is performed on the power of the absolute value of the corresponding raw impulse response function to produce the time-dependent energy metric associated with the corresponding raw impulse response functions of the microphone and the speaker.
8. The method according to claim 7, wherein, Performing the smoothing operation includes: Within a specified time range, a moving average of the power of the absolute value of the corresponding original impulse response function is generated.
9. The method according to claim 7, wherein, Generating the normalization factor includes: A second power of the time-dependent energy metric is generated, which is the reciprocal of the first power, and is associated with the corresponding raw impulse response function of each microphone in the microphone array and each speaker in the speaker array.
10. The method according to claim 6, wherein, Generating the corresponding impulse response function includes: A segment corresponding to the attenuation normalized impulse response function of the microphone and the speaker is obtained as the corresponding impulse response function of each microphone in the microphone array and each speaker in the speaker array, the segment starting at a first time and ending at a second time.
11. The method according to claim 10, wherein, The first time is the minimum distance of the echo travel based on the reverberant sound field.
12. The method according to claim 10, wherein, The second time is an estimate based on the length of time it takes for the corresponding original impulse response function to decay to the background noise associated with the measurement of the original impulse response function.
13. A method for calibrating a loudspeaker, comprising: The reverberant sound field is received via each microphone in the microphone array based on the audio signal generated by each speaker in the speaker array; For each microphone in the microphone array and each speaker in the speaker array, a corresponding power spectral density is generated for the microphone and the speaker based on the corresponding reverberant sound field generated by the speaker and received at the microphone; Based on the ratio of the average power spectral density on the speaker array and the microphone array to the average power spectral density on the microphone array, a corresponding calibration filter is generated for each speaker in the speaker array; as well as The acoustic signal is generated by the loudspeaker array using the corresponding calibration filter generated for each loudspeaker in the loudspeaker array, and each loudspeaker in the loudspeaker array generates the same spectrum of the acoustic signal in response to the same output stimulus using the corresponding calibration filter.
14. A computer program product including a non-transitory storage medium, the computer program product including code that, when executed by processing circuitry of a computing device, causes the processing circuitry to perform a method, the method comprising: The reverberant sound field is received via each microphone in the microphone array based on the audio signal generated by each speaker in the speaker array; For each microphone in the microphone array and each speaker in the speaker array, a corresponding power spectral density is generated for the microphone and the speaker based on the corresponding reverberant sound field generated by the speaker and received at the microphone; Based on the ratio of the average power spectral density on the speaker array and the microphone array to the average power spectral density on the speaker array, a corresponding calibration filter is generated for each microphone in the microphone array. as well as The microphone array records acoustic signals from the user using the corresponding calibration filter generated for each microphone in the microphone array, with each microphone in the microphone array recording the same spectrum of the acoustic signals using the corresponding calibration filter.
15. The computer program product according to claim 14, wherein, Generating the corresponding power spectral density for each microphone in the microphone array and each speaker in the speaker array includes: Based on the corresponding reverberant sound field generated by the loudspeaker and received at the microphone, a corresponding impulse response function is generated for the microphone and the loudspeaker.
16. The computer program product according to claim 15, wherein, Generating the corresponding power spectral density for each microphone in the microphone array and each speaker in the speaker array further includes: Perform autocorrelation on the corresponding impulse response functions of the microphone and speaker to generate autocorrelation impulse response functions; and A frequency space transformation is performed on the autocorrelation impulse response function to generate the power spectral density of the microphone and speaker.
17. The computer program product according to claim 16, wherein, Performing the transformation of the autocorrelation impulse response function to the frequency space includes: Generate a window function that is constant during a specified time interval and equal to zero outside the specified time interval; and Perform a Fourier transform operation on the product of the window function and the autocorrelation impulse response function.
18. The computer program product according to claim 15, wherein, The method further includes, before generating the corresponding impulse response function: A swept sine chirp signal having a frequency between a first frequency and a second frequency is generated at the speaker as the audio signal, and the swept sine chirp signal is received at the microphone.
19. The computer program product according to claim 15, wherein, Generating the corresponding impulse response function includes: For each microphone in the microphone array and each speaker in the speaker array: Measure the raw impulse response function corresponding to the microphone and the speaker; and Generate a time-dependent energy metric associated with the original impulse response function; A normalization factor is generated based on the average of the corresponding time-related energy metrics associated with each microphone in the microphone array and each speaker in the speaker array across the microphone array and the speaker array; and The original impulse response function corresponding to each microphone in the microphone array and each speaker in the speaker array is divided by the normalization factor to produce the attenuated normalized impulse response function corresponding to that microphone and that speaker.
20. An electronic device, the electronic device comprising: Memory; as well as A control circuit coupled to the memory, the control circuit being configured to: The reverberant sound field is received via each microphone in the microphone array based on the audio signal generated by each speaker in the speaker array; For each microphone in the microphone array and each speaker in the speaker array, a corresponding power spectral density is generated for the microphone and the speaker based on the corresponding reverberant sound field generated by the speaker and received at the microphone; Based on the ratio of the average power spectral density on the speaker array and the microphone array to the average power spectral density on the speaker array, a corresponding calibration filter is generated for each microphone in the microphone array. as well as The microphone array records acoustic signals from the user using the corresponding calibration filter generated for each microphone in the microphone array, with each microphone in the microphone array recording the same spectrum of the acoustic signals using the corresponding calibration filter.
Citation Information
Patent Citations
Adaptive self-calibration of small microphone array by soundfield approximation and frequency domain magnitude equalization
US20130170666A1
Sound level estimation
US20170127206A1