Automatic calibration of microphone arrays for telepresence conferencing

Automatic calibration of microphones and speakers in telepresence systems using power spectral densities from reverberant sound fields addresses the inefficiencies of conventional methods, achieving precise and consistent audio performance.

JP7724317B2Active Publication Date: 2025-08-15GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024017209
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-08-15
Estimated Expiration
2040-10-30

AI Technical Summary

Technical Problem

Conventional methods for calibrating microphones and speakers in telepresence conferencing systems are cumbersome, prone to errors, and require manual setup, making them inefficient and inaccurate.

Method used

Generate calibration filters for microphones and speakers by deriving power spectral densities from reverberant sound fields, using a computer to measure impulse response functions and normalize them, allowing automatic calibration without external hardware.

Benefits of technology

The solution provides accurate and consistent audio calibration insensitive to room configuration and hardware positioning, ensuring high-quality directional audio signals and realistic spatialized output without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007724317000011
    Figure 0007724317000011
  • Figure 0007724317000012
    Figure 0007724317000012
  • Figure 0007724317000013
    Figure 0007724317000013
Patent Text Reader

Abstract

To provide a method that calibrates a microphone and a speaker used in applications such as representation conferencing.SOLUTION: A method includes receiving, via each microphone of a microphone array, a reverberant sound field on the basis of an audio signal generated by each speaker of a speaker array, generating respective power spectral densities for the microphone and the speakers on the basis of each reverberant sound field received by the microphone, generating each calibration filter for each microphone of the microphone array as a ratio of the power spectral density averaged for the speaker array and the microphone array to the power spectral density averaged for the speaker array, and recording an acoustic signal from a user by the microphone array using the respective calibration filters generated for the respective microphone arrays.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This specification relates to the calibration of microphones and speakers used in applications such as telepresence conferencing. [Background technology]

[0002] background A telepresence conferencing system may include a large number of microphones for detecting directional audio signals from a user and multiple speakers for providing directional audio signals to the user. Summary of the Invention [Means for solving the problem]

[0003] overview In one general aspect, the method may include receiving, via each microphone of the microphone array, a reverberant sound field based on an audio signal generated by each speaker of the speaker array. The method may also include generating, for each microphone of the microphone array and for each speaker of the speaker array, a power spectral density for that microphone and for that speaker based on each reverberant sound field generated by that speaker and received by that microphone. The method may further include generating, for each microphone of the microphone array, a respective calibration filter as a ratio of the power spectral density averaged over the speaker array and the microphone array to the power spectral density averaged over the speaker array. The method may further include recording an acoustic signal from a user with the microphone array using the respective calibration filter generated for each microphone array, wherein each of the microphone arrays using the respective calibration filter records essentially the same spectrum of the acoustic signal.

[0004] In another general aspect, a computer program product includes a non-transitory storage medium, the computer program product including code that, when executed by a processing circuit of a computing device, causes the processing circuit to perform a method. The method may include receiving, via each microphone of a microphone array, a reverberant sound field based on an audio signal generated by each speaker of a speaker array. The method may also include generating, for each microphone of the microphone array and for each speaker of the speaker array, a power spectral density for that microphone and for that speaker based on the respective reverberant sound fields generated by that speaker and received by that microphone. The method may further include generating, for each microphone of the microphone array, a respective calibration filter as a ratio of a power spectral density averaged over the speaker array and the microphone array to a power spectral density averaged over the speaker array. The method may further include recording an acoustic signal from a user with the microphone array using the respective calibration filter generated for each microphone array, wherein each of the microphone arrays using the respective calibration filter records essentially the same spectrum of the acoustic signal.

[0005] In another general aspect, an electronic device includes a memory and a control circuit coupled to the memory. The control circuit can be configured to receive a reverberant sound field based on an audio signal generated by each speaker of the speaker array via each microphone of the microphone array. The control circuit also controls, for each microphone of the microphone array and for each speaker of the speaker array, each reverberant sound field generated by the speaker and received by the microphone. The control circuitry may be configured to generate a power spectral density for each of the microphones and the speaker based on the acoustic sound field. The control circuitry may be further configured to generate a respective calibration filter for each microphone of the microphone array as a ratio of the power spectral density averaged over the speaker array and the microphone array to the power spectral density averaged over the speaker array. The control circuitry may be further configured to record an acoustic signal from the user with the microphone array using the respective calibration filter generated for each microphone array, wherein each of the microphone arrays using the respective calibration filter records essentially the same spectrum of the acoustic signal.

[0006] In another general aspect, a method may include receiving, via each microphone of a microphone array, a reverberant sound field based on an audio signal generated by each speaker of a speaker array. The method may also include generating, for each microphone of the microphone array and for each speaker of the speaker array, a power spectral density for that microphone and for that speaker based on the reverberant sound field generated by that speaker and received by that microphone. The method may further include generating, for each microphone of the microphone array, a respective calibration filter as a ratio of the power spectral density averaged over the microphone array and the speaker array to the power spectral density averaged over the speaker array. The method may further include generating an acoustic signal by the speaker array using each calibration filter generated for each speaker of the speaker array, wherein each of the speaker arrays using the respective calibration filters generates essentially the same spectrum of the acoustic signal in response to the same output stimulus.

[0007] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0008] [Figure 1A] FIG. 1 illustrates an exemplary electronic environment for implementing the technical solutions described herein. [Figure 1B] 1B illustrates an exemplary configuration of microphones and speakers within the electronic environment shown in FIG. 1A. [Figure 1C] FIG. 1 illustrates an exemplary configuration of microphones and speakers in a telepresence system. [Figure 2] 1B is a flowchart illustrating an exemplary method of implementing a technical solution within the electronic environment shown in FIG. 1A. [Figure 3] 1 is a flowchart illustrating an exemplary process for calibrating microphones of a microphone array according to the technical solution. [Figure 4A] 1B is a plot illustrating an exemplary raw impulse response function from two speakers to four microphones in the electronic environment shown in FIG. 1A. [Figure 4B] 4B is a plot illustrating an exemplary time-dependent energy metric associated with the raw impulse response function of FIG. 4A. [Figure 4C] 4C is a plot illustrating the example time-dependent energy metric of FIG. 4B averaged over all speakers and microphones. [Figure 4D] 1 is a plot showing example decay normalized impulse response functions corresponding to four microphones and raw impulse response functions of two speakers. [Figure 4E] 4E is a plot illustrating an exemplary sub-segment of the decaying normalized impulse response function shown in FIG. 4D. [Figure 4F] 4F is a plot illustrating an autocorrelation function of an exemplary multi-channel white noise derived from a sub-segment of the decaying normalized impulse response function shown in FIG. 4E. [Figure 4G] 4F is a plot illustrating an exemplary power spectral density corresponding to the autocorrelation function of the multi-channel white noise shown in FIG. 4F. [Figure 4H]4H is a plot illustrating the example power spectral density of FIG. 4G averaged over the speaker. [Figure 4I] 4H is a plot illustrating an example microphone calibration filter derived from the power spectral density averaged for the speaker of FIG. 4H. [Figure 4J] 4H is a plot illustrating the example power spectral density of FIG. 4G averaged over the microphone. [Figure 4K] 4J is a plot illustrating an example speaker calibration filter derived from the power spectral densities averaged for the microphone of FIG. 4J. [Figure 5] 10A-10C illustrate examples of computing devices and mobile computing devices that can be used with the circuits described herein. DETAILED DESCRIPTION OF THE INVENTION

[0009] Detailed Description To accurately capture signals from the microphones that can be used to generate high-quality, directional audio signals, each microphone (e.g., microphone gain) in the array may be calibrated relative to the other microphones. Additionally, to accurately render realistic spatialized output in a telepresence system, each speaker (e.g., speaker gain) must also be calibrated relative to the other speakers. Conventional approaches to performing such calibration require the use of external hardware, e.g., sound sources and microphones placed at the expected locations of users / speakers in the telepresence conferencing system.

[0010] However, for such telepresence systems, a technical challenge with the above-described conventional approaches to calibrating microphones and speakers is that the equipment is cumbersome to use and store, requires manual setup and teardown, and is prone to error if the hardware is not precisely positioned relative to the location of the actual user of the system. The equipment may also be prone to error if hardware such as volume knobs or equalizer controls are not precisely configured.

[0011] In contrast to conventional approaches to solving the above-mentioned technical problem, one technical solution to the above-mentioned technical problem includes generating calibration filters for microphones and / or speakers by deriving a power spectral density at each microphone in response to a signal generated by each speaker. For example, a computer in an improved telepresence system can measure a raw impulse response function corresponding to each channel, i.e., each speaker / microphone pair. In some embodiments, the computer normalizes the raw impulse response function based on the contribution of various reverberant reflections to the reverberant sound field energy. The computer then extracts a subsegment of each impulse response function between a start time and an end time that is later than the time the signal was generated by the speaker. The computer then generates a white noise power spectral density for each channel based on the subsegment. In this case, the microphone calibration function is based on the inverse of the power spectral density averaged over the speakers. In this case, the speaker calibration function is based on the inverse of the power spectral density averaged over the microphones.

[0012] The technical advantages of the above technical solution are that it is insensitive to the room configuration and can be performed automatically without human intervention. Also, it is insensitive to the hardware configuration, e.g., the position of the microphone and speaker relative to each other. Furthermore, the technical solution does not require any external hardware beyond that already present in the telepresence system. Essentially, the user can generate the calibration filter by simply flipping a switch.

[0013] In some embodiments, the computer normalizes the raw impulse response functions of all channels using the impulse response energy averaged over the speakers and microphones. In some embodiments, the start time is based on the distance traveled by the reverberant wave in the reverberant sound field. In some embodiments, the end time is based on the noise floor associated with the measurement process. In some embodiments, the white noise power spectral density of a channel is based on the Fourier transform of the white noise autocorrelation of a subsegment of that channel. In some embodiments, the Fourier transform replaces a windowed version of the white noise autocorrelation function.

[0014] 1A is a diagram illustrating an exemplary electronic environment 100 in which the above-described improvement techniques may be implemented. As shown in FIG. 1A, the exemplary electronic environment 100 includes a computer 120.

[0015] The computer 120 includes a network interface 122, one or more processing units 124, and a memory 126. For example, the network interface 122 may include an Ethernet adapter or the like for converting electrical and / or optical signals received from a network into an electronic format for use by the computer 120. The processing unit 124 includes one or more processing chips and / or assemblies. The memory 126 includes both volatile memory (e.g., RAM) and non-volatile memory, such as one or more ROMs, disk drives, solid-state drives, etc. The processing unit 124 and the memory 126 together form a control circuit. The control circuit is configured and arranged to perform the various methods and functions described herein.

[0016] In some embodiments, one or more of the components of computer 120 may comprise a processor (e.g., processing unit 124) configured to process instructions stored in memory 126. Such instructions, as illustrated in Figure 1, include a reverberant sound field manager 130, an impulse response manager 140, a power spectral density manager 150, and a calibration filter manager 160. Additionally, as shown in Figure 1, memory 126 is configured to store various data, as described for each manager that uses such data.

[0017] The reverberant sound field manager 130 is configured to generate reverberant sound field data 132. The reverberant sound field data 132 represents the reverberant sound field produced by the speakers and used to measure the impulse response at the microphones. The sound field is reverberant because, when converted to an audio signal at the speakers, the audio signal may be reflected by nearby walls, ceiling, floor, and objects in the room containing the computer 120, speakers, and microphones.

[0018]

number

[0019]

number

[0020]

number

[0021]

number

[0022]

number

[0023]

number

[0024]

number

[0025]

number

[0026]

number

[0027] These calibration filters are used to calibrate microphones to record the same spectrum in response to the same input stimulus and to calibrate speakers to produce the same spectrum in response to the same input stimulus.

[0028] 1B illustrates an exemplary configuration 170 of microphones 172 and speakers 174 and a computer 120 capable of performing calibration of the microphones 172 and speakers 174. In the configuration 170 shown in FIG. 1B, there are 16 microphones and two speakers. Microphones 172 that may be used in configuration 170 include the Invensense ICS-52000 TDM microphone, etc. Speakers 174 that may be used in configuration 170 include the Tymphany TC5FC07-04, etc. However, any number of microphones and speakers may be contemplated.

[0029] 1C illustrates an exemplary configuration 180 of microphones and speakers in a telepresence system. The telepresence system 180 can be utilized by multiple users to conduct video conference communications (e.g., telepresence sessions) in 3D, for example. Typically, the system 180 illustrated in FIG. 1C would be used to capture video and / or images of users during a 2D or 3D video conference.

[0030] As shown in Figure 1C, a telepresence system 180 is in use by a first user 182 and a second user 182'. 1 and 2. Users 182 and 182′ are participating in a 3D telepresence session using presence system 180. In such an example, telepresence system 180 enables each of users 182 and 182′ to see a highly realistic, visually consistent image of the other, facilitating users to interact in a manner similar to if they were physically present.

[0031] Telepresence system 180 can include one or more 2D or 3D displays. Here, user 182 is provided with 3D display 190, and user 182′ is provided with 3D display 192. 3D displays 190, 192 can utilize any of a number of types of 3D display technologies to provide an autostereoscopic view for each viewer (here, user 102 or user 104, etc.). In some implementations, 3D displays 190, 192 can be standalone units (e.g., freestanding or wall-mounted). In some implementations, displays 190, 192 can be 2D displays.

[0032] In general, displays such as displays 190, 192 can provide images that approximate the 3D optical properties of actual objects in the real world without the use of a head-mounted display (HMD) device. Generally, the displays described herein include flat panel displays, lenticular lenses (e.g., microlens arrays), and / or parallax barriers to redirect images to multiple different viewing regions associated with the displays.

[0033] In some exemplary displays, there may be only one location that provides a 3D view of the image content (e.g., a user, an object, etc.) that such a display provides. A user can sit in this one location and experience realistic 3D images with correct parallax and minimal distortion. If the user moves to a different physical location (or changes head position or eye gaze position), the image content (e.g., the user, objects worn by the user, and / or other objects) will begin to appear less realistic, 2D, and / or distorted. The systems and techniques described herein can reconstruct the image content projected from the display, ensuring that the user can continue to experience realistic, parallax-correct, low-distortion 3D images in real time, even while moving. Thus, the systems and techniques described herein have the advantage of maintaining and providing 3D image content and objects for display to the user regardless of user movement that occurs while the user is viewing the 3D display.

[0034] 1, telepresence system 180 may comprise one or more networks. Network 198 may be a publicly available network (e.g., the Internet) or a private network, to name two examples. Network 198 may be wired, wireless, or a combination of the two. Network 198 may comprise or utilize one or more other devices or systems, including, but not limited to, one or more servers (not shown).

[0035] The telepresence system 180 further comprises a microphone array 172 and a speaker array 174 for the user 182, and a similar microphone array 172' and a speaker array 174' for the user 182'. These components are arranged in a state of normal operation to provide the most realistic audio experience for the users 182 and 182'. The speaker arrays 174 and 174' can provide 3D audio signals locally. The microphone arrays 172 and 172' can be used to detect the 3D audio signals from the users. The audio signals are then encoded and used to render a 3D sound field representing the sound in the telepresence system 180. The request may be sent to a remote user for processing.

[0036] 2 is a flowchart illustrating an example method 200 for calibrating a microphone and speaker. Method 200 may be performed by a software entity described with respect to FIG. 1 resident in memory 126 of user device computer 120 and executed by processing unit 124. Alternatively, method 200 may be performed by a software entity resident in memory of a computing device different from (e.g., remote from) user device computer 120.

[0037]

number

[0038] At 204, the power spectral density manager 150 generates, for each microphone in the microphone array and for each speaker in the speaker array, a respective power spectral density (e.g., power spectral density data 154) for that microphone and that speaker based on the respective reverberant sound field generated by that speaker and received by that microphone. The generation of power spectral densities is described in more detail in FIG. 3.

[0039] At 206, the calibration filter manager 160 generates, for each microphone of the microphone array, a respective calibration filter (e.g., microphone calibration data 162) that is based on the ratio of the power spectral density averaged over the speaker array and microphone array (Equation (7)) to the power spectral density averaged over the speaker array (Equation (5)).

[0040] These calibration filters are used to calibrate microphones to record the same spectrum in response to the same input stimulus, and to calibrate speakers to produce the same spectrum in response to the same input stimulus. Also, because the calibration filters are based on the reverberant sound field rather than the direct sound field, the calibration coefficients are not significantly affected by the geometry of the environment in which the speakers and microphones are located, nor by any nodes in the direct sound signal, and the calibration filters allow the microphones to produce high quality, highly directional sound signals.

[0041] At 208, the computer 120 records acoustic signals from the user with the microphone array using respective calibration filters generated for each microphone in the microphone array. Each microphone in the microphone array records essentially the same spectrum of the acoustic signal. The signals recorded from the calibrated microphones can be processed to generate spatial audio signals representative of sounds in the microphone array's environment (e.g., utterances made by one or more speakers using a telepresence system with the microphone array), and the generated spatial audio signals can be transmitted to a sound rendering system (e.g., a remote telepresence system) for rendering.

[0042] 3 is a flowchart illustrating an example process 300 for calibrating microphones of a microphone array. Process 300 may be performed by a software entity described with respect to FIG. 1 resident in memory 126 of user device computer 120 and executed by processing unit series 124. Alternatively, process 300 may be performed by a software entity resident in memory of a computing device different from (e.g., remote from) user device computer 120.

[0043] At 301, the impulse response manager 140 measures the reverberant (raw) impulse response from each channel, i.e., each microphone / speaker pair. As previously mentioned, the impulse response may be derived from reflections from walls, ceilings, floors, or objects received at the microphones resulting from the swept sine chirp generated by the reverberant sound field manager 130. The actual recording of the signal occurs at a start time, which occurs a sufficient time after the direct audio signal is received at the microphone. Thus, the reverberant sound field measured at the microphone contains only signals reflected off boundaries and obstacles.

[0044] An exemplary raw impulse response function for two speakers and four microphones is shown in FIG. 4A, resulting in eight raw impulse response functions. In some embodiments, each of the raw impulse response functions is measured based on the reverberant sound field from one speaker. In some embodiments, each speaker generates a swept sine chirp at a separate time. In some embodiments, these raw impulse response functions are measured at the microphone array one at a time. In some embodiments, the raw impulse response functions are measured at the microphone array one at a time.

[0045] At 302, the attenuation normalization manager 141 estimates the impulse response energy as a function of time for each channel. Figure 4B shows an example time-dependent energy metric associated with the raw impulse response function of Figure 4A. Note that the Y-axis coordinate value of the plot in Figure 4B is the square root of the energy.

[0046] At 303, the attenuation normalization manager 141 averages the impulse response energies across microphones and speakers to generate an average impulse response energy according to equation (3). Figure 4C shows the example time-dependent energy metric of Figure 4B averaged over all speakers and microphones.

[0047] At 304, the decay normalization manager 141 normalizes the raw impulse responses for each channel using the average impulse response energy to generate decay normalized impulse response functions. Figure 4D shows an example of a decay normalized impulse response function corresponding to the raw impulse response functions of four microphones and two speakers.

[0048] At 305, the subsegment manager 142 generates a subsegment of the damped normalized impulse response function. The sub-segment is extracted over a fixed time interval (i.e., from a first time to a second time) to generate a time-based sub-segment. Figure 4E shows an example sub-segment of the decaying normalized impulse response function shown in Figure 4D.

[0049] At 306, the convolution manager 151 generates white noise autocorrelation functions from the individual subsegments. Figure 4F shows an example multi-channel white noise autocorrelation function derived from the subsegments of the decaying normalized impulse response function shown in Figure 4E.

[0050] At 307, the transform manager 152 performs a Fourier transform on the white noise autocorrelation function over a short time window to generate a power spectral density for each channel according to equation (4). Figure 4G shows an example power spectral density corresponding to the autocorrelation function of the multi-channel white noise shown in Figure 4F.

[0051] At 308, the calibration filter manager 160 generates an average of the power spectral densities for the speakers to generate a speaker-averaged power spectral density. Figure 4H shows the example power spectral density of Figure 4G averaged over the speakers.

[0052] At 309, the calibration filter manager 160 generates a microphone calibration filter as the ratio of the power spectral density averaged over the microphone and speaker to the power spectral density averaged over the speaker. Figure 4I shows an example microphone calibration filter derived from the speaker-averaged power spectral density of Figure 4H.

[0053] Note that steps 308 and 309 can also be applied to generating speaker calibration filters. Figure 4J shows the example power spectral densities of Figure 4G averaged over the microphones. Figure 4K shows an example speaker calibration filter derived from the power spectral densities averaged over the microphones of Figure 4J.

[0054] FIG. 5 is a diagram illustrating an example of a generic computing device 500 and a generic mobile computing device 550 that may be used with the techniques described herein.

[0055] 5, computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants (PDAs), servers, blade servers, mainframes, and other suitable computers. Computing device 550 is intended to represent various forms of mobile devices, such as PDAs, cell phones, smartphones, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and do not limit the scope of the invention(s) described and / or claimed herein.

[0056] Computing device 500 includes a processor 502, memory 504, storage device 506, a high-speed interface 508 connected to memory 504 and a high-speed expansion port 510, and a low-speed interface 512 connected to a low-speed bus 514 and storage device 506. Each of components 502, 504, 506, 508, 510, and 512 are connected to one another using various buses and may be implemented on a common motherboard or otherwise implemented as appropriate. Processor 502 is capable of processing instructions for execution within computing device 500. The instructions may be used to display graphics for a GUI on an external input / output device, such as a display 516 coupled to high-speed interface 508. The computing device 500 includes instructions stored in memory 504 or stored on storage device 506 for displaying the network information. In other implementations, multiple processors and / or multiple buses may be utilized, along with multiple memories and types of memory, as appropriate. Also, multiple computing devices 500 may be connected, each providing a portion of the required operations (e.g., as a server bank, a cluster of blade servers, or a multi-processor system).

[0057] The memory 504 stores information within the computing device 500. In one implementation, the memory 504 is one or more volatile storage devices. In another implementation, the memory 504 is one or more non-volatile storage devices. The memory 504 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.

[0058] Storage device 506 can provide mass storage for computing device 500. In one embodiment, storage device 506 can be or include a computer-readable medium, such as a floppy disk drive, hard disk drive, optical disk drive, or tape drive, a flash memory or other similar solid-state memory device, or an array of devices, including devices included in a storage area network or other configuration. A computer program product can be tangibly embodied on an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 504, storage device 506, or memory on processor 502.

[0059] High-speed controller 508 manages operations requiring more bandwidth for computing device 500, while low-speed controller 512 manages operations requiring more lower bandwidth. This allocation of functionality is merely exemplary. In one embodiment, high-speed controller 508 is coupled to memory 504 (e.g., through a graphics processor or accelerator), display 516, and high-speed expansion port 510. High-speed expansion port 510 may accept various expansion cards (not shown). In this embodiment, low-speed controller 512 is coupled to storage device 506 and low-speed expansion port 514. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, etc., or to a network device, such as a switch or router, for example, through a network adapter.

[0060] Computing device 500 may be implemented in a number of different forms as shown. For example, it may be implemented as a standard server 520, multiple times in a cluster of such servers, or as part of a rack server system 524. Additionally, it may be implemented as a personal computer, such as a laptop computer 522. Alternatively, the components of computing device 500 may be combined with other components of a mobile device (not shown), such as device 550. Each such device may include one or more of computing devices 500, 550, and the entire system may consist of multiple computing devices 500, 550 in communication with each other.

[0061] Computing device 550 includes, among other things, a processor 552, memory 564, input / output devices such as a display 554, a communication interface 566, and a transceiver 568. Storage devices, such as a microdrive or other device, may be provided on device 550. Each of the components 550, 552, 564, 554, 566, and 568 are connected to one another using various buses, and some of these components may be implemented on a common motherboard or in other suitable manners.

[0062] The processor 552 can execute instructions (including instructions stored in memory 564) within the computing device 450. The processor may be implemented as a chipset of chips including separate analog and digital processors. The processor enables the coordination of other components of the device 550, for example, controlling the user interface, controlling applications run by the device 550, and controlling wireless communications by the device 550.

[0063] Processor 552 may communicate with a user through a control interface 558 and a display interface 556 coupled to a display 554. Display 554 may be, for example, a TFT LCD (thin film transistor liquid crystal display) or an OLED (organic light emitting diode) display, or other suitable display technology. Display interface 556 may include appropriate circuitry for driving display 554 to present graphical and other information to the user. Control interface 558 may receive commands from the user and convert them for execution by processor 552. Additionally, an external interface 562 may be provided in communication with processor 552 to enable device 550 to communicate with other devices over short distances. For example, external interface 562 may enable wired communication in some implementations and wireless communication in other implementations, and multiple interfaces may be used.

[0064] Memory 564 stores information within computing device 550. Memory 564 may be embodied as one or more computer-readable media, one or more volatile storage devices, or one or more nonvolatile storage devices. Additionally, expansion memory 574 may be provided and connected to device 550 through expansion interface 572. Expansion interface 572 may include, for example, a Single In Line Memory Module (SIMM) card interface. Such expansion memory 574 may provide additional storage space for device 550 or may store applications or other information for device 550. Specifically, expansion memory 574 may include instructions for performing or assisting in the above-described processes and may also include secured information. Thus, for example, expansion memory 574 may be provided as a security module for device 550, or instructions enabling secured use of device 550 may be programmed into expansion memory 574. Additionally, secured applications may be provided via SIMM cards, along with additional information, such as placing identifying information on the SIMM card in an unhackable manner.

[0065] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier may be a computer-readable or machine-readable medium, such as memory 564, expansion memory 574, or memory on processor 552, and may be received, for example, via transceiver 568 or external interface 562.

[0066] Device 550 may communicate wirelessly through communication interface 566. The interface may include digital signal processing circuitry, if necessary. The communication interface may enable communication under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communication may occur, for example, through radio frequency transceiver 568. Additionally, short-range communication may occur, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). Additionally, GPS (Global Positioning System) may be used. System) receiver module 570 may provide additional navigation-related or location-related wireless data to device 550. The additional navigation-related or location-related wireless data may be utilized by applications executing on device 550, as appropriate.

[0067] Device 550 may also communicate via voice using audio codec 560. Audio codec 560 may accept voice information from a user and convert it into usable digital information. Similarly, audio codec 560 may generate sounds for the user, such as through a speaker in a handset of device 550. Such sounds may include audio from a voice telephone call, recorded audio (e.g., voice messages, music files, etc.), and sounds generated by applications running on device 550.

[0068] The computing device 550 may be implemented in a number of different forms, as shown, for example, as a mobile phone 580, or as part of a smartphone 582, personal digital assistant, or other similar mobile device.

[0069] Various implementations of the systems and techniques described herein may be realized in digital electronic circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable on a programmable system. The programmable system includes at least one programmable processor, which may be a special-purpose processor or a general-purpose processor, coupled to a storage system, at least one input device, and at least one output device to receive and transmit data and instructions.

[0070] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that accept machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0071] To enable user interaction, the systems and techniques described herein include a computer with a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) for allowing the user to provide input to the computer. Other types of devices may be used to enable interaction with a user; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual, auditory, or haptic feedback), and input from the user may be accepted in any form, such as acoustic, speech, or tactile input.

[0072] The systems and techniques described herein may be implemented in a computer system with back-end components (e.g., a data server), a computer system with middleware components (e.g., an application server), a computer system with front-end components (e.g., a client computer having a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or any combination of such back-end, middleware, and front-end components. These components of the system may be connected to each other by any form or medium of digital data communication (e.g., a communication network). Communication networks include LANs (local area networks), WANs (wide area networks), the Internet, and the like.

[0073] A computer system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server exists by virtue of computer programs running on the respective computers and having a client-server relationship.

[0074] Returning to FIG. 1 , in some embodiments, memory 126 may be any type of memory, such as RAM, disk drive memory, flash memory, etc. In some embodiments, memory 126 may be implemented as two or more memory components (e.g., two or more RAM components or disk drive memory) associated with components of compression computer 120. In some embodiments, memory 126 may be database memory. In some embodiments, memory 126 may be or include non-local memory. For example, memory 126 may be or include memory shared by multiple devices (not shown). In some embodiments, memory 126 may be associated with a server device (not shown) in a network and configured to provide components of compression computer 120.

[0075] The components of the compression computer 120 (e.g., modules, processing unit 124) may be configured to operate based on one or more platforms (e.g., one or more similar or different platforms), which may include one or more types of hardware, software, firmware, operating systems, runtime libraries, etc. In some implementations, the components of the compression computer 120 may be configured to operate within a collection of devices (e.g., a server farm). In such implementations, the functionality and processing of the components of the compression computer 120 may be distributed across several devices included in the collection of devices.

[0076] The components of computer 120 may be or include any type of hardware and / or software configured to process attributes. In some implementations, one or more of the components shown in the components of computer 120 in FIG. 1 may be hardware-based modules (e.g., DSPs (digital signal processors), FPGAs (field programmable gate arrays), memory), firmware modules, and / or software-based modules (e.g., modules of computer code, 1. The computer 120 may be a hardware-based program (a series of computer-readable instructions that can be executed by a processor) or may include such hardware-based modules. For example, in some embodiments, one or more of the components of the computer 120 may be or include software modules configured to be executed by at least one processor (not shown). In some embodiments, the functionality of the components may be included in modules and / or components different from those shown in FIG. 1.

[0077] Although not shown, in some embodiments, components of computer 120 (or portions thereof) may be configured to operate within, for example, a data center (e.g., a cloud computing environment), a computer system, one or more server / host devices, etc. In some embodiments, components of computer 120 (or portions thereof) may be configured to operate within a network. Thus, components of computer 120 (or portions thereof) may be configured to function within various types of network environments that may include one or more devices and / or one or more server devices. For example, the network may be or include a local area network (LAN), a wide area network (WAN), etc. The network may be or include a wireless network and / or a wired network implemented using, for example, gateway devices, bridges, switches, etc. The network may include one or more segments and / or portions of segments based on various protocols, such as Internet Protocol (IP) and / or proprietary protocols. The network may include at least a portion of the Internet.

[0078] In some embodiments, one or more of the components of computer 120 may be or include a processor configured to process instructions stored in memory. For example, depth image manager 130 (and / or portions thereof), viewpoint manager 140 (and / or portions thereof), ray casting manager 150 (and / or portions thereof), SDV manager 160 (and / or portions thereof), aggregation manager 170 (and / or portions thereof), route finding manager 180 (and / or portions thereof), and depth image generation manager 190 (and / or portions thereof) may be a combination of a processor and memory configured to execute instructions related to processing to implement one or more functions.

[0079] While several embodiments have been described, it will be understood that various modifications may be made without departing from the spirit and scope of the present specification.

[0080] When an element is referred to as being provided on, connected to, electrically connected to, coupled to, or electrically coupled with another element, it will be understood that the element may be directly provided on, connected to, or coupled with the other element, or that one or more intermediate elements may be present. In contrast, when an element is referred to as being directly provided on, connected to, or directly coupled with another element, there are no intermediate elements present. Although the terms "directly provided on," "directly connected to," or "directly coupled to" may not be used throughout the detailed description, elements shown as being directly provided on, directly connected to, or directly coupled with may be so referred to. The claims of this application may be amended to describe examples of relationships described or illustrated in the specification.

[0081] While certain features of the above-described embodiments have been illustrated as described herein, many variations, substitutions, changes, and equivalents will now occur to those skilled in the art. Therefore, it is to be understood that the claims are intended to encompass within the scope of the embodiments all such variations and modifications. These have been presented by way of example only, and not limitation, and it should be understood that various changes in form and detail may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination except mutually exclusive combinations. The embodiments described herein may include various combinations and / or subcombinations of the functions, components, and / or features of the different embodiments described.

[0082] Additionally, the illustrated logic flows do not necessarily require the particular order shown, or the sequence shown, to achieve desirable results. Additionally, other steps may be provided or eliminated from the described flows, and other components may be added or removed from the described systems. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. receiving, via a microphone array, a reverberant sound field based on audio signals generated by the speaker array; generating power spectral densities for microphones of the microphone array and speakers of the speaker array based on the reverberant sound fields generated by the speakers and received at the microphones; generating a calibration filter for the microphone based on the power spectral densities averaged over the speaker array and the microphone array.

2. The method of claim 1 , further comprising recording acoustic signals from a user with the microphone array using the respective calibration filters generated by the microphones.

3. generating the calibration filters for the microphones of the microphone array based on the power spectral densities averaged over the speaker array and the microphone array, 3. The method of claim 1, comprising generating the calibration filter for the microphone as a ratio of the power spectral density averaged over the speaker array and the microphone array to the power spectral density of the microphone averaged over the speaker array.

4. Generating the respective power spectral densities of the microphone and the speaker comprises:

4. The method of claim 1, further comprising generating respective impulse response functions for the microphone and the loudspeaker based on the respective reverberant sound fields generated by the loudspeaker and received at the microphone.

5. Generating the respective power spectral densities of the microphone and the speaker comprises: performing an autocorrelation of the impulse response functions of the microphone and the speaker to generate an autocorrelated impulse response function; 5. The method of claim 4, further comprising: performing a transform on the autocorrelation impulse response function to frequency space to generate the power spectral densities of the microphone and the speaker.

6. Performing a transformation of the autocorrelation impulse response function to a frequency space includes: generating a window function that is equal to a constant within a specified time interval and equal to zero outside said specified time interval; and performing a Fourier transform operation on the product of the window function and the autocorrelation impulse response function.

7. Before generating each of the impulse response functions, 7. The method of claim 4, further comprising generating, as the audio signal, a swept sine chirp signal having a frequency between a first frequency and a second frequency at the speaker, wherein the swept sine chirp signal is received at the microphone.

8. Generating each of the impulse response functions includes: measuring a raw impulse response function corresponding to the microphone and the speaker; generating a time-dependent energy metric associated with the RAW impulse response function; generating a normalization factor for the microphone array and the speaker array based on an average of the time-dependent energy metrics associated with the microphones and the speakers; and dividing the raw impulse response functions corresponding to the microphone and the loudspeaker by the normalization factor to generate damped normalized impulse response functions corresponding to the microphone and the loudspeaker.

9. Generating the time-dependent energy metric associated with each of the RAW impulse response functions corresponding to the microphone and the speaker comprises: generating a first power of the absolute value of the raw impulse response function corresponding to the microphone and the speaker; 9. The method of claim 8, comprising performing a smoothing operation on the first power of the absolute value of the raw impulse response function to generate the time-dependent energy metric associated with the raw impulse response function corresponding to the microphone and the speaker.

10. performing the smoothing operation 10. The method of claim 9, comprising generating a moving average of the first power of the absolute value of each of the raw impulse response functions during a specified time period.

11. generating the normalization factor 11. The method of claim 9, comprising generating a second power of the time-dependent energy metric associated with the RAW impulse response functions corresponding to the microphone and the speaker, the second power being the reciprocal of the first power.

12. Generating each of the impulse response functions includes:

12. The method of claim 8, further comprising obtaining a sub-segment of the damped normalized impulse response function corresponding to the microphone and the loudspeaker as the respective impulse response functions corresponding to the microphone and the loudspeaker, the sub-segment starting at a first time and ending at a second time.

13. The method of claim 12 , wherein the first time period is based on the shortest distance traveled by echo waves in the reverberant sound field.

14. A computer program comprising code which, when executed by processing circuitry of a computing device, causes the processing circuitry to perform the method of any of claims 1 to 13.

15. a memory for storing the computer program of claim 14; and a control circuit coupled to the memory.

Citation Information

Patent Citations

  • Voice recognition robot

    JP2008164747A

  • Pickup signal processing apparatus, method, and program

    JP2010232717A

  • Adaptive self-calibration of small microphone array by soundfield approximation and frequency domain magnitude equalization

    US20130170666A1

  • Conferencing Device Self Test

    US20150049583A1

  • Sound level estimation

    US20170127206A1