Techniques for performing positional calibration for a multi-channel speaker system using directional microphone arrays
The use of directional microphones within audio output devices for automated speaker calibration and channel assignment addresses the challenges of wireless multi-channel speaker systems, ensuring accurate placement and high-quality audio performance.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HARMAN INT IND INC
- Filing Date
- 2023-09-25
- Publication Date
- 2026-04-23
AI Technical Summary
The transition to wireless connectivity in multi-channel speaker systems necessitates effective techniques for accurate speaker position calibration and channel assignment, as existing methods using external microphones or omni-directional microphones are prone to misalignment and incorrect placement, leading to compromised audio quality and reversed sound images.
A computer-implemented method using directional microphones within each audio output device to capture a test audio object, determine room impulse responses, and automatically assign channels based on computed positions, ensuring accurate speaker placement and channel assignment.
This method provides automated, accurate, and robust calibration even in compact structures, eliminating the risk of reversed sound images and compromised audio quality, while aligning with sleek speaker aesthetics and facilitating a user-friendly setup process.
Smart Images

Figure CN2023121054_23042026_PF_FP_ABST
Abstract
Description
TECHNIQUES FOR PERFORMING POSITIONAL CALIBRATION FOR A MULTI-CHANNEL SPEAKER SYSTEM USING DIRECTIONAL MICROPHONE ARRAYSBACKGROUND
[0001] Field of the Various Embodiments
[0002] The various embodiments relate generally to audio output devices and, more specifically to techniques for performing positional calibration for a multi-channel speaker system using directional microphone arrays.
[0003] Description of the Related Art
[0004] The field of integrated home entertainment systems has witnessed a surge in popularity with the advent of multi-channel speaker arrangements. These systems offer immersive audio experiences, enhancing the enjoyment of movies, music, virtual reality, and gaming. To cater to the growing demand for seamless and clutter-free solutions, several companies have introduced operating systems enabling users to establish wireless connections among a designated group of speakers, creating a cohesive multi-channel speaker configuration. This wireless approach eliminates the need for extensive cabling, promoting both convenience and aesthetic appeal in interior spaces. However, the shift towards wireless connectivity introduces a novel challenge -the accurate detection and assignment of speaker channels during the setup process.
[0005] The accurate calibration of speaker position and assignment of speaker channels is crucial for achieving optimal audio performance in a multi-channel speaker system. Incorrectly assigned channels can lead to skewed audio localization and compromised surround sound quality. Traditional wired systems have relied on manual calibration procedures to determine speaker positions and assign corresponding channels. Yet, the transition to wireless connectivity necessitates innovative methods for channel calibration.
[0006] Acoustic channel position calibration is a pivotal requirement to mitigate the potential for reversed sound image and subpar audio quality. Many existing commercial speaker calibration systems on the market use an external microphone for accurate calibration. In addition to the added clutter caused by external microphones, one significant drawback of relying on an external microphone for speaker calibration is the potential for misalignment or incorrect placement. Users might not position the external microphone accurately or consistently across different calibration sessions. This variability in microphone placement could lead to inconsistent calibration results, negatively impacting the overall audio quality and performance of the multi-channel speaker system. In other commercial systems, calibration involves in-situ measurements facilitated by inserting microphones into the speakers themselves. While this type of calibration is more user-friendly, these systems do not perform automatic speaker assignment correction. This deficiency can result in an unsatisfactory auditory experience for users, manifested as a reversed sound image when the left and right speakers are inadvertently switched. Certain other existing speaker systems have utilized omni-directional microphones for speaker position calibration using direction of arrival (DOA) information. However, this approach faces limitations, particularly when dealing with compact structures that lack ample spacing for microphone placement. For instance, closely spaced omnidirectional microphones prevent accurate estimation of the time difference of arrival (TDOA) between various acoustic signals. The absence of systems providing automated speaker channel assignment presents a notable inconvenience to users and an increased risk of incorrect channel assignments.
[0007] As the foregoing illustrates, what is needed are more effective techniques for performing positional calibration for a plurality of audio output devices and seamless automatic channel assignment in a multi-device system comprising the plurality of audio output devices, such as compact audio output devices.SUMMARY
[0008] In various embodiments, a computer-implemented method of performing positional calibration for a multi-channel audio system includes, capturing, using a plurality of directional microphones within a second audio output device, a test audio object emitted by a first audio output device; determining, for each of the plurality of directional microphones configured within the second audio output device, a respective room impulse response (RIR) associated with the test audio signal captured by the direction microphone; determining a position of the second audio output device relative to the first audio output device based on the RIRs; assigning a respective channel in the multi-channel audio system to each of the first audio output device and the second audio output device based on the computed position; and emitting audio from the first audio output device based on the respective channel assigned to the first audio output device.
[0009] At least one technical advantage of the disclosed techniques relative to the prior art is that, with the disclosed techniques, an automated calibration of speaker positions in a wireless multi-channel speaker system can be performed providing for more accurate assignment of each speaker to its fitting channel within the multi-channel arrangement. Using the calibration for channel assignment in the system ensures that the sound image is reproduced correctly. In addition, the disclosed calibration techniques can determine the location of each audio output device relative to the other audio output devices automatically and accurately, as compared with other methods that involve performing in-situ measurements at the factory or using external microphones. Unlike the limitations associated with external microphones or closely spaced omni-directional microphones, the disclosed techniques ensure robust and dependable calibration outcomes even within compact structures. Consequently, the disclosed techniques not only eliminate the risk of reversed sound images and compromised audio quality resulting from manual or inconsistent calibration approaches but also offers a streamlined and user-friendly setup process to perform auto-channel assignment in a multi-channel system. These technical advantages provide one or more technological improvements over prior art approaches.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] So that the manner in which the above recited features of the various embodiments can be understood in detail, a more particular description of the inventive concepts, briefly summarized above, may be had by reference to various embodiments, some of which are illustrated in the appended drawings. It is to be noted, however, that the appended drawings illustrate only typical embodiments of the inventive concepts and are therefore not to be considered limiting of scope in any way, and that there are other equally effective embodiments.
[0011] Figure 1 illustrates an audio output device configured according to various embodiments;
[0012] Figures 2A and 2B are illustrations of a relative position determination of a first audio output device by a second output device of Figure 1, according to various embodiments;
[0013] Figure 3 is an illustration of a Room Impulse Response (RIR) signal computed for each microphone in an audio output device of Figure 1 with two directional microphones, according to various embodiments;
[0014] Figures 4A, 4B and 4C are illustrations of arrangements of directional microphones in exemplary directional microphone arrays incorporated within an audio output device of Figure 1, according to various embodiments;
[0015] Figure 5 is an illustration of calibration signals between pairs of audio output devices from the plurality of audio output devices represented in Figure 1, according to various embodiments;
[0016] Figure 6 is an illustration of calibration signals between the plurality of audio output devices in Figure 1, according to various embodiments;
[0017] Figure 7 is an illustration of a process for computing a RIR function by an audio output device of Figure 1, according to various embodiments;
[0018] Figures 8A and 8B provide an illustration of the manner in which a relative angular position of an audio output device is determined by the device of Figure 1, according to various embodiments;
[0019] Figure 9 illustrates a flow diagram of method steps for performing positional calibration for the plurality of audio output devices of Figure 1, according to various embodiments;
[0020] Figure 10 illustrates a flow diagram of method steps for determining an angular position for an audio output device represented in Figure 1, according to various embodiments.DETAILED DESCRIPTION
[0021] In the following description, numerous specific details are set forth to provide a more thorough understanding of the various embodiments. However, it will be apparent to one skilled in the art that the inventive concepts may be practiced without one or more of these specific details.
[0022] Figure 1 illustrates an audio output device 110A configured to implement one or more aspects of the various embodiments. As shown, the audio output device 110A includes, without limitation, a processor 102, memory 104, storage 106, an interconnect bus 108, a speaker 192 and an array of N directional microphones (including microphones 126A, 126B …126N) . In various embodiments, the positioning of the N directional microphones 126 within audio output device 110A allows for flexibility, with the exception of directly in front of the speaker 192. For example, these N directional microphones 126 could be located above, below, or behind the speaker 192.
[0023] As shown, the memory 104 includes, without limitation, a position calibration engine 114, device positions for each of M-1 audio outputted devices (e.g., device positions 140B, 140C …140M corresponding to respective positions of audio output devices 110B, 110C …110M relative to audio output device 110A) , a channel assignment module 128 and an audio object 124. For explanatory purposes, multiple instances of like objects are denoted with reference numbers identifying the object and letters identifying the instance, where needed. Within the framework of the position calibration engine 114, the memory 104 can encompass additional components in various embodiments. These can include an RIR peak amplitude module 152 and a RIR time of arrival module 154.
[0024] The processor 102 can be any suitable processor, such as a central processing unit (CPU) , a graphics processing unit (GPU) , an application-specific integrated circuit (ASIC) , a field programmable gate array (FPGA) , a digital signal processor (DSP) , and / or any other type of processing unit, or a combination of different processing units, such as a CPU configured to operate in conjunction with a GPU. In general, the processor 102 can be any technically feasible hardware unit capable of processing data and / or executing software applications.
[0025] Memory 104 can include a random-access memory (RAM) module, a flash memory unit, or any other type of memory unit or combination thereof. The processor 102 is configured to read data from and write data to memory 104. Memory 104 includes various software programs (e.g., an operating system, one or more applications) that can be executed by the processor 102 and application data associated with the software programs. Storage 106 can include non-volatile storage for applications and data and can include fixed or removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-Ray, HD-DVD, or other magnetic, optical, or solid-state storage devices. The interconnect bus 108 connects the processor 102, the memory 104, the storage 106, and any other components of the audio output device 110A.
[0026] The audio output device 110A is coupled to a plurality of M-1 audio output devices 110 (e.g., audio output devices 110B, 110C …110M) located in a physical space 112. In total, M audio output devices, which include audio output device 110A and the M-1 audio output devices 110 coupled to audio output device 110A, are depicted in Figure 1. In various embodiments, M is any number (e.g., two, three, five, and / or six or more) . Each of the M-1 audio output devices 110 coupled to audio output device 110A operate substantially similarly to audio output device 110A and are internally configured with substantially similar components.
[0027] As depicted in Figure 1, each of the M-1 audio output devices situated in the physical space 112 can establish a communication link with the audio output device 110A through wireless techniques. These wireless techniques can involve technologies such as Wi-Fi, Bluetooth, or other comparable wireless methods, and the communication itself occurs via a network 130. Moreover, each of the M-1 audio output devices is capable of emitting an audio signal 180, which becomes discernible to the audio output device 110A when positioned in close enough vicinity to the audio output device 110A. Furthermore, reciprocally, the audio output device 110A is also capable of emitting an audio signal 180, which becomes perceptible to the audio output devices 110 if positioned adequately close to audio output device 110A.
[0028] In some embodiments, the plurality of audio output devices 110 includes, for example, a set of speakers in a home theater system. In some embodiments, the combined configuration of audio output device 110A and the M-1 audio output devices 110 constitutes a multi-channel speaker system. In some audio systems (e.g., a home theater system) that are capable of rendering spatial audio, each audio output device 110 corresponds to a particular channel corresponding to a certain location within the physical space 112, such as a front left channel, a front right channel, a rear left channel, and a rear right channel.
[0029] As shown, the position calibration engine 114 is a program stored in the memory 104 and executed by the processor 102 to calibrate a relative position (e.g., including a relative angular position) for each of the M-1 audio output devices 110 in the multi-channel speaker system. Similarly, each of the M-1 audio output devices 110 can also be configured with a substantially similar position calibration engine (not shown) that calibrates a relative position for the audio output device 110A (or any of the other audio output devices 110) . Positional calibration is performed for each of the M audio output devices 110 because the effectiveness of the rendered spatial audio (e.g., the clarity with which a listener perceives that an audio signal is positioned at a particular location within the physical space 112) is related to the accuracy with which the relative locations of the audio output devices 110 match the actual relative locations of the audio output devices 110 within the physical space 112. Rather than defining each of the M audio output devices 110 as rendering a channel associated with a fixed location (e.g., a front left speaker that is expected to be positioned in a front left corner of the room) , the audio output device 110A performs a calibration to detect the position (or angle) of each M-1 audio output device 110 relative to the position of the audio output device 110A within the physical space 112. Based on the calibration, the audio output device 110A stores a relative device position 140 of each of the M-1 audio output devices 110 (e.g., device positions 140B, 140C …140M corresponding to respective positions of audio output devices 110B, 110C …110M relative to audio output device 110A) . As mentioned above, each of the M-1 audio output devices 110 can be configured to perform a similar calibration to determine relative positional information for the audio output device 110A (or even each of the other audio output devices 110) .
[0030] In some embodiments, the audio signal 180 is generated by the audio output device 110A for each of the plurality of audio output devices 110 to calibrate a respective position for each of the plurality of audio output devices 110 relative to audio output device 110A. For example, in audio systems in which a first audio output device 110A comprises a front left speaker, the position calibration engine 114, using speaker 192, transmits, to the second audio output device 110B, which can comprise a front right speaker, the audio signal 180, which includes the audio object 124. Notably, the audio object 124 can, in some embodiments, be a stored chirp signal residing in memory 104. Alternatively, the position calibration engine 114 could generate the chirp signal in real-time during the calibration in other embodiments.
[0031] While not shown in Figure 1, in various embodiments, the second audio output device 110B comprises N directional microphones (which operate substantially similarly to directional microphones 126 configured in audio output device 110A) , where N is any number (e.g., two, three, five, and / or six or more) . As noted above, each of the M-1 audio output devices 110 is also configured with a position calibration engine substantially similar to the position calibration engine 114 shown in Figure 1. The position calibration engine in audio output device 110B, for example, also comprises an RIR module, which operates substantially similarly to the RIR peak amplitude module 152 depicted in Figure 1, and a RIR time of arrival module 154, which operates substantially similarly to the RIR time of arrival module 154 depicted in Figure 1.
[0032] Continuing with the example from above, during the calibration, every directional microphone within the second audio output device 110B (the front right speaker) captures and processes the audio object 124, enabling the computation of a room impulse response (RIR) function corresponding to an audio signal (based on the audio object 124) received at each of the directional microphones. The RIR module within the second audio output device 110B (not shown) identifies a directional microphone from the set of N directional microphones that registers a highest peak value of the computed RIR function. In this scenario, consider a left-side directional microphone (from the N microphones) positioned within the second audio output device 110B (where the second audio output device 110B corresponds to the front right speaker and the first audio output device 110A corresponds to the front left speaker, as mentioned above) . The left-side directional microphone, oriented towards the incoming audio signal 180 from the first audio output device 110 (which comprises the front left speaker) , can, for example, register the highest peak RIR value. This outcome is attributed to the fact that the audio signal 180 reaches the left-side directional microphone through the most direct path (as will be discussed in further detail in relation to Figures 2A and 2B) . Based on an evaluation of the direction in which the left-side directional microphone faces, a position (or angle) of the audio output device 110B relative to the audio output device 110A can be determined. In some embodiments, using the channel assignment module 128, a channel in the multi-speaker channel system is assigned to each of the first audio output device 110A and the second audio output device 110B based on the determined relative positions. In some embodiments, the relative position (or angle) of the second audio output device 110B relative to the first audio output device 110A is stored in memory at the second audio output device 110B (e.g., in memory dedicated to device positions substantially similar to device positions 140 shown in Figure 1) . In some embodiments, the relative position can also be transmitted to the first audio output device 110A and stored in device position 140B.
[0033] In some embodiments, a respective channel is assigned to each of the M audio output devices 110 using channel assignment module 128 in a multi-channel speaker system subsequent to the complete computation of the entire speaker arrangement's geometry in the multi-channel speaker system. For example, continuing with the example from above, after determining a relative position of the second audio output device 110B relative to the first audio output device 110A using the directional microphones in the second audio output device 110B, a comparable calibration for determining the relative positions of audio output devices 110C to 110M relative to audio output device 110A can be conducted. Here, the first audio output device 110A transmits the audio object 124 to each of the other audio output devices 110C to 110M. Each of the audio output devices 110C to 110M then computes and stores a relative position (or angle) of the respective audio output device relative the first audio output device 110A by performing the calibration detailed above. Notably, it’s unnecessary for a separate or distinct transmission of the audio object 124 to each audio output device. In fact, each of the M-1 audio output devices (e.g., audio output devices 110B, 110C …110M) can compute the relative position for the first audio output device 110A utilizing the same audio object 124 transmission. In some embodiments, the relative position (or angle) of each of the M-1 audio output devices 110 relative the first audio output device 110A can also be transmitted over network 130 to the first audio output device and stored in device positions 140.
[0034] In some embodiments, once the relative position (or angle) of each of the other audio output devices (e.g., audio output device 110B, 110C …110M) relative the first audio output device 110A is computed, the computation of the entire speaker arrangement’s geometry in the multi-channel speaker system is considered complete and a respective channel can be assigned to each of the M audio output devices using the channel assignment module 128. In other words, enough information is available after computing the relative position (or angle) of each of the other audio output devices (e.g., audio output device 110B, 110C …110M) relative to the first audio output device 110A to assign an appropriate channel to each of the M audio output devices. Accordingly, no other calibrations are needed to assign the appropriate channels.
[0035] In other embodiments, however, a confirmation process is undertaken by the position calibration engine 114 whereby the computed relative positions (or angles) of each of the other audio output devices (e.g., audio output device 110B, 110C …110M) relative the first audio output device 110A are confirmed by computing, at the first audio output device 110A, a relative position (angle) of the first audio output device 110 relative to each of the other audio output devices (e.g., audio output device 110B, 110C …110M) . Accordingly, a relative position for the first audio output device 110A relative to each of the other M-1 audio output devices 110 is determined using the directional microphones housed in the first audio output device 110A. During discrete time intervals, each of the M-1 audio output devices 110 transmits an audio output object, such as a chirp signal (or an analogous audio output object substantially similar to audio object 124) to the first audio output device 110A. Employing the directional microphones within the first audio output device 110A, a calibration is employed to compute and store a relative position for the first audio output device 110A relative to each one of the other M-1 audio output devices 110 at the first audio output device 110A.
[0036] Where the first audio output device 110A is configured as the principal entity (or controlling entity) in the system, the relative positions computed and stored at each of the other M-1 audio output devices 110 are relayed to the first audio output device 110A over network 130. These relative positions are compared with respective relative positions computed by and stored at the first audio output device 110A. When the transmitted relative positions from the other audio output devices correspond to the respective positions determined by the first audio output device 110A, a comprehensive establishment of the complete speaker arrangement geometry within the multi-channel system can be confirmed. This confirmation is performed by the position calibration engine 114 and the confirmed positions can be stored in memory 104, specifically, using device positions 140 in memory 104. Thereafter, a respective channel can be assigned to each of the M audio output devices 110 in the multi-channel speaker system.
[0037] In some embodiments, depending on the number and configuration of the directional microphones within audio output device 110B, the relative position can include an angular position associated with audio output device 110A. In general, the M audio output devices 110 will typically be positioned at the corners of a regular polygon that might correspond to the expected locations of the channels. However, in certain situations, rather than being positioned at the vertices of a square or a rectangle, the M audio output devices 110 can be positioned at the vertices of a less predictable quadrilateral shape. Whether the M audio output devices are positioned at the corners of a regular polygon or a less predictable quadrilateral shape, an array of N directional microphones, where N is greater than 2, can be used to determine an angular position of an audio output device for which a calibration is being performed. In such instances, an angular orientation of the directional microphone in the array registering the highest value of the computed RIR function can be used to determine the relative angular position of the audio output device.
[0038] In some embodiments, to obtain a more precise angular position of the target audio output device, an interpolation process might be employed. This involves determining angular orientations for directional microphones neighboring the directional microphone which corresponds to the highest registered RIR values, along with their respective RIR calculations. Subsequently, an interpolation is performed using the angular orientations and RIR values for the neighboring directional microphones and the angular orientation and RIR value for the directional microphone registering the highest peak RIR amplitude value. This technique yields a more refined angular position estimate for the audio output device position being calibrated.
[0039] In some embodiments, in addition to the RIR peak amplitude module 152, the position calibration engine 114 further comprises a RIR time of arrival module 154, which determines a relative time of arrival of an RIR peak between one or more microphones. In some embodiments, the RIR time of arrival module 154 module is supplementary to the RIR peak amplitude module 152 and collaborates with the RIR peak amplitude module 152 to aid in the identification of the directional microphone within an N directional microphone array that most directly captures the audio signal 180 transmitted by the audio output device undergoing calibration. To illustrate, in the context mentioned earlier, one of the directional microphones in the N directional microphone array situated at the second audio output device 110B typically records the highest peak RIR value. The RIR time of arrival module 154 can be used to determine which of the directional microphones in the N directional microphone array receives the audio signal 180 ahead of the other directional microphones. In typical circumstances, the directional microphone registering the peak RIR value would also encounter the audio signal 180 ahead of the other directional microphones within the N directional microphone array. In such situations, the RIR time of arrival module 154 can be used to confirm that the directional microphone that registers the highest peak RIR value also receives the audio signal 180 ahead of other modules.
[0040] Nevertheless, the subtle variations in time of arrival among the N directional microphones could limit the efficacy of relying solely on time of arrival to accurately ascertain which directional microphone holds a more direct orientation toward the audio output device generating the audio signal 180. One challenge in relying exclusively on time of arrival calculations is the inherent imperfection of microphone directionality. Sounds originating from the rear of the directional microphone can inadvertently affect the outcomes. Moreover, the timing of the RIR peak can be further complicated by the influence of room reverberation. As a result, in several scenarios, the outcomes derived from the RIR time of arrival module 154 are employed to reinforce the conclusions drawn by the RIR peak amplitude module 152 152. This composite information is then employed to ascertain a relative position, inclusive of angular position, for the device undergoing calibration.
[0041] The disclosed techniques perform positional calibration leveraging arrays of directional microphones to achieve precise calibration of acoustic channel positions. Unlike existing calibration systems available in the market, the disclosed techniques result in compact speaker systems. The utilization of directional microphones allows for a streamlined design that aligns with the modern trend towards sleek and minimalistic speaker aesthetics. This compact design holds significant advantages, particularly for slim or small speakers that have limited space available for the mounting of additional components like external microphones. By integrating the channel assignment system within such space-constrained speaker configurations, the disclosed techniques present a breakthrough in achieving accurate channel calibration without compromising on the overall design and form factor of the speakers. Accordingly, the disclosed techniques lay the groundwork for a more efficient, convenient, and visually pleasing approach to setting up multi-channel speaker systems in contemporary home entertainment environments. The disclosed techniques align seamlessly with the contemporary trend towards wireless and integrated entertainment systems, providing an elegant and effective solution to the challenges faced in multi-channel speaker configurations.
[0042] Figures 2A and 2B are illustrations of a relative position determination of a first audio output device by a second output device of Figure 1, according to various embodiments. Figure 2A illustrates a configuration where two audio output devices are oriented in a same direction with the primary audio output device 210A positioned directly to the right of the secondary audio output device 210B. Both audio output devices of Figure 2A can operate substantially similarly to the audio output devices 110 of Figure 1. The primary audio output device 210A comprises an array of two directional microphones, 204A and 204B mounted directly behind the main lobe of a speaker 230A. Alternatively, the directional microphones 204A and 204B can be mounted above or below the speaker 230A. The secondary audio output device 210B comprises a speaker 230B. Note that the example of Figure 2A includes only two audio output devices and, accordingly, it is not necessary for both audio output devices to be equipped with directional microphones because the positional calibration can be performed using the primary audio output device. However, the secondary audio output device 210B would need to be equipped with at least a speaker to be able to transmit the audio object 124 (e.g., a chirp signal) .
[0043] When performing positional calibration, sound signals can be emitted from the speaker 230B associated with the secondary audio output device 210B while the speaker 230A is muted. In the illustrated scenario depicted in Figure 2A, the directional microphone 204A is oriented towards the source of the incoming audio signal 220. As a result of this alignment, the directional microphone 204A captures the incoming audio signal 220 directly. At the same time, most signals that arrive at directional microphone 204B are reflected audio signals 222 from the wall 210. Accordingly, the direct sound is received by directional microphone 204A at an angle of 0 degrees, whereas the reflected sound approaches from the opposite angle at directional microphone 204B, measuring 180 degrees.
[0044] The sound energy, typically computed by performing a RIR computation, experienced at directional microphone 204A will be higher (or, at least, have a higher peak) as compared to the sound energy experienced at directional microphone 204B. Furthermore, the direct audio signal 220 will be experienced at directional microphone 204A earlier than the reflected audio signal 222 will be experienced at directional microphone 204B. If both directional microphones 204A and 204B perform an RIR computation using the audio signals (e.g., a chirp signal) received at the respective microphones, the RIR computed for the audio signals received at directional microphone 204A would register higher peak values than the RIR computed for audio signals received at directional microphone 204B. Furthermore, the direct audio signal 220 would be received at directional microphone 204A earlier than the reflected audio signals would be received at directional microphone 204B.
[0045] Figure 3 is an illustration of a RIR function computed for each directional microphone in an audio output device of Figure 1 with two directional microphones, according to various embodiments. As discussed in connection with Figure 2A, in an audio output device 210A (which is substantially similar to audio output device 110A of Figure 1) , the two directional microphones 204A and 204B will receive audio signals transmitted from speaker 230B with different levels of intensity. As shown in Figure 3, signal 310 can be associated with the RIR computations performed for audio signals received at directional microphone 204A while signal 320 can be associated with RIR computations performed for reflected audio signals received at directional microphone 204B. As depicted in Figure 3, the signal310 exhibits a more pronounced peak value, and the peak reaches directional microphone 204A at a time duration ΔT 330 prior to the arrival of the reflected audio signal at directional microphone 204B.
[0046] Referring back to Figure 2B, a similar configuration to Figure 2A is illustrated with substantially similar components, but where the primary audio output device 210A is rotated such that the incoming audio signals from speaker 230B do not align with the directivity of directional microphones 204A and 204B. Figure 2B exemplifies the challenge posed by incoming audio signals arriving at oblique angles in the context of two directional microphones. In this configuration, the direct sound is received by directional microphone 204A at an angle of 100 degrees, whereas the reflected sound approaches directional microphone 204B from the opposite angle, measuring 280 degrees. In this scenario, the direct audio signal 220 would, in most cases, register a higher RIR peak value at directional microphone 204A and arrive earlier as compared to the reflected audio signal 222 received at directional microphone 204B. Nevertheless, the energy contrast between the direct and reflected audio signals might not be substantial enough for accurate calibration under specific rotation angles. This issue of misalignment between incoming audio signals and microphone directivity becomes especially prominent within a multi-speaker, multi-channel system. For those situations, additional microphones can be employed for accurate calibration.
[0047] Figures 4A, 4B and 4C are illustrations of arrangements of directional microphones in exemplary directional microphone arrays incorporated within an audio output device of Figure 1, according to various embodiments. In some embodiments, an N directional microphone arrangement, where N is greater than 2, can be necessary to deal with the challenge of irregular wall reflection and audio signals arriving at oblique angles from multiple audio output devices positioned at various angles relative to the audio output device comprising the directional microphone array. Figure 4A illustrates an exemplary arrangement with four directional microphones (e.g., directional microphones 420A, 420B, 420C and 420D) where the directional microphones are distributed along a circle in a circular formation and each microphone faces outward. In this setup, the potential of the directional microphones 420 extends beyond merely calibrating the position of an audio output device emitting a direct audio signal 450 from the leftward direction relative to the microphone arrangement. The configuration depicted in Figure 4A can also effectively calibrate the orientation of audio output devices situated in various positions, including, but not limited to, behind, in front of, or to the right of the audio output device comprising the directional microphone arrangement.
[0048] Figure 4B illustrates an exemplary arrangement with four directional microphones (e.g., directional microphones 420A, 420B, 420C and 420D) comprising two pairs of two back-to-back microphones. Note that the arrangement is not limited to two pairs of directional microphones, but any number of pairs of directional microphones can be envisioned. The arrangement of Figure 4B can be particularly useful in instances where pairs of audio output devices are calibrated individually prior to calibrating the relative positions of all the pairs in the system (as will be discussed further in connection with Figure 5) .
[0049] Figure 4C illustrates an exemplary arrangement with four directional microphones (e.g., directional microphones 420A, 420B, 420C, 420D, 420E, 420F, 420G, and 420H) where the directional microphones are distributed along a circle and each directional microphone faces outward. The arrangement of directional microphones in Figure 4C is similar to the exemplary arrangement of Figure 4A but utilizes additional directional microphones. With more microphones utilized, in addition to identifying the left and right channels seamlessly, the direction of arrival of sound can also be estimated and an angular position associated with one or more other audio output devices in the multi-channel system can be determined. For example, in the configuration of Figure 4C, each of the directional microphones is associated with a respective angular orientation. An angular orientation of the directional microphone registering the highest RIR peak amplitude value in this arrangement can be used to compute a relative angular position of the audio output device being calibrated.
[0050] Figure 5 is an illustration of calibration signals between pairs of audio output devices from the plurality of audio output devices represented in Figure 1, according to various embodiments. The audio device arrangement of Figure 5 contains multiple audio output devices (e.g., audio output devices 502A, 502B, 502C and 502D) , where each audio output device comprises a respective directional microphone array (e.g., directional microphone arrays 530A, 530B, 530C and 530D) . Within the context of the audio output devices shown in Figure 5, the configuration of directional microphone arrays can be chosen from a range of options, including those depicted in Figures 4A, 4B, or 4C, or similar alternatives. However, when considering a situation where individual calibration of two speaker pairs precedes the alignment of all pairs within the system, the arrangement presented in Figure 4B emerges as a suitable choice. Alternatively, the arrangement shown in Figure 4A can also be used.
[0051] While the setups depicted in Figures 2A and 2B necessitate just a single calibration due to the involvement of only two audio output devices, configurations featuring more than two audio output devices require multiple calibration processes to establish relative positions for all the audio output devices within the system. For example, the speaker system of Figure 5, where each audio output device comprises at least four directional microphones (e.g., arranged similar to the directional microphones in Figures 4A or 4B) a four-channel calibration would be needed, where each audio output device is dedicated to one of the four channels.
[0052] In some embodiments, each of the audio output device pairs (e.g., the audio output device pair towards the back comprising audio output devices 502A and 502B, the audio output device pair towards the front comprising audio output devices 502C and 502D) can first be calibrated separately. The relative position of audio output device 502 (a) can be determined by emitting an audio object 124 (e.g., a chirp signal) from audio output device 502A. Using the directional microphone array in audio output device 502B, an RIR function can be computed for each of the directional microphones in the directional microphone array 530B. The position of audio output device 502B relative to audio output device 502A is determined using an orientation of the directional microphone in the directional microphone array 530B that registers the highest RIR peak value. Having calibrated the back pair of audio output devices comprising audio output devices 502A and 502B, the front pair comprising audio output devices 502C and 502D can be similarly calibrated.
[0053] After calibrating the left and right channels for each pair, the relative position of the two pairs can be calibrated using the same procedures. For example, either audio output device 502A or audio output device 502B can transmit an audio object 124, where the audio object 124 is received by directional microphone arrays in either the audio output device 502C or the audio output device 502D. Using RIR functions computed by the directional microphone arrays in either of the audio output devices 502C and 502D, the position of the front pair comprising audio output devices 502C and 502D relative to either of the devices in the back pair can be determined. Alternatively, one of the devices in the front pair can transmit the audio object 124 while directional microphone arrays in either of the devices of the back pair can be used to determine a relative position of the front pair devices. Thereafter, an appropriate channel can be assigned to each of the devices in the four-channel system.
[0054] In some embodiments, instead of calibrating pairs of devices separately, a relative position, including an angular position, of each of the devices in the multi-channel system relative to one of the primary or controller devices can be determined. Figure 6 is an illustration of calibration signals between the plurality of audio output devices in Figure 1, according to various embodiments. The audio device arrangement of Figure 6 contains multiple audio output devices (e.g., audio output devices 602A, 602B, 602C and 602D) , where each audio output device comprises a respective directional microphone array (e.g., directional microphone arrangements 630A, 630B, 630C and 630D) .
[0055] Within the context of the audio output devices shown in Figure 6, the configuration of directional microphone arrays can be chosen from a range of options, including those depicted in Figures 4A, 4B, or 4C, or similar alternatives. However, when considering a situation where a relative position, including an angular position, of each of the devices in the multi-channel system relative to one of the primary or controller devices (e.g., primary audio output device 602A) is to be determined, the arrangement presented in Figure 4C emerges as a suitable choice. As the setup in Figure 4C involves multiple directional microphones, each audio output device can derive a more precise angular position concerning the primary device. For instance, if the directional microphone configuration from Figure 4C is adopted in the audio output devices illustrated in Figure 6, during the calculation of the relative position between audio output device 502D and audio output device 502A, the orientation of a specific directional microphone (e.g., directional microphone 420A from Figure 4C) can be leveraged to enhance the accuracy in determining the relative angular positioning of audio output device 502D with respect to audio output device 502A.
[0056] To determine the relative angular position of each of the audio output devices 602B, 602C and 602D relative to the primary audio output device 602A, a chirp signal (or any other analogous signal) can be transmitted from audio output device 602A to each of the other audio output devices 602B, 602C and 602D. Based on the received audio signals at each of the other audio output devices 602B, 602C and 602D, an RIR function is computed for each of directional microphones in the respective directional microphone arrangements 630B, 630C and 630D. Based on the orientation of the directional microphone associated with the peak RIR values, an angular position for each of the other audio output devices relative to the primary audio output device 602A is computed. As mentioned previously, the angular positions can be stored at the respective audio output devices. In some embodiments, when employing the channel assignment module 128 from Figure 1, configuring the primary audio output device 602A to execute channel assignments based on the calculated relative angular positions can also involve transmitting these computed angular positions to the primary audio output device 602A. In some embodiments, these positions can be stored within the device positions 140 at the primary audio output device 602A as depicted in Figure 1.
[0057] In some embodiments, however, an additional confirmation process is undertaken, whereby, chirp signals are transmitted during discrete time intervals from each of the other audio output devices 602B, 602C and 602D to the primary audio output device 602A. The primary audio output device 602A, using directional microphone arrangement 630A and the calibration process described above, determines relative angular positions associated with each of the other audio output devices. The primary audio output device 602A then corroborates the determined relative angular positions with the previously computed angular positions transmitted to the audio output device 602A from each of the other audio output devices. Channel assignments are performed once each of the computed relative angular positions are confirmed with previously computed angular positions by each of the other audio output devices. Thereafter, each of the audio output devices 602A, 602B, 602C and 602D can be grouped into the system with a channel assigned using the correct angular position.
[0058] Figure 7 is an illustration of a process for computing a RIR function by a directional microphone integrated in an audio output device of Figure 1, according to various embodiments. The RIR computation at each directional microphone can be performed by a position calibration engine (e.g., position calibration engine 114 shown in Figure 1) at the respective audio output device. As shown in Figure 7, a chirp signal 710 is emitted by a speaker 780 of an audio output device and is received at a directional microphone of an audio output device. Note that in some embodiments, audio object 124 referred to in Figure 1 is not limited to being a chirp signal but can be any other type of signal as well. Using the chirp signal 710, a reverse signal is determined. In some embodiments, the reverse signal can be a reverse chirp signal 712. Thereafter, an impulse function h (n) 720 is derived for the chirp signal 710 as received by a directional microphone. The position calibration engine at the respective audio output device then computes a Fast Fourier Transform (FFT) 716 using the impulse function h (n) 720 derived from the chirp signal 710. The position calibration engine also computes an FFT 714 from the reverse chirp signal 712. The FFT 714 of the reverse chirp signal is then multiplied with the FFT associated with the chirp signal 710. An Inverse Fast Fourier Transform 718 is computed using the results of the multiplication to generate the RIR signal 722.
[0059] As mentioned previously, in some embodiments, to obtain a more precise angular position of the target audio output device, an interpolation process might be employed. Figures 8A and 8B provide an illustration of the manner in which a relative angular position of an audio output device is determined by the device of Figure 1, according to various embodiments.
[0060] . Referring to the directional microphone arrangement of Figure 8A comprising eight directional microphones (e.g., directional microphones 820A, 820B …820H) , in certain circumstances, an audio signal 880 received at the directional microphone arrangement 800 may not align precisely with the directivity of any of the directional microphones in the directional microphone arrangement. In the example of Figure 8A, a highest peak RIR value can be registered by either directional microphone 820A or the directional microphone 820B. While the angular position of an audio output device (now shown in Figure 8A) from which the audio signal 880 emanates can be estimated by using the angular orientation of either of the directional microphones 820A or 820B, in various embodiments, a more precise angular position can be computed by performing an interpolation. In some embodiments, an interpolation is performed using RIR values and angular orientations associated with the directional microphone receiving the highest RIR peak value and one or more directional microphones neighboring the directional microphone receiving the highest RIR peak value.
[0061] Figure 8B illustrates the manner in which a more precise angular position for an audio output device can be computed by performing an interpolation operation. In a representative microphone setup, an audio signal might exhibit its maximum intensity across three adjacent directional microphones within a multi-microphone configuration. However, deducing the angular source position of the audio signal by solely relying on the angular orientation of the directional microphone linked to the highest RIR value might not yield a highly accurate angular determination. Therefore, to achieve precise angular calibration, it is necessary for the position calibration engine 114 to consider the angular orientations of the neighboring directional microphones as well. As shown in Figure 8B, in a directional microphone arrangement of, for example, six microphones, the highest intensity of an incoming sound signal can be experienced across three adjacent directional microphones, with the highest RIR value experienced at directional microphone N (not shown in Figure 8B) , which is associated with angular orientation αN 804. Simply using the angular orientation 804 to determine the angular position of the audio source will in certain scenarios not be precise enough. Accordingly, the angular orientations αN-1 802 associated with a directional microphone N-1 and αN+1 806 associated with a directional microphone N+1 will need to be used to determine a more precise angular position.
[0062] As shown in Figure 8B, the RIR peak values associated with each of the directional microphones N-1, N and N+1 (not shown in Figure 8B) can be plotted on a y-axis with the associated angular orientations for each of the directional microphones plotted on the x-axis. A straight line 840 can be drawn that passes through the highest RIR point associated with directional microphone N and the lowest RIR point associated with directional microphone N+1. Using the slope of line 840, another line 820 is also drawn through the peak RIR value registered at directional microphone N-1. The intersection of lines 820 and 840 can be used to determine the actual angular position αACTUAL 808 of the audio source. This technique yields a more refined angular position estimate for the audio output device position being calibrated.
[0063] Figure 9 illustrates a flow diagram of method steps for performing positional calibration for the plurality of audio output devices of Figure 1, according to various embodiments. The method steps of Figure 9 can be applied at least in part, by, for example, respective position calibration engines (substantially similar to the position calibration engine 114 of audio output device 110) in each of the audio output devices 110 of Figure 1. Although the method steps of Figure 9 are described with respect to the audio output device 110 of Figure 1 and the techniques illustrated in Figures 2-8, many systems configured to perform the method steps, in any order, can fall within the scope of the various embodiments.
[0064] As shown, a method 900 begins with a step 902 where a position calibration engine 114 of an audio output device 110A (which can be the controller audio output device in a multi-channel system) causes the audio output device to output an audio object 124, which, in various embodiments, can be a chirp signal. However, in other embodiments, the audio output device can output a tone of a given frequency, a frequency sweep over a portion of a human-audible frequency range and / or a human-inaudible frequency range, white or pink noise, or the like.
[0065] Step 904 is performed for each other audio output device of the plurality of audio output devices. For example, as discussed in connection with Figure 1, when audio output device 110A transmits the audio object 124, a calibration process can be performed at each of the other M-1 audio output devices (e.g., 110B, 110C …110M) to determine a relative position of the audio output device 110A.
[0066] At step 906, a position calibration engine in each of the other audio output devices determines a first relative angular position associated with the audio output device emitting the chirp signal (as will be discussed in further detail in connection with Figure 10) . As discussed in connection with Figure 1, during the calibration process, every directional microphone within each of the other audio output devices (e.g., the second audio output device 110B) captures and processes the audio object 124, enabling the computation of a RIR function corresponding to the audio object 124 received at each of the directional microphones. The RIR peak amplitude module within the second audio output device 110B, for example, identifies a directional microphone from the set of N directional microphones that registers a peak value of the computed RIR function. Based on an evaluation of the direction in which the directional microphone registering the peak RIR values is oriented, a position (or angle) of the audio output device 110B relative to the audio output device 110A can be determined. A similar calibration process is conducted for each of the other audio output devices that are part of the multi-channel system. As discussed in connection with Figures 8A and 8B, step 906 can also include determining a more precise angular position for an audio output device by performing an interpolation operation.
[0067] At step 908, a second relative angular position is determined by the audio output device (e.g., the primary audio output device 110A) for each of the other M-1 audio output devices and a comparison is performed between a first relative angular position determined by each of the M-1 audio output devices and a second relative angular position determined by the primary audio output device 110 for each of the M-1 audio output devices. As noted previously in connection with Figure 1, a confirmation process is undertaken by the position calibration engine 114 whereby the computed relative positions (or angles) of each of the other audio output devices (e.g., audio output device 110B, 110C …110M) relative to the audio output device 110A are confirmed by computing, at the first audio output device 110A, a relative position (angle) of the controlling audio output device 110 relative to each of the other audio output devices (e.g., audio output device 110B, 110C …110M) . Accordingly, a relative position for the controlling audio output device 110A relative to each of the other M-1 audio output devices 110 is determined using the directional microphones housed in the first audio output device 110A. During discrete time intervals, each of the M-1 audio output devices 110 transmits an audio output object, such as a chirp signal (or an analogous audio output object substantially similar to audio object 124) to the audio output device 110A. Employing the directional microphones within the first audio output device 110A, a calibration (as describe above) is employed to compute a relative position for the audio output device 110A relative to each one of the other M-1 audio output devices 110 at the audio output device 110A. Where the controlling audio output device 110A is configured as the principal entity (or controlling entity) in the system, the first relative positions computed and stored at each of the other M-1 audio output devices 110 are relayed to the first audio output device 110A over network 130. These relative positions are compared with respective second relative positions computed by and stored at the audio output device 110A.
[0068] Responsive to a determination that the second relative angular positions determined by the audio output device 110A for each of the other audio output devices correlates with the respective first relative angular positions, at step 910, a relative position for each other audio output device is stored at the primary audio output device 110A. With the relative positions of each of the other M-1 audio output devices stored at audio output device 110, a comprehensive establishment of the complete speaker arrangement geometry within the multi-channel system can be confirmed.
[0069] At step 912, a channel can be assigned to each of the audio output devices in the system of Figure 1 based on the determined relative positions for each of the plurality of audio output devices. As discussed in connection with Figure 1, after the relative positions for each of the other M-1 audio output devices is stored at audio output device 110A, each of the plurality of M audio output devices can be assigned a respective channel of the plurality of channels in the multi-channel system.
[0070] At step 914, audio is outputted at each audio output device of the plurality of audio output devices in accordance with the respective channel assignment. For example, a multi-channel audio is received (e.g., at a primary or controlling audio output device) and distributed to the other audio output devices based on the assignments. Each of the audio output devices then output the respective one or more channels assigned to the audio output device.
[0071] Figure 10 illustrates a flow diagram 1000 of method steps for determining an angular position for an audio output device represented in Figure 1, according to various embodiments. The method steps of Figure 10 can be applied, e.g., by the position calibration engine 114 of Figure 1. Although the method steps of Figure 10 are described with respect to the audio output device 110A of Figure 1, many systems configured to perform the method steps, in any order, can fall within the scope of the various embodiments. Note that the steps of flow diagram 1000 can be performed as part of step 906 of the method 900 illustrated in Figure 9.
[0072] At step 1002, a chirp signal is received from an audio output device. For example, referring to Figure 6, a chirp signal transmitted by audio output device 602A can be received at audio output device 602D.
[0073] At step 1004, a reverse chirp signal is determined from the chirp signal. For example, a reverse chirp signal 712 can be computed from the chirp signal 710 as shown in Figure 7.
[0074] Step 1006 is performed for each directional microphone of a plurality of directional microphones. The plurality of directional microphones can, for example, be configured in a given arrangement (e.g., in exemplary arrangements shown in Figures 4A, 4B and 4C) at audio output device 602D.
[0075] At step 1008, an FFT of the chirp signal is computed for each directional microphone. For example, as shown in Figure 7, an FFT 716 of the chirp signal is computed.
[0076] At step 1010, an FFT of the reverse chirp signal is also computed for each directional microphone. For example, as shown in Figure 7, an FFT 714 of the reverse chirp signal is computed.
[0077] At step 1012, the FFT of the chirp signal is multiplied with the FFT of the reverse chirp signal. For example, as shown in Figure 7, the FFT 714 of the reverse chirp signal is multiplied with the FFT 716 of the chirp signal.
[0078] At step 1014, an IFFT of the result of the multiplication is computed to determine a RIR associated with the received chirp signal. Referring again to Figure 7, for example, an RIR signal 722 is computed using the IFFT 718 obtained by multiplying the FFT 714 with the FFT 716.
[0079] At step 1016, a directional microphone from the plurality of directional microphones associated with a highest peak RIR value is determined. As mentioned previously, typically, the directional microphone with the highest peak RIR value will also receive the audio signal earlier than the other directional microphones and be associated with the earliest time of arrival for the audio signal amongst the plurality of directional microphones.
[0080] At step 1018, an angular position for the audio output device is determined. Referring to Figure 6, for example, based on an orientation of the directional microphone that register a peak RIR value from the directional microphone arrangement 630D at audio output device 602D, a relative angular position of audio output device 602D relative to audio output device 602A can be determined. As discussed in connection with Figures 8A and 8B, this angular position can also be obtained by performing an interpolation between the angular position associated with the directional microphone with the highest RIR value and one or more other directional microphones neighboring the directional microphone with the highest RIR value from the plurality of directional microphones in the microphone arrangement 630D.
[0081] In sum, techniques for performing positional calibration for a multi-channel speaker system includes causing a first audio output device of a plurality of audio output devices to output a chirp signal. The chirp signal is captured at a second audio output device of the plurality of audio devices using a plurality of directional microphones. For each of the plurality of directional microphones included in the second audio output device, a RIR signal associated with the chirp signal is computed. Further, a directional microphone from the plurality of directional microphones that registers a highest value of the computed RIR signal (and / or receives the chirp signal at the earliest time) is identified. A relative position of the first audio device is derived by evaluating a direction in which the identified directional microphone with the highest value of the computed RIR signal is oriented. In some embodiments, an angular relative position of the first audio device can also be obtained by performing an interpolation between the angular position associated with the directional microphone with the highest RIR value and one or more other directional microphones neighboring the directional microphone with the highest RIR value from the plurality of directional microphones in the microphone arrangement of the multi-channel speaker system. Thereafter, a channel in the multi-channel speaker system is assigned to the first audio device based on the relative position and audio is transmitted from the first audio device over the assigned channel.
[0082] At least one technical advantage of the disclosed techniques relative to the prior art is that, with the disclosed techniques, an automated calibration of speaker positions in a wireless multi-channel speaker system can be performed coupled with the precise assignment of each speaker to its fitting channel within the multi-channel arrangement. The disclosed calibration technique leverages the relative positions of the speakers, thereby offering a seamless and efficient method of assigning channels to each speaker in a multi-channel audio environment. In addition, the disclosed calibration techniques can determine the location of each audio output device relative to the other audio output devices automatically and accurately, as compared with other methods that involve performing in-situ measurements at the factory or external microphones. Unlike the limitations associated with external microphones or closely spaced omni-directional microphones, the disclosed techniques ensure robust and dependable calibration outcomes even within compact structures. Consequently, the disclosed techniques not only eliminate the risk of reversed sound images and compromised audio quality resulting from manual or inconsistent calibration approaches but also offers a streamlined and user-friendly setup process to perform auto-channel assignment in a multi-channel system. These technical advantages provide one or more technological improvements over prior art approaches.
[0083] 1. According to some embodiments, a computer-implemented method of performing positional calibration for a multi-channel audio system comprises capturing, using a plurality of directional microphones within a first audio output device, a test audio object emitted by a second audio output device; determining, for each of the plurality of directional microphones, a respective room impulse response (RIR) associated with the test audio object captured by the directional microphone; determining a position of the second audio output device relative to the first audio output device based on the RIRs; assigning a respective channel in the multi-channel audio system to each of the first audio output device and the second audio output device based on the position; and emitting audio from the first audio output device based on the respective channel assigned to the fist audio output device.
[0084] 2. The computer-implemented method according to clause 1, wherein the test audio object is a chirp signal.
[0085] 3. The computer-implemented method according to clauses 1-2, wherein determining the position of the second audio output device relative to the first audio output device comprises identifying a first directional microphone from the plurality of directional microphones that registers a highest peak value of the respective RIR; and computing the position of the second audio output device using an angular orientation of the first directional microphone.
[0086] 4. The computer-implemented method according to clauses 1-3, wherein determining the position of the second audio output device relative to the first audio output device further comprises evaluating a direction in which one or more other directional microphones from the plurality of directional microphones adjacent to the first directional microphone are oriented and performing an interpolation.
[0087] 5. The computer-implemented method according to clauses 1-4, wherein determining the position of the second audio output device relative to the first audio output device further comprises computing, for each of the plurality of directional microphones, a time of arrival of the test audio object; identifying a first directional microphone from the plurality of directional microphones that registers a highest peak value of the respective RIR and an earliest time of arrival of the test audio object; and computing a position of the second audio output device relative to the first audio output device by evaluating a direction in which the first directional microphone associated with the highest peak value of the respective RIR and the earliest time of arrival of the test audio object is oriented.
[0088] 6. The computer-implemented method according to clauses 1-5, wherein the plurality of directional microphones are arranged in a circular formation.
[0089] 7. The computer-implemented method according to clauses 1-6, wherein the plurality of directional microphones are arranged as pairs of back-to-back microphones.
[0090] 8. The computer-implemented method according to clauses 1-7, wherein determining the respective RIR for a first directional microphone of the plurality of directional microphones comprises computing a reverse signal from the test audio object; computing a Fast Fourier Transform (FFT) of the test audio object as captured by the first directional microphone; computing a Fast Fourier Transform (FFT) of the reverse signal; multiplying the FFT of the test audio object as captured by the first directional microphone; with the FFT of the reverse signal; and computing the respective RIR as an Inverse Fast Fourier Transform (IFFT) of a result of the multiplication.
[0091] 9. The computer-implemented method according to clauses 1-8, further comprising capturing, using the plurality of directional microphones within the first audio output device, another test audio object emitted by a third audio output device; determining, for each of the plurality of directional microphones, a RIR associated with the another test audio object captured by the directional microphone; determining a position of the third audio output device relative to the first audio output device based on the RIRs; assigning a respective channel in the multi-channel audio system to each of the first audio output device, the second audio output device and the third audio output device based on the determined positions; and emitting audio from the first audio output device, the second audio output device and the third audio output device based on the respective channels assigned to the first audio output device, the second audio output device, and the third audio output device.
[0092] 10. The computer-implemented method according to clauses 1-9, further comprising causing the first audio output device to emit another test audio object; receiving from the second audio output device a position of the first audio output device relative to the second audio output device computed at the second audio output device using the another test audio object; comparing the received position of the first audio output device relative to the second audio output device to the position of the second audio output device relative to the first audio output device; and responsive to a determination that the received position of the first audio output device relative to the second audio output device correlates with the position of the second audio output device relative to the first audio output device, assigning a respective channel in the multi-channel audio system to each of the first audio output device and the second audio output device.
[0093] 11. According to some embodiments, one or more non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to perform the steps of capturing, using a plurality of directional microphones within a first audio output device, a test audio object emitted by a second audio output device; determining, for each of the plurality of directional microphones, a respective room impulse response (RIR) associated with the test audio object captured by the directional microphone; determining a position of the second audio output device relative to the first audio output device based on the RIRs; assigning a respective channel in a multi-channel audio system to each of the first audio output device and the second audio output device based on the position; and emitting audio from the first audio output device based on the respective channel assigned to the fist audio output device.
[0094] 12. The one or more non-transitory computer readable media according to clause 11, wherein the test audio object is a chirp signal.
[0095] 13. The one or more non-transitory computer readable media according to clauses 11-12, wherein determining the position of the second audio output device relative to the first audio output device comprises identifying a first directional microphone from the plurality of directional microphones that registers a highest peak value of the respective RIR; and computing the position of the second audio output device using an angular orientation of the first directional microphone.
[0096] 14. The one or more non-transitory computer readable media according to clauses 11-13, wherein determining the position of the second audio output device relative to the first audio output device further comprises evaluating a direction in which one or more other directional microphones from the plurality of directional microphones adjacent to the first directional microphone are oriented and performing an interpolation.
[0097] 15. The one or more non-transitory computer readable media according to clauses 11-14, wherein determining the position of the second audio output device relative to the first audio output device further comprises computing, for each of the plurality of directional microphones, a time of arrival of the test audio object; identifying a first directional microphone from the plurality of directional microphones that registers a highest peak value of the respective RIR and an earliest time of arrival of the test audio object; and computing a position of the second audio output device relative to the first audio output device by evaluating a direction in which the first directional microphone associated with the highest peak value of the respective RIR and the earliest time of arrival of the test audio object is oriented.
[0098] 16. The one or more non-transitory computer readable media according to clauses 11-15, further comprising capturing, using the plurality of directional microphones within the first audio output device, another test audio object emitted by a third audio output device; determining, for each of the plurality of directional microphones, a respective RIR associated with the another test audio object captured by the directional microphone; determining a position of the third audio output device relative to the first audio output device based on the RIRs; assigning a respective channel in the multi-channel audio system to each of the first audio output device, the second audio output device and the third audio output device based on the determined positions; and emitting audio from the first audio output device, the second audio output device and the third audio output device based on the respective channels assigned to the first audio output device, the second audio output device, and the third audio output device.
[0099] 17. The one or more non-transitory computer readable media according to clauses 11-16, further comprising causing the first audio output device to emit another test audio object; receiving from the second audio output device a position of the first audio output device relative to the second audio output device computed at the second audio output device using the another test audio object; comparing the received position of the first audio output device relative to the second audio output device to the position of the second audio output device relative to the first audio output device; and responsive to a determination that the received position of the first audio output device relative to the second audio output device correlates with the position of the second audio output device relative to the first audio output device, assigning a respective channel in the multi-channel audio system to each of the first audio output device and the second audio output device.
[0100] 18. According to some embodiments, an audio output device comprises a memory storing instructions, and one or more processors that execute the instructions to perform steps comprising capturing, using a plurality of directional microphones within the audio output device, a test audio object emitted by another audio output device; determining, for each of the plurality of directional microphones, a respective room impulse response (RIR) associated with the test audio object captured by the directional microphone; determining a position of the another audio output device relative to the audio output device based on the RIRs; assigning a respective channel in a multi-channel audio system to each of the audio output device and the another audio output device based on the position; and emitting audio from the audio output device based on the respective channel assigned to the audio output device.
[0101] 19. The audio output device according to clause 18, wherein the plurality of directional microphones are arranged in a circular formation or as pairs of back-to-back microphones.
[0102] 20. The audio output device according to clauses 18-19, wherein the plurality of directional microphones are located above or below a speaker in the audio output device.
[0103] The descriptions of the various embodiments have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0104] Aspects of the present embodiments may be embodied as a system, method or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, etc. ) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a “module, ” a “system, ” or a “computer. ” In addition, any hardware and / or software technique, process, function, component, engine, module, or system described in the present disclosure may be implemented as a circuit or set of circuits. Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer readable medium (s) having computer readable program code embodied thereon.
[0105] Any combination of one or more computer readable medium (s) may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0106] Aspects of the present disclosure are described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine. The instructions, when executed via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / acts specified in the flowchart and / or block diagram block or blocks. Such processors may be, without limitation, general purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.
[0107] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function (s) . It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0108] While the preceding is directed to embodiments of the present disclosure, other and further embodiments of the disclosure may be devised without departing from the basic scope thereof, and the scope thereof is determined by the claims that follow.
Claims
1.A computer-implemented method of performing positional calibration for a multi-channel audio system, the method comprising:capturing, using a plurality of directional microphones within a first audio output device, a test audio object emitted by a second audio output device;determining, for each of the plurality of directional microphones, a respective room impulse response (RIR) associated with the test audio object captured by the directional microphone;determining a position of the second audio output device relative to the first audio output device based on the RIRs;assigning a respective channel in the multi-channel audio system to each of the first audio output device and the second audio output device based on the position; andemitting audio from the first audio output device based on the respective channel assigned to the first audio output device.2.The computer-implemented method of claim 1, wherein the test audio object is a chirp signal.3.The computer-implemented method of claim 1, wherein determining the position of the second audio output device relative to the first audio output device comprises:identifying a first directional microphone from the plurality of directional microphones that registers a highest peak value of the respective RIR; andcomputing the position of the second audio output device using an angular orientation of the first directional microphone.4.The computer-implemented method of claim 3, wherein determining the position of the second audio output device relative to the first audio output device further comprises evaluating a direction in which one or more other directional microphones from the plurality of directional microphones adjacent to the first directional microphone are oriented and performing an interpolation.5.The computer-implemented method of claim 1, wherein determining the position of the second audio output device relative to the first audio output device further comprises:computing, for each of the plurality of directional microphones, a time of arrival of the test audio object;identifying a first directional microphone from the plurality of directional microphones that registers a highest peak value of the respective RIR and an earliest time of arrival of the test audio object; andcomputing a position of the second audio output device relative to the first audio output device by evaluating a direction in which the first directional microphone associated with the highest peak value of the respective RIR and the earliest time of arrival of the test audio object is oriented.6.The computer-implemented method of claim 1, wherein the plurality of directional microphones are arranged in a circular formation.7.The computer-implemented method of claim 1, wherein the plurality of directional microphones are arranged as pairs of back-to-back microphones.8.The computer-implemented method of claim 1, wherein determining the respective RIR for a first directional microphone of the plurality of directional microphones comprises:determining a reverse signal from the test audio object;computing a Fast Fourier Transform (FFT) of the test audio object as captured by the first directional microphone;computing a Fast Fourier Transform (FFT) of the reverse signal;multiplying the FFT of the test audio object as captured by the first directional microphone; with the FFT of the reverse signal; andcomputing the respective RIR as an Inverse Fast Fourier Transform (IFFT) of a result of the multiplication.9.The computer-implemented method of claim 1, further comprising:capturing, using the plurality of directional microphones within the first audio output device, another test audio object emitted by a third audio output device;determining, for each of the plurality of directional microphones, a respective RIR associated with the another test audio object captured by the directional microphone;determining a position of the third audio output device relative to the first audio output device based on the RIRs;assigning a respective channel in the multi-channel audio system to each of the first audio output device, the second audio output device and the third audio output device based on the determined positions; andemitting audio from the first audio output device, the second audio output device and the third audio output device based on the respective channels assigned to the first audio output device, the second audio output device, and the third audio output device.10.The computer-implemented method of claim 1, further comprising:causing the first audio output device to emit another test audio object;receiving from the second audio output device a position of the first audio output device relative to the second audio output device computed at the second audio output device using the another test audio object;comparing the received position of the first audio output device relative to the second audio output device to the position of the second audio output device relative to the first audio output device; andresponsive to a determination that the received position of the first audio output device relative to the second audio output device correlates with the position of the second audio output device relative to the first audio output device, assigning a respective channel in the multi-channel audio system to each of the first audio output device and the second audio output device.11.A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to perform the steps of:capturing, using a plurality of directional microphones within a first audio output device, a test audio object emitted by a second audio output device;determining, for each of the plurality of directional microphones, a respective room impulse response (RIR) associated with the test audio object captured by the directional microphone;determining a position of the second audio output device relative to the first audio output device based on the RIRs;assigning a respective channel in a multi-channel audio system to each of the first audio output device and the second audio output device based on the position; andemitting audio from the first audio output device based on the respective channel assigned to the first audio output device.12.The non-transitory computer readable medium of claim 11, wherein the test audio object is a chirp signal.13.The non-transitory computer readable medium of claim 11, wherein determining the position of the second audio output device relative to the first audio output device comprises:identifying a first directional microphone from the plurality of directional microphones that registers a highest peak value of the respective RIR; andcomputing the position of the second audio output device using an angular orientation of the first directional microphone.14.The non-transitory computer readable medium of claim 13, wherein determining the position of the second audio output device relative to the first audio output device further comprises evaluating a direction in which one or more other directional microphones from the plurality of directional microphones adjacent to the first directional microphone are oriented and performing an interpolation.15.The non-transitory computer readable medium of claim 11, wherein determining the position of the second audio output device relative to the first audio output device further comprises:computing, for each of the plurality of directional microphones, a time of arrival of the test audio object;identifying a first directional microphone from the plurality of directional microphones that registers a highest peak value of the respective RIR and an earliest time of arrival of the test audio object; andcomputing a position of the second audio output device relative to the first audio output device by evaluating a direction in which the first directional microphone associated with the highest peak value of the respective RIR and the earliest time of arrival of the test audio object is oriented.16.The non-transitory computer readable medium of claim 11, further comprising:capturing, using the plurality of directional microphones within the first audio output device, another test audio object emitted by a third audio output device;determining, for each of the plurality of directional microphones, a respective RIR associated with the another test audio object captured by the directional microphone;determining a position of the third audio output device relative to the first audio output device based on the RIRs;assigning a respective channel in the multi-channel audio system to each of the first audio output device, the second audio output device and the third audio output device based on the determined positions; andemitting audio from the first audio output device, the second audio output device and the third audio output device based on the respective channels assigned to the first audio output device, the second audio output device, and the third audio output device.17.The non-transitory computer readable medium of claim 11, further comprising:causing the first audio output device to emit another test audio object;receiving from the second audio output device a position of the first audio output device relative to the second audio output device computed at the second audio output device using the another test audio object;comparing the received position of the first audio output device relative to the second audio output device to the position of the second audio output device relative to the first audio output device; andresponsive to a determination that the received position of the first audio output device relative to the second audio output device correlates with the position of the second audio output device relative to the first audio output device, assigning a respective channel in the multi-channel audio system to each of the first audio output device and the second audio output device.18.An audio output device comprising:a plurality of directional microphones;one or more speakers;a memory storing instructions, andone or more processors that execute the instructions to perform steps comprising:capturing, using the plurality of directional microphones, a test audio object emitted by another audio output device;determining, for each of the plurality of directional microphones, a respective room impulse response (RIR) associated with the test audio object captured by the directional microphone;determining a position of the another audio output device relative to the audio output device based on the RIRs;assigning a respective channel in a multi-channel audio system to each of the audio output device and the another audio output device based on the position; andemitting, using the one or more speakers, audio based on the respective channel assigned to the audio output device.19.The audio output device of claim 18, wherein the plurality of directional microphones are arranged in a circular formation or as pairs of back-to-back microphones.20.The audio output device of claim 18, wherein the plurality of directional microphones are located above or below a speaker in the audio output device.