Auto-location of audio devices
The method addresses the limitations of existing audio device localization by using DOA and TOA data to accurately locate audio devices in irregular environments, enhancing localization accuracy and flexibility.
Patent Information
- Application Number
- JP2023533781
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-07-22
- Filing Date
- 2021-12-02
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2041-12-02
AI Technical Summary
Existing audio device localization methods require synchronized microphones and speakers, assume standard layouts, and are not robust to measurement errors, especially in environments with irregularly distributed and heterogeneous audio devices.
The method utilizes direction of arrival (DOA) and time of arrival (TOA) data between pairs of audio devices to minimize a nonlinear optimization problem, allowing for the automatic localization of audio devices in irregular environments without requiring synchronization or known test stimuli.
Enables accurate localization of audio devices in complex environments by leveraging DOA and TOA data, overcoming synchronization constraints and measurement errors, and accommodating irregular device layouts.
Smart Images

Figure 0007797506000007 
Figure 0007797506000008 
Figure 0007797506000009
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to Spanish Patent Application No. P202031212, filed December 3, 2021, and No. P202130458, filed May 20, 2021, and U.S. Provisional Application No. 63 / 155369, filed March 2, 2021, No. 63 / 203403, filed July 21, 2021, and No. 63 / 224778, filed July 22, 2021, all of which are incorporated herein by reference in their entirety.
[0002] Technical Field The present disclosure relates to systems and methods for automatically locating audio devices. [Background technology]
[0003] Audio devices, including but not limited to smart audio devices, are being widely deployed and are becoming a common fixture in many homes. While existing systems and methods for locating audio devices provide benefits, improved systems and methods would be desirable.
[0004] Notation and Name Throughout this disclosure, including the claims, the terms "speaker," "loudspeaker," and "audio reproduction transducer" are used interchangeably to refer to any sound-emitting transducer (or collection of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., woofers and tweeters) that may be driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feeds may receive different processing in different circuit branches coupled to different transducers.
[0005] Throughout this disclosure, including the claims, the phrase performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing the operation directly on the signal or data, or to performing the operation on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing before performing the operation).
[0006] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.
[0007] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, video, or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipeline processing on audio or other audio data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.
[0008] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean a direct or indirect connection. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.
[0009] As used herein, a "smart device" is an electronic device generally configured to communicate with one or more other devices (or networks) via various wireless protocols, such as Bluetooth, Zigbee, near-field communications, Wi-Fi, light fidelity (Li-Fi), 3G, 4G, 5G, etc., and capable of operating interactively and / or autonomously to some degree. Some notable types of smart devices are smartphones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smart watches, smart bands, smart keychains, and smart audio devices. The term "smart device" can also refer to devices that exhibit some characteristics of ubiquitous computing, such as artificial intelligence.
[0010] As used herein, the phrase "smart audio device" refers to a smart device that is either a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of virtual assistant functionality). A single-purpose audio device is a device (e.g., a television (TV)) that includes or is coupled to at least one microphone (and optionally includes or is coupled to at least one speaker and / or at least one camera) and is designed largely or primarily to achieve a single purpose. For example, while televisions are typically capable of (and are considered capable of) playing audio from program material, modern televisions most often run some kind of operating system on which applications, including television viewing applications, run locally. In this sense, a single-purpose audio device with a speaker and microphone is often configured to run local applications and / or services that directly use the speaker and microphone. Several single-purpose audio devices can be configured to group together to achieve audio playback over a zone or user-configured area.
[0011] One common type of multipurpose audio device is an audio device that implements at least some aspects of virtual assistant functionality, although other aspects of the virtual assistant functionality may be implemented by one or more other devices, such as one or more servers with which the multipurpose audio device is configured to communicate. Such multipurpose audio devices are sometimes referred to herein as "virtual assistants." A virtual assistant is a device (e.g., a smart speaker or voice-assistant-integrated device) that includes or is coupled to at least one microphone (and, optionally, includes or is coupled to at least one speaker and / or at least one camera). In some examples, a virtual assistant may provide the ability to utilize multiple devices (different from the virtual assistant) for applications that are in some way cloud-enabled or otherwise not entirely implemented within or on the virtual assistant itself. In other words, at least some aspects of the virtual assistant functionality, e.g., speech recognition functionality, may be implemented (at least in part) by one or more servers or other devices with which the virtual assistant can communicate over a network such as the Internet. Virtual assistants may sometimes cooperate, for example, in a discrete, conditionally defined manner. For example, two or more virtual assistants can cooperate in the sense that one of them, e.g., the virtual assistant that is most confident that it heard the wake word, will respond to that word. Connected virtual assistants can, in some implementations, form a kind of constellation, which may be managed by one main application that may be (or may implement) a virtual assistant.
[0012] Here, "wake word" is used broadly to mean any sound (e.g., a word spoken by a human being, or some other sound) that the smart audio device is configured to wake up in response to detecting ("listening") for that sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "awakening" refers to the device entering a state in which it waits for (i.e., listens for) a voice command. In some instances, what may be referred to herein as a "wake word" may include multiple words, e.g., a phrase.
[0013] Here, the term "wake word detector" refers to a device (or software including instructions for configuring a device) configured to continuously look for alignment between real-time audio (e.g., speech) features and a trained model. Typically, a wake word event is triggered whenever the wake word detector determines that the probability that the wake word has been detected exceeds a predetermined threshold. For example, the threshold may be a predetermined threshold adjusted to provide a reasonable compromise between false accept and false reject rates. Following a wake word event, the device may enter a state (which may be referred to as an "awake" or "attentive" state) in which it listens for commands and passes received commands to a larger, more computationally intensive recognizer.
[0014] As used herein, the terms "program stream" and "content stream" refer to a collection of one or more audio signals, and possibly video signals, that are intended to be listened to at least in part together. Examples include selections of music, movie soundtracks, movies, television programs, audio portions of television programs, podcasts, live voice calls, synthesized voice responses from smart assistants, etc. In some cases, a content stream may contain multiple versions of at least a portion of an audio signal, e.g., the same dialogue in multiple languages. In such cases, only one version of the audio data or portion thereof (e.g., a version corresponding to a single language) is intended to be played at a time. Summary of the Invention [Means for solving the problem]
[0015] At least some aspects of the present disclosure may be implemented via methods. Some such methods may involve localizing an audio device. For example, some methods may involve localizing an audio device in an audio environment. Some such methods may involve obtaining, by a control system, direction of arrival (DOA) data corresponding to sound emitted by at least a first smart audio device in the audio environment. In some implementations, the first smart audio device may include a first audio transmitter and a first audio receiver. In some examples, the DOA data may correspond to sound received by at least a second smart audio device in the audio environment. In some instances, the second smart audio device may include a second audio transmitter and a second audio receiver. In some examples, the DOA data may also correspond to sound emitted by at least a second smart audio device and received by at least the first smart audio device.
[0016] Some such methods may involve receiving, by a control system, configuration parameters. In some examples, the configuration parameters may correspond to an audio environment and / or one or more audio devices in the audio environment. Some such methods may involve, by the control system, minimizing a cost function based at least in part on the DOA data and the configuration parameters to estimate a position and / or orientation of at least a first smart audio device and a second smart audio device.
[0017] According to some examples, the DOA data may also correspond to sounds received by one or more passive audio receivers of the audio environment. In some examples, each of the one or more passive audio receivers may include a microphone array, but in some cases may lack an audio emitter. In some such examples, minimizing a cost function may also provide an estimated position and orientation of each of the one or more passive audio receivers.
[0018] In some examples, the DOA data may also correspond to sounds emitted by one or more audio emitters of the audio environment. In some cases, each of the one or more audio emitters may include at least one sound-emitting transducer, but in some cases may lack a microphone array. In some such examples, minimizing a cost function may also provide an estimated position of each of the one or more audio emitters.
[0019] In some implementations, the DOA data may also correspond to sounds emitted by a third through Nth smart audio device in the audio environment, where N corresponds to the total number of smart audio devices in the audio environment. In some examples, the DOA data may also correspond to sounds received by each of the first through Nth smart audio devices from all other smart audio devices in the audio environment. In some such examples, minimizing the cost function may involve estimating the positions and / or orientations of the third through Nth smart audio devices.
[0020] According to some examples, the configuration parameters may include the number of audio devices in the audio environment, one or more dimensions of the audio environment, and / or one or more constraints on audio device position and / or orientation. In some instances, the configuration parameters may include disambiguation data for rotation, translation, and / or scaling.
[0021] Some methods may involve receiving, by a control system, a seed layout for the cost function. The seed layout, in some examples, may specify a correct number of audio transmitters and receivers in the audio environment and optional positions and orientations for each of the audio transmitters and receivers in the audio environment.
[0022] Some methods may involve receiving, by a control system, a weighting factor associated with one or more elements of DOA data, which may indicate, for example, availability and / or reliability of the one or more elements of DOA data.
[0023] Some methods may involve the control system obtaining one or more elements of DOA data using beamforming methods, steered power response methods, time difference of arrival methods, structured signal methods, or combinations thereof.
[0024] Some methods may involve receiving, by a control system, time of arrival (TOA) data corresponding to sound emitted by at least one audio device in the audio environment and received by at least one other audio device in the audio environment. In some such examples, a cost function may be based at least in part on the TOA data. Some such methods may involve estimating at least one playback latency and / or estimating at least one recording latency. In some examples, the cost function may operate with respect to rescaled positions, rescaled latencies, and / or rescaled arrival times.
[0025] According to some examples, the cost function may include a first term that depends only on the DOA data. In some such examples, the cost function may include a second term that depends only on the TOA data. In some such examples, the first term may include a first weighting factor and the second term may include a second weighting factor. In some instances, one or more TOA elements of the second term may have a TOA element weighting factor that indicates the availability and / or reliability of each of the one or more TOA elements.
[0026] In some examples, the configuration parameters may include playback latency data, recording latency data, data for disambiguating latency symmetry, rotation disambiguation data, translation disambiguation data, scaling disambiguation data, and / or one or more combinations thereof.
[0027] Some other aspects of the present disclosure may be implemented via methods. Some such methods may involve localizing a device. For example, some methods may involve localizing a device in an audio environment. Some such methods may involve obtaining, by a control system, direction of arrival (DOA) data corresponding to transmissions of at least a first transceiver of a first device in the environment. The first transceiver, in some examples, may include a first transmitter and a first receiver. In some cases, the DOA data may correspond to transmissions received by at least a second transceiver of a second device in the environment. In some examples, the second transceiver may include a second transmitter and a second receiver. In some cases, the DOA data may correspond to transmissions from at least a second transceiver received by the at least first transceiver.
[0028] In some examples, the first device and the second device may be audio devices, and the environment may be an audio environment. According to some such examples, the first transmitter and the second transmitter may be audio transmitters. In some such examples, the first receiver and the second receiver may be audio receivers. In some implementations, the first transceiver and the second transceiver may be configured to transmit and receive electromagnetic waves.
[0029] Some such methods may involve receiving, by a control system, configuration parameters. In some cases, the configuration parameters may correspond to an environment and / or to one or more devices in the environment. Some such methods may involve minimizing, by the control system, a cost function based at least in part on the DOA data and the configuration parameters to estimate a position and / or orientation of at least a first device and a second device.
[0030] In some examples, the DOA data may also correspond to transmissions received by one or more passive receivers in the environment, each of which may, for example, include a receiver array but lack a transmitter. In some such examples, minimizing a cost function may also provide an estimated position and / or orientation of each of the one or more passive receivers.
[0031] According to some examples, the DOA data may also correspond to transmissions from one or more transmitters in the environment. In some cases, each of the one or more transmitters may lack a receiver array. In some such examples, minimizing the cost function may also provide an estimated location of each of the one or more transmitters.
[0032] In some examples, the DOA data may also correspond to transmissions emitted by third through Nth transceivers of third through Nth devices in the environment, where N corresponds to the total number of transceivers in the environment. In some such examples, the DOA data may also correspond to transmissions received by each of the first through Nth transceivers from all other transceivers in the environment. In some such examples, minimizing the cost function may involve estimating the positions and / or orientations of the third through Nth transceivers.
[0033] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. Thus, some innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media storing software.
[0034] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, the apparatus may include an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof. However, in some implementations, the apparatus may be another type of device, such as a mobile device, laptop, server, etc.
[0035] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Please note that the relative dimensions of the following figures may not be drawn to scale. [Brief explanation of the drawings]
[0036] [Figure 1] Shows an example of the geometric relationships between four audio devices in an environment. [Figure 2] 2 illustrates an audio emitter located within the audio environment of FIG. 1; [Figure 3] 2 illustrates an audio receiver located within the audio environment of FIG. 1. [Figure 4] 11 is a flow diagram outlining an example of a method that may be performed by a control system of an apparatus such as that shown in FIG. 10. [Figure 5] FIG. 10 is a flow diagram outlining another example of a method for automatically estimating a device's position and orientation based on DOA data. [Figure 6]FIG. 1 is a flow diagram outlining an example of a method for automatically estimating a device's position and orientation based on DOA and TOA data. [Figure 7] FIG. 10 is a flow diagram outlining another example method for automatically estimating a device's position and orientation based on DOA and TOA data. [Figure 8A] 1 shows an example of an audio environment. [Figure 8B] 10 illustrates an additional example of determining listener angular orientation data. [Figure 8C] 10 illustrates an additional example of determining listener angular orientation data. [Figure 8D] FIG. 8C illustrates an example of determining the appropriate rotation for audio device coordinates according to the method described above. [Figure 9A] FIG. 1 is a flow diagram outlining an example of a localization method. [Figure 9B] FIG. 10 is a flow diagram outlining another example of a localization method. [Figure 10] FIG. 1 is a block diagram illustrating example components of an apparatus capable of implementing various aspects of the disclosure. [Figure 11] This example shows an example of a floor plan for an audio environment, which is a living space.
[0037] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION
[0038] In addition to existing audio devices, including televisions and soundbars, the emergence of smart speakers incorporating multiple driver units and microphone arrays, as well as connected devices with new microphone and speaker capabilities, such as light bulbs and microwave ovens, creates the problem that dozens of microphones and speakers need to be located relative to each other to achieve coordination. Audio devices cannot be assumed to be in a standard layout (such as a discrete Dolby 5.1 loudspeaker layout). In some cases, audio devices within an environment may be randomly located, or at least distributed within the environment in an irregular and / or asymmetric manner.
[0039] Furthermore, audio devices cannot be assumed to be homogeneous or synchronous. As used herein, audio devices may be referred to as "synchronous" or "synchronized" if sound is detected or emitted by those audio devices according to the same or synchronized sample clocks. For example, a first synchronized microphone of a first audio device in the environment may digitally sample audio data according to a first sample clock, and a second microphone of a second synchronized audio device in the environment may digitally sample audio data according to the first sample clock. Alternatively or additionally, a first synchronized speaker of a first audio device in the environment may emit sound according to a speaker setup clock, and a second synchronized speaker of a second audio device in the environment may emit sound according to said speaker setup clock.
[0040] Some previously disclosed methods for automatic speaker localization require synchronized microphones and / or speakers. For example, some existing tools for device localization rely on sample synchrony between all microphones in the system and require known test stimuli and passing full-bandwidth audio data between sensors.
[0041] The assignee has produced several speaker localization techniques for cinemas and homes that are excellent solutions for the use cases for which they were designed. Some such methods are based on time-of-flight derived from the impulse response between the sound source and a microphone approximately co-located with each loudspeaker. System latency in the record and playback chain can also be estimated, but requires sample synchrony between clocks and the need for known test stimuli to estimate the impulse response.
[0042] Recent examples of sound source localization in this context relax the constraints by requiring intra-device microphone synchronization but not inter-device synchronization. Additionally, some such methods forgo the need to pass audio between sensors through low-bandwidth message passing, such as through detection of the time of arrival (TOA, also known as "time of flight") of direct (unreflected) sound or detection of the dominant direction of arrival (DOA) of direct sound. Each approach has several potential advantages and disadvantages. For example, some previously deployed TOA methods can determine device geometry excluding unknown translations, rotations, and reflections around one of three axes. With only one microphone per device, the rotation of individual devices is also unknown. Some previously deployed DOA methods can determine device geometry excluding unknown translations, rotations, and scaling. While some such methods can produce satisfactory results under ideal conditions, their robustness to measurement errors has not been demonstrated.
[0043] Some embodiments disclosed herein allow for localization of a collection of smart audio devices based on 1) the DOAs between each pair of audio devices in an audio environment and 2) minimization of a nonlinear optimization problem designed for input of data type 1). Other embodiments disclosed herein allow for localization of a collection of smart audio devices based on 1) the DOAs between each pair of audio devices in a system, 2) the TOAs between each pair of devices, and 3) minimization of a nonlinear optimization problem designed for input of data types 1) and 2).
[0044] FIG. 1 illustrates an example of the geometric relationships between four audio devices in an environment. In this example, audio environment 100 is a room that includes television 101 and audio devices 105a, 105b, 105c, and 105d. According to this example, audio devices 105a-105d are located at positions 1-4 in audio environment 100, respectively. As with other examples disclosed herein, the types, numbers, locations, and orientations of elements shown in FIG. 1 are intended to be exemplary only. Other implementations may have different types, numbers, and arrangements of elements, such as more or fewer audio devices, audio devices in different locations, audio devices with different capabilities, etc.
[0045] In this implementation, each of audio devices 105a-105d is a smart speaker that includes a microphone system and a speaker system including at least one speaker. In some implementations, each microphone system includes an array of at least three microphones. According to some implementations, television 101 may include a speaker system and / or a microphone system. In some such implementations, an auto-localization method may be used to automatically localize television 101, or a portion of television 101 (e.g., television speakers, television transceiver, etc.). This is described below, for example, with reference to audio devices 105a-105d.
[0046] Some of the embodiments described in this disclosure allow for automatic localization of a set of audio devices, such as audio devices 105a-105d shown in FIG. 1, based on the direction of arrival (DOA) between each pair of audio devices, the time of arrival (TOA) of the audio signals between each pair of devices, or both the DOA and TOA of the audio signals between each pair of devices. In some cases, as in the example shown in FIG. 1, each of the audio devices is enabled with at least one driver unit and a microphone array, and the microphone array is capable of providing the direction of arrival of incoming sound. According to this example, double arrows 110a-b represent sound transmitted by audio device 105a and received by audio device 105b, as well as sound transmitted by audio device 105b and received by audio device 105a. Similarly, double-headed arrows 110ac, 110ad, 110bc, 110bd, and 110cd represent sounds transmitted and received by audio device 105a and audio device 105c, sounds transmitted and received by audio device 105a and audio device 105d, sounds transmitted and received by audio device 105b and audio device 105c, sounds transmitted and received by audio device 105b and audio device 105d, and sounds transmitted and received by audio device 105c and audio device 105d, respectively.
[0047] In this example, each of audio devices 105a-105d has an orientation represented by arrows 115a-115d, which may be defined in various ways. For example, the orientation of an audio device with a single loudspeaker may correspond to the direction in which the single loudspeaker is pointing. In some examples, the orientation of an audio device with multiple loudspeakers pointing in different directions may be indicated by the direction in which one of the loudspeakers is pointing. In other examples, the orientation of an audio device with multiple loudspeakers pointing in different directions may be indicated by the direction of a vector corresponding to the sum of the audio output in the different directions in which each of the multiple loudspeakers is pointing. In the example shown in FIG. 1, the orientation of arrows 115a-115d is defined with reference to a Cartesian coordinate system. In other examples, the orientation of arrows 115a-115d may be defined with reference to another type of coordinate system, such as a spherical or cylindrical coordinate system.
[0048] In this example, the television 101 includes an electromagnetic interface 103 configured to receive electromagnetic waves. In some examples, the electromagnetic interface 103 may be configured to transmit and receive electromagnetic waves. According to some implementations, at least two of the audio devices 105a-105d may include an antenna system configured as a transceiver. The antenna system may be configured to transmit and receive electromagnetic waves. In some examples, the antenna system includes an antenna array having at least three antennas. Some of the embodiments described in this disclosure enable automatic localization of a set of devices, such as the audio devices 105a-105d and / or the television 101 shown in FIG. 1, based at least in part on the DOAs of the electromagnetic waves transmitted between the devices. Thus, the double-headed arrows 110ab, 110ac, 110ad, 110bc, 110bd, and 110cd may also represent electromagnetic waves transmitted between the audio devices 105a, 105d.
[0049] According to some examples, the antenna system of a device (such as an audio device) may be co-located with a loudspeaker of the device, e.g., adjacent to the loudspeaker. In some such examples, the antenna system orientation may correspond to the loudspeaker orientation. Alternatively or additionally, the antenna system of the device may have a known or predetermined orientation relative to one or more loudspeakers of the device.
[0050] In this example, audio devices 105a-105d are configured to wirelessly communicate with each other and with other devices. In some examples, audio devices 105a-105d may include a network interface configured for communication between audio devices 105a-105d and other devices over the Internet. In some implementations, the auto-localization process disclosed herein may be performed by a control system of one of audio devices 105a-105d. In other examples, the auto-localization process may be performed by another device in audio environment 100, such as what may be referred to as a smart home hub, configured for wireless communication with audio devices 105a-105d. In other examples, the auto-localization process may be performed at least in part by a device external to audio environment 100, such as a server, based on information received from one or more of audio devices 105a-105d and / or the smart home hub.
[0051] FIG. 2 illustrates audio emitters located within the audio environment of FIG. 1. Some implementations provide automatic localization of one or more audio emitters, such as person 205 of FIG. 2. In this example, person 205 is at position 5. Here, sound emitted by person 205 and received by audio device 105a is represented by single-sided arrow 210a. Similarly, sound emitted by person 205 and received by audio devices 105b, 105c, and 105d are represented by single-sided arrows 210b, 210c, and 210d. Audio emitters may be localized based on the DOA of the audio emitter sounds as captured by audio devices 105a-105d and / or television 101, based on the difference in TOA of the audio emitter sounds as measured by audio devices 105a-105d and / or television 101, or based on both the difference in DOA and TOA.
[0052] Alternatively or additionally, some implementations may provide for automatic location of one or more electromagnetic wave emitters. Some of the embodiments described in this disclosure allow for automatic location of one or more electromagnetic wave emitters based at least in part on the DOA of the electromagnetic waves transmitted by the one or more electromagnetic wave emitters. If the electromagnetic wave emitters were at position 5, the electromagnetic waves emitted by the electromagnetic wave emitters and received by audio devices 105a, 105b, 105c, and 105d may also be represented by single arrows 210a, 210b, 210c, and 210d.
[0053] FIG. 3 illustrates audio receivers located within the audio environment of FIG. 1. In this example, the microphone of smartphone 305 is enabled, but the speaker of smartphone 305 is not currently emitting sound. Some embodiments provide automatic localization of one or more passive audio receivers, such as smartphone 305 of FIG. 3, when smartphone 305 is not emitting sound. Here, sound emitted by audio device 105a and received by smartphone 305 is represented by single arrow 310a. Similarly, sound emitted by audio devices 105b, 105c, and 105d and received by smartphone 305 are represented by single arrows 310b, 310c, and 310d.
[0054] If the audio receiver includes a microphone array and is configured to determine the DOA of the received sound, the audio receiver may be localized based at least in part on the DOA of the sound emitted by the audio devices 105a-105d and captured by the audio receiver. In some examples, the audio receiver may be localized based at least in part on the difference in the TOA of the smart audio device captured by the audio receiver, regardless of whether the audio receiver includes a microphone array. Still other embodiments may allow automatic localization of a set of a smart audio device, one or more audio emitters, and one or more receivers based on DOA alone, or DOA and TOA, by combining the methods described above.
[0055] Arrival direction localization Figure 4 is a flow diagram outlining one example of a method that may be performed by a control system of an apparatus such as that shown in Figure 10. The blocks of method 400, as well as other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described.
[0056] Method 400 is an example of an audio device localization process. In this example, method 400 involves determining the position and orientation of two or more smart audio devices, each smart audio device including a loudspeaker system and an array of microphones. According to this example, method 400 involves determining the position and orientation of a smart audio device based at least in part on audio emitted by all smart audio devices and captured by all other smart audio devices according to a DOA estimate. In this example, initial blocks of method 400 rely on the control system of each smart audio device to extract the DOA from input audio captured by that smart audio device's microphone array, for example, by using the time difference of arrival between individual microphone capsules of the microphone array.
[0057] In this example, block 405 involves acquiring audio emitted by all smart audio devices in the audio environment and captured by all other smart audio devices in the audio environment. In some such examples, block 405 may involve causing each smart audio device to emit a sound, which in some instances may be a sound having a predetermined duration, frequency content, etc. This predetermined type of sound may be referred to herein as a structured source signal. In some implementations, the smart audio devices may be or include audio devices 105a-105d of FIG. 1.
[0058] In some such examples, block 405 may involve a sequential process of causing a single smart audio device to emit sound while other smart audio devices "listen" for sound. 1, block 405 may include: (a) causing audio device 105a to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 105b-105d; then (b) causing audio device 105b to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 105a, 105c, and 105d; then (c) causing audio device 105c to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 105a, 105b, and 105d; and then (d) causing audio device 105d to emit a sound and receiving microphone data corresponding to the emitted sound from the microphone arrays of audio devices 105a, 105b, and 105c. These emitted sounds may or may not be the same, depending on the particular implementation.
[0059] In another example, block 405 may be involved in a simultaneous process of causing all smart audio devices to emit sound while other smart audio devices "listen" for sound. For example, block 405 may include the following steps: (1) causing audio device 105a to emit a first sound and receiving microphone data corresponding to the emitted first sound from the microphone arrays of audio devices 105b-105d; (2) causing audio device 105b to emit a second sound different from the first sound and receiving microphone data corresponding to the emitted second sound from the microphone arrays of audio devices 105a, 105c, and 105d; (3) causing audio device 105b to emit a second sound different from the first sound and receiving microphone data corresponding to the emitted second sound from the microphone arrays of audio devices 105a, 105c, and 105d; (4) causing audio device 105c to emit a third sound, different from the first sound and the second sound, and receiving microphone data corresponding to the emitted third sound from the microphone arrays of audio devices 105a, 105b, 105d; and (5) causing audio device 105d to emit a fourth sound, different from the first sound, the second sound, and the third sound, and receiving microphone data corresponding to the emitted fourth sound from the microphone arrays of audio devices 105a, 105b, 105c.
[0060] In this example, block 410 involves a process of pre-processing the audio signal acquired via the microphone. Block 410 may involve, for example, applying one or more filters, noise or echo suppression processes, etc. Some additional pre-processing examples are described below.
[0061] According to this example, block 415 involves determining DOA candidates from the pre-processed audio signal resulting from block 410. For example, if block 405 involved emitting and receiving a structured source signal, block 415 may involve one or more deconvolution methods to yield impulse responses and / or "pseudoranges" from which the time differences of arrival of dominant peaks can be used in conjunction with the known microphone array geometry of the smart audio device to estimate DOA candidates.
[0062] However, not all implementations of method 400 involve obtaining microphone signals based on predetermined sound emissions. Thus, some examples of block 415 include “blind” methods, such as steered response power, receiver-side beamforming, or other similar methods, applied to any audio signal from which one or more DOAs may be extracted by peak picking. Some examples are described below. While DOA data may be determined via blind methods or using structured source signals, it will be understood that in most cases, TOA data may only be determined using structured source signals. Furthermore, more accurate DOA information may generally be obtained using structured source signals.
[0063] According to this example, block 420 involves selecting one DOA that corresponds to the sound emitted by each of the other smart audio devices. In many cases, the microphone array may detect both directly arriving sound and reflected sound transmitted by the same audio device. Block 420 may involve selecting the audio signal that most likely corresponds to the directly transmitted sound. Several additional examples of determining candidate DOA and selecting a DOA from two or more candidate DOAs are described below.
[0064] In this example, block 425 involves receiving the DOA information resulting from each smart audio device's implementation of block 420 (in other words, receiving a set of DOAs corresponding to sounds transmitted from all smart audio devices in the audio environment to all other smart audio devices) and performing a localization method based on the DOA information (e.g., implementing a localization algorithm via a control system). In some disclosed implementations, block 425 involves minimizing a cost function, possibly subject to several constraints and / or weights, as described below, e.g., with reference to FIG. 5. In some such examples, the cost function receives as input data DOA values from all smart audio devices to all other smart audio devices and returns as output an estimated position and an estimated orientation of each smart audio device. In the example shown in FIG. 4, block 430 represents the estimated smart audio device positions and estimated smart audio device orientations generated in block 425.
[0065] 5 is a flow diagram outlining another example of a method for automatically estimating the position and orientation of a device based on DOA data. Method 500 may be performed, for example, by implementing a localization algorithm via a control system of an apparatus such as that shown in FIG. 10. The blocks of method 500, as well as other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described.
[0066] According to this example, DOA data is acquired in block 505. According to some implementations, block 505 may involve acquiring acoustic DOA data, for example, as described above with reference to blocks 405-420 of Figure 4. Alternatively or additionally, block 505 may involve acquiring DOA data corresponding to electromagnetic waves transmitted and received by each of a plurality of devices in the environment.
[0067] In this example, the localization algorithm receives as input the DOA data obtained in block 505 from every smart device to every other smart device in the audio environment, along with any configuration parameters 510 specified for the audio environment. In some examples, optional constraints 525 may be applied to the DOA data. The configuration parameters 510, minimization weights 515, optional constraints 525, and seed layout 530 may be retrieved from memory by, for example, a control system running software to implement cost function 520 and nonlinear search algorithm 535. The configuration parameters 510 may include, for example, data corresponding to maximum room dimensions, loudspeaker layout constraints, external inputs for setting global translation (e.g., two parameters), global rotation (one parameter), and global scale (one parameter), etc.
[0068] According to this example, configuration parameters 510 are provided to a cost function 520 and a nonlinear search algorithm 535. In some examples, configuration parameters 510 are provided to optional constraints 525. In this example, cost function 520 takes into account the difference between the measured DOA and the DOA estimated by the optimizer's localization solution.
[0069] In some embodiments, optional constraints 525 impose restrictions on the possible audio device positions and / or orientations, such as imposing a condition that audio devices be a certain minimum distance from each other. Alternatively or additionally, optional constraints 525 may impose restrictions on dummy minimization variables that are introduced for convenience, for example, as described below.
[0070] In this example, the non-linear search algorithm 535 is also provided with minimization weights 515. Some examples are described below.
[0071] According to some implementations, the nonlinear search algorithm 535 is an algorithm that can find a local solution to a continuous optimization problem of the form:
number
[0072] The nonlinear search algorithm 535 may vary according to the particular implementation. Examples of nonlinear search algorithms 535 include gradient descent, Broyden-Fletchers-Goldfarb-Shanno (BFGS), Interior Point Optimization (IPOPT), etc. Some nonlinear search algorithms only require values of the cost function and constraints, while some other methods may require first derivatives (gradients, Jacobians) of the cost function and constraints, and some other methods may require second derivatives (Hessians) of the same functions. When derivatives are required, they can be provided explicitly, or they can be calculated automatically using automatic or numerical differentiation techniques.
[0073] Some nonlinear search algorithms require seed point information to begin the minimization, as suggested by seed layout 530 provided to nonlinear search algorithm 535 in FIG. 5. In some examples, the seed point information may be provided as a layout of the same number of smart audio devices with corresponding positions and orientations (in other words, the same number as the actual number of smart audio devices for which DOA data is acquired). The positions and orientations may be arbitrary and need not be actual or approximate positions and orientations of the smart audio devices. In some examples, the seed point information may indicate smart audio device positions along an axis of the audio environment or another arbitrary line, a circle, rectangle, or other geometric shape within the audio environment, or the like. In some examples, the seed point information may indicate an arbitrary smart audio device orientation, which may be a predetermined smart audio device or a random starting audio device orientation.
[0074] In some embodiments, the cost function 520 can be formulated in terms of complex plane variables as follows:
number
[0075] According to this example, the result of the minimization is device location data 540 x 200, which indicates the 2D location of the smart device. k (representing two real unknowns per device) and device orientation data 545 z indicating the orientation vector of the smart device. k (representing two additional real variables per device). From the orientation vector, the smart device orientation angle α k Only ������ is significant for the problem (one real unknown per device). Therefore, in this example, there are three significant unknowns per smart device.
[0076] In some examples, the result evaluation block 550 involves calculating a residual of a cost function at the resulting position and orientation. A relatively lower residual indicates a relatively more accurate device localization value. According to some implementations, the result evaluation block 550 may involve a feedback process. For example, some such examples may implement a feedback process that involves comparing the residual of a given DOA candidate combination with another DOA candidate combination. This is described, for example, in the discussion of DOA robustness metrics below.
[0077] As noted above, in some implementations, block 505 may involve determining DOA candidates and acquiring acoustic DOA data as described above with reference to blocks 405-420 of Figure 4, which involve selecting DOA candidates. Thus, Figure 5 includes a dashed line from result evaluation block 550 to block 505 to represent one flow of an optional feedback process. Additionally, Figure 4 includes a dashed line from block 430 (which may involve result evaluation in some examples) to DOA candidate selection block 420 to represent another flow of an optional feedback process.
[0078] In some embodiments, the nonlinear search algorithm 535 may not accept complex-valued variables. In such cases, all complex-valued variables may be replaced by a pair of real variables.
[0079] In some implementations, there may be additional prior information regarding the availability or reliability of each DOA measurement. In some such examples, the loudspeaker may be localized using only a subset of all possible DOA elements. Missing DOA elements may be masked, for example, with a corresponding zero weight in the cost function. In some such examples, the weights w nm may be either 0 or 1, e.g., 0 for measurements that are missing or considered not sufficiently reliable, and 1 for measurements that are reliable. In some other embodiments, the weight w nmmay have continuous values between 0 and 1 as a function of the reliability of the DOA measurements. In embodiments where no a priori information is available, the weights w nm may simply be set to 1.
[0080] In some implementations, the condition |z k |=1 (one condition per smart audio device) may be added as a constraint to ensure normalization of the vectors indicating the orientation of smart audio devices. In other examples, these additional constraints may not be required, and the vectors indicating the orientation of smart audio devices may be left unnormalized. Other implementations may add conditions regarding the proximity of smart audio devices as constraints. This may be the case, for example, when |x n -x m | ≥ D, where D is the minimum distance between smart audio devices.
[0081] Minimizing the above cost function does not completely determine the absolute positions and orientations of the smart audio devices. According to this example, the cost function remains invariant under global rotation (one independent parameter), global translation (two independent parameters), and global rescaling (one independent parameter) that simultaneously affect all smart device positions and orientations. The global rotation, translation, and rescaling cannot be determined from minimizing the cost function. Different layouts related by symmetry transformations are completely indistinguishable in this framework and are said to belong to the same equivalence class. Therefore, the configuration parameters should provide a criterion that allows uniquely defining a smart audio device layout that represents the entire equivalence class. In some embodiments, it may be advantageous to select a criterion such that this smart audio device layout defines a frame of reference that is close to the frame of reference of a listener near the reference listening position. Examples of such criteria are provided below. In some other examples, the criterion may be purely mathematical and decoupled from a realistic frame of reference.
[0082] Symmetry disambiguation criteria may include a reference position that fixes global translational symmetry (e.g., smart audio device 1 should be at the origin of the coordinate system); a reference orientation that fixes two-dimensional rotational symmetry (e.g., smart device 1 should be oriented toward an area of the audio environment designated as the front, such as where television 101 is located in Figures 1-3); and a reference distance that fixes global scaling symmetry (e.g., smart device 2 should be a unit distance from smart device 1). In total, there are four parameters in this example that cannot be determined from the minimization problem and must be provided as external inputs. Thus, in this example, there are 3N-4 unknowns that can be determined from the minimization problem.
[0083] As described above, in some examples, in addition to a set of smart audio devices, there may be one or more passive audio receivers with microphone arrays and / or one or more audio emitters. In such cases, the localization process may use techniques to determine the positions and orientations of the smart audio devices, the positions of the emitters, and the positions and orientations of the passive receivers from the audio emitted by all the smart audio devices and all the emitters and captured by all the other smart audio devices and all the passive receivers based on DOA estimation.
[0084] In some such instances, the localization process may proceed in a manner similar to that described above. In some instances, the localization process may be based on the same cost function as above, which is provided below for the convenience of the reader.
number
[0085] However, when the localization process involves passive audio receivers and / or audio emitters that are not audio receivers, the variables in the above equations need to be interpreted slightly differently, where N represents the total number of devices, and the breakdown of devices is N smart N smart audio devices rec N passive audio receivers and emit emitters, so N=N smart +N rec +N emit In some examples, the weights w nm DOA may have a sparse structure to mask missing data due to passive receiver or emitter-only devices (or other audio sources without receivers, such as humans), so that if device n is an audio emitter without a receiver, then for all m, nm DOA = 0 and device m is an audio receiver, then w for all n nm DOA = 0. For both smart audio devices and passive receivers, both the position and the angle can be determined, while for audio emitters, only the position is available. The total number of unknowns is 3N smart +3N rec +2N emit It is -4.
[0086] Combined time-of-arrival and direction-of-arrival localization The following discussion highlights the differences between the DOA-based localization process described above and the combined DOA and TOA localization in this section, and those details not explicitly given can be assumed to be the same as in the DOA-based localization process described above.
[0087] 6 is a flow diagram outlining an example of a method for automatically estimating the position and orientation of a device based on DOA and TOA data. Method 600 may be performed, for example, by implementing a localization algorithm via a control system of an apparatus such as that shown in FIG. 10. The blocks of method 600, as well as other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described.
[0088] According to this example, DOA data is obtained in blocks 605-620. According to some implementations, blocks 605-620 may involve obtaining acoustic DOA data from multiple smart audio devices, for example, as described above with reference to blocks 405-420 of Figure 4. In some alternative implementations, blocks 605-620 may involve obtaining DOA data corresponding to electromagnetic waves transmitted and received by each of multiple devices in an environment.
[0089] However, in this example, block 605 also involves obtaining TOA data. According to this example, the TOA data includes measured TOA of audio emitted and received by all smart audio devices in the audio environment (e.g., all pairs of smart audio devices in the audio environment). In some embodiments involving emitting a structured source signal, the audio used to extract the TOA data may be the same as that used to extract the DOA data. In other embodiments, the audio used to extract the TOA data may be different from the audio used to extract the DOA data.
[0090] According to this example, block 616 involves detecting TOA candidates in the audio data, and block 618 involves selecting a single TOA from among the TOA candidates for each smart audio device pair. Some examples are described below.
[0091] Various techniques can be used to obtain TOA data. One method is to use a room calibration audio sequence, such as a sweep (e.g., a logarithmic sine tone) or a maximum length sequence (MLS). Optionally, either of the aforementioned sequences may be used with band-limiting to the near-ultrasonic audio frequency range (e.g., 18 kHz to 24 kHz). In this audio frequency range, most standard audio equipment can emit and record sounds, but such signals cannot be perceived by humans because they lie beyond the capabilities of normal human hearing. Some alternative implementations may involve recovering TOA components from hidden signals in the primary audio signal, such as direct sequence spread spectrum signals.
[0092] Given a set of DOA data from every smart audio device to every other smart audio device, and a set of TOA data from every pair of smart audio devices, the localization method 625 of FIG. 6 may be based on minimizing a cost function, possibly subject to some constraints. In this example, the localization method 625 of FIG. 6 receives the above-mentioned DOA and TOA values as input data and outputs estimated position and orientation data 630 corresponding to the smart audio devices. In some examples, the localization method 625 may also output playback and recording latencies of the smart audio devices, for example, up to some global symmetry that cannot be determined from the minimization problem. Some examples are described below.
[0093] 7 is a flow diagram outlining another example of a method for automatically estimating the position and orientation of a device based on DOA and TOA data. Method 700 may be performed, for example, by implementing a localization algorithm via a control system of an apparatus such as that shown in FIG. 10. The blocks of method 700, as well as other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described.
[0094] Except as described below, in some examples, blocks 705, 710, 715, 720, 725, 730, 735, 740, 745, and 750 may be as described above with reference to blocks 505, 510, 515, 520, 525, 530, 535, 540, 545, and 550 of FIG. 5 . However, in this example, cost function 720 and nonlinear optimization method 735 are modified relative to cost function 520 and nonlinear optimization method 535 of FIG. 5 to operate on both DOA data and TOA data. The TOA data of block 708 may, in some examples, be obtained as described above with reference to FIG. 6 . Another difference compared to the process of FIG. 5 is that in this example, nonlinear optimization method 735 also outputs record and playback latency data 747 corresponding to a smart audio device, e.g., as described below. Thus, in some implementations, result evaluation block 750 may involve evaluating both DOA data and / or TOA data. In some such examples, the operations of block 750 may include a feedback process involving DOA data and / or TOA data. For example, some such examples may implement a feedback process involving comparing the residual of a given TOA / DOA candidate combination with another TOA / DOA candidate combination, as described, for example, in the discussion of TOA / DOA robustness measurements below.
[0095] In some examples, the result evaluation block 750 involves calculating a residual of a cost function at the resulting position and orientation. A relatively lower residual typically indicates a relatively more accurate device localization value. According to some implementations, the result evaluation block 750 may involve a feedback process. For example, some such examples may implement a feedback process that involves comparing the residual of a given TOA / DOA candidate combination with another TOA / DOA candidate combination. This is described, for example, in the discussion of TOA and DOA robustness measurements below.
[0096] Thus, FIG. 6 includes dashed lines from block 630 (which may involve result evaluation in some examples) to DOA candidate selection block 620 and TOA candidate selection block 618 to represent the flow of an optional feedback process. In some implementations, block 705 may involve acquiring acoustic DOA data as described above with reference to blocks 605-620 of FIG. 6, which involves determining and selecting DOA candidates. In some examples, block 708 may involve acquiring acoustic TOA data as described above with reference to blocks 605-618 of FIG. 6, which involves determining and selecting TOA candidates. Although not shown in FIG. 7, some optional feedback processes may involve returning from result evaluation block 750 to block 705 and / or block 708.
[0097] According to this example, the localization algorithm proceeds by minimizing a cost function, possibly subject to some constraints, and can be written as follows: In this example, the localization algorithm receives as input DOA data 705 and TOA data 708, along with configuration parameters 710 specified for the listening environment and possibly some optional constraints 725. In this example, the cost function takes into account the difference between the measured and estimated DOAs, and the difference between the measured and estimated TOAs. In some embodiments, the constraints 725 impose limits on possible device positions, orientations, and / or latencies, such as imposing a condition that audio devices be a certain minimum distance from each other and / or imposing a condition that some device latencies should be zero.
[0098] In some implementations, the cost function can be formulated as follows:
number
number
[0099] There are up to five real unknowns per smart audio device: device position x n (two real unknowns per device), device orientation α n (one real unknown per device) and record and playback latency n and k n (Two additional unknowns per device). From these, only device location and latency are significant for the TOA portion of the cost function. If there are links or constraints between latencies that are known a priori, some implementations can reduce the number of effective unknowns.
[0100] In some examples, there may be additional prior information, for example, regarding the availability or reliability of each TOA measurement. In some of these examples, the weights w nm TOA can be 0 or 1, e.g., 0 for measurements that are not available (or deemed not sufficiently reliable) and 1 for measurements that are reliable. In this way, device localization may be estimated using only a subset of all possible DOA and / or TOA factors. In some other implementations, the weights may have continuous values from 0 to 1, e.g., as a function of the reliability of the TOA measurements. In some examples where no prior reliability information is available, the weights may simply be set to 1.
[0101] According to some implementations, one or more additional constraints may be imposed on the possible values of the latencies and / or the relationships between different latencies.
[0102] In some examples, the location of an audio device may be measured in standard units of length, such as meters, and the latency and arrival times may be given in standard units of time, such as seconds. However, nonlinear optimization methods often work better when the scales of variation of the different variables used in the minimization process are on the same order of magnitude. Thus, some implementations may involve rescaling location measurements so that the range of variation of smart device location is between -1 and 1, and also rescaling latency and arrival times so that these values are also between -1 and 1.
[0103] Minimizing the above cost function does not completely determine the absolute location and orientation or latency of smart audio devices. TOA information provides an absolute distance scale, which means that the cost function is no longer invariant under scale transformations, but remains invariant under global rotations and translations. Furthermore, latency is subject to an additional global symmetry: if the same global amount is added simultaneously to all playback and recording latencies, the cost function remains invariant. These global transformations cannot be determined from minimizing the cost function. Similarly, configuration parameters should provide a criterion that allows uniquely defining a device layout that represents the entire equivalence class.
[0104] In some examples, the symmetry disambiguation criteria may include a reference position that fixes global translational symmetry (e.g., smart device 1 should be at the origin of the coordinate system); a reference orientation that fixes two-dimensional rotational symmetry (e.g., smart device 1 should be oriented toward the front); and a reference latency (e.g., the recording latency for device 1 should be 0). In total, in this example, there are four parameters that cannot be determined from the minimization problem and must be provided as external inputs. Thus, there are 5N-4 unknowns that can be determined from the minimization problem.
[0105] In some implementations, in addition to the set of smart audio devices, there may be one or more passive audio receivers and / or one or more audio emitters, which may not have a functioning microphone array. Including latency as a minimization variable allows some disclosed methods to localize receivers and emitters whose emission and reception times are not precisely known. In some such implementations, the TOA cost function described above may be implemented. This cost function is reproduced below for the reader's convenience.
number
[0106] As mentioned above with reference to the DOA cost function, the cost function variables need to be interpreted slightly differently when the cost function is used for localization estimation involving passive receivers and / or emitters, where N represents the total number of devices, and the breakdown of devices is N smart N smart audio devices rec N passive audio receivers and emit emitters, so N=N smart +N rec +N emit The weight w nm DOA may have a sparsity structure to mask missing data due to passive receivers or dedicated emitters, so that, for example, if device n is an audio emitter, then w nm DOA = 0 and device m is an audio receiver, then w for all n nm DOA= 0. According to some implementations, for smart audio devices, location, orientation, and recording and playback latency must be determined; for passive receivers, location, orientation, and recording latency must be determined; and for audio emitters, location and playback latency must be determined. Thus, according to some such examples, the total number of unknowns is 5N smart +4N rec +3N emit It is -4.
[0107] Global translation and rotation disambiguation Solutions to both the DOA-only problem and the combined TOA and DOA problem are subject to global translational and rotational ambiguities. In some instances, the translational ambiguity can be resolved by treating the emitter-only source as the listener and translating all devices so that the listener is located at the origin.
[0108] Rotational ambiguity can be resolved by imposing additional constraints on the solution. For example, some multi-loudspeaker environments may include a television (TV) loudspeaker and a sofa positioned for TV viewing. After locating the loudspeakers in the environment, some methods may involve finding a vector connecting the listener to the TV viewing direction. Some such methods may then involve causing the TV to emit sound from its loudspeaker and / or prompting the user to walk to the TV and locating the user's speech. Some implementations may involve rendering an audio object that pans around the environment. A user may provide user input (e.g., say "stop") indicating when the audio object is at one or more predetermined locations within the environment, such as the front of the environment, the TV location in the environment, etc. Some implementations include a mobile phone app with an inertial measurement unit that prompts the user to point the mobile phone in two defined directions, the first direction being the direction of a particular device (e.g., the device with the illuminated LED), and the second direction being the user's desired viewing direction, such as the front of the environment, the TV location in the environment, etc. Some detailed disambiguation examples will now be described with reference to Figures 8A-8D.
[0109] 8A illustrates an example audio environment. According to some examples, audio device position data output by one of the disclosed localization methods may include an estimate of the audio device position for each of audio devices 1-5 relative to audio device coordinate system 807. In this implementation, audio device coordinate system 807 is a Cartesian coordinate system with the position of the microphone for audio device 2 as its origin. Here, the x-axis of audio device coordinate system 807 corresponds to line 803 between the position of the microphone for audio device 2 and the position of the microphone for audio device 1.
[0110] In this example, listener location is determined by prompting listener 805, depicted as sitting on couch 103 (e.g., via audio prompts from one or more loudspeakers in environment 800a), to make one or more utterances 827 and estimating listener location according to time-of-arrival (TOA) data. The TOA data corresponds to microphone data acquired by multiple microphones in the environment. In this example, the microphone data corresponds to detection of the one or more utterances 827 by at least some (e.g., three, four, or all five) of the microphones in audio devices 1-5.
[0111] Alternatively or additionally, the listener position may be estimated according to DOA data provided by at least some (e.g., two, three, four, or all five) microphones of audio devices 1 through 5. According to some such examples, the listener position may be determined according to the intersection of lines 809a, 809b, etc., corresponding to the DOA data.
[0112] According to this example, the listener position corresponds to the origin of listener coordinate system 820. In this example, listener angular orientation data is represented by the y'-axis of listener coordinate system 820, which corresponds to line 813a between listener's head 810 (and / or listener's nose 825) and soundbar 830 of television 101. In the example shown in FIG. 8A, line 813a is parallel to the y'-axis. Thus, angle Θ represents the angle between the y-axis and the y'-axis. In this example, block 1225 of FIG. 12 may involve rotating audio device coordinates by angle Θ about the origin of listener coordinate system 820. Thus, although the origin of audio device coordinate system 807 is shown in FIG. 8A to correspond to audio device 2, some implementations involve coordinating the origin of audio device coordinate system 807 with the origin of listener coordinate system 820 before rotating the audio device coordinates by angle Θ about the origin of listener coordinate system 820. This coordinating may be performed by a coordinate transformation from audio device coordinate system 807 to listener coordinate system 820.
[0113] The location of the soundbar 830 and / or television 101 may, in some examples, be determined by causing the soundbar to emit sound and estimating the location of the soundbar according to DOA and / or TOA data that may correspond to detection of the sound by at least some (e.g., three, four, or all five) microphones of audio devices 1-5. Alternatively or additionally, the location of the soundbar 830 and / or television 101 may be determined by prompting a user to walk to the television and localizing the user's speech according to DOA and / or TOA data that may correspond to detection of the sound by at least some (e.g., three, four, or all five) microphones of audio devices 1-5. Some such methods may involve, for example, applying a cost function, as described above. Some such methods may involve triangulation. Such examples may be useful in situations where the soundbar 830 and / or television 101 do not have associated microphones.
[0114] In some other examples where the soundbar 830 and / or television 101 have associated microphones, the location of the soundbar 830 and / or television 101 may be determined according to a TOA and / or DOA method, such as those disclosed herein. According to some such methods, the microphones may be co-located with the soundbar 830.
[0115] According to some implementations, the soundbar 830 and / or the television 101 may have an associated camera 811. The control system may be configured to capture an image of the listener's head 810 (and / or the listener's nose 825). In some such examples, the control system may be configured to determine a line 813a between the listener's head 810 (and / or the listener's nose 825) and the camera 811. The listener angular orientation data may correspond to the line 813a. Alternatively or additionally, the control system may be configured to determine an angle Θ between the line 813a and the y-axis of the audio device coordinate system.
[0116] FIG. 8B shows an additional example of determining listener angular orientation data. According to this example, the listener position has already been determined in block 1215 of FIG. 12. Here, a control system controls the loudspeakers of environment 800b to render audio object 835 at various positions within environment 800b. In some such examples, the control system may cause the loudspeakers to render audio object 835 so that audio object 835 appears to rotate around listener 805, for example, by rendering audio object 835 so that audio object 835 appears to rotate around the origin of listener coordinate system 820. In this example, curved arrow 840 indicates a portion of the trajectory of audio object 835 as it rotates around listener 805.
[0117] According to some such examples, listener 805 may provide user input indicating when audio object 835 is in the direction listener 805 is facing (e.g., say "stop"). In some such examples, control system may be configured to determine line 813b between the listener position and the position of audio object 835. In this example, line 813b corresponds to the y' axis of the listener coordinate system, which indicates the direction listener 805 is facing. In alternative implementations, listener 805 may provide user input indicating when audio object 835 is in front of the environment, at the TV position of the environment, at the audio device position, etc.
[0118] FIG. 8C shows an additional example of determining listener angular orientation data. According to this example, the listener position has already been determined in block 1215 of FIG. 12. Here, listener 805 uses handheld device 845 to provide input regarding listener's 805 viewing direction by pointing handheld device 845 toward television 101 or soundbar 830. The dashed outline of handheld device 845 and listener's arm indicates that, in this example, listener 805 was pointing handheld device 845 toward audio device 2 at a time prior to when listener 805 was pointing handheld device 845 toward television 101 or soundbar 830. In other examples, listener 805 may be pointing handheld device 845 toward another audio device, such as audio device 1. According to this example, handheld device 845 is configured to determine an angle α between audio device 2 and television 101 or soundbar 830, where angle α approximates the angle between audio device 2 and the viewing direction of listener 805.
[0119] Handheld device 845, in some examples, may be a cellular phone including an inertial sensor system and a wireless interface configured to communicate with a control system controlling the audio devices of environment 800c. In some examples, handheld device 845 may be running an application or “app” configured to control handheld device 845 to perform a desired function, such as by providing user prompts (e.g., via a graphical user interface), by receiving input indicating that handheld device 845 is pointing in a desired direction, storing corresponding inertial sensor data, and / or by transmitting corresponding inertial sensor data to a control system controlling the audio devices of environment 800c.
[0120] According to this example, a control system (which may be a control system of handheld device 845, a control system of a smart audio device in environment 800c, or a control system controlling an audio device in environment 800c) is configured to determine the orientation of lines 813c and 850 according to inertial sensor data, for example, according to gyroscope data. In this example, line 813c is parallel to axis y' and may be used to determine the listener angular orientation. According to some examples, the control system may determine an appropriate rotation of audio device coordinates around the origin of listener coordinate system 820 according to angle α between audio device 2 and the viewing direction of listener 805.
[0121] FIG. 8D shows an example of determining an appropriate rotation of audio device coordinates according to the method described with reference to FIG. 8C. In this example, the origin of audio device coordinate system 807 is co-located with the origin of listener coordinate system 820. Co-locating the origin of audio device coordinate system 807 with the origin of listener coordinate system 820 is possible after the listener position is determined. Co-locating the origin of audio device coordinate system 807 with the origin of listener coordinate system 820 may include transforming the audio device position from audio device coordinate system 807 to listener coordinate system 820. Angle α was determined as described above with reference to FIG. 8C. Thus, angle α corresponds to the desired orientation of audio device 2 in listener coordinate system 820. In this example, angle β corresponds to the orientation of audio device 2 in audio device coordinate system 807. The angle Θ, which in this example is β-α, indicates the rotation required to align the y-axis of the audio device coordinate system 807 with the y'-axis of the listener coordinate system 820.
[0122] DOA robustness index As discussed above with reference to Figure 4, some examples use "blind" methods applied to any signal, including steered response power, beamforming, or other similar methods, and a robustness measure may be added to improve accuracy and stability. Some implementations include time integration of the beamformer steered response to filter out transient components and detect only persistent peaks, as well as to average out random errors and fluctuations in those persistent DOA. Other examples may use only a limited frequency band as input, which may be tailored to the room or signal type for better performance.
[0123] For example, when using a "supervised" method involving the use of a structured source signal and deconvolution methods to generate impulse responses, preprocessing strategies can be implemented to increase the accuracy and salience of the DOA peaks. In some examples, such preprocessing may include truncation using an amplitude window of some time width beginning at the onset of the impulse response on each microphone channel. Such examples may incorporate an impulse response onset detector so that each channel onset can be found independently.
[0124] In some examples based on either the "blind" or "supervised" methods described above, further processing may be added to improve DOA accuracy. It is important to note that DOA selection based on peak detection (e.g., during steered-response power (SRP) or impulse response analysis) is sensitive to environmental acoustics. Environmental acoustics can cause capture of non-primary path signals due to reflections and device occlusions, which attenuate both received and transmitted energy. These occurrences can reduce the accuracy of device-pair DOAs and introduce errors into the optimizer's localization solution. Therefore, it is prudent to consider all peaks within a predetermined threshold as candidates for the ground truth DOA. One example of a predetermined threshold is the requirement that the peak be greater than the mean steered-response power (SRP). For all detected peaks, saliency thresholding and removal of candidates below the mean signal level have proven to be simple yet effective initial filtering techniques. As used herein, "prominence" is a measure of how large a local peak is compared to its neighboring minima, which differs from thresholding based solely on power. An example of a prominence threshold is the requirement that the power difference between a peak and its neighboring minima be greater than or equal to a threshold. Retaining promising candidates improves the likelihood that a device pair contains a usable DOA in their set (within an acceptable error tolerance from the correct answer). However, if the signal is corrupted by strong reflections / obscurations, the device pair may not contain a usable DOA. In some examples, a selection algorithm may be implemented to do one of the following: 1) select the best usable DOA candidate for each device pair; 2) determine that none of the candidates are usable and therefore null the optimization contribution for that pair using a cost function weighting matrix; or 3) select the best inferred candidate, but apply non-binary weighting to the DOA contribution if it is difficult to unambiguously determine the amount of error the best candidate introduces.
[0125] After initial optimization using the best inferred candidates, in some examples, the localization solution may be used to calculate the residual cost contribution of each DOA. Outlier analysis of the residual cost may provide evidence of the DOA pairs that are most impacting the localization solution, with extreme outliers flagging those DOAs as potentially inaccurate or suboptimal. Recursive optimization of outlier DOA pairs based on their residual cost contributions using the remaining candidates and weightings applied to the contributions of that device pair may then be used for candidate processing according to one of the three options previously mentioned. This is an example of a feedback process, as described above with reference to Figures 4-7. According to some implementations, repeated optimization and processing decisions may be performed until all detected candidates have been evaluated and the residual cost contributions of the selected DOAs are balanced.
[0126] The drawback of candidate selection based on optimizer evaluation is that it is computationally intensive and sensitive to candidate traversal order. An alternative technique with less computational weight involves determining all permutations of candidates in a set and performing a triangle alignment method for device localization on these candidates. A related triangle alignment method is disclosed in U.S. Patent No. 6,275,999, which is incorporated herein by reference for all purposes. The localization results can then be evaluated by calculating the total and residual costs they incur with respect to the DOA candidates used in triangulation. Decision logic for parsing these metrics can be used to determine the best candidates and their respective weightings to be fed into the nonlinear optimization problem. When the list of candidates is large, and therefore the number of permutations is large, filtering and intelligent traversal through the permutation list can be applied. [Patent Document 1] U.S. Provisional Patent Application No. 62 / 992,068, filed March 19, 2020, entitled "Audio Device Auto-Location"
[0127] TOA robustness index As discussed above with reference to Figure 6, the use of multiple candidate TOA solutions adds robustness compared to systems utilizing a single or minimal TOA value, ensuring that the impact of errors on finding the optimal speaker layout is minimized. Once the system's impulse response is obtained, in some instances, each of the TOA matrix elements can be reconstructed by searching for the peak corresponding to the direct sound. Under ideal conditions (e.g., no noise, no obstructions in the direct path between the sound source and receiver, and the speaker pointing directly toward the microphone), this peak is easily identifiable as the largest peak in the impulse response. However, in the presence of noise, obstructions, or misalignment of the speaker and microphone, the peak corresponding to the direct sound does not necessarily correspond to a maximum value. Furthermore, in such conditions, the peak corresponding to the direct sound may be difficult to isolate from other reflections and / or noise. Direct sound identification can be a challenging process in some instances. Inaccurate identification of the direct sound degrades (and in some cases completely destroys) the automatic localization process. Therefore, when there is a possibility of error in the direct sound identification process, it can be effective to consider multiple candidates for the direct sound. In some such cases, the peak selection process may include two parts: (1) a direct sound search algorithm that searches for suitable peak candidates, and (2) a peak candidate evaluation process to increase the probability of selecting the correct TOA matrix element.
[0128] In some implementations, the process of searching for direct sound candidate peaks may include a method for identifying significant candidates for the direct sound. Some such methods may be based on the following steps: (1) identifying one first reference peak (e.g., the maximum absolute value of the impulse response (IR)), the “first peak,” (2) evaluating the noise level around (before and after) this first peak, (3) searching for alternative peaks before (and possibly after) the first peak that are above the noise level, (4) ranking the found peaks according to their probability of corresponding to the correct TOA, and, optionally, (5) grouping nearby peaks (to reduce the number of candidates).
[0129] Once direct sound candidate peaks are identified, some implementations may involve a multiple peak evaluation step. The direct sound candidate peak search results, in some instances, in one or more candidate values for each TOA matrix element, ranked according to their estimated probability. Multiple TOA matrices can be formed by selecting among different candidate values. To evaluate the likelihood of a given TOA matrix, a minimization process (such as the one described above) may be implemented. This process can generate a residual of the minimization, which is a good estimate of the internal coherence of the TOA and DOA matrices. A perfect noiseless TOA matrix will yield a zero residual, while a TOA matrix with inaccurate matrix elements will yield a large residual. In some implementations, the method searches for a set of candidate TOA matrix elements that produces a TOA matrix with the smallest residual. This is an example of the evaluation process described above with reference to FIGS. 6 and 7, which may include a result evaluation block 750. In one example, the evaluation process may involve the following steps: (1) selecting an initial TOA matrix, (2) evaluating the initial matrix using a residual from the minimization process, (3) changing one matrix element of the TOA matrix from a list of TOA candidates, (4) re-evaluating the matrix using the residual from the minimization process, (5) accepting the change if the residual is smaller, and not accepting the change otherwise, and (6) sequentially repeating steps 3 through 5. In some examples, the evaluation process may stop when all TOA candidates have been evaluated or when a predetermined maximum number of iterations has been reached.
[0130] Examples of localization methods FIG. 9A is a flow diagram outlining an example of a localization method. The blocks of method 900, as well as other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described. In this implementation, method 900 involves estimating the position and orientation of audio devices within an environment. The blocks of method 900 may be performed by one or more devices, which may be (or may include) apparatus 1000 shown in FIG. 10.
[0131] In this example, block 905, by a control system, obtains direction-of-arrival (DOA) data corresponding to sound emitted by at least a first smart audio device in the audio environment. The control system may be, for example, control system 1010 described below with reference to FIG. 10. According to this example, the first smart audio device includes a first audio transmitter and a first audio receiver, and the DOA data corresponds to sound received by at least a second smart audio device in the audio environment, where the second smart audio device includes a second audio transmitter and a second audio receiver. In this example, the DOA data also corresponds to sound emitted by the at least second smart audio device and received by the at least first smart audio device. In some examples, the first and second smart audio devices may be two of audio devices 105a-105d shown in FIG. 1.
[0132] The DOA data may be obtained in various manners depending on the particular implementation. In some cases, determining the DOA data may involve one or more of the DOA-related methods described above with reference to FIG. 4 and / or in the "DOA Robustness Metrics" section. Some implementations may involve the control system obtaining one or more elements of the DOA data using beamforming methods, steered powered response methods, time difference of arrival methods, and / or structured signal methods.
[0133] According to this example, block 910 involves receiving, by the control system, configuration parameters. In this implementation, the configuration parameters correspond to the audio environment itself, one or more audio devices in the audio environment, or both the audio environment and one or more audio devices in the audio environment. According to some examples, the configuration parameters may indicate the number of audio devices in the audio environment, one or more dimensions of the audio environment, one or more constraints on audio device position or orientation, and / or disambiguation data for at least one of rotation, translation, or scaling. In some examples, the configuration parameters may include playback latency data, recording latency data, and / or data for disambiguating latency symmetry.
[0134] In this example, block 915 involves minimizing a cost function based at least in part on the DOA data and configuration parameters to estimate the position and orientation of at least the first smart audio device and the second smart audio device by the control system.
[0135] According to some examples, the DOA data may also correspond to sounds emitted by a third through Nth smart audio device in the audio environment, where N corresponds to the total number of smart audio devices in the audio environment. In such examples, the DOA data may also correspond to sounds received by each of the first through Nth smart audio devices from all other smart audio devices in the audio environment. In such instances, minimizing the cost function may involve estimating the positions and / or orientations of the third through Nth smart audio devices.
[0136] In some examples, the DOA data may also correspond to sounds received by one or more passive audio receivers of the audio environment. Each of the one or more passive audio receivers may include a microphone array but may lack an audio emitter. Minimizing the cost function may also provide an estimated position and orientation of each of the one or more passive audio receivers. According to some examples, the DOA data may also correspond to sounds emitted by one or more audio emitters of the audio environment. Each of the one or more audio emitters may include at least one sound-emitting transducer but may lack a microphone array. Minimizing the cost function may also provide an estimated position of each of the one or more audio emitters.
[0137] In some examples, method 900 may involve receiving, by a control system, a seed layout for the cost function. The seed layout may specify, for example, the correct number of audio transmitters and receivers in the audio environment and optional positions and orientations for each of the audio transmitters and receivers within the audio environment.
[0138] According to some examples, method 900 may involve receiving, by a control system, a weighting factor associated with one or more elements of DOA data, which may indicate, for example, availability and / or reliability of the one or more elements of DOA data.
[0139] In some examples, method 900 may involve receiving, by a control system, time-of-arrival (TOA) data corresponding to sound emitted by at least one audio device in the audio environment and received by at least one other audio device in the audio environment. In some such examples, a cost function may be based at least in part on the TOA data. Some such methods may involve estimating at least one playback latency and / or at least one recording latency. According to some examples, the cost function may operate with respect to a rescaled position, a rescaled latency, and / or a rescaled time of arrival.
[0140] In some examples, the cost function may include a first term that depends only on DOA data and a second term that depends only on TOA data. In some such examples, the first term may include a first weighting factor and the second term may include a second weighting factor. According to some such examples, one or more TOA elements in the second term may have a TOA element weighting factor that indicates the availability or reliability of each of the one or more TOA elements.
[0141] 9B is a flow diagram outlining another example of a localization method. The blocks of method 950, as with other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described. In this implementation, method 950 involves estimating the position and orientation of a device within an environment. The blocks of method 950 may be performed by one or more devices, which may be (or may include) apparatus 1000 shown in FIG. 10.
[0142] In this example, block 955, by a control system, obtains direction-of-arrival (DOA) data corresponding to transmissions of at least a first transceiver of a first device in the environment. The control system may be, for example, control system 1010 described below with reference to FIG. 10. According to this example, the first transceiver includes a first transmitter and a first receiver, and the DOA data corresponds to transmissions received by at least a second transceiver of a second device in the environment, and the second transceiver also includes a second transmitter and a second receiver. In this example, the DOA data also corresponds to transmissions from at least the second transceiver received by the at least the first transceiver. According to some examples, the first transceiver and the second transceiver may be configured to transmit and receive electromagnetic waves. In some examples, the first and second smart audio devices may be two of audio devices 105a-105d shown in FIG. 1.
[0143] The DOA data may be obtained in various manners depending on the particular implementation. In some cases, determining the DOA data may involve one or more of the DOA-related methods described above with reference to FIG. 4 and / or in the "DOA Robustness Metrics" section. Some implementations may involve the control system obtaining one or more elements of the DOA data using beamforming methods, steered powered response methods, time difference of arrival methods, and / or structured signal methods.
[0144] According to this example, block 960 involves receiving, by the control system, configuration parameters. In this implementation, the configuration parameters correspond to the environment itself, one or more devices in the audio environment, or both the environment and one or more audio devices in the audio environment. According to some examples, the configuration parameters may indicate the number of audio devices in the environment, one or more dimensions of the environment, one or more constraints on device position or orientation, and / or disambiguation data for at least one of rotation, translation, or scaling. In some examples, the configuration parameters may include playback latency data, recording latency data, and / or data for disambiguating latency symmetry.
[0145] In this example, block 965 involves the control system minimizing a cost function based at least in part on the DOA data and configuration parameters to estimate the position and orientation of at least the first device and the second device.
[0146] According to some implementations, the DOA data may also correspond to transmissions emitted by third through Nth transceivers of third through Nth devices in the environment, where N corresponds to the total number of transceivers in the environment. The DOA data also corresponds to transmissions received by each of the first through Nth transceivers from all other transceivers in the environment. In some such implementations, minimizing the cost function may involve estimating the positions and / or orientations of the third through Nth transceivers.
[0147] In some examples, the first device and the second device may be smart audio devices, and the environment may be an audio environment. In some such examples, the first transmitter and the second transmitter may be audio transmitters. In some such examples, the first receiver and the second receiver may be audio receivers. According to some such examples, the DOA data may also correspond to sounds emitted by a third through Nth smart audio device in the audio environment, where N corresponds to the total number of smart audio devices in the audio environment. In such examples, the DOA data may also correspond to sounds received by each of the first through Nth smart audio devices from all other smart audio devices in the audio environment. In such cases, minimizing the cost function may involve estimating the positions and orientations of the third through Nth smart audio devices. Alternatively and / or additionally, in some examples, the DOA data may correspond to electromagnetic waves emitted and received by devices in the environment.
[0148] In some examples, the DOA data may also correspond to sounds received by one or more passive receivers in the environment. Each of the one or more passive receivers may include a receiver array but may lack a transmitter. Minimizing the cost function may also provide an estimated position and orientation of each of the one or more passive receivers. According to some examples, the DOA data may also correspond to transmissions from one or more transmitters in the environment. In some such examples, each of the one or more transmitters may lack a receiver array. Minimizing the cost function may also provide an estimated position of each of the one or more transmitters.
[0149] In some examples, method 950 may involve receiving, by a control system, a seed layout for the cost function. The seed layout may specify, for example, the correct number of transmitters and receivers in the audio environment and optional positions and orientations for each of the transmitters and receivers within the audio environment.
[0150] According to some examples, method 950 may involve receiving, by a control system, a weighting factor associated with one or more elements of DOA data, which may indicate, for example, availability and / or reliability of the one or more elements of DOA data.
[0151] In some examples, method 950 may involve receiving, by a control system, time-of-arrival (TOA) data corresponding to sound emitted by at least one audio device in the audio environment and received by at least one other audio device in the audio environment. In some such examples, a cost function may be based at least in part on the TOA data. Some such methods may involve estimating at least one playback latency and / or at least one recording latency. According to some such examples, the cost function may operate with respect to a rescaled position, a rescaled latency, and / or a rescaled time of arrival.
[0152] In some examples, the cost function may include a first term that depends only on DOA data and a second term that depends only on TOA data. In some such examples, the first term may include a first weighting factor and the second term may include a second weighting factor. According to some such examples, one or more TOA elements in the second term may have a TOA element weighting factor that indicates the availability or reliability of each of the one or more TOA elements.
[0153] FIG. 10 is a block diagram illustrating example components of an apparatus capable of implementing various aspects of the present disclosure. The apparatus 1000 may be configured to perform, for example, the methods described above with reference to FIG. 9A and / or FIG. 9B . According to some examples, the apparatus 1000 may be or include a smart audio device (such as a smart speaker) configured to perform at least some of the methods disclosed herein. In other implementations, the apparatus 1000 may be or include another device configured to perform at least some of the methods disclosed herein. In some such implementations, the apparatus 1000 may be or include a smart home hub or server.
[0154] In this example, the device 1000 includes an interface system 1005 and a control system 1010. In some implementations, the interface system 1005 may be configured to receive input from each of multiple microphones in the environment. The interface system 1005 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 1005 may include one or more wireless interfaces. The interface system 1005 may include one or more devices for implementing a user interface, such as one or more microphones, one or more loudspeakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 1005 may include one or more interfaces between the control system 1010 and a memory system, such as the optional memory system 1015 shown in FIG. 10 . However, the control system 1010 may include a memory system.
[0155] The control system 1010 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components. In some implementations, the control system 1010 may reside in more than one device. For example, in some implementations, a portion of the control system 1010 may reside in a device within the audio environment 100 depicted in FIG. 1 (e.g., one of the audio devices 105a-105d or a smart home device), and another portion of the control system 1010 may reside in a device outside the audio environment 100, such as a server, a mobile device (e.g., a smartphone or tablet computer), or the like. The interface system 1005 may also reside in more than one device in some such examples.
[0156] In some implementations, the control system 1010 may be configured, at least in part, to perform the methods disclosed herein. According to some examples, the control system 1010 may be configured to implement the methods described above, for example, with reference to Figures 4-9B.
[0157] In some examples, device 1000 may include an optional microphone system 1020 shown in FIG. 10. Microphone system 1020 may include one or more microphones. In some examples, microphone system 1020 may include an array of microphones. In some examples, device 1000 may include an optional loudspeaker system 1025 shown in FIG. 10. Loudspeaker system 1025 may include one or more loudspeakers. In some examples, microphone system 1020 may include an array of loudspeakers. In some such examples, device 1000 may be or include an audio device. For example, device 1000 may be or include one of audio devices 105a-105d shown in FIG. 1.
[0158] In some examples, the apparatus 1000 may include an optional antenna system 1030 shown in FIG. 10 . According to some examples, the antenna system 1030 may include an array of antennas. In some examples, the antenna system 1030 may be configured to transmit and / or receive electromagnetic waves. According to some implementations, the control system 1010 may be configured to estimate a distance between two audio devices in the environment based on antenna data from the antenna system 1030. For example, the control system 1010 may be configured to estimate a distance between two audio devices in the environment according to a direction of arrival of the antenna data and / or a received signal strength of the antenna data.
[0159] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. For example, some or all of the methods described herein may be performed by control system 1010 according to instructions stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may reside, for example, in optional memory system 1015 shown in FIG. 10 and / or in control system 1010. Thus, various inventive aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media having software stored thereon. The software may include, for example, instructions for controlling at least one device to process audio data. The software may be executable by one or more components of a control system, such as control system 1010 of FIG. 10.
[0160] Figure 11 shows an example floor plan of an audio environment, in this example a residential space. As with other figures provided herein, the types and number of elements shown in Figure 11 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements.
[0161] According to this example, environment 1100 includes living room 1110 at the upper left, kitchen 1115 at the lower center, and bedroom 1122 at the lower right. Squares and circles distributed throughout the living space represent a set of loudspeakers 1105a-1105h that are conveniently positioned in the space but do not conform to a standardized layout (arbitrarily placed). At least some of these loudspeakers may be smart speakers in some implementations. In some examples, television 1130 may be configured, at least in part, to implement one or more disclosed embodiments. In this example, environment 1100 includes cameras 1111a-1111e distributed throughout the environment. In some implementations, one or more smart audio devices in environment 1100 may also include one or more cameras. The one or more smart audio devices may be single-purpose audio devices or virtual assistants. In some such examples, one or more cameras of optional sensor system 130 may be present in or on television 1130, in a mobile phone, or in a smart speaker such as one or more of loudspeakers 1105b, 1105d, 1105e, or 1105h. Cameras 1111a-1111e are not shown in all illustrations of environments 1100 presented in this disclosure, but each of environments 1100 may nevertheless include one or more cameras in some implementations.
[0162] Some aspects of the present disclosure include systems or devices configured (e.g., programmed) to perform one or more examples of the disclosed methods, and tangible computer-readable media (e.g., disks) storing code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system including an input device, memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data presented thereto.
[0163] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on an audio signal, including performing one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and memory) that is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device.
[0164] Another aspect of the present disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) storing code (e.g., executable code) for performing one or more examples of the disclosed methods or steps thereof.
[0165] While particular embodiments of and applications of the present disclosure have been described herein, it will be apparent to those skilled in the art that many variations of the embodiments and applications described herein are possible without departing from the scope of the present disclosure.
Claims
1. 1. A method for localizing an audio device in an audio environment, the method comprising: acquiring, by a control system, direction of arrival (DOA) data corresponding to sound emitted by at least a first smart audio device of the audio environment, the first smart audio device including a first audio transmitter and a first audio receiver, the DOA data corresponding to sound received by at least a second smart audio device of the audio environment, the second smart audio device including a second audio transmitter and a second audio receiver, the DOA data also corresponding to sound emitted by at least the second smart audio device and received by at least the first smart audio device; receiving, by the control system, configuration parameters corresponding to the audio environment, corresponding to one or more audio devices of the audio environment, or corresponding to both the audio environment and the one or more audio devices of the audio environment; and minimizing, by the control system, a cost function based at least in part on the DOA data and the configuration parameters to estimate positions and orientations of at least the first smart audio device and the second smart audio device. method.
2. 2. The method of claim 1 , wherein the DOA data corresponds to sounds received by one or more passive audio receivers of the audio environment, each of the one or more passive audio receivers including a microphone array but lacking an audio emitter, and wherein minimizing the cost function also provides an estimated position and orientation of each of the one or more passive audio receivers.
3. 3. The method of claim 1, wherein the DOA data also corresponds to sounds emitted by one or more audio emitters of the audio environment, each of the one or more audio emitters including at least one sound emitting transducer but lacking a microphone array, and wherein minimizing the cost function also provides an estimated position of each of the one or more audio emitters.
4. 4. The method of claim 1, wherein the DOA data also corresponds to sounds emitted by a third through an Nth smart audio device of the audio environment, where N corresponds to a total number of smart audio devices in the audio environment, and the DOA data also corresponds to sounds received by each of the first through Nth smart audio devices from all other smart audio devices in the audio environment, and minimizing the cost function involves estimating the positions and orientations of the third through Nth smart audio devices.
5. 5. The method of claim 1, wherein the configuration parameters include the number of audio devices in the audio environment, one or more dimensions of the audio environment, one or more constraints on audio device position or orientation, or disambiguation data for at least one of rotation, translation, or scaling.
6. 6. The method of claim 1, further comprising receiving, by the control system, a seed layout for the cost function, the seed layout specifying a correct number of audio transmitters and receivers in the audio environment and arbitrary positions and orientations for each of the audio transmitters and receivers in the audio environment.
7. 7. The method of claim 1, further comprising receiving, by the control system, weighting factors associated with one or more elements of the DOA data, the weighting factors indicating at least one of availability or reliability of the one or more elements.
8. 8. The method of claim 1, further comprising: acquiring, by the control system, one or more elements of the DOA data using at least one of a beamforming method, a steered powered response method, a time difference of arrival method, or a structured signal method.
9. 9. The method of claim 1, further comprising receiving, by the control system, time-of-arrival (TOA) data corresponding to sounds emitted by at least one audio device of the audio environment and received by at least one other audio device of the audio environment, wherein the cost function is based at least in part on the TOA data.
10. 10. The method of claim 9, further comprising estimating at least one playback latency, estimating at least one recording latency, or estimating at least one playback latency and at least one recording latency.
11. The method of claim 10 , wherein the cost function operates on at least one of a rescaled position, a rescaled latency, or a rescaled arrival time.
12. 12. The method of claim 9, wherein the cost function includes a first term that depends only on the DOA data and a second term that depends only on the TOA data.
13. The method of claim 12 , wherein the first term includes a first weighting factor and the second term includes a second weighting factor.
14. The method of claim 12 , wherein one or more TOA elements of the second term have a TOA element weight factor indicating availability or reliability of each of the one or more TOA elements.
15. 15. The method of claim 1, wherein the configuration parameters include at least one of playback latency data, recording latency data, data for disambiguating latency symmetry, rotation disambiguation data, translation disambiguation data, or scaling disambiguation data.
16. Apparatus configured to carry out the method of any one of claims 1 to 15.
17. A system configured to carry out the method of any one of claims 1 to 15.
18. One or more non-transitory media storing software including instructions for controlling one or more devices to perform the method of any one of claims 1 to 15.
19. 1. A method for localizing a device in an environment, the method comprising: acquiring, by a control system, direction of arrival (DOA) data corresponding to transmissions of at least a first transceiver of a first device in the environment, the first transceiver including a first transmitter and a first receiver, the DOA data corresponding to transmissions received by at least a second transceiver of a second device in the environment, the second transceiver including a second transmitter and a second receiver, the DOA data also corresponding to transmissions from at least the second transceiver received by at least the first transceiver; receiving, by the control system, configuration parameters corresponding to the environment, corresponding to one or more devices of the environment, or corresponding to both the environment and the one or more devices of the environment; and minimizing, by the control system, a cost function based at least in part on the DOA data and the configuration parameters to estimate positions and orientations of at least the first device and the second device. method.
20. 20. The method of claim 19, wherein the DOA data also corresponds to transmissions received by one or more passive receivers in the environment, each of the one or more passive receivers including a receiver array but lacking a transmitter, and wherein minimizing the cost function also provides an estimated position and orientation of each of the one or more passive receivers.
21. 21. The method of claim 19 or 20, wherein the DOA data also corresponds to transmissions from one or more transmitters in the environment, each of the one or more transmitters lacking a receiver array, and wherein minimizing the cost function also provides an estimated position of each of the one or more transmitters.
22. 22. The method of claim 19, wherein the DOA data also corresponds to transmissions emitted by third through Nth transceivers of third through Nth devices in the environment, where N corresponds to a total number of transceivers in the environment, and the DOA data also corresponds to transmissions received by each of the first through Nth transceivers from all other transceivers in the environment, and wherein minimizing the cost function includes estimating the positions and orientations of the third through Nth transceivers.
23. 23. The method of any one of claims 19 to 22, wherein the first device and the second device are audio devices and the environment is an audio environment.
24. the first transmitter and the second transmitter are audio transmitters; the first receiver and the second receiver are audio receivers; 24. The method of claim 23.
25. 24. The method of any one of claims 19 to 23, wherein the first transceiver and the second transceiver are configured to transmit and receive electromagnetic waves.
26. Apparatus configured to carry out the method of any one of claims 19 to 25.
27. A system configured to carry out the method of any one of claims 19 to 25.
28. 26. One or more non-transitory media storing software including instructions for controlling one or more devices to perform the method of any one of claims 19 to 25.
Citation Information
Patent Citations
Wireless coordination of audio source
JP2020184788A
Automatic discovery and localization of speaker locations in surround sound systems
US20190253801A1
Acoustic environment mapping
WO2018029341A1