Distributed Audio Device Ducking
A control system optimizes audio processing for multiple devices by considering device location and echo management to enhance SER and speech recognition in multi-device audio environments.
Patent Information
- Application Number
- JP2024527308
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-04-12
- Filing Date
- 2022-11-04
- Publication Date
- 2025-12-22
- Estimated Expiration
- 2042-11-04
AI Technical Summary
Existing audio devices struggle with managing the signal-to-echo ratio (SER) in multi-device environments, where echoes from other devices limit the effectiveness of speech recognition, particularly when multiple audio devices are present in a single acoustic space.
Implementing a control system that adjusts audio processing for multiple audio devices based on factors like device location, echo management system performance, and user position to optimize the SER by modifying audio rendering and echo cancellation techniques.
Enhances the signal-to-echo ratio, improving speech recognition performance by reducing echoes and maintaining a desirable listening experience across multiple audio devices.
Smart Images

Figure 0007789915000094 
Figure 0007789915000095 
Figure 0007789915000096
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 278,003, filed November 10, 2021, U.S. Provisional Patent Application No. 63 / 362,842, filed April 12, 2022, and European Patent Application No. 22167857.6, filed April 12, 2022, all of which are incorporated herein by reference in their entireties. [Technical field]
[0002] The present disclosure relates to systems and methods for orchestrating and implementing audio devices, such as smart audio devices, and for controlling the speech-to-echo ratio (SER) in such audio devices. [Background technology]
[0003] Audio devices, including but not limited to smart audio devices, are being widely deployed and are becoming a common feature in many homes. While existing systems and methods for controlling audio devices provide benefits, improved systems and methods are desirable.
[0004] Notation and Nomenclature Throughout this disclosure, including the claims, the terms "speaker," "loudspeaker," and "audio-reproducing transducer" are used interchangeably to refer to any sound-emitting transducer (or set of transducers). A typical headphone set includes two speakers. A speaker may be implemented to include multiple transducers (e.g., woofers and tweeters) that may be driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may receive different processing in different circuit branches coupled to different transducers.
[0005] Throughout this disclosure, including the claims, the phrase performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to a signal or data) is used broadly to indicate performing an operation on the signal or data directly, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone pre-filtering or pre-processing prior to performing the operation on it).
[0006] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.
[0007] Throughout this disclosure, including the claims, the term "processor" is used broadly to denote a system or device that is programmable or, in some cases, configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipeline processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.
[0008] Throughout this disclosure, including the claims, the terms "couples" or "coupled" are used to mean either a direct connection or an indirect connection. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.
[0009] As used herein, a "smart device" is generally an electronic device configured to communicate with one or more other devices (or networks) via various wireless protocols, such as Bluetooth, Zigbee, near field communication, Wi-Fi, Light Fidelity (Li-Fi), 3G, 4G, 5G, etc., and capable of operating interactively and / or autonomously to some degree. Some notable types of smart devices include smartphones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smart watches, smart bands, smart keychains, and smart audio devices. The term "smart device" may also refer to devices that exhibit some properties of ubiquitous computing, such as artificial intelligence.
[0010] The term "smart audio device" is used herein to refer to a smart device that is either a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of virtual assistant functionality). A single-purpose audio device is a device (e.g., a television (TV)) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera) and is designed largely or primarily to achieve a single purpose. For example, while TVs are typically capable of (and are considered capable of) playing audio from program material, modern TVs most often run some kind of operating system on which applications, including television viewing applications, run locally. In this sense, single-purpose audio devices having speaker(s) and microphone(s) are often configured to run local applications and / or services to directly use the speaker(s) and microphone(s). Some single-purpose audio devices may be configured to be grouped to achieve audio playback across a zone or user-configured area.
[0011] One common type of multipurpose audio device is an audio device that implements at least some aspects of virtual assistant functionality, while other aspects of the virtual assistant functionality may be implemented by one or more other devices, such as one or more servers with which the multipurpose audio device is configured to communicate. Such multipurpose audio devices are sometimes referred to herein as "virtual assistants." A virtual assistant is a device (e.g., a smart speaker or voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera). In some examples, a virtual assistant may be cloud-enabled in some sense or otherwise provide the ability to utilize multiple devices (different from the virtual assistant) for applications not fully implemented in or on the virtual assistant itself. In other words, at least some aspects of the virtual assistant functionality, e.g., speech recognition functionality, may be implemented (at least in part) by one or more servers or other devices with which the virtual assistant may communicate over a network such as the Internet. Virtual assistants may also operate together, for example, in a discrete and conditionally defined manner. For example, two or more virtual assistants may work together in the sense that one of them, for example, the one most confident that it has heard the wake word, will respond to the wake word. In some implementations, connected virtual assistants may form a kind of constellation that may be managed by one main application that may be (or may implement) the virtual assistant.
[0012] As used herein, "wake word" is used broadly to refer to any sound (e.g., a word uttered by a human being, or some other sound), and the smart audio device is configured to awaken in response to detecting ("listening") the sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "awakening" indicates that the device enters a state in which it waits for (or is listening for) a sound command. In some instances, what may be referred to herein as a "wake word" may include more than one word, e.g., a phrase.
[0013] As used herein, the term "wake word detector" refers to a device (or software including instructions for configuring a device) configured to continuously search for a match between real-time sound (e.g., speech) features and a trained model. Typically, a wake word event is triggered whenever the wake word detector determines that the probability that a wake word has been detected exceeds a predefined threshold. For example, the threshold may be a predetermined threshold adjusted to provide a reasonable compromise between false acceptance and false rejection rates. Following a wake word event, the device may enter a state (sometimes referred to as an "awake" or "attentive" state) in which it listens for commands and passes received commands to a larger, more computationally intensive recognizer.
[0014] As used herein, the terms "program stream" and "content stream" refer to a collection of one or more audio signals, and in some cases video signals, at least some of which are intended to be listened to together. Examples include selections of music, movie soundtracks, movies, television programs, audio portions of television programs, podcasts, live voice calls, synthesized voice responses from smart assistants, etc. In some cases, a content stream may include multiple versions of at least a portion of an audio signal, e.g., the same dialogue in two or more languages. In such cases, only one version of the audio data or portion thereof (e.g., the version corresponding to a single language) is intended to be played at a time. Summary of the Invention
[0015] At least some aspects of the present disclosure may be implemented via one or more audio processing methods. In some cases, the method(s) may be implemented, at least in part, by a control system and / or via instructions (e.g., software) stored on one or more non-transitory media. Some such methods may include receiving, by the control system, an output signal from one or more microphones in an audio environment. The output signal, in some cases, may include a signal corresponding to a person's current speech. In some such examples, the current speech may be or may include a wake word utterance.
[0016] Some such methods may include determining, by the control system, in response to the output signals, one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for two or more audio devices in the audio environment based at least in part on the audio device location information and the echo management system information. In some examples, the audio processing modifications may include reducing loudspeaker playback levels for one or more loudspeakers in the audio environment. Some such methods may include causing, by the control system, one or more types of audio processing modifications to be applied.
[0017] In some examples, at least one of the audio processing modifications may correspond to an increase in the signal-to-echo ratio. According to some such examples, the echo management system information may include a model of echo management system performance. For example, the model of echo management system performance may include an acoustic echo canceller (AEC) performance matrix. In some examples, the model of echo management system performance may include a measure of expected echo return loss enhancement provided by the echo management system.
[0018] According to some examples, determining one or more types of audio processing modifications may be based at least in part on optimizing a cost function. Alternatively or additionally, in some examples, the one or more types of audio processing modifications may be based at least in part on acoustic models of inter-device echo and intra-device echo. Alternatively or additionally, in some examples, the one or more types of audio processing modifications may be based at least in part on the interaudibility of audio devices in the audio environment, for example, on an interaudibility matrix.
[0019] In some examples, one or more types of audio processing modifications may be based at least in part on an estimated location of the person. In some such examples, the estimated location of the person may be based at least in part on output signals from multiple microphones in the audio environment. According to some such examples, the audio processing modifications may include altering a rendering process to warp the rendering of the audio signal away from the estimated location of the person.
[0020] Alternatively or additionally, in some examples, one or more types of audio processing modifications may be based at least in part on listening intent, which in some such examples may include spatial components, frequency components, or both spatial and frequency components.
[0021] Alternatively or additionally, in some examples, the one or more types of audio processing modifications may be based at least in part on one or more constraints. In some such examples, the one or more constraints may be based at least in part on a perceptual model. Alternatively or additionally, the one or more constraints may be based at least in part on audio content energy preservation, audio spatiality preservation, an audio energy vector, a regularization constraint, or a combination thereof. Alternatively or additionally, some examples may include updating an acoustic model of the audio environment, a model of echo management system performance, or both after applying the one or more types of audio processing modifications.
[0022] In some examples, the one or more types of audio processing modifications may include spectral modifications, hi some such examples, the spectral modifications may include reducing the level of the audio data in a frequency band from 500 Hz to 3 KHz.
[0023] Aspects of some disclosed implementations include a control system configured (e.g., programmed) to perform one or more of the disclosed methods or steps thereof, and a tangible, non-transitory computer-readable medium (e.g., a disk or other tangible storage medium) implementing non-transitory storage of data that stores code (e.g., executable code) for performing one or more of the disclosed methods or steps thereof. For example, some disclosed embodiments may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of various operations on data, including one or more of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system that includes an input device, a memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more of the disclosed methods (or steps thereof) in response to data asserted thereto.
[0024] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices, such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure may be implemented in software-stored non-transitory media.
[0025] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Note that the relative dimensions of the following figures may not be drawn to scale. [Brief explanation of the drawings]
[0026] [Figure 1A] FIG. 1 is a block diagram illustrating example components of an apparatus capable of implementing various aspects of the present disclosure. [Figure 1B] An example of an audio environment is shown below. [Figure 2] 1C illustrates the echo paths between three of the audio devices of FIG. 1B. [Figure 3] FIG. 1 is a system block diagram illustrating components of an audio device according to an example. [Figure 4] 1 illustrates elements of a ducking module according to an example. [Figure 5] FIG. 1 is a block diagram illustrating an example of an audio device including a ducking module. [Figure 6] FIG. 10 is a block diagram illustrating an alternative example of an audio device including a ducking module. [Figure 7] FIG. 1 is a flow diagram outlining an example of a method for determining a ducking solution. [Figure 8] FIG. 10 is a flow diagram outlining another example of a method for determining a ducking solution. [Figure 9] FIG. 1 is a flow diagram outlining an example of the disclosed method. [Figure 10] 1B is a flow diagram outlining an example of a method that may be performed by an apparatus such as that shown in FIG. 1A. [Figure 11] FIG. 2 is a block diagram of elements of an example embodiment configured to implement a zone classifier. [Figure 12] 1B is a flow diagram outlining an example of a method that may be performed by an apparatus such as apparatus 150 of FIG. 1A. [Figure 13] 1B is a flow diagram outlining another example of a method that may be performed by an apparatus such as apparatus 150 of FIG. 1A. [Figure 14] 1B is a flow diagram outlining another example of a method that may be performed by an apparatus such as apparatus 150 of FIG. 1A. [Figure 15]FIG. 1 illustrates an exemplary set of speaker activation and object rendering positions. [Figure 16] FIG. 1 illustrates an exemplary set of speaker activation and object rendering positions. [Figure 17] 1B is a flow diagram outlining an example of a method that may be performed by an apparatus or system such as that shown in FIG. 1A. [Figure 18] 10 is a graph of speaker activation in an exemplary embodiment. [Figure 19] 10 is a graph of object rendering positions in an exemplary embodiment. [Figure 20] 10 is a graph of speaker activation in an exemplary embodiment. [Figure 21] 10 is a graph of object rendering positions in an example embodiment. [Figure 22] 10 is a graph of speaker activation in an exemplary embodiment. [Figure 23] 10 is a graph of object rendering positions in an example embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0027] Some embodiments are configured to implement a system including coordinated audio devices, also referred to herein as orchestrated audio devices. In some implementations, the orchestrated audio devices may include smart audio devices. According to some such implementations, two or more of the smart audio devices may be wake word detectors or may be configured to implement wake word detectors. Thus, in such examples, multiple microphones (e.g., asynchronous microphones) may be present in the audio environment.
[0028] Currently, designers generally view audio devices as a single interface point for audio that may be a mix of entertainment, communication, and information services. Using audio for notifications and voice control has the advantage of avoiding visual or physical intrusion. In all forms of interactive audio, the problem of increasing full-duplex (input and output) audio capabilities remains a challenge. When there is audio output in a room that is not related to in-room transmission or information-based capture, it is desirable to remove this audio from the captured signal (e.g., by echo cancellation and / or echo suppression).
[0029] Some disclosed embodiments provide techniques for managing the listener or “user” experience to improve a key criterion for successful full duplex on one or more audio devices. This criterion is known herein as the signal-to-echo ratio (SER), also referred to as the voice-to-echo ratio, and may be defined as the ratio between the audio signal or other desired signal to be captured in an audio environment (e.g., a room) via one or more microphones and the “echo” presented at the audio device, including signals from the one or more microphones corresponding to output program content, interactive content, etc. being played by one or more loudspeakers in the audio environment. Those skilled in the art will recognize that, in this context, “echo” is not necessarily reflected before being captured by the microphone.
[0030] Such an embodiment may be useful in situations where two or more audio devices are present within acoustic range of a user, each capable of presenting audio program material at an appropriate volume at the user's location for a desired entertainment, communication, or information service. The value of such an embodiment may be particularly high when three or more audio devices are also present near the user. If an audio device is closer to the user, the audio device may be more advantageous in terms of its ability to accurately localize sounds or convey specific audio signaling and imaging to the user. However, if these audio devices include one or more microphones, one or more of these audio devices may also have a preferred microphone system for picking up the user's voice.
[0031] Audio devices often need to respond to a user's voice commands while the audio device is playing content, causing the audio device's microphone system to detect the content played by the audio device. In other words, the audio device hears its own "echo." The special properties of wake word detectors may enable such devices to outperform more common speech recognition engines in the presence of this echo. A common mechanism implemented in these audio devices, commonly referred to as "ducking," involves reducing the audio device's playback level after detecting the wake word so that the audio device can better recognize post-wake word commands uttered by the user. Such ducking generally results in an improvement in SER, a common metric for predicting speech recognition performance.
[0032] In the context of distributed, orchestrated audio devices, where multiple audio devices are located in a single acoustic space (also referred to herein as an "audio environment"), ducking only the playback of a single audio device may not be an optimal solution. This may be true in part because "echoes" (detected audio playback) from other unducked audio devices in the audio environment may limit the maximum achievable SER by ducking only the playback of a single audio device.
[0033] Thus, some disclosed embodiments may cause audio processing changes for two or more audio devices of an audio environment to increase the SER at one or more microphones of the audio environment. In some examples, the audio processing change(s) may be determined according to the results of an optimization process. According to some examples, the optimization process may include trading off objective sound capture performance goals against constraints that preserve one or more aspects of a user's listening experience. In some examples, the constraints may be perceptual constraints, objective constraints, or a combination thereof. Some disclosed examples include implementing a model that describes the echo management signal chain, sometimes referred to herein as a "capture stack," the acoustic space, and the perceptual impact of the audio processing change(s), and explicitly trading them off (e.g., finding a solution that takes all such factors into account).
[0034] According to some examples, the process may include, for example, a closed-loop system in which the acoustic and capture stack models are updated after each audio processing change (such as each change of one or more rendering parameters). Some such examples may include iteratively improving the performance of the audio system over time.
[0035] Some disclosed implementations may be based at least in part on one or more of the following factors, or a combination thereof: A model of the acoustic environment (in other words, the acoustics of the audio environment) and the echo management signal chain that can predict the SER achieved for a given configuration and output solution; Constraints that limit the output solution according to both objective and subjective metrics, including: - Content energy conservation; - Spatiality preservation; - energy vector; and / or - Regularization such as: Level 1 regularization (L1), Level 2 regularization (L2), or any level of regularization (LN) distortion of two-dimensional or three-dimensional arrays of loudspeaker activation, sometimes referred to herein as "waffle"; and / or ◇ L1, L2 or LN regularization of ducking gain; · Listening objectives that may determine both the spatial and level components of the target solution.
[0036] According to some examples, the ducking solution may be based at least in part on one or more of the following factors, or a combination thereof: In some cases, a simple gain may be applied to already rendered audio content; These gains may be full-band or frequency dependent, depending on the particular implementation; Input to a renderer that the renderer may use in combination with Waffle in some instances; and / or Inputs to a waffle-generating device, module, etc., sometimes referred to herein as a waffle maker, which in some instances may be a component of a renderer. The waffle maker, in some instances, may use such inputs to generate ducked waffles.
[0037] In some implementations, such input to the waffle maker and / or renderer may be used to modify audio playback such that audio objects may appear to be "moved away" from the location where the wake word was detected. Some such implementations may include determining relative activation of a set of loudspeakers in an audio environment by optimizing a cost that is a function of (a) a model of the perceived spatial location of the audio signal as played when played on a set of loudspeakers in the audio environment, (b) a measure of proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker in the set of loudspeakers, and (c) one or more additional dynamically configurable functions. Some such implementations are described in more detail below.
[0038] FIG. 1A is a block diagram illustrating example components of a device capable of implementing various aspects of the present disclosure. As with other figures provided herein, the types and number of elements shown in FIG. 1A are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 150 may be configured to perform at least some of the methods disclosed herein. In some implementations, device 150 may be or include one or more components of an audio system. For example, device 150 may be an audio device, such as a smart audio device, in some implementations. In other examples, device 150 may be a mobile device (such as a mobile phone), a laptop computer, a tablet device, a television, or another type of device.
[0039] According to some alternative implementations, apparatus 150 may be or include a server. In some such examples, apparatus 150 may be or include an encoder. Thus, in some instances, apparatus 150 may be a device configured for use within an audio environment, such as a home audio environment, while in other instances, apparatus 150 may be a device configured for use in the “cloud,” e.g., a server.
[0040] In this example, device 150 includes interface system 155 and control system 160. Interface system 155, in some implementations, may be configured for communication with one or more other devices in an audio environment. The audio environment, in some examples, may be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. Interface system 155, in some implementations, may be configured to exchange control information and associated data with audio devices in the audio environment. The control information and associated data, in some examples, may relate to one or more software applications that device 150 is executing.
[0041] The interface system 155, in some implementations, may be configured to receive or provide a content stream. The content stream may include audio data. The audio data may include, but is not limited to, an audio signal. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. The metadata may be provided, for example, by what may be referred to herein as an “encoder.” In some examples, the content stream may include video data and audio data corresponding to the video data.
[0042] Interface system 155 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, interface system 155 may include one or more wireless interfaces. Interface system 155 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, interface system 155 may include one or more interfaces between control system 160 and a memory system, such as optional memory system 165 shown in FIG. 1A . However, in some cases, control system 160 may include the memory system. Interface system 155, in some implementations, may be configured to receive input from one or more microphones in the environment.
[0043] Control system 160 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0044] In some implementations, control system 160 may reside in more than one device. For example, in some implementations, a portion of control system 160 may reside in a device within one of the environments shown herein, and another portion of control system 160 may reside in a device outside the environment, such as a server, a mobile device (e.g., a smartphone, or a tablet computer). In other examples, a portion of control system 160 may reside in a device within one of the environments shown herein, and another portion of control system 160 may reside in one or more other devices in the environment. For example, control system functionality may be distributed across multiple smart audio devices in the environment, or may be shared by an orchestration device (such as what may be referred to herein as a smart home hub) and one or more other devices in the environment. In other examples, a portion of control system 160 may reside in a device implementing a cloud-based service, such as a server, and another portion of control system 160 may reside in another device implementing the cloud-based service, such as another server, a memory device, or the like. Interface system 155 may also reside in more than one device in some examples.
[0045] In some implementations, control system 160 may be configured to at least partially perform the methods disclosed herein. According to some examples, control system 160 may be configured to determine and cause audio processing changes for two or more audio devices of the audio environment to increase SER at one or more microphones of the audio environment. In some examples, the audio processing change(s) may be based at least in part on audio device location information and echo management system information. According to some examples, the audio processing change(s) may be responsive to microphone output signals corresponding to a person's current speech, such as the utterance of a wake word. In some examples, the audio processing change(s) may be determined according to the results of an optimization process. According to some examples, the optimization process may include trading off objective sound capture performance goals against constraints that preserve one or more aspects of the user's listening experience. In some examples, the constraints may be perceptual constraints, objective constraints, or a combination thereof.
[0046] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices, such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may reside, for example, in optional memory system 165 and / or control system 160 shown in FIG. 1A. Accordingly, various innovative aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media having software stored thereon. The software may include, for example, instructions for controlling at least one device to perform some or all of the methods disclosed herein. The software may be executable by one or more components of a control system, such as, for example, control system 160 of FIG. 1A.
[0047] In some examples, device 150 may include optional microphone system 170 shown in FIG. 1A . Optional microphone system 170 may include one or more microphones. According to some examples, optional microphone system 170 may include an array of microphones. In some examples, the array of microphones may be configured to determine direction of arrival (DOA) and / or time of arrival (TOA) information, for example, according to instructions from control system 160. The array of microphones may, in some cases, be configured for receive-side beamforming, for example, according to instructions from control system 160. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker of a speaker system, a smart audio device, or the like. In some examples, device 150 may not include microphone system 170. However, in some such implementations, device 150 may nevertheless be configured to receive microphone data for one or more microphones in the audio environment via interface system 160. In some such implementations, a cloud-based implementation of device 150 may be configured to receive microphone data, or data corresponding to microphone data, from one or more microphones in the audio environment via interface system 160.
[0048] According to some implementations, device 150 may include optional loudspeaker system 175 shown in FIG. 1A. Optional loudspeaker system 175 may include one or more loudspeakers, sometimes referred to herein as "speakers," or more generally, "audio reproduction transducers." In some examples (e.g., cloud-based implementations), device 150 may not include loudspeaker system 175.
[0049] In some implementations, device 150 may include optional sensor system 180 shown in FIG. 1A . Optional sensor system 180 may include one or more touch sensors, gesture sensors, motion detectors, etc. According to some implementations, optional sensor system 180 may include one or more cameras. In some implementations, the camera may be a freestanding camera. In some examples, one or more cameras of optional sensor system 180 may reside within a smart audio device, which may be configured to at least partially implement a virtual assistant in some examples. In some such examples, one or more cameras of optional sensor system 180 may reside within a television, a mobile phone, or a smart speaker. In some examples, device 150 may not include sensor system 180. However, in some such implementations, device 150 may nevertheless be configured to receive sensor data for one or more sensors in the audio environment via interface system 160.
[0050] In some implementations, device 150 may include optional display system 185 shown in FIG. 1A . Optional display system 185 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some cases, optional display system 185 may include one or more organic light-emitting diode (OLED) displays. In some examples, optional display system 185 may include one or more displays of a smart audio device. In other examples, optional display system 185 may include a television display, a laptop display, a mobile device display, or another type of display. In some examples in which device 150 includes display system 185, sensor system 180 may include a touch sensor system and / or a gesture sensor system proximate one or more displays of display system 185. According to some such implementations, control system 160 may be configured to control display system 185 to present one or more graphical user interfaces (GUIs).
[0051] According to some such examples, device 150 may be or include a smart audio device. In some such implementations, device 150 may be or include a wake word detector. For example, device 150 may be or include a virtual assistant.
[0052] Figure 1B shows an example of an audio environment. As with other figures provided herein, the types, numbers, and arrangements of elements shown in Figure 1B are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements, different arrangements of elements, etc.
[0053] According to this example, audio environment 100 includes audio devices 110A, 110B, 110C, 110D, and 110E. Audio devices 110A-110E may, in some examples, be instances of apparatus 150 of FIG. 1A. In this example, each audio device 110A-110E includes at least one respective microphone 120A, 120B, 120C, 120D, and 120E and at least one respective loudspeaker 121A, 121B, 121C, 121D, and 121E. In this example, individual instances of microphones 120A-120E and loudspeakers 121A-121E are shown. However, one or more of audio devices 110A-110E may include a microphone system including multiple microphones and / or a loudspeaker system including multiple loudspeakers. According to some examples, each of audio devices 110A-110E may be a smart audio device, such as a smart speaker.
[0054] In some examples, some or all of audio devices 110A-110E may be orchestrated audio devices that operate (at least in part) according to instructions from an orchestration device. According to some such examples, the orchestration device may be one of audio devices 110A-110E. In other examples, the orchestration device may be another device, such as a smart home hub.
[0055] In this case, persons 101A and 101B are in an audio environment. In this example, an acoustic event is caused by speaker 101A speaking near audio device 110A. Element 102 is intended to represent the speech of person 101A. In this example, speech 102 corresponds to the speech of a wake word by person 101A.
[0056] Figure 2 shows the echo paths between three of the audio devices of Figure 1B. Elements of Figure 2 that are not described with reference to Figure 1B are as follows: 200AA: echo path from device 110A to device 110A (from loudspeaker 121A to microphone 120A); 200AB: echo path from device 110A to device 110B (loudspeaker 121A to microphone 120B); 200AC: echo path from device 110A to device 110C (loudspeaker 121A to microphone 120C); 200BA: echo path from device 110B to device 110A (from loudspeaker 121B to microphone 120A); 200BB: echo path from device 110B to device 110B (from loudspeaker 121B to microphone 120B); 200BC: echo path from device 110B to device 110C (from loudspeaker 121B to microphone 120C); 200CA: echo path from device 110C to device 110A (from loudspeaker 121C to microphone 120A); 200CB: echo path from device 110C to device 110C (from loudspeaker 121C to microphone 120B); 200CC: Echo path from device 110C to device 110C (from loudspeaker 121C to microphone 120C).
[0057] These echo paths indicate the effect that each audio device's reproduced audio, or "echo," has on other audio devices. This effect is sometimes referred to herein as the "mutual audibility" of the audio devices. Mutual audibility depends on various factors, including the position and orientation of each audio device within the audio environment, the playback level of each audio device, the loudspeaker capabilities of each audio device, etc. Some implementations may include constructing a more accurate representation of the mutual audibility of audio devices within an audio environment, such as an audibility matrix A, which represents the energy of echo paths 200AA-200CC. For example, each column of the audibility matrix may represent an audio device loudspeaker, and each row of the audibility matrix may represent an audio device microphone, or vice versa. In some such audibility matrices, the diagonal of the audibility matrix may represent the echo path from an audio device's loudspeaker(s) to the same audio device's microphone(s).
[0058] 2, it can be seen that if echo paths 200CA and 200BA from audio devices 110C and 110B, respectively, to audio device 100A have strong coupling, the echoes from audio devices 110C and 110B will become significant in the echo (residual) of audio device 110A. In such a case, if the audio system's response to detecting voice 102 (which corresponds to the wake word in this example) is simply to lower the volume of the nearest loudspeaker, this would involve ducking only playback from loudspeaker(s) 121A. If this were the audio system's only response to detecting the wake word, potential SER improvement could be significantly limited. Accordingly, some disclosed examples may include other responses to detecting the wake word in such situations.
[0059] 3 is a system block diagram illustrating components of an audio device according to an example. In FIG. 3, the block representing audio device 110A includes loudspeaker 121A and microphone 120A. In some examples, loudspeaker 121A may be one of multiple loudspeakers in a loudspeaker system, such as loudspeaker system 175 of FIG. 1A. Similarly, according to some implementations, microphone 120A may be one of multiple microphones in a microphone system, such as microphone system 170 of FIG. 1A.
[0060] In this example, audio device 110A includes renderer 201A, echo management system (EMS) 203A, and audio processor / communications block 240A. In this example, EMS 203A may be or include an acoustic echo canceller (AEC), an acoustic echo suppressor (AES), or both an AEC and an AES. According to this example, renderer 201A is configured to render audio data 301 received by or stored in audio device 110A for playback on loudspeaker 121A. In some examples, the audio data may include one or more audio signals and associated spatial data. The spatial data may indicate, for example, an intended perceived spatial location corresponding to an audio signal. In some examples, the spatial data may be or include spatial metadata corresponding to an audio object. In this example, the renderer output 220A is provided to the loudspeaker 121A for playback, and the renderer output 220A is also provided to the EMS 203A as a reference for echo cancellation.
[0061] In addition to receiving renderer output 220A, in this example EMS 203A also receives microphone signal 223A from microphone 120A. In this example, EMS 203A processes microphone signal 223A and provides an echo-canceled residual 224A (sometimes referred to herein as "residual output 224A") to audio processor / communications block 240A.
[0062] In some implementations, the voice processor / communications block 240A may be configured for voice recognition functions. In some examples, the voice processor / communications block 240A may be configured to provide telecommunications services such as telephone calls, video conferencing, etc. Although not shown in FIG. 3 , the voice processor / communications block 240A may be configured to communicate with one or more networks, loudspeaker 121A, and / or microphone 120A, for example, via an interface system. The one or more networks may include, for example, a local Wi-Fi network, one or more types of telephone networks, etc.
[0063] 4 shows elements of a ducking module according to one example. In this implementation, the ducking module 400 is implemented by an instance of the control system 160 of FIG. 1A. In this example, the elements of FIG. 4 are as follows: 401: Acoustic models of inter-device and intra-device echoes, including, in some examples, an acoustic model of user speech. According to some examples, the acoustic model 401 may be or include a model of how playback from each audio device in an audio environment appears as echo detected by the microphones of all audio devices (its own and others) in the audio environment. In some examples, the acoustic model 401 may be based at least in part on an audio environment impulse response estimate, characteristics of the impulse response such as peak amplitude, decay time, etc. In some examples, the acoustic model 401 may be based at least in part on an audibility estimate. In some such examples, the audibility estimate may be based on microphone measurements. Alternatively or additionally, the audibility estimate may be inferred depending on audio device position, e.g., based on the fact that echo power is inversely proportional to the distance between audio devices. In some examples, the acoustic model 401 may be based at least in part on long-term estimates of AEC / AES filter taps. In some examples, the acoustic model 401 may be based at least in part on waffles, which may include information about the capabilities (loudness) of the loudspeaker(s) of each audio device; 452: Spatial information including information regarding the position of each of a plurality of audio devices within the audio environment. In some examples, the spatial information 452 may include information regarding the orientation of each of the plurality of audio devices. According to some examples, the spatial information 452 may include information regarding the position of one or more people within the audio environment. In some instances, the spatial information 452 may include information regarding the impulse response of at least a portion of the audio environment; 402: A model of EMS performance, which may represent the performance of the EMS 203A of FIG. 3 or another EMS (such as that of FIG. 5 or 6). In this example, the EMS performance model 402 predicts how well an EMS (AEC, AES, or both) will perform. In some examples, the EMS performance model 402 may predict how well an EMS will perform given the current audio environment impulse response, the current noise level(s) of the audio environment, the type of algorithm(s) being used to implement the EMS, the type of content being played in the audio environment, the number of echo criteria being fed to the EMS algorithm(s), the capabilities / quality of the loudspeakers in the audio environment (nonlinearities in loudspeakers may impose an upper limit on expected performance), or a combination thereof. According to some examples, the EMS performance model 402 may be based at least in part on empirical observations, for example, by observing how an EMS performs under various conditions, storing data points based on such observations, and building a model based on these data points (e.g., by curve fitting). In some examples, the EMS performance model 402 may be based at least in part on machine learning, such as by training a neural network based on empirical observations of EMS performance. Alternatively or additionally, the EMS performance model 402 may be based at least in part on a theoretical analysis of the algorithm(s) used by the EMS. In some examples, the EMS performance model 402 may indicate the ERLE (Echo Return Loss Enhancement) caused by the operation of the EMS, which is a useful metric for evaluating EMS performance. The ERLE may indicate, for example, the amount of additional signal loss applied by the EMS between each audio device; 403: Information about one or more current listening objectives. In some examples, the listening objective information 403 may set a goal, such as a SER goal or a SER improvement goal, for the ducking module to achieve. According to some examples, the listening objective information 403 may include both a spatial component and a level component; · 450: Target-related factors that may be used to determine the target, such as external triggers, acoustic events, mode indications, etc.; 404: One or more constraints to be applied during the process of determining a ducking solution, such as a constraint that trades off improved listening performance (in other words, improved ability of one or more microphones to capture audio, such as human speech) against other metrics (such as degradation of the listening experience of people in the audio environment). For example, in one example, a constraint may prevent the ducking module 400 from reducing loudspeaker playback levels for some or all loudspeakers in the audio environment to an unacceptably low level, such as 0 dBFS (decibels relative to full scale); 451: Metadata about the current audio content, which may include spatiality metadata, level metadata, content type metadata, etc. Such metadata may provide information (directly or indirectly) about the impact that ducking one or more loudspeakers will have on the listening experience of a person in the audio environment. For example, if the spatiality metadata indicates that a "loud" audio object is being played by multiple loudspeakers in the audio environment, ducking one of those loudspeakers may not have an adverse impact on the listening experience. As another example, if the content metadata indicates that the content is a podcast, in some instances a monologue or dialogue in the podcast may be played by multiple loudspeakers in the audio environment, and therefore ducking one of those loudspeakers may not have an adverse impact on the listening experience. However, if the content metadata indicates that the audio content corresponds to a movie or television program, the dialogue of such content may be reproduced primarily or entirely through particular loudspeakers (e.g., "front" loudspeakers), and ducking those loudspeakers may have an undesirable effect on the listening experience; 405: The model used to derive the perceptually driven constraints. Some detailed examples are given elsewhere in this specification; 406: An optimization algorithm, which may vary depending on the particular implementation. In some examples, the optimization algorithm 406 may be or include a closed-form optimization algorithm. In some cases, the optimization algorithm 406 may be or include an iterative process. Some detailed examples are described elsewhere herein; 480: Ducking solution output by ducking module 400. As described in more detail elsewhere herein (e.g., with reference to FIGS. 5 and 6), ducking solution 480 may differ according to various factors, including whether ducking solution 480 is provided to a renderer or whether ducking solution 480 is provided for audio data output from a renderer.
[0064] According to some examples, the AEC model 402 and the acoustic model 401 may provide a means for estimating or predicting the SER, for example, as described below.
[0065] The constraint(s) 404 and the perceptual model 405 may be used to ensure that the ducking solution 480 output by the ducking module 400 is neither degenerate nor trivial. An example of a trivial solution is to globally set the playback level to 0. The constraint(s) 404 may be perceptual and / or objective. According to some examples, the constraint(s) 404 may be based at least in part on a perceptual model, such as a model of human hearing. In some examples, the constraint(s) 404 may be based at least in part on audio content energy preservation, audio spatiality preservation, an audio energy vector, or one or more combinations thereof. According to some examples, the constraint(s) 404 may be or include a regularization constraint. The listening intent information 403 may determine, for example, a current target SER improvement to be made by distributed ducking (in other words, ducking two or more audio devices in an audio environment).
[0066] Global Optimization In some examples, the selection of audio device(s) for ducking involves using estimated SER and / or wake word information obtained when the wake word is detected to select the audio device to listen for the next utterance. If this audio device selection is incorrect, it is highly unlikely that the best listening device will be able to understand the command spoken after the wake word. This is because automatic speech recognition (ASR) is more difficult than wake word detection (WWD), which is one of the motivating factors for ducking. If the best listening device is not ducked, ASR is likely to fail for all of the audio devices. Therefore, in some such examples, the ducking method includes optimizing the ASR stage by using the previous estimate (from WWD) to duck the closest (or best estimated) audio device(s).
[0067] Some ducking implementations therefore include using previous estimates when determining a ducking solution. However, in implementations such as that shown in FIG. 4, listening objectives and constraints can be applied to achieve more robust ASR performance. In some such examples, the ducking method may include configuring the ducking algorithm so that the SER improvement is significant at all potential user locations in the acoustic space. In this way, it can be ensured that at least one of the microphones in the room has sufficient SER for robust ASR performance. Such implementations may be advantageous when the talker location is unknown or there is uncertainty regarding the talker location. Some such examples may include accounting for variance in talker and / or audio device position estimates by spatially expanding the SER improvement zone to account for one or more uncertainties.
[0068] Some such examples may include the use of the δ parameter discussed below, or a similar parameter. Other examples may include multi-parameter models that describe or accommodate uncertainty in talker position and / or audio device position estimates.
[0069] In some embodiments, the ducking method may be performed in the context of one or more user zones, a set of acoustic features, as described in more detail later in this document,
number
[0070] In some examples, the ducking method may use quantities related to one or more user zones in the process of selecting an audio device for ducking or other audio processing modification. If both z and p are available, an exemplary audio device selection decision may be made according to the following formula:
number
[0071] In some implementations, the acoustic features W(j) may be used directly in the ducking method. For example, if the wake word confidence score associated with utterance j is w(j), then n If (j), then audio device selection may be performed according to the following formula:
number
[0072] 5 is a block diagram illustrating an example of an audio device including a ducking module. Renderer 201A, audio processor / communications block 240A, EMS 203A, loudspeaker(s) 121A, and microphone(s) 120A may function substantially as described with reference to FIG. 3, except as described below. In this example, renderer 201A, audio processor / communications block 240A, EMS 203A, and ducking module 400 are implemented by an instance of control system 160 described with reference to FIG. 1A. Ducking module 400 may be, for example, an instance of ducking module 400 described with reference to FIG. 4. Thus, ducking module 400 may be configured to determine one or more types of audio processing modifications (indicated by or corresponding to ducking solution 480) to apply to rendered audio data (e.g., audio data rendered into loudspeaker feed signals) for at least audio device 110A. The audio processing modification may be or may include a reduction in loudspeaker playback levels of one or more loudspeakers in the audio environment. In some examples, the ducking module 400 may be configured to determine one or more types of audio processing modification to apply to rendered audio data for two or more audio devices in the audio environment.
[0073] In this example, the renderer output 220A and the ducking solution 480 are provided to a gain multiplier 501. In some examples, the ducking solution 480 includes a gain for the gain multiplier 501 to apply to the renderer output 220A to generate processed audio data 502. According to this example, the processed audio data 502 is provided to the EMS 203A as a local reference for echo cancellation, where the processed audio data 502 is also provided to the loudspeaker(s) 121A for playback.
[0074] In some examples, ducking module 400 may be configured to determine ducking solution 480 as described below with reference to Figure 8. According to some examples, ducking module 400 may be configured to determine ducking solution 480 as described in the "Optimizing for a Specific Device" section below.
[0075] 6 is a block diagram illustrating an alternative example of an audio device including a ducking module. Renderer 201A, audio processor / communications block 240A, EMS 203A, loudspeaker(s) 121A, and microphone(s) 120A may function substantially as described with reference to FIG. 3, except as described below. Ducking module 400 may be, for example, an instance of ducking module 400 described with reference to FIG. 4. In this example, renderer 201A, audio processor / communications block 240A, EMS 203A, and ducking module 400 are implemented by an instance of control system 160 described with reference to FIG. 1A.
[0076] According to this example, ducking module 400 is configured to provide ducking solution 480 to renderer 201A. In some such examples, ducking solution 480 may cause renderer 201A to implement one or more types of audio processing modifications, which may include reducing loudspeaker playback levels, during the process of rendering received audio data 301 (or during the process of rendering audio data stored in memory of audio device 110A). In some examples, ducking module 400 may be configured to determine ducking solution 480 for implementation of one or more types of audio processing modifications via one or more instances of renderer 201A in one or more other audio devices in the audio environment.
[0077] In this example, the renderer output 201A outputs processed audio data 502. According to this example, the processed audio data 502 is provided to the EMS 203A as a local reference for echo cancellation, where the processed audio data 502 is also provided to the loudspeaker(s) 121A for playback.
[0078] In some examples, the ducking solution 480 may include one or more penalties implemented by a flexible rendering algorithm, for example, as described below. In some such examples, the penalty may be a loudspeaker penalty estimated to cause the desired SER improvement. According to some examples, determining one or more types of audio processing modifications may be based on optimizing a cost function by the ducking module 400 or by the renderer 201A.
[0079] FIG. 7 is a flow diagram outlining an example of a method for determining a ducking solution. In some examples, method 720 may be performed by an apparatus such as that shown in FIG. 1A, FIG. 5, or FIG. 6. In some examples, method 720 may be performed by a control system of an orchestration device, which may, in some cases, be an audio device. In some examples, method 720 may be performed, at least in part, by a ducking module, such as ducking module 400 of FIG. 4, FIG. 5, or FIG. 6. According to some examples, method 720 may be performed, at least in part, by a renderer. The blocks of method 720, as well as other methods described herein, are not necessarily performed in the order presented. Furthermore, such methods may include more or fewer blocks than those illustrated and / or described.
[0080] In this example, the process begins at block 725. In some cases, block 725 may correspond to a boot-up process or a time when the boot-up process is complete and the device configured to perform method 720 is ready to function.
[0081] According to this example, block 730 includes waiting for a wake word to be detected. If method 720 is being performed by an audio device, block 730 may also include playing rendered audio data corresponding to received or stored audio content, such as music content, a podcast, an audio soundtrack for a movie or television program, etc.
[0082] In this example, upon detecting the wake word (e.g., at block 730), the SER of the wake word is estimated at block 735. S(a) is an estimate of the speech-to-echo ratio at device a. By definition, the speech-to-echo ratio in dB is given by:
number
[0083] In the above formula,
number
number
number
number
number
number
[0084] According to this example, block 740 includes obtaining a target SER (in this case from block 745) and calculating a target SER improvement. In some implementations, the desired SER improvement (SERI) may be determined as follows:
number
[0085] In the above equation, m represents the device / microphone location for which SER is being improved, and TargetSER represents a threshold that, in some examples, may be set according to the application in use. For example, a wake word detection algorithm may tolerate a lower operating SER than a command detection algorithm, which may tolerate a lower operating SER than a large vocabulary speech recognizer. Typical values for TargetSER are on the order of -6 dB to 12 dB. In some cases, if S(m) is not known or not easily estimated, a preset value based on offline measurements of recorded speech and echo in a typical echoic room or setting may be sufficient. Some embodiments may determine the audio device for which audio processing (e.g., rendering) should be modified by specifying f_n ranging from 0 to 1. Other embodiments may use a speech-to-echo ratio improvement in decibels, s, potentially calculated according to: n (also denoted herein as s_n) may include specifying the degree to which audio processing (e.g., rendering) should be modified:
number
[0086] Some embodiments may calculate f_n directly from the device geometry, for example, as follows:
number
[0087] In the above equation, m represents the index of the device to be selected for maximum audio processing (e.g., rendering) modification, and H(m,i) represents the approximate physical distance between device m and device i. Other implementations may include other choices of relaxation or smoothing functions over device geometries.
[0088] Thus, H is a property of the physical location of an audio device within an audio environment. H may be determined or estimated according to various methods, depending on the particular implementation. Various examples of methods for estimating the location of an audio device within an audio environment are described below.
[0089] In this example, block 750 includes calculating what may be referred to herein as a "ducking solution." The ducking solution may include determining a reduction in loudspeaker playback levels for one or more loudspeakers in the audio environment, although the ducking solution may also include one or more other audio processing modifications, such as those disclosed herein. The ducking solution determined in block 750 is an example of ducking solution 480 of FIGS. 4, 5, and 6. Thus, block 750 may be performed by ducking module 400.
[0090] According to this example, the ducking solution determined in block 750 is based (at least in part) on the target SER, the ducking constraints (represented by block 755), the AEC model (represented by block 765), and the acoustic model (represented by block 760). The acoustic model may be an instance of the acoustic model 401 of FIG. 4. The acoustic model may be based at least in part on inter-device audibility, for example, sometimes referred to herein as inter-device echo or interaudibility. The acoustic model, in some examples, may be based at least in part on intra-device echo. In some instances, the acoustic model may be based at least in part on an acoustic model of a user's speech, such as the acoustic characteristics of typical human speech, the acoustic characteristics of human speech previously detected in the audio environment, etc. The AEC model, in some examples, may be an instance of the AEC model 402 of FIG. 4. The AEC model may be indicative of the performance of the EMS 203A of FIG. 5 or FIG. 6. In some examples, the EMS performance model 402 may be indicative of actual or expected ERLE (echo return loss enhancement) caused by the operation of the AEC. The ERLE may indicate, for example, the amount of additional signal loss applied by the AEC between each audio device. According to some examples, the EMS performance model 402 may be based at least in part on an expected ERLE for a given number of echo criteria. In some examples, the EMS performance model 402 may be based at least in part on an estimated ERLE calculated from actual microphone and residual signals.
[0091] In some examples, the ducking solution determined in block 750 may be an iterative solution, while in other examples, the ducking solution may be a closed-form solution. Examples of both iterative and closed-form solutions are disclosed herein.
[0092] In this example, block 770 includes applying the ducking solution determined in block 750. In some examples, the ducking solution may be applied to rendered audio data, as shown in Figure 5. In other examples, the ducking solution determined in block 750 may be provided to a renderer, as shown in Figure 6. The ducking solution may be applied as part of the process of rendering audio data input to the renderer.
[0093] In this example, block 775 includes detecting another utterance, which in some examples may be a command uttered after the wake word. According to this example, block 780 includes estimating the SER of the utterance detected in block 775. According to this example, block 785 includes updating the AEC model and the acoustic model based at least in part on the SER estimated in block 780. According to this example, the process of block 785 occurs after a ducking solution is applied. In a perfect system, the actual SER improvement and the actual SER would be exactly what was targeted. In real-world systems, the actual SER improvement and the actual SER are likely to differ from what was targeted. In such an implementation, method 720 includes using at least the SER to update the information and / or models used to calculate the ducking solution. According to this example, the ducking solution is based at least in part on the acoustic model of block 760. For example, the acoustic model of block 760 may have indicated very strong acoustic coupling between audio device X and microphone Y, and as a result, the ducking solution may have included significant ducking of the signal from microphone Y. However, after estimating the SER of the detected speech in block 775 while the ducking solution was applied, the control system may have determined that the actual SER and / or SERI were not as expected. If so, block 785 may include updating the acoustic model accordingly (in this example, by reducing the acoustic coupling estimate between audio device X and microphone Y). According to this example, the process then returns to block 730.
[0094] FIG. 8 is a flow diagram outlining another example of a method for determining a ducking solution. In some examples, method 800 may be performed by an apparatus such as that shown in FIG. 1A, FIG. 5, or FIG. 6. In some examples, method 800 may be performed by a control system of an orchestration device, which may, in some cases, be an audio device. In some examples, method 800 may be performed, at least in part, by a ducking module, such as ducking module 400 of FIG. 4, FIG. 5, or FIG. 6. According to some examples, method 800 may be performed, at least in part, by a renderer. The blocks of method 800, as well as other methods described herein, are not necessarily performed in the order presented. Furthermore, such methods may include more or fewer blocks than those illustrated and / or described.
[0095] In this example, the process begins at block 805. In some cases, block 805 may correspond to a boot-up process or a time when the boot-up process is complete and a device configured to perform method 800 is ready to function.
[0096] According to this example, block 810 includes estimating a current echo level without application of a ducking solution. In this example, block 810 includes estimating the current echo level based (at least in part) on an acoustic model (represented by block 815) and an AEC model (represented by block 820). Block 810 may include estimating a current echo level resulting from a current ducking candidate solution. The estimated current echo level may, in some examples, be combined with a current speech level to generate a current estimated SER improvement.
[0097] The acoustic model may be an instance of acoustic model 401 of FIG. 4. The acoustic model may be based, for example, at least in part on inter-device audibility, sometimes referred to herein as inter-device echo or inter-audibility. The acoustic model, in some examples, may be based at least in part on intra-device echo. In some instances, the acoustic model may be based at least in part on an acoustic model of a user's speech, such as acoustic characteristics of typical human speech, acoustic characteristics of human speech previously detected in the audio environment, etc.
[0098] The AEC model may, in some examples, be an instance of the AEC model 402 of FIG. 4. The AEC model may represent the performance of the EMS 203A of FIG. 5 or FIG. 6. In some examples, the EMS performance model 402 may represent the actual or expected ERLE (Echo Return Loss Enhancement) caused by the operation of the AEC. The ERLE may represent, for example, the amount of additional signal loss applied by the AEC between each audio device. According to some examples, the EMS performance model 402 may be based at least in part on the expected ERLE for a given number of echo criteria. In some examples, the EMS performance model 402 may be based at least in part on estimated ERLE calculated from actual microphone and residual signals.
[0099] According to this example, block 825 includes obtaining a current ducking solution (represented by block 850) and estimating the SER based on applying the current ducking solution. In some examples, the ducking solution may be determined as described with reference to block 750 of FIG. 7.
[0100] In this example, block 830 includes calculating the difference or "error" between the current estimate of the SER improvement and the target SER improvement (represented by block 835). In some alternative examples, block 830 may include calculating the difference between the current estimate of the SER and the target SER.
[0101] Block 840, in this example, includes determining whether the difference or "error" calculated in block 830 is sufficiently small. For example, block 840 may include determining whether the difference calculated in block 830 is equal to or less than a threshold value. The threshold value may be in the range of 0.1 dB to 1.0 dB, such as 0.1 dB, 0.2 dB, 0.3 dB, 0.4 dB, 0.5 dB, 0.6 dB, 0.7 dB, 0.8 dB, 0.9 dB, or 1.0 dB, in some examples. In such an example, if block 840 determines that the difference calculated in block 830 is equal to or less than the threshold value, the process ends at block 845. The current ducking solution may be output and / or applied.
[0102] However, if it is determined in block 840 that the difference calculated in block 830 is not sufficiently small (e.g., not equal to or less than a threshold), then in this example, the process proceeds to block 855. According to this example, the ducking solution is or includes a ducking vector. In this example, block 855 includes calculating gradients of the cost function and constraint functions with respect to the ducking vector. The cost function may correspond to (or describe) the error between the estimated SER improvement and the target SER improvement, for example, as determined in block 830.
[0103] In some implementations, the constraint function may penalize the influence of the ducking vector on one or more objective functions (e.g., an audio energy preservation function), one or more subjective functions (such as one or more perceptually-based functions), or a combination thereof. In some such examples, one or more of the constraints may be based on a perceptual model of human hearing. According to some examples, one or more of the constraints may be based on audio spatiality preservation.
[0104] In some examples, block 855 may include optimizing a cost that is a function of a model of the perceived spatial location of the audio signal as played back on a set of loudspeakers in the environment and a measure of proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker in the set of loudspeakers. In some such examples, the cost may be a function of one or more additional dynamically configurable features. In some such examples, at least one of the one or more additional dynamically configurable features corresponds to echo cancellation performance. According to some such examples, at least one of the one or more additional dynamically configurable features corresponds to interaudibility of loudspeakers in the audio environment. Detailed examples are provided below. However, other implementations may not include these types of cost functions.
[0105] According to this example, block 865 includes updating the current ducking solution using gradients and one or more types of optimizers, such as the following algorithm, stochastic gradient descent, or another known optimizer.
[0106] In this example, block 870 includes evaluating a change in the ducking solution from a previous ducking solution. According to this example, if block 870 determines that the change in the ducking solution from the previous solution is less than a threshold value, the process ends at block 875. According to some examples, the threshold value may be expressed in decibels. According to some such examples, the threshold value may be in a range of 0.1 dB to 1.0 dB, such as 0.1 dB, 0.2 dB, 0.3 dB, 0.4 dB, 0.5 dB, 0.6 dB, 0.7 dB, 0.8 dB, 0.9 dB, or 1.0 dB. In some examples, if block 870 determines that the change in the ducking solution from the previous solution is less than or equal to a threshold value, the process ends.
[0107] However, in this example, if it is determined in block 870 that the change in the ducking solution from the previous solution is greater than or equal to a threshold, the ducking solution of block 850 is updated to become the current ducking solution, and the process returns to block 825. In some examples, method 800 may continue until block 845 or block 875 is reached. According to some examples, method 800 may end if block 845 or block 875 is not reached within a certain time interval or within a number of iterations.
[0108] After method 800 completes, the resulting ducking solution may be applied. In some examples, the ducking solution may be applied to rendered audio data, as shown in FIG. 5. In other examples, the ducking solution determined via method 800 may be provided to a renderer, as shown in FIG. 6. The ducking solution may be applied as part of the process of rendering audio data input to the renderer. However, as noted elsewhere herein, in some implementations, method 800 may be performed at least in part by a renderer. According to some such implementations, the renderer may both determine and apply the ducking solution.
[0109] The following algorithm is one example of obtaining a ducking solution. In some examples, the ducking solution may be, include, or indicate a gain to be applied to the rendered audio data. Thus, in some such examples, the ducking solution may be appropriate for ducking module 400, as shown in FIG. 5.
[0110] The following symbols are defined as follows: A represents the interaudibility matrix (audibility between each audio device); · P represents the nominal (unducked) playback level vector (across the audio devices); D represents a ducking solution vector (across devices), which may correspond to the ducking solution 480 output by the ducking module 400; · C represents the AEC performance matrix, which in this example indicates the ERLE (Echo Return Loss Enhancement), which is the amount of additional signal loss applied by the AEC between each audio device.
[0111] The net echo at the microphone feed of audio device i can be expressed as:
number
[0112] In Equation 1, J represents the number of audio devices in the room, and A j i represents the audibility of audio device j relative to audio device i, and P j represents the playback level of audio device j. Then, taking into account the effect of ducking any of the audio devices, the echo in the microphone feed can be expressed as:
number
[0113] In Equation 2, D j denotes the ducking gain. Considering also a naive model of the AEC, where the nominal ERLE is applied independently to each loudspeaker rendering, the echo power in the residual can be expressed as:
number
[0114] In Equation 3, C j i represents the ability of audio device i to cancel echo from audio device j. In some examples, Cj i is assumed to be either -20dB or 0dB.
[0115] Generating C can be as trivial as setting an entry to be the nominal cancellation performance value when a particular device is performing cancellation for that particular non-local ("far") device entry in the C matrix. More complex models may account for audio environment adaptive noise (any noise in the adaptive filter process) and inter-channel correlation. For example, such a model may predict how the AEC will perform in the future if a given ducking solution is applied. Some such models may be based at least in part on the echo and noise levels of the audio environment.
[0116] In this example, the distributed ducking problem is formulated as an optimization problem that involves minimizing the total echo power in the AEC residual by varying the ducking vector, which can be expressed as:
number
number
number
number
[0117] Iterative Ducking Solution In one example, a gradient-based iterative solution to a distributed ducking optimization problem may take the following form:
number
[0118] In Equation 8, F represents the cost function that describes the distributed ducking problem, and D n denotes the ducking vector at the nth iteration. In this form, D(x)∈[0,1] is constrained. However, adding a regularization term to Equation 8 does not guarantee that D(x)∈[0,1] without adding some heuristics and / or hard constraints. Another approach involves formulating a gradient-based iterative ducking solution as follows:
number
[0119] In some examples, Z∈[0,1]. However, if one wishes to improve the SER at audio device i by reducing the echo power in the microphone feed of audio device i while also maintaining (at least to some extent) the perceptual quality of the audio content being rendered, it may prove sufficient to ensure D≧0 and Z≧0. This means allowing some audio devices to increase their full-band volume in order to maintain, at least to some extent, the quality of the rendered content. Thus, in some examples, Z may be defined as:
number
[0120] In Equation 10, F represents a cost function that describes the distributed ducking problem, and R is a regularization term that aims to maintain the quality of the rendered audio content. In some examples, R can be an energy conservation constraint. The regularization term R can also find a sensible solution for D.
[0121] Considering T as the target SER improvement at audio device i by ducking the audio device in the audio environment, F can be defined as:
number
[0122] In Equation 11, E res,i n represents the echo in the residual for device i at the nth iteration evaluated using Equation 3, and E res,i 0 represents the echo in the residual for device i when D is all ones. However, by defining F as shown in Equation 11, the step size becomes a function of the error and not just T. In some instances, F can be reformulated to remove its dependence on the target SER as follows:
number
[0123] It is potentially advantageous to adjust each element of the ducking vector proportionally to the sensitivity of E to each element. Additionally, it is potentially advantageous to allow the user some control over the step size. With these goals in mind, in some instances, F may be reformulated as follows:
number
[0124] In Equation 13, M scales the individual contribution of each device to the echo in the residual for audio device i. In some examples, M may be expressed as:
number
[0125] In Equation 14,
number
number
[0126] In Equation 15, λ represents the Lagrange multiplier. Equation 15 allows the control system to find a simple solution for D. However, in some instances, a ducking solution that maintains an acceptable listening experience and conserves total energy in the audio environment can be determined by defining the regularization term R as follows:
number
[0127] Ducking Solver Algorithm With the above discussion in mind, the following ducking solver algorithm can be easily understood: The following algorithm is an example of method 800 of FIG.
number
[0128] According to some such examples, THRESH may be 6 dB, 8 dB, 10 dB, 12 dB, 14 dB, and so on.
[0129] FIG. 9 is a flow diagram outlining an example of the disclosed method. In some examples, method 900 may be performed by an apparatus such as that shown in FIG. 1A, FIG. 5, or FIG. 6. In some examples, method 900 may be performed by a control system of an orchestration device, which may, in some cases, be an audio device. In some examples, method 900 may be performed, at least in part, by a ducking module, such as ducking module 400 of FIG. 4, FIG. 5, or FIG. 6. According to some examples, method 900 may be performed, at least in part, by a renderer. The blocks of method 900, as well as other methods described herein, are not necessarily performed in the order presented. Furthermore, such methods may include more or fewer blocks than those illustrated and / or described.
[0130] In this example, block 905 includes receiving, by the control system, an output signal from one or more microphones in the audio environment, where the output signal includes a signal corresponding to a person's current speech. In some instances, the current speech may be or may include a wake word utterance.
[0131] According to this example, block 910 includes determining, by the control system, in response to the output signal, one or more audio processing modifications to apply to audio data being rendered to loudspeaker feed signals for two or more audio devices in the audio environment based at least in part on the audio device location information and the echo management system information. In this example, the audio processing modifications include reducing loudspeaker playback levels for one or more loudspeakers in the audio environment. Thus, the audio processing modifications may include or be indicated by what is referred to herein as a ducking solution. According to some examples, at least one of the one or more types of audio processing modifications may correspond to an increase in signal-to-echo ratio.
[0132] However, in some examples, audio processing modifications may include or involve modifications other than reducing loudspeaker playback levels. For example, audio processing modifications may include shaping the spectrum of the output of one or more loudspeakers, sometimes referred to herein as “spectral modification” or “spectral shaping.” Some such examples may include shaping the spectrum using a substantially linear equalization (EQ) filter designed to produce an output that differs from the spectrum of the audio desired to be detected. In some examples, if the output spectrum is being shaped to detect human voices, the filter may reduce frequencies within a range of approximately 500 to 3 kHz (e.g., plus or minus 5% or 10% at each end of the frequency range). Some examples may include shaping loudness to emphasize low and high frequencies and leave space in the mid-band (e.g., within a range of approximately 500 to 3 kHz).
[0133] Alternatively or additionally, in some examples, audio processing modifications may include modifying output ceilings or peaks to lower peak levels and / or reduce distortion products that may further degrade the performance of any echo cancellation that is part of the overall system creating the achieved SER for audio detection, for example, via a time-domain dynamic range compressor or a multi-band frequency-dependent compressor. Such audio signal modifications can effectively reduce the amplitude of the audio signal and can help limit loudspeaker excursion.
[0134] Alternatively or additionally, in some examples, the audio processing modifications may include the system (e.g., the audio processing manager) spatially steering the audio in a manner that tends to reduce the energy or coupling of the output of one or more loudspeakers to one or more microphones enabling higher SER. Some such implementations may include the "warping" examples described herein.
[0135] Alternatively or additionally, in some examples, audio processing modifications may include energy conservation and / or creating continuity across a specific or broad set of listening locations. In some examples, energy removed from one loudspeaker may be compensated for by providing additional energy to another loudspeaker. In some cases, the overall loudness may remain the same, or essentially the same. While this is not a required feature, it may be an effective means of enabling more severe changes to the audio processing of the "closest" device, or set of closest devices, without loss of content. However, continuity and / or energy conservation may be particularly important when dealing with complex audio outputs and audio scenes.
[0136] Alternatively or additionally, in some examples, audio processing changes may include an activation time constant. For example, changes to audio processing may be applied a little faster (e.g., 100-200 ms) than they are restored to their normal state (e.g., 1000-10000 ms), so that the change(s) in audio processing appear intentional if noticeable, but the subsequent restoration from the change(s) may not appear to be related to any actual event or change (from the user's perspective) and, in some cases, may be so slow as to be barely noticeable.
[0137] In this example, block 915 includes causing the control system to apply one or more types of audio processing modifications. In some examples, such as shown in FIG. 5, the audio processing modifications may be applied to the rendered audio data according to the ducking solution 480 from the ducking module 400. According to some examples, such as shown in FIG. 5, the audio processing modifications may be applied by a renderer. In some such examples, the one or more types of audio processing modifications may include modifying the rendering process to warp the rendering of the audio signal away from the estimated location of the person. However, in some such examples, such audio processing modifications may nevertheless be based at least in part on the ducking solution 480 from the ducking module 400.
[0138] According to some examples, the echo management system information may include a model of echo management system performance. In some examples, the model of echo management system performance may be or include an acoustic echo canceller (AEC) performance matrix. In some examples, the model of echo management system performance may be or include a measure of expected echo return loss enhancement provided by the echo management system.
[0139] In some examples, one or more types of audio processing modifications may be based at least in part on acoustic models of inter-device echo and intra-device echo. Alternatively or additionally, in some examples, one or more types of audio processing modifications may be based at least in part on an interaudibility matrix.
[0140] Alternatively or additionally, in some examples, one or more types of audio processing modifications may be based at least in part on an estimated location of the person. In some examples, the estimated location may correspond to a point, while in other examples, the estimated location may correspond to an area, such as a user zone. According to some such examples, the user zone may be a portion of the audio environment, such as a sofa area, a table area, a chair area, etc. In some examples, the estimated location may correspond to an estimated location of the person's head. According to some examples, the estimated location of the person may be based at least in part on output signals from multiple microphones in the audio environment.
[0141] In some implementations, one or more types of audio processing modifications may be based at least in part on listening intent, which may include, for example, spatial content, frequency content, or both.
[0142] According to some examples, the one or more types of audio processing modifications may be based at least in part on one or more constraints. The one or more constraints may be based on a perceptual model, such as a model of human hearing. Alternatively or additionally, the one or more constraints may include audio content energy preservation, audio spatiality preservation, an audio energy vector, a regularization constraint, or a combination thereof.
[0143] In some examples, the method 900 may include updating an acoustic model of the audio environment, a model of echo management system performance, or both, after applying one or more types of audio processing modifications.
[0144] In some examples, determining one or more types of audio processing modifications may be based at least in part on optimizing a cost function. According to some such examples, the cost function may correspond to or be similar to one of the cost functions of Equations 10-13. Other examples of audio processing modifications based at least in part on optimizing a cost function are described in detail below.
[0145] Audio Device Location Method Example 9 and elsewhere herein, in some examples, audio processing modifications (such as those corresponding to ducking solutions) may be based at least in part on audio device location information. The locations of audio devices within the audio environment may be determined or estimated by various methods, including but not limited to those described in the following paragraphs.
[0146] Some such methods may include receiving direct instructions by a user, for example, using a smartphone or tablet device to mark or indicate the approximate location of a device on a floor plan or similar graphical representation of the environment. Such digital interfaces are already commonplace in managing the configuration, grouping, names, purposes, and identities of smart home devices. For example, such direct instructions may be provided via an Amazon Alexa smartphone application, a Sonos S2 controller application, or a similar application.
[0147] Some examples may include using measured signal strengths (sometimes referred to as Received Signal Strength Indicator or RSSI) of common wireless communication technologies such as Bluetooth, Wi-Fi, ZigBee, etc. to solve a basic triangulation problem to generate estimates of physical distance between devices, for example, as disclosed in J. Yang and Y. Chen, “Indoor Localization Using Improved RSS-Based Lateration Methods,” GLOBECOM 2009 - 2009 IEEE Global Telecommunications Conference, Honolulu, HI, 2009, pp. 1-6, doi: 10.1109 / GLOCOM.2009.5425237, and / or Mardeni, Mardeni, R. & Othman, Shaifull & Nizam (2010), “Node Positioning in ZigBee Network Using Trilateration Method Based on the Received Signal Strength Indicator (RSSI),” 46, both of which are incorporated herein by reference.
[0148] U.S. Pat. No. 10,779,084, entitled "Automatic Discovery and Localization of Speaker Locations in Surround Sound Systems," which is incorporated herein by reference, describes a system that can automatically locate loudspeakers and microphones in a listening environment by acoustically measuring the time of arrival (TOA) between each speaker and microphone.
[0149] International Publication No. 2021 / 127286 A1, entitled "Audio Device Auto-Location," which is incorporated herein by reference, discloses methods for estimating an audio device location, a listener position, and a listener location in an audio environment. Some disclosed methods include estimating an audio device location in the environment via direction-of-arrival (DOA) data and by determining interior angles for each of a plurality of triangles based on the DOA data. In some examples, each triangle has a vertex corresponding to an audio device location. Some disclosed methods include determining a side length of each side of each of the triangles and performing a forward alignment process to align each of the plurality of triangles to generate a forward alignment matrix. Some disclosed methods include determining to perform a backward alignment process to align each of the plurality of triangles in reverse sequence to generate a backward alignment matrix. A final estimate of each audio device location may be based at least in part on values of the forward alignment matrix and the backward alignment matrix.
[0150] Other disclosed methods in International Publication No. WO 2021 / 127286 A1 include estimating listener location, and in some cases, including listener location. Some such methods include prompting a listener to make one or more utterances (e.g., via audio prompts from one or more loudspeakers in the environment) and estimating the listener location according to DOA data. The DOA data may correspond to microphone data acquired by multiple microphones in the environment. The microphone data may correspond to detection of one or more utterances by the microphones. At least some of the microphones may be co-located with the loudspeakers. According to some examples, estimating the listener location may include a triangulation process. Some such examples include triangulating the user's voice by finding intersections between DOA vectors passing through the audio device. Some disclosed methods of determining a listener's orientation include prompting a user to identify one or more loudspeaker locations. Some such examples include prompting a user to identify one or more loudspeaker locations by moving next to the loudspeaker location(s) and making an utterance. Another example includes prompting a user to identify one or more loudspeaker locations by pointing at each of the one or more loudspeaker locations using a handheld device such as a mobile phone that includes an inertial sensor system and a wireless interface configured to communicate with a control system controlling audio devices in the audio environment (such as a control system of an orchestration device). Some disclosed methods include determining a listener's orientation by having loudspeakers render audio objects so that the audio objects appear to rotate around the listener, and prompting the listener to make an utterance (such as "Stop!") when the listener perceives the audio object to be within a location such as a loudspeaker location, a television location, or the like.Some disclosed methods include determining the location and / or orientation of the listener via camera data, e.g., by determining the relative location of the listener to one or more audio devices of the audio environment according to the camera data, by determining the orientation of the listener relative to one or more audio devices of the audio environment according to the camera data (e.g., according to the direction the listener is facing), etc.
[0151] Shi, Guangi et al., "Spatial Calibration of Surround Sound Systems including Listener Position Estimation," AES 137 thConvention, October 2014 describes a system in which a single linear microphone array associated with a playback system component with a predictable location, such as a sound bar or front-center speaker, measures the time difference of arrival (TDOA) of both satellite loudspeakers and the listener to determine the location of both the loudspeakers and the listener. In this case, the listening orientation is essentially defined as the line connecting the detected listening position to the playback system component containing the linear microphone array, such as a sound bar collocated with a television (mounted directly above or below the television). Because the sound bar's location is predictably located directly above or below the video screen, the geometry of the measured distances and angles of incidence can be converted to absolute positions relative to any point in front of that reference sound bar location using simple trigonometric principles. The distance between the loudspeakers and microphones in a linear microphone array can be estimated by playing a test signal and measuring the time of flight (TOF) between the emitting loudspeaker and the receiving microphone. For this purpose, the time delay of the direct component of the measured impulse response can be used. The impulse response between the loudspeaker and the microphone array elements can be obtained by playing a test signal through the loudspeaker under analysis. For example, either a maximum length sequence (MLS) or a chirp signal (also known as a logarithmic sine sweep) can be used as the test signal. The room impulse response can be obtained by calculating the circular cross-correlation between the captured signal and the MLS input. Figure 2 in this reference shows the echo impulse response obtained using the MLS input. This impulse response is said to resemble measurements made in a typical office or living room. The delay of the direct component is used to estimate the distance between the loudspeaker and the microphone array elements. For loudspeaker distance estimation, any loopback latency of the audio device used to play the test signal should be calculated and removed from the measured TOF estimate.
[0152] Example of estimating location and orientation of a person in an audio environment The location and orientation of a person in an audio environment may be determined or estimated by a variety of methods, including but not limited to those described in the following paragraphs.
[0153] Hess, Wolfgang, "Head-Tracking Techniques for Virtual Acoustic Applications" (AES 133rd Convention, October 2012), incorporated herein by reference, presents a number of commercially available techniques for tracking both the position and orientation of a listener's head in the context of a spatial audio playback system. One specific example described is the Microsoft Kinect. Its depth sensing and standard camera, along with publicly available software (Windows® Software Development Kit (SDK)), allow for the simultaneous tracking of the head position and orientation of several listeners in space using a combination of skeletal tracking and facial recognition. Kinect for Windows has been discontinued, but the Azure Kinect Developer Kit (DK), which implements Microsoft's next-generation depth sensor, is now available.
[0154] U.S. Patent No. 10,779,084, entitled "Automatic Discovery and Localization of Speaker Locations in Surround Sound Systems," which is incorporated herein by reference, describes a system that can automatically locate loudspeakers and microphones in a listening environment by acoustically measuring the time of arrival (TOA) between each speaker and microphone. A listening location can be detected by arranging and positioning a microphone (e.g., a microphone in a mobile phone held by the listener) at a desired listening position, and an associated listening orientation can be defined by placing another microphone, e.g., on a TV, at a point in the listener's viewing direction. Alternatively, a listening orientation can be defined by positioning a loudspeaker, such as a loudspeaker on a TV, in the viewing direction.
[0155] International Publication No. 2021 / 127286 A1, entitled "Audio Device Auto-Location," which is incorporated herein by reference, discloses methods for estimating an audio device location, a listener position, and a listener location in an audio environment. Some disclosed methods include estimating an audio device location in the environment via direction-of-arrival (DOA) data and by determining interior angles for each of a plurality of triangles based on the DOA data. In some examples, each triangle has a vertex corresponding to an audio device location. Some disclosed methods include determining a side length of each side of each of the triangles and performing a forward alignment process to align each of the plurality of triangles to generate a forward alignment matrix. Some disclosed methods include determining to perform a backward alignment process to align each of the plurality of triangles in reverse sequence to generate a backward alignment matrix. A final estimate of each audio device location may be based at least in part on values of the forward alignment matrix and the backward alignment matrix.
[0156] Other disclosed methods in International Publication No. WO 2021 / 127286 A1 include estimating listener location, and in some cases, including listener location. Some such methods include prompting a listener to make one or more utterances (e.g., via audio prompts from one or more loudspeakers in the environment) and estimating the listener location according to DOA data. The DOA data may correspond to microphone data acquired by multiple microphones in the environment. The microphone data may correspond to detection of one or more utterances by the microphones. At least some of the microphones may be co-located with the loudspeakers. According to some examples, estimating the listener location may include a triangulation process. Some such examples include triangulating the user's voice by finding intersections between DOA vectors passing through the audio device. Some disclosed methods of determining a listener's orientation include prompting a user to identify one or more loudspeaker locations. Some such examples include prompting a user to identify one or more loudspeaker locations by moving next to the loudspeaker location(s) and making an utterance. Another example includes prompting a user to identify one or more loudspeaker locations by pointing at each of the one or more loudspeaker locations using a handheld device such as a mobile phone that includes an inertial sensor system and a wireless interface configured to communicate with a control system controlling audio devices in the audio environment (such as a control system of an orchestration device). Some disclosed methods include determining a listener's orientation by having loudspeakers render audio objects so that the audio objects appear to rotate around the listener, and prompting the listener to make an utterance (such as "Stop!") when the listener perceives the audio object to be within a location such as a loudspeaker location, a television location, or the like.Some disclosed methods include determining the location and / or orientation of the listener via camera data, e.g., by determining the relative location of the listener to one or more audio devices of the audio environment according to the camera data, by determining the orientation of the listener relative to one or more audio devices of the audio environment according to the camera data (e.g., according to the direction the listener is facing), etc.
[0157] Shi, Guangi et al., "Spatial Calibration of Surround Sound Systems including Listener Position Estimation," AES 137 thConvention, October 2014 describes a system in which a single linear microphone array associated with a playback system component with a predictable location, such as a sound bar or front-center speaker, measures the time difference of arrival (TDOA) of both satellite loudspeakers and the listener to determine the location of both the loudspeakers and the listener. In this case, the listening orientation is essentially defined as the line connecting the detected listening position to the playback system component containing the linear microphone array, such as a sound bar collocated with a television (mounted directly above or below the television). Because the sound bar's location is predictably located directly above or below the video screen, the geometry of the measured distances and angles of incidence can be converted to absolute positions relative to any point in front of that reference sound bar location using simple trigonometric principles. The distance between the loudspeakers and microphones in a linear microphone array can be estimated by playing a test signal and measuring the time of flight (TOF) between the emitting loudspeaker and the receiving microphone. For this purpose, the time delay of the direct component of the measured impulse response can be used. The impulse response between the loudspeaker and the microphone array elements can be obtained by playing a test signal through the loudspeaker under analysis. For example, either a maximum length sequence (MLS) or a chirp signal (also known as a logarithmic sine sweep) can be used as the test signal. The room impulse response can be obtained by calculating the circular cross-correlation between the captured signal and the MLS input. Figure 2 in this reference shows the echo impulse response obtained using the MLS input. This impulse response is said to resemble measurements made in a typical office or living room. The delay of the direct component is used to estimate the distance between the loudspeaker and the microphone array elements. For loudspeaker distance estimation, any loopback latency of the audio device used to play the test signal should be calculated and removed from the measured TOF estimate.
[0158] Estimating people's locations according to user zones In some examples, the estimated location of a person in an audio environment may correspond to a user zone. This section describes a method for estimating the user zone in which a person is located, based at least in part on a microphone signal.
[0159] FIG. 10 is a flow diagram outlining an example of a method that may be performed by an apparatus as shown in FIG. 1A. The blocks of method 1000 are not necessarily performed in the order shown, similar to other methods described herein. Further, such a method may include more or fewer blocks than illustrated and / or described. In this implementation, method 1000 includes estimating the location of a user in an environment.
[0160] In this example, block 1005 includes receiving an output signal from each of a plurality of microphones in an environment. In this case, each of the plurality of microphones is present at a microphone location in the environment. According to this example, the output signal corresponds to the user's current utterance. In some examples, the current utterance may be or include a wake word utterance. Block 1005 may include, for example, a control system (such as control system 120 in FIG. 1A) receiving an output signal from each of a plurality of microphones in an environment via an interface system (such as interface system 205 in FIG. 1A).
[0161] In some examples, at least some of the microphones in the environment may provide output signals that are asynchronous with respect to output signals provided by one or more other microphones. For example, a first microphone of the plurality of microphones may sample audio data according to a first sample clock, and a second microphone of the plurality of microphones may sample audio data according to a second sample clock. In some instances, at least one of the microphones in the environment may be included in or configured to communicate with a smart audio device.
[0162] According to this example, block 1010 includes determining a plurality of current acoustic features from the output signal of each microphone. In this example, the “current acoustic features” are acoustic features derived from the “current utterance” of block 1005. In some implementations, block 1010 may include receiving the plurality of current acoustic features from one or more other devices. For example, block 1010 may include receiving at least some of the plurality of current acoustic features from one or more wake word detectors implemented by the one or more other devices. Alternatively or additionally, in some implementations, block 1010 may include determining the plurality of current acoustic features from the output signal.
[0163] Regardless of whether the acoustic features are determined by a single device or multiple devices, the acoustic features may be determined asynchronously. When the acoustic features are determined by multiple devices, the acoustic features will generally be determined asynchronously unless the devices are configured to coordinate the process of determining the acoustic features. When the acoustic features are determined by a single device, in some implementations, the acoustic features may nevertheless be determined asynchronously because the single device may receive the output signals of each microphone at different times. In some examples, the acoustic features may be determined asynchronously because at least some of the microphones in the environment may provide output signals that are asynchronous with respect to the output signals provided by one or more other microphones.
[0164] In some examples, the acoustic features may include a wake word confidence metric, a wake word duration metric, and / or at least one reception level metric. The reception level metric may indicate a reception level of the sound detected by the microphone and may correspond to a level of the microphone output signal.
[0165] Alternatively or additionally, the acoustic features may include one or more of the following: Average state entropy (purity) for each wake word state along its 1-best (Viterbi) alignment with the acoustic model. Connectionist Temporal Classification Loss (CTC loss) for the acoustic model of the wake word detector. The wake word detector may be trained to provide an estimate of the talker's distance from the microphone and / or an RT60 estimate in addition to the wake word confidence. The distance estimate and / or the RT60 estimate may be acoustic features. Instead of or in addition to the broadband received level / power at the microphone, the acoustic feature may be the received level in several logarithmic / Mel / Bark-spaced frequency bands, which may vary according to the particular implementation (e.g., 2 frequency bands, 5 frequency bands, 20 frequency bands, 50 frequency bands, 1 octave frequency band, or 1 / 3 octave frequency band). · Cepstral representation of the spectral information at preceding points, calculated by taking the DCT (Discrete Cosine Transform) of the logarithm of the band power. Band power in frequency bands weighted for human speech. For example, the acoustic signature may be based only on a specific frequency band (e.g., 400 Hz to 1.5 kHz). In this example, higher and lower frequencies may be ignored. · Confidence of the voice activity detector per band or per bin. The acoustic signature may be based at least in part on a long-term noise estimate so as to ignore microphones with insufficient signal-to-noise ratio. Kurtosis as a measure of the "peakiness" of speech. Kurtosis can be an indicator of smearing due to long reverberation tails. Estimated onset time of the wake word. Onset and duration are expected to be equal across all microphones within a frame or so. Outliers may give clues to unreliable estimates. This assumes a level of synchrony, not necessarily across samples, but across frames of, say, tens of milliseconds.
[0166] According to this example, block 1015 includes applying a classifier to the plurality of current acoustic features. In some such examples, applying the classifier may include applying a model trained on previously determined acoustic features derived from a plurality of previous utterances made by a user in a plurality of user zones in the environment. Various examples are provided herein.
[0167] In some examples, the user zones may include a sink area, a food preparation area, a refrigerator area, a dining area, a sofa area, a television area, a bedroom area, and / or an entryway area. According to some examples, one or more of the user zones may be predetermined user zones. In some such examples, the one or more predetermined user zones may have been selected by a user during a training process.
[0168] In some implementations, applying the classifier may include applying a Gaussian mixture model trained on previous utterances. According to some such implementations, applying the classifier may include applying a Gaussian mixture model trained on one or more of a normalized wake word confidence, a normalized average received level, or a maximum received level of the previous utterances. However, in alternative implementations, applying the classifier may be based on a different model, such as one of the other models disclosed herein. In some instances, the model may be trained using training data labeled with user zones. However, in some examples, applying the classifier includes applying a model trained using unlabeled training data that is not labeled with user zones.
[0169] In some examples, the previous utterance may have been or included a wake word utterance, and in some such examples, the previous utterance and the current utterance may have been utterances of the same wake word.
[0170] In this example, block 1020 includes determining an estimate of a user zone in which the user is currently located based at least in part on the output from the classifier. In some such examples, the estimate may be determined without reference to the geometric locations of the microphones. For example, the estimate may be determined without reference to the coordinates of individual microphones. In some examples, the estimate may be determined without estimating the geometric location of the user.
[0171] Some implementations of method 1000 may include selecting at least one speaker according to the estimated user zone. Some such implementations may include controlling the at least one selected speaker to provide sound to the estimated user zone. Alternatively or additionally, some implementations of method 1000 may include selecting at least one microphone according to the estimated user zone. Some such implementations may include providing a signal output by the at least one selected microphone to a smart audio device.
[0172] 11 is a block diagram of elements of an example embodiment configured to implement a zone classifier. According to this example, a system 1100 includes multiple loudspeakers 1104 distributed throughout at least a portion of an environment (e.g., an environment such as that shown in FIG. 1A or 1B). In this example, the system 1100 includes a multi-channel loudspeaker renderer 1101. According to this implementation, the output of the multi-channel loudspeaker renderer 1101 serves as both a loudspeaker drive signal (a speaker feed for driving the speakers 1104) and an echo reference. In this implementation, the echo reference is provided to an echo management subsystem 1103 via multiple loudspeaker reference channels 1102 that include at least some of the speaker feed signals output from the renderer 1101.
[0173] In this implementation, the system 1100 includes multiple echo management subsystems 1103. According to this example, the echo management subsystems 1103 are configured to implement one or more echo suppression processes and / or one or more echo cancellation processes. In this example, each of the echo management subsystems 1103 provides a corresponding echo management output 1103A to one of the wake word detectors 1106. The echo management output 1103A has an attenuated echo relative to the input to the associated one of the echo management subsystems 1103.
[0174] According to this implementation, the system 1100 includes N microphones 1105 (N is an integer) distributed throughout at least a portion of an environment (e.g., the environment shown in FIG. 1A or FIG. 1B). The microphones may include array microphones and / or spot microphones. For example, one or more smart audio devices located within the environment may include an array of microphones. In this example, the outputs of the microphones 1105 are provided as inputs to the echo management subsystems 1103. According to this implementation, each of the echo management subsystems 1103 captures the output of an individual microphone 1105 or an individual group or subset of the microphones 1105.
[0175] In this example, the system 1100 includes multiple wake word detectors 1106. According to this example, each of the wake word detectors 1106 receives audio output from one of the echo management subsystems 1103 and outputs multiple acoustic features 1106A. The acoustic features 1106A output from each echo management subsystem 1103 may include (but are not limited to) wake word confidence, wake word duration, and receive level measurements. Although three arrows representing three acoustic features 1106A are shown as being output from each echo management subsystem 1103, more or fewer acoustic features 1106A may be output in alternative implementations. Furthermore, although these three arrows impinge on the classifier 1107 along approximately vertical lines, this does not indicate that the classifier 1107 necessarily receives acoustic features 1106A from all of the wake word detectors 1106 simultaneously. As noted elsewhere herein, the acoustic features 1106A may, in some cases, be determined and / or provided to the classifier asynchronously.
[0176] According to this implementation, the system 1100 includes a zone classifier 1107, sometimes referred to as a classifier 1107. In this example, the classifier receives multiple features 1106A from multiple wake word detectors 1106 for multiple (e.g., all) microphones 1105 in the environment. According to this example, the output 1108 of the zone classifier 1107 corresponds to an estimate of a user zone in which the user is currently located. According to some such examples, the output 1108 may correspond to one or more posterior probabilities. The estimate of the user zone in which the user is currently located may be or correspond to a maximum posterior probability according to Bayesian statistics.
[0177] Next, we will describe an exemplary implementation of a classifier that may, in some examples, correspond to the zone classifier 1107 of FIG. i Let (n) be the i-th microphone signal (i={1...N}) at discrete time n (i.e., microphone signal x i(n) are the outputs of the N microphones 1105. The N signals x i (n) process produces a "clean" microphone signal e at discrete time n. i (n), where i={1...N}. The clean signal e, designated 1103A in FIG. i (n) are fed to wake word detectors 1106 in this example, where each wake word detector 1106 receives a vector of features w i (j), where j={1...J} is the index corresponding to the jth wake word utterance. In this example, the classifier 1107 takes as input the following aggregate feature set:
number
[0178] According to some implementations, a set of zone labels C k (k={1...K}) may correspond to the number K of different user zones in the environment. For example, the user zones may include a sofa zone, a kitchen zone, a reading chair zone, etc. Some examples may define more than one zone within a kitchen or other room. For example, a kitchen area may include a sink zone, a food preparation zone, a refrigerator zone, and a dining zone. Similarly, a living room area may include a sofa zone, a TV zone, a reading chair zone, one or more entry / exit zones, etc. Zone labels for these zones may be selectable by the user, for example, during a training phase.
[0179] In some implementations, the classifier 1107 determines the posterior probability p(C k |W(j)) with probability p(C k |W(j)) is the user's zone C k The probability of being in each of the zones (zone C k(For each of each and each utterance, the j-th utterance and the k-th zone) are shown and are an example of the output 1108 of the classifier 1107.
[0180] According to some examples, the training data can be collected by prompting the user to select or define a zone, such as a sofa zone (e.g., for each user zone). The training process can include prompting the user to make training utterances, such as wake words, near the selected or defined zone. In the example of the sofa zone, the training process can include prompting the user to make training utterances at the center and the extreme ends of the sofa. The training process can include prompting the user to repeat the training utterances several times at each location within the user zone. The user can then be prompted to move to another user zone and continue until all specified user zones are covered.
[0181] FIG. 12 is a flowchart outlining an example of a method that can be performed by an apparatus such as apparatus 200 of FIG. 1A. The blocks of method 1200 are not necessarily executed in the order shown, similar to other methods described herein. Further, such a method can include more or fewer blocks than those illustrated and / or described. In this implementation, method 1200 includes training a classifier to estimate the location of a user within an environment.
[0182] In this example, block 1205 includes prompting the user to make at least one training utterance at each of a plurality of locations within a first user zone of the environment. The training utterance(s) can, in some examples, be one or more instances of a wake word utterance. According to some implementations, the first user zone can be any user zone selected and / or defined by the user. In some cases, the control system has a corresponding zone label (e.g., the zone label C described above) k, a corresponding instance of a zone label of one of the zones ( ) and associate the zone label with the training data obtained for the first user zone.
[0183] An automated prompt system may be used to collect these training data. As described above, the interface system 205 of the device 200 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. For example, the device 200 may provide the user with the following prompts on the screen of the display system, or the user may hear them announced via one or more speakers during the training process: · "Please move to the couch." "Say the wake word 10 times while moving your head." "Move to a position halfway between the sofa and the reading chair and say the wake word 10 times." "Stand in the kitchen as if you were cooking and say the wake word 10 times."
[0184] In this example, block 1210 includes receiving a first output signal from each of a plurality of microphones in the environment. In some examples, block 1210 may include receiving the first output signal from all of the active microphones in the environment, while in other examples, block 1210 may include receiving the first output signal from a subset of all of the active microphones in the environment. In some examples, at least some of the microphones in the environment may provide output signals that are asynchronous with respect to output signals provided by one or more other microphones. For example, a first microphone of the plurality of microphones may sample audio data according to a first sample clock, and a second microphone of the plurality of microphones may sample audio data according to a second sample clock.
[0185] In this example, each microphone of the plurality of microphones is present at a microphone location in the environment. In this example, the first output signal corresponds to a detected instance of training utterance received from a first user zone. Because block 1205 includes prompting a user to make at least one training utterance at each of a plurality of locations within the first user zone of the environment, in this example, the term "first output signal" refers to the set of all output signals corresponding to the training utterances for the first user zone. In other examples, the term "first output signal" may refer to a subset of all output signals corresponding to the training utterances for the first user zone.
[0186] According to this example, block 1215 includes determining one or more first acoustic features from each of the first output signals. In some examples, the first acoustic features may include a wake word confidence metric and / or a received level metric. For example, the first acoustic features may include a normalized wake word confidence metric, a normalized average received level indication, and / or a maximum received level indication.
[0187] As described above, block 1205 includes prompting a user to produce at least one training utterance at each of a plurality of locations within a first user zone of the environment, so in this example, the term "first output signals" refers to the set of all output signals corresponding to the training utterances for the first user zone. Therefore, in this example, the term "first acoustic features" refers to a set of acoustic features derived from the set of all output signals corresponding to the training utterances for the first user zone. Therefore, in this example, the set of first acoustic features is at least as large as the set of first output signals. For example, if two acoustic features are determined from each of the output signals, the set of first acoustic features will be twice as large as the set of first output signals.
[0188] In this example, block 1220 includes training a classifier model to create a correlation between the first user zone and the first acoustic feature. The classifier model may be, for example, any of those disclosed herein. According to this implementation, the classifier model is trained without reference to the geometric locations of the multiple microphones. In other words, in this example, data regarding the geometric locations of the multiple microphones (e.g., microphone coordinate data) is not provided to the classifier model during the training process.
[0189] FIG. 13 is a flowchart outlining another example of a method that can be performed by an apparatus such as apparatus 200 of FIG. 1A. The blocks of method 1300 are not necessarily performed in the order shown, similar to other methods described herein. For example, in some implementations, at least a portion of the acoustic feature determination process of block 1325 may be performed before block 1315 or block 1320. Further, such a method may include more or fewer blocks than illustrated and / or described. In this implementation, method 1300 includes training a classifier to estimate the location of a user within an environment. Method 1300 provides an example of extending method 1200 to multiple user zones of an environment.
[0190] In this example, block 1305 includes prompting the user to make at least one training utterance at a location within a user zone of the environment. In some cases, block 1305 may be performed in the manner described above while referring to block 1205 of FIG. 12, except that block 1305 relates to a single location within the user zone. The training utterance(s) can, in some examples, be one or more instances of wake word utterances. According to some implementations, the user zone can be any user zone selected and / or defined by the user. In some cases, the control system can create a corresponding zone label (e.g., a corresponding instance of one of the zone labels C described above) and associate the zone label with the training data obtained for the user zone. k of one of the zone labels) and associate the zone label with the training data obtained for the user zone.
[0191] According to this example, block 1310 is performed substantially as described above with reference to block 1210 of FIG. 12 . However, in this example, the process of block 1310 is generalized to any user zone, not necessarily the first user zone from which training data is obtained. Thus, the output signals received in block 1310 are “output signals from each of a plurality of microphones in the environment, each of the plurality of microphones being present at a microphone location in the environment, and the output signals corresponding to instances of detected training utterances received from the user zone.” In this example, the term “output signals” refers to the set of all output signals corresponding to one or more training utterances at a user zone location. In other examples, the term “output signals” may refer to a subset of all output signals corresponding to one or more training utterances at a user zone location.
[0192] According to this example, block 1315 includes determining whether sufficient training data has been obtained for the current user zone. In some such examples, block 1315 may include determining whether output signals corresponding to a threshold number of training utterances have been obtained for the current user zone. Alternatively or additionally, block 1315 may include determining whether output signals corresponding to training utterances in a threshold number of locations within the current user zone have been obtained. If not, method 1300 returns to block 1305 in this example, and the user is prompted to make at least one additional utterance at a location within the same user zone.
[0193] However, if in block 1315 it is determined that sufficient training data has been acquired for the current user zone, then in this example the process proceeds to block 1320. According to this example, block 1320 includes determining whether to acquire training data for additional user zones. According to some examples, block 1320 may include determining whether training data has been acquired for each user zone previously identified by the user. In other examples, block 1320 may include determining whether training data has been acquired for a minimum number of user zones. The minimum number may be user-selectable. In other examples, the minimum number may be a recommended minimum number per environment, a recommended minimum number per room of the environment, etc.
[0194] If, in block 1320, it is determined that training data should be obtained for additional user zones, in this example, the process continues to block 1322, which includes prompting the user to move to another user zone of the environment. In some examples, the next user zone may be selectable by the user. According to this example, the process continues to block 1305 after the prompt of block 1322. In some such examples, the user may be prompted to confirm that the user has arrived in the new user zone after the prompt of block 1322. According to some such examples, the user may be required to confirm that the user has arrived in the new user zone before the prompt of block 1305 is provided.
[0195] If, at block 1320, it is determined that training data should not be obtained for additional user zones, in this example, the process proceeds to block 1325. In this example, method 1300 includes obtaining training data for K user zones. In this implementation, block 1325 includes determining first through second acoustic features from first through second output signals corresponding to the first through second user zones, respectively, for which training data was obtained. In this example, the term "first output signal" refers to the set of all output signals corresponding to training utterances for the first user zone, and the term "Hth output signal" refers to the set of all output signals corresponding to training utterances for the Kth user zone. Similarly, the term "first acoustic feature" refers to the set of acoustic features determined from the first output signal, and the term "Gth acoustic feature" refers to the set of acoustic features determined from the Hth output signal.
[0196] According to these examples, block 1330 includes training a classifier model to create correlations between the first through Kth user zones and the first through Kth acoustic features, respectively. The classifier model may be, for example, any of the classifier models disclosed herein.
[0197] In the previous example, the user zone (e.g., the zone label C described above) k , according to a corresponding instance of a zone label of one of the zones. However, the model may be trained according to either labeled user zones or unlabeled user zones, depending on the particular implementation. If labeled, each training utterance may be paired with a label corresponding to the user zone, for example, as follows:
number
[0198] Training a classifier model may include determining a best fit to the labeled training data. Without loss of generality, suitable classification techniques for the classifier model may include: Bayesian classifiers with per-class distributions described, for example, by multivariate normal distributions, full-covariance Gaussian mixture models, or diagonal-covariance Gaussian mixture models; · Vector quantization; · Nearest neighbor (k-means); A neural network with one output for each class but with a SoftMax output layer; Support Vector Machines (SVM); and / or Boosting techniques such as Gradient Boosting Machines (GBM).
[0199] In one example of implementing the unlabeled case, the data may be automatically divided into K clusters, where K may also be unknown. The unlabeled automatic division may be performed, for example, by using classical clustering techniques, such as the k-means algorithm or Gaussian mixture modeling.
[0200] To improve robustness, regularization may be applied to the classifier model training, and the model parameters may be updated over time as new utterances are presented.
[0201] Further aspects of the embodiment will now be described.
[0202] An exemplary acoustic feature set (e.g., acoustic feature 1106A of FIG. 11) may include the likelihood of wake word confidence, the average received level over the estimated duration of the most confident wake word, and the maximum received level over the duration of the most confident wake word. The features may be normalized with respect to their maximum values for each wake word utterance. Training data may be labeled and a full covariance Gaussian mixture model (GMM) may be trained to maximize the expected value of the training labels. The estimated zone may be the class that maximizes the posterior probability.
[0203] The above description of some embodiments discusses learning an acoustic zone model from a set of training data collected during a prompted collection process. In that model, the training time (or configuration mode) and the runtime (or normal mode) may be considered two separate modes in which a microphone system may be placed. An extension of this approach is online learning in which some or all of the acoustic zone model is learned or adapted online (e.g., at runtime or in normal mode). In other words, in some implementations, the process of training the classifier may continue even after the classifier has been applied in the "runtime" process to estimate the user zone in which the user is currently located (e.g., according to method 1000 of FIG. 10).
[0204] FIG. 14 is a flowchart outlining another example of a method that may be performed by an apparatus such as apparatus 200 of FIG. 1A. The blocks of method 1400 are not necessarily performed in the order shown, as with other methods described herein. Further, such a method may include more or fewer blocks than those illustrated and / or described. In this implementation, method 1400 includes ongoing training of a classifier during a "runtime" process of estimating the location of a user within an environment. Method 1400 is an example of what is referred to herein as an online learning mode.
[0205] In this example, block 1405 of method 1400 corresponds to blocks 1005-1020 of method 1000. Here, block 1405 includes providing an estimate of a user zone in which the user is currently located based at least in part on output from the classifier. According to this implementation, block 1410 includes obtaining implicit or explicit feedback regarding the estimate of block 1405. In block 1415, the classifier is updated according to the feedback received in block 1405. Block 1415 may include, for example, one or more reinforcement learning methods. As suggested by the dashed arrow from block 1415 to block 1405, in some implementations, method 1400 may include returning to block 1405. For example, method 1400 may include providing a future estimate of a user zone in which the user will be located at a future time based on applying the updated model.
[0206] Explicit techniques for obtaining feedback may include: Use a voice user interface (UI) to ask the user if the prediction was correct. For example, a sound may be provided to the user indicating: "I think you're sitting on the couch, so please say 'correct' or 'wrong'." Use a voice UI to inform the user that they can always correct an incorrect prediction (e.g., a sound could be provided to the user indicating: "I can now predict where you will be when you talk to me. If the prediction is wrong, say something like, 'Amanda, I'm not sitting on the couch. I'm sitting in a reading chair.'" Use a voice UI to inform the user that a correct prediction can always be rewarded (e.g., a sound could be provided to the user indicating: "I'm now able to predict where you will be when you talk to me. If my prediction is correct, please help me improve my predictions by saying something like, 'Amanda, that's right. I'm sitting on the couch.'"). Includes physical buttons or other UI elements that the user can interact with to provide feedback (e.g., thumbs-up and / or thumbs-down buttons on the physical device or within a smartphone app).
[0207] The goal of predicting the user zone in which the user is located may be to inform microphone selection or adaptive beamforming schemes that attempt to more effectively pick up sound from the user's acoustic zone, for example, to better recognize commands following a wake word. In such a scenario, implicit techniques for obtaining feedback on the quality of the zone prediction may include: Penalizing predictions that result in misrecognition of the command following the wake word. Proxies that may indicate misrecognition may include a user interrupting the voice assistant's response to a command by, for example, uttering a counter-command such as "Amanda, stop!"; Penalizing predictions that give the speech recognizer low confidence that it successfully recognized the command. Many automatic speech recognition systems have the ability to return a confidence level with their results that can be used for this purpose; Penalizing predictions that result in the second-pass wake word detector being unable to retroactively detect the wake word with high confidence; and / or Enhanced predictions that result in highly confident recognition of wake words and / or correct recognition of user commands.
[0208] The following is an example in which a second-pass wake word detector cannot retroactively detect the wake word with high confidence. After obtaining output signals corresponding to the current utterance from microphones in the environment and determining acoustic features based on the output signals (e.g., via multiple first-pass wake word detectors configured to communicate with the microphones), assume that the acoustic features are provided to a classifier. In other words, the acoustic features are presumed to correspond to the detected wake word utterance. Further, assume that the classifier determines that the person who made the current utterance is most likely located in Zone 3, which in this example corresponds to a reading chair. For example, there may be a learned microphone or combination of microphones known to be best at hearing a person's voice when the person is located in Zone 3, to send to a cloud-based virtual assistant service for voice command recognition.
[0209] Further, assume that after determining which microphone(s) will be used for speech recognition, but before the person's voice is actually sent to the virtual assistant service, a second-pass wake word detector operates on the microphone signal corresponding to the speech detected by the selected microphone(s) for Zone 3 being submitted for command recognition. If that second-pass wake word detector does not match the first-pass wake word detectors where the wake word actually was uttered, it is likely because the classifier incorrectly predicted the zone. Therefore, this classifier should be penalized.
[0210] Techniques for post-updating the zone mapping model after one or more wake words have been spoken may include the following. Maximum a posteriori (MAP) adaptation of Gaussian mixture models (GMM) or nearest neighbor models; and / or Reinforcement learning, for example in neural networks, by associating appropriate "one-hot" (for correct predictions) or "one-cold" (for incorrect predictions) ground truth labels with SoftMax outputs and applying online backpropagation to determine new network weights.
[0211] Some examples of MAP adaptation in this context may include adjusting the mean in the GMM each time a wake word is spoken. In this way, the mean may more closely match the acoustic characteristics observed when a subsequent wake word is spoken. Alternatively or additionally, such examples may include adjusting the variance / covariance or mixture weight information in the GMM each time a wake word is spoken.
[0212] For example, a MAP adaptation scheme may be as follows:
number
[0213] In the above formula, μ i,old where σ represents the mean of the i-th Gaussian in the mixture, α represents a parameter that controls how aggressive the MAP adaptation should be (α can be in the range [0.9, 0.999]), and x represents the feature vector of the new wake word utterance. The index "i" corresponds to the mixture element that returns the highest a priori probability of containing the speaker's location at the time of the wake word.
[0214] Alternatively, each of the mixture elements may be adjusted according to their a priori probability of containing the wake word, for example, as follows:
number
[0215] In the above formula, β i =α*(1−P(i)), where P(i) represents the a priori probability that observation x is due to mixture element i.
[0216] In one reinforcement learning example, there may be three user zones. Suppose, for a particular wake word, the model predicts the probabilities for the three user zones to be [0.2, 0.1, 0.7]. If a second source of information (e.g., a second-pass wake word detector) confirms that zone 3 was correct, the ground truth label may be [0, 0, 1] ("one hot"). Posterior updates to the zone mapping model may involve backpropagating error through the neural network, effectively meaning that the neural network will predict zone 3 more strongly if the same input is presented again. Conversely, if the second source of information indicates that zone 3 was an inaccurate prediction, the ground truth label may be [0.5, 0.5, 0.0], in one example. Backpropagating error through the neural network will make the model less likely to predict zone 3 if the same input is presented in the future.
[0217] Further examples of audio processing modifications including cost function optimization As noted elsewhere herein, in various disclosed examples, one or more types of audio processing modifications may be based on optimizing a cost function. Some such examples include flexible rendering.
[0218] Flexible rendering allows spatial audio to be rendered across any number of arbitrarily arranged speakers. Given the widespread deployment of audio devices, including but not limited to smart audio devices (e.g., smart speakers) in the home, there is a need to implement flexible rendering techniques that enable consumer products to perform flexible rendering of audio and playback of the audio so rendered.
[0219] Several techniques have been developed to implement flexible rendering. They cast the rendering problem as one of cost function minimization, where the cost function consists of two terms: a first term that models the desired spatial impression the renderer is trying to achieve, and a second term that assigns costs to speaker activations. Until now, this second term has focused on creating a sparse solution, where only speakers in close proximity to the desired spatial location of the audio being rendered are activated.
[0220] Spatial audio playback in consumer environments is typically tied to a predetermined number of loudspeakers arranged in predetermined locations, e.g., 5.1 and 7.1 surround sound. In these cases, content is authored specifically for the associated loudspeakers and encoded as one discrete channel per loudspeaker (e.g., Dolby Digital, Dolby Digital Plus, etc.). More recently, immersive object-based spatial audio formats (Dolby Atmos) have been introduced that break this association between content and specific loudspeaker locations. Instead, content may be described as a collection of individual audio objects, each with potentially time-varying metadata that describes the desired perceptual location of the audio object in three-dimensional space. During playback, the content is converted into loudspeaker feeds by a renderer that adapts to the number and locations of the loudspeakers in the playback system. However, many such renderers still constrain the location of a set of loudspeakers to be one of a set of pre-defined layouts (e.g., for Dolby Atmos, 3.1.2, 5.1.2, 7.1.4, 9.1.6, etc.).
[0221] Moving beyond such constrained rendering, methods have been developed that allow object-based audio to be flexibly rendered across a truly arbitrary number of loudspeakers placed in arbitrary positions. These methods require the renderer to have knowledge of the number and physical locations of the loudspeakers in the listening space. For such systems to be practical for the average consumer, an automated method for locating the loudspeakers is desirable. One such method relies on the use of multiple microphones, possibly co-located with the loudspeakers. By playing an audio signal through the loudspeakers and recording with the microphones, the distance between each loudspeaker and microphone is estimated. From these distances, the locations of both the loudspeakers and the microphones are then estimated.
[0222] Concurrent with the introduction of object-based spatial audio in consumer spaces has been the rapid adoption of so-called “smart speakers,” such as the Amazon Echo product line. While the immense popularity of these devices can be attributed to the simplicity and convenience offered by their wireless connectivity and integrated voice interface (e.g., Amazon's Alexa), the acoustic capabilities of these devices are generally limited, particularly with regard to spatial audio. In most cases, these devices are constrained to mono or stereo playback. However, combining the flexible rendering and automatic localization techniques described above with multiple orchestrated smart speakers can result in a system with highly sophisticated spatial playback capabilities that remains extremely simple for consumers to set up. Wireless connectivity allows consumers to place as many speakers as they need wherever convenient, without the need to run speaker wires, and the built-in microphones can be used to automatically localize the speakers for their associated flexible renderers.
[0223] Conventional flexible rendering algorithms are designed to achieve a particular desired perceived spatial impression as closely as possible. In an orchestrated smart speaker system, sometimes maintaining this spatial impression may not be the most important or desired objective. For example, if someone is simultaneously attempting to speak to an integrated voice assistant, it may be desirable to momentarily modify the spatial rendering in a manner that reduces the relative playback level on speakers near a particular microphone in order to increase the signal-to-noise ratio and / or signal-to-echo ratio (SER) of the microphone signal containing the detected speech. Some embodiments described herein may be implemented as modifications to existing flexible rendering methods to enable such dynamic modifications to the spatial rendering, for example, to achieve one or more additional objectives.
[0224] Existing flexible rendering techniques include Center of Mass Amplitude Panning (CMAP) and Flexible Virtualization (FV). Broadly speaking, both of these techniques render a set of one or more audio signals, each with an associated desired perceptual spatial location, for playback on a set of two or more loudspeakers, with the relative activation of the loudspeakers in the set being a function of a model of the perceptual spatial locations of the audio signals played on the loudspeakers and the proximity of the desired perceptual spatial locations of the audio signals to the loudspeaker locations. The model ensures that the audio signals are heard by the listener near their intended spatial locations, and a proximity term controls which loudspeakers are used to achieve this spatial impression. In particular, the proximity term favors the activation of loudspeakers that are near the desired perceptual spatial location of the audio signals. For both CMAP and FV, this functional relationship is conveniently derived from the following cost function, written as the sum of two terms, one for spatial aspect and one for proximity:
number
number
number
number
[0225] In a particular definition of the cost function, g opts Although the relative levels of the components of g are appropriate, it is difficult to control the absolute level of optimal activation resulting from the above minimization. opt Subsequent normalization of σ may be performed to control the absolute level of activation. For example, it may be desirable to normalize the vectors to have unit length, which follows the commonly used low-power panning rule:
number
[0226] The exact behavior of the flexible rendering algorithm depends on the two terms in the cost function, C spatial and C proximity It is determined by the specific construction of the CMAP. spatial represents the perceived spatial location of audio signals played from a set of loudspeakers in terms of their associated activation gains gi It is derived from a model that places the loudspeakers at the center of mass at their positions weighted by (elements of vector g):
number
number
number
number
number
number
number
number
number
number
[0227] For this purpose, the second term of the cost function, C proximity can be defined as the distance-weighted sum of the absolute squares of the speaker activations, which can be compactly represented in matrix form as follows:
number
number
number
number
[0228] Combining the two terms of the cost function defined in equations 8 and 9a gives the overall cost function:
number
[0229] Setting the derivative of this cost function with respect to g equal to zero and solving for g gives the optimal speaker activation solution:
number
[0230] In general, an optimal solution to Equation 27 may result in speaker activations with negative values. For the CMAP construction of a flexible renderer, such negative activations may be undesirable, so Equation 27 may be minimized subject to the condition that all activations remain positive.
[0231] 15 and 16 are diagrams illustrating an example set of speaker activation and object rendering positions. In these examples, the speaker activation and object rendering positions correspond to speaker positions of 4, 64, 165, -87, and -4 degrees. FIG. 15 shows speaker activations 1505a, 1510a, 1515a, 1520a, and 1525a, which contain the optimal solution of Equation 11 for these specific speaker positions. FIG. 16 plots the individual speaker positions as dots 1605, 1610, 1615, 1620, and 1625, which correspond to speaker activations 1505a, 1510a, 1515a, 1520a, and 1525a, respectively. FIG. 16 also shows ideal object positions (in other words, the positions where audio objects should be rendered) for a number of possible object angles as dots 1630a, and the corresponding actual rendering positions for those objects as dots 1635a connected to the ideal object positions by dotted lines 1640a.
[0232] One class of embodiments includes a method for rendering audio for playback by at least one (e.g., all or some) of a plurality of coordinated (orchestrated) smart audio devices. For example, a set of smart audio devices present in a user's home (in a system) may be orchestrated to handle a variety of simultaneous use cases, including (in one embodiment) flexible rendering of audio for playback by all or some of the smart audio devices (i.e., by speaker(s) of all or some of the smart audio devices). Many interactions with the system are possible that require dynamic modifications to the rendering. Such modifications may, but are not necessarily, focused on spatial fidelity.
[0233] Some embodiments are methods for rendering audio for playback by at least one (e.g., all or some) of a set of smart audio devices (or for playback by at least one (e.g., all or some) of a speaker of another set of speakers). The rendering may include minimizing a cost function, the cost function including at least one dynamic speaker activation term. Examples of such dynamic speaker activation terms include (but are not limited to): · The proximity of the speaker to one or more listeners; · The proximity of the speaker to the attractive or repulsive forces; · Audibility of the loudspeaker relative to a certain location (e.g., listener position or baby room); · Speaker capabilities (e.g. frequency response and distortion); · Synchronicity of speakers relative to other speakers; · Wake word performance; Echo cancellation performance.
[0234] The dynamic speaker activation term(s) may enable at least one of a variety of behaviors, including warping the spatial presentation of audio away from a particular smart audio device so that its microphone can better hear the talker, or so that the smart audio device's speaker(s) can better hear a secondary audio stream.
[0235] Some embodiments implement rendering for playback over the speaker(s) of multiple coordinated (orchestrated) smart audio devices, while other embodiments implement rendering for playback over the speaker(s) of a separate set of speakers.
[0236] Pairing a flexible rendering method (implemented according to some embodiments) with a set of wireless smart speakers (or other smart audio devices) can result in an extremely capable and easy-to-use spatial audio rendering system. Considering interactions with such a system, it becomes apparent that dynamic modifications to the spatial rendering may be desirable to optimize for other objectives that may arise during use of the system. To achieve this goal, one class of embodiments augments existing flexible rendering algorithms (where speaker activation is a function of the spatial and proximity terms previously disclosed) with one or more additional, dynamically configurable features that depend on one or more properties of the audio signal being rendered, the set of speakers, and / or other external inputs. According to some embodiments, the existing flexible rendering cost function given in Equation 1 is augmented with these one or more additional dependencies according to the following equation:
number
[0237] In Equation 28, the term
number
number
number
number
number
number
number
number
[0238]
number
[0239]
number
[0240]
number
[0241] Using the new cost function defined in Equation 28, the optimal set of activations can be found by minimizing and possibly post-regularizing with respect to g as specified previously in Equations 28a and 28b.
[0242] In some cases, the cost term C j by a ducking module, such as ducking module 400 of FIG. 6,
number
[0243] FIG. 17 is a flowchart outlining an example of a method that may be performed by an apparatus or system such as that shown in FIG. 1A. The blocks of method 1700 are not necessarily performed in the order shown, similar to other methods described herein. Further, such a method may include more or fewer blocks than those illustrated and / or described. The blocks of method 1700 may be performed by one or more devices that may be (or include) a control system such as control system 160 shown in FIG. 1A).
[0244] In this implementation, block 1705 includes receiving audio data by the control system via an interface system. In this example, the audio data includes one or more audio signals and associated spatial data. According to this implementation, the spatial data indicates an intended perceived spatial position corresponding to the audio signal. In some cases, the intended perceived spatial position may be explicit, such as indicated by position metadata such as Dolby Atmos position metadata. In other examples, the intended perceived spatial position may be implicit; for example, the intended perceived spatial position may be an assumed location associated with a channel according to a Dolby 5.1, Dolby 7.1, or another channel-based audio format. In some examples, block 1705 includes the rendering module of the control system receiving audio data via the interface system.
[0245] According to this example, block 1710 includes rendering, by the control system, the audio data for playback through a set of loudspeakers in the environment to generate a rendered audio signal. In this example, rendering each of the one or more audio signals included in the audio data includes determining the relative activation of a set of loudspeakers in the environment by optimizing a cost function. According to this example, the cost is a function of a model of the perceived spatial location of the audio signal when played on the set of loudspeakers in the environment. In this example, the cost is also a function of a measure of proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker in the set of loudspeakers. In this implementation, the cost is also a function of one or more additional dynamically configurable features. In this example, the dynamically configurable functionality is based on one or more of the following: the proximity of the loudspeaker to one or more listeners, the proximity of the loudspeaker to an attractive force location (where attractive force is a factor that favors relatively higher loudspeaker activation closer to the attractive force location), the proximity of the loudspeaker to a repulsive force location (where repulsive force is a factor that favors relatively lower loudspeaker activation closer to the repulsive force location), the performance of each loudspeaker relative to other loudspeakers in the environment, the synchronicity of the loudspeaker relative to other loudspeakers, wake word performance, or echo cancellation performance.
[0246] In this example, block 1715 includes providing the rendered audio signals to at least some of the loudspeakers of the set of loudspeakers of the environment via an interface system.
[0247] According to some examples, the model of perceptual spatial location may generate binaural responses corresponding to audio object locations at the left and right ears of a listener. Alternatively or additionally, the model of perceptual spatial location may place the perceptual spatial location of an audio signal reproduced from a set of loudspeakers at the center of gravity of the positions of the set of loudspeakers weighted by the associated activation gains of the loudspeakers.
[0248] In some examples, the one or more additional dynamically configurable features may be based at least in part on the level of the one or more audio signals. In some instances, the one or more additional dynamically configurable features may be based at least in part on the spectrum of the one or more audio signals.
[0249] Some examples of method 1700 include receiving loudspeaker layout information. In some examples, the one or more additional dynamically configurable features may be based at least in part on the location of each of the loudspeakers in the environment.
[0250] Some examples of method 1700 include receiving loudspeaker specification information. In some examples, the one or more additional dynamically configurable features may be based at least in part on the capabilities of each loudspeaker, which may include one or more of a frequency response, a playback level limit, or parameters of one or more loudspeaker dynamics processing algorithms.
[0251] According to some examples, the one or more additional dynamically configurable features may be based at least in part on measurements or estimates of acoustic transmission from each loudspeaker to the other loudspeakers. Alternatively or additionally, the one or more additional dynamically configurable features may be based at least in part on one or more human listener or speaker locations in the environment. Alternatively or additionally, the one or more additional dynamically configurable features may be based at least in part on measurements or estimates of acoustic transmission from each loudspeaker to the listener or speaker location. The acoustic transmission estimate may be based at least in part on walls, furniture, or other objects that may be between each loudspeaker and the listener or speaker location, for example.
[0252] Alternatively or additionally, the one or more additional dynamically configurable features may be based at least in part on the object locations of one or more non-loudspeaker objects or landmarks in the environment. In some such implementations, the one or more additional dynamically configurable features may be based at least in part on measurements or estimates of acoustic transmission from each loudspeaker to the object or landmark locations.
[0253] By employing one or more well-defined additional cost terms to implement flexible rendering, numerous new and useful behaviors can be achieved. All example behaviors listed below are cast in terms of penalizing certain loudspeakers under certain conditions that are deemed undesirable. As a result, these loudspeakers are less activated in the spatial rendering of a set of audio signals. While in many of these cases it might be conceivable to simply lower the volume of the undesired loudspeakers independently of any modifications to the spatial rendering, such a strategy could significantly degrade the overall balance of the audio content. For example, certain components of the mix could become completely inaudible. In contrast, the disclosed embodiments integrate these penalties into the core optimization of the rendering, allowing the rendering to adapt and perform the best possible spatial rendering using the remaining speakers that are less penalized. This is a much more elegant, adaptable, and effective solution.
[0254] Example use cases include, but are not limited to: Providing a more balanced spatial presentation around the listening area - It has been found that spatial audio is best presented across loudspeakers that are approximately the same distance from the intended listening area. Costs can be structured so that loudspeakers that are significantly closer or farther than the average distance from the loudspeakers to the listening area are penalized, reducing their activation; Moving the audio away from or towards the listener or talker - When a user of the system intends to speak to a smart voice assistant of the system or associated with the system, it may be beneficial to create a cost that penalizes loudspeakers that are closer to the talker, in this way these loudspeakers will be activated less, which allows their associated microphones to hear the talker better; - To provide a single listener with a more intimate experience that minimizes playback levels for others in the listening space, speakers far from the listener's location can be heavily penalized so that only the speakers closest to the listener are activated most significantly; Moving the audio away from or towards a landmark, zone, or area Certain locations near listening spaces may be considered sensitive, such as baby rooms, infant beds, offices, reading areas, study areas, etc. In such cases, a cost may be established that penalizes the use of speakers near this location, zone or area. Alternatively, for the same (or similar) case as above, a system of speakers may have generated measurements of sound transmission from each speaker into the baby room, especially if one of the speakers (with attached or associated microphone) is in the baby room itself. In this case, rather than using the physical proximity of the speakers to the baby room, a cost may be constructed that penalizes the use of speakers with a high measured sound transmission into this room; and / or Optimal use of speaker capabilities The capabilities of different loudspeakers can vary significantly. For example, one common smart speaker may only include a single 1.6" full-range driver with limited low-frequency capabilities, while another smart speaker includes a much more capable 3" woofer. These capabilities are generally reflected in the frequency response of the speaker, and thus the set of responses associated with a speaker may be utilized in cost terms. Speakers with lower capabilities at certain frequencies compared to other speakers, as measured by their frequency response, may be penalized and therefore activated less. In some implementations, such frequency response values may be stored with the smart loudspeaker and then reported to a computational unit responsible for optimizing the flexible rendering; Many speakers include two or more drivers, each responsible for reproducing a different frequency range. For example, one common smart speaker is a two-way design, including a woofer for lower frequencies and a tweeter for higher frequencies. Typically, such speakers include crossover circuitry to split the full-range audio signal into appropriate frequency ranges and send them to the respective drivers. Alternatively, such speakers may provide the Flexible Renderer playback access to each individual driver, as well as information about each individual driver's capabilities, such as frequency response. By applying cost terms such as those described immediately above, in some examples, the Flexible Renderer may automatically construct a crossover between two drivers based on their relative capabilities at different frequencies; The above examples using frequency response focus on the inherent capabilities of the speaker, but may not accurately reflect the capabilities of the speaker as placed in the listening environment. In certain cases, the measured frequency response of the speaker at the intended listening position may be available through some calibration procedure. Such measurements may be used instead of pre-calculated responses to better optimize the use of the speaker. For example, a particular speaker may be inherently very competent at certain frequencies, but due to its placement (e.g., behind a wall or furniture), may produce a very limited response at the intended listening position. Capturing this response and feeding the measurement into the appropriate cost term can prevent significant activation of such a speaker; Frequency response is only one aspect of a loudspeaker's reproduction capabilities. Many small loudspeakers begin to distort as the reproduction level increases, especially for low frequencies, and then reach their excursion limits. To reduce such distortion, many loudspeakers implement dynamics processing that constrains the reproduction level below some limiting thresholds that may be variable over frequency. If a speaker is near or at these thresholds, while other speakers participating in flexible rendering are not, it makes sense to reduce the signal level in the limiting speaker and redirect this energy to other, less burdened speakers. Such behavior can be achieved automatically according to some embodiments by appropriately configuring the associated cost terms. Such cost terms may include one or more of the following: Monitoring the global playback volume in relation to a loudspeaker's limit threshold, e.g. loudspeakers whose volume level is closer to their limit threshold may be penalized more; Monitoring dynamic signal levels, possibly varying over frequency, in relation to loudspeaker limiting thresholds, also possibly varying over frequency. For example, loudspeakers whose monitored signal levels are closer to their limiting thresholds may be penalized more; Direct monitoring of loudspeaker dynamics processing parameters, such as limiting gain. In some such instances, loudspeakers whose parameters indicate more limiting may be penalized more; Monitoring the actual instantaneous voltage, current, and power being supplied by the amplifier to the loudspeaker to determine whether the loudspeaker is operating within its linear range. For example, loudspeakers that are operating less linearly may be penalized more; - Smart speakers with integrated microphones and interactive voice assistants typically employ some type of echo cancellation to reduce the level of the audio signal played from the speaker that is picked up by the recording microphone. The greater this reduction, the more likely the speaker is to hear and understand talkers in the space. If the echo canceller residual is consistently high, this may be an indication that the speaker is being driven into a nonlinear region where the echo path becomes difficult to predict. In such cases, it may make sense to redirect signal energy away from the speaker, and therefore a cost term that takes echo canceller performance into account may be beneficial. Such a cost term may assign a higher cost to speakers with poor performance for their associated echo canceller; To achieve predictable imaging when rendering spatial audio across multiple loudspeakers, it is generally required that playback across a set of loudspeakers be reasonably synchronized over time. While this is naturally the case for wired loudspeakers, with a large number of wireless loudspeakers, synchronization can be difficult and the final results may be variable. In such cases, each loudspeaker may be able to report its relative degree of synchronization with the target, and this degree may then be reflected in the synchronization cost term. In some such examples, loudspeakers that are less synchronized may be penalized more and therefore excluded from rendering. Additionally, strict synchronization may not be required for certain types of audio signals, e.g., components of an audio mix intended to be diffuse or omnidirectional. In some implementations, the components may be tagged as such in metadata, and the synchronization cost term may be modified so that the penalty is reduced.
[0255] Further examples of embodiments are now described. Similar to the proximity costs defined in Equations 25a and 25b, the new cost function term
number
number
number
number
[0256] Combining Equations 29a and 29b with matrix-quadratic versions of the CMAP and FV cost functions given in Equation 26 yields a potentially useful implementation of the general extended cost function (in some embodiments) given in Equation 28:
number
[0257] With this definition of the new cost function terms, the overall cost function remains a matrix quadratic equation, and the optimal set of activations g opt , we can obtain:
number
[0258] Weight term w ij , for each of the loudspeakers, a given consecutive penalty value
number
number
number
number
[0259] If all loudspeakers are penalized, it is often convenient to subtract a minimum penalty from all weight terms in post-processing, so that at least one of the speakers is not penalized.
number
[0260] As noted above, there are many possible use cases that can be realized using the new cost function terms described herein (and similar new cost function terms employed according to other embodiments). More specific details will now be provided using three examples: moving audio toward a listener or talker, moving audio away from a listener or talker, and moving audio away from a landmark.
[0261] In a first example, what is referred to herein as a "gravitational force" is used to attract audio to a location, which in some examples may be a listener or talker location, a landmark location, a furniture location, etc. This location may be referred to herein as a "gravitational force" or "attractor location." As used herein, a "gravitational force" is a factor that favors relatively higher loudspeaker activation in proximity to the attractional force. According to this example, a weight w ij is the fixed attractor location
number
number
[0262] To demonstrate the use case of "pulling" audio to a listener or talker, j = 20, β j Set =3,
number
number
number
[0263] In the second and third examples, a "repulsion force" is used to "keep" audio away from a location, which may be a person's location (e.g., a listener location, a talker location, etc.) or another location, such as a landmark location, a furniture location, etc. In some examples, a repulsion force may be used to keep audio away from an area or zone of a listening environment, such as an office area, a reading area, a bed or sleeping area (e.g., a crib or bedroom). According to some such examples, a specific location may be used to represent a zone or area. For example, a location representing a crib may be an estimated location of an infant's head, an estimated sound source location corresponding to the infant, etc. This location may be referred to herein as a "repulsion force location" or "repulsion location." As used herein, a "repulsion force" is a factor that favors relatively low loudspeaker activation in proximity to a repulsion force location. According to this example, similar to the attractive forces of Equations 35a and 35b, a fixed repulsion location
number
number
[0264] To demonstrate the use case of keeping audio away from the listener or talker, an example is j = 5, β j Set =2,
number
number
number
[0265] A third example use case is to "steer" audio away from acoustically sensitive landmarks, such as the door to a sleeping infant's room.
number
[0266] In a further example of method 1700 of FIG. 17, a use case is responding to a selection of two or more audio devices in an audio environment for ducking or other audio processing corresponding to a ducking solution, and applying a penalty to the two or more audio devices corresponding to the ducking solution. Following the previous example, the selection of the two or more audio devices is, in some embodiments, based on a value f , which is a unitless parameter that controls the degree to which audio processing changes occur on audio device i. i Many combinations are possible. In one simple example, the penalty assigned to audio device i for ducking purposes is w ij =f i, may be directly selected as weights. In some examples, one or more of such weights may be determined by a ducking module, such as ducking module 400 of FIG. 6. In some such examples, ducking solution 480 provided to a renderer may include one or more of such weights. In other examples, these weights may be determined by the renderer. In some such examples, one or more of such weights may be determined by the renderer in response to ducking solution 480. According to some examples, one or more of such weights may be determined according to an iterative process, such as method 800 of FIG. 8.
[0267] In addition to the previous example of determining weights, in some implementations, the weights may be determined as follows:
number
[0268] In the above formula, α j , β j , τ j represent adjustable parameters that indicate the overall strength of the penalty, the steepness of the onset of the penalty, and the extent of the penalty, respectively, as explained above with reference to Equation 33.
[0269] In the previous example, the unitless parameter f that describes the ducking solution i As an alternative to , the improvement in the speech-to-echo ratio at audio device i is expressed directly in decibels. i was also introduced. Expressing the solution in this way, the penalty can also be determined as follows:
number
[0270] where f i is the dB term s i is replaced by converting it to a value in the range 0 to infinity. iAs becomes more negative, the penalty increases, thereby moving more audio away from,device i.
[0271] In some examples (such as the previous two equations), the adjustable parameter α j , β j , τ j One or more of may be determined by a ducking module, such as ducking module 400 of FIG. 6. In some such examples, ducking solution 480 provided to a renderer may include one or more of such adjustable parameters. In other examples, the one or more adjustable parameters may be determined by the renderer. In some such examples, the one or more adjustable parameters may be determined by the renderer in response to ducking solution 480. According to some examples, one or more of such adjustable parameters may be determined according to an iterative process, such as method 800 of FIG. 8.
[0272] The aforementioned ducking penalty can be understood as part of a combination of multiple penalty conditions resulting from multiple simultaneous use cases. For example, audio can be "pushed away" from sensitive landmarks using penalties such as those described in Equations 35c-d, while the term f as determined by the decision aspect i Or it can still be "away" from microphone locations where it is desired to use s1 to improve SER.
[0273] Aspects of some disclosed implementations include a system or device configured (e.g., programmed) to perform one or more of the disclosed methods and a tangible computer-readable medium (e.g., disc) storing code for implementing one or more of the disclosed methods, or steps thereof. For example, the system may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including one or more of the disclosed methods, or steps thereof. Such a general-purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more of the disclosed methods (or steps thereof) in response to data asserted thereto.
[0274] Some disclosed embodiments are implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform the required processing on audio signal(s), including implementing one or more disclosed methods. Alternatively, some embodiments (or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) programmed and / or otherwise configured with software or firmware to perform any of the various operations, including one or more disclosed methods or steps thereof. Alternatively, elements of some disclosed embodiments are implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform one or more disclosed methods or steps thereof, the system also including other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to perform one or more disclosed methods or steps thereof would typically be coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device.
[0275] Another aspect of some disclosed implementations is a computer-readable medium (e.g., a disc or other tangible storage medium) that stores code (e.g., executable code) for performing any embodiment of one or more of the disclosed methods or steps thereof.
[0276] While particular embodiments and applications are described herein, it will be apparent to those skilled in the art that many variations to the embodiments and applications described herein are possible without departing from the scope of the material described and claimed herein. While particular implementations have been shown and described, it should be understood that the present disclosure should not be limited to the specific embodiments described and shown or to the specific methods described.
[0277] Various aspects of the invention can be understood from the following enumerated exemplary embodiments (EEE).
[0278] <eee1>1. A method of audio processing, comprising: receiving, by a control system, output signals from one or more microphones in the audio environment, the output signals including signals corresponding to a person's current speech; determining, by the control system, in response to the output signals, based at least in part on the audio device location information and the echo management system information, one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for two or more audio devices in the audio environment, wherein the audio processing modifications include reducing loudspeaker playback levels for one or more loudspeakers in the audio environment; applying, by a control system, one or more types of audio processing modifications; A method comprising:
[0279] <eee2>The method of EEE1, wherein at least one of the one or more types of audio processing modifications corresponds to an increase in signal-to-echo ratio.
[0280] <eee3>The method of any one of EEE1 to EEE2, wherein the echo management system information includes a model of echo management system performance.
[0281] <eee4>The method of EEE3, wherein the model of echo management system performance includes an acoustic echo canceller (AEC) performance matrix.
[0282] <eee5>The method according to EEE3 or EEE4, wherein the model of the echo management system performance includes a measure of expected echo return loss enhancement provided by the echo management system.
[0283] <eee6>The method of any one of EEE1 to 5, wherein the one or more types of audio processing modifications are based at least in part on acoustic models of inter-device echo and intra-device echo.
[0284] <eee7>The method of any one of EEE1 to 6, wherein the one or more types of audio processing modifications are based at least in part on an interaudibility matrix.
[0285] <eee8>The method of any one of EEE1 to 7, wherein the one or more types of audio processing modifications are based at least in part on the estimated location of the person.
[0286] <eee9>The method according to EEE8, wherein the estimated location of the person is based at least in part on output signals from a plurality of microphones in the audio environment.
[0287] <eee10>The method of EEE8 or EEE9, wherein the one or more types of audio processing modifications include modifying a rendering process to warp the rendering of the audio signal away from the estimated location of the person.
[0288] <eee11>The method of any one of EEE1 to 10, wherein the one or more types of audio processing modifications are based at least in part on listening intent.
[0289] <eee12>The method of claim 3, wherein the listening objective includes at least one of a spatial component or a frequency component.
[0290] <eee13>13. The method of any one of EEE1 to 12, wherein the one or more types of audio processing modifications are based at least in part on one or more constraints.
[0291] <eee14>The method according to EEE13, wherein the one or more constraints are based on a perceptual model.
[0292] <eee15>The method of any one of EEE13 to EEE14, wherein the one or more constraints include one or more of an audio content energy preservation, an audio spatiality preservation, an audio energy vector, or a regularization constraint.
[0293] <eee16>16. The method of any one of EEE1 to 15, further comprising updating at least one of an acoustic model of the audio environment or a model of echo management system performance after applying one or more types of audio processing modifications.
[0294] <eee17>17. The method of any one of EEE1 to 16, wherein determining one or more types of audio processing modifications is based on optimizing a cost function.
[0295] <eee18>31. The method of any one of EEE1 to 17, wherein the one or more types of audio processing modifications include spectral modifications.
[0296] <eee19>8. The method of claim 8, wherein the spectral modification comprises reducing the level of the audio data in a frequency band from 500 Hz to 3 KHz.
[0297] <eee20>The method of any one of EEE1 to 19, wherein the current utterance comprises a wake word utterance.
[0298] <eee21>An apparatus configured to perform the method according to any one of EEE1 to 20.
[0299] <eee22>A system configured to perform the method described in any one of EEE1 to 20.
[0300] <eee23> One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the methods described in any one of EEE1 to 20.< / eee23>
Claims
1. 1. A method of audio processing, comprising: receiving, by a control system, output signals from one or more microphones in an audio environment, the output signals including signals corresponding to a person's current speech; determining, by the control system, in response to the output signals, based at least in part on audio device location information and echo management system information, one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for two or more audio devices in the audio environment, the audio processing modifications comprising reducing loudspeaker playback levels for one or more loudspeakers in the audio environment; and causing the control system to apply the one or more types of audio processing modifications; reducing loudspeaker playback levels of one or more loudspeakers in the audio environment is formulated as an optimization problem based on a model of echo management system performance and an interaudibility matrix; the optimization problem includes minimizing total echo power in a residual of the model of echo management system performance by varying loudspeaker playback level reductions of one or more loudspeakers in the audio environment. method.
2. The method of claim 1 , wherein at least one of the one or more types of audio processing modifications corresponds to an increase in signal-to-echo ratio.
3. The method of claim 1 , wherein the model of echo management system performance includes an acoustic echo canceller (AEC) performance matrix.
4. The method of claim 1 , wherein the model of echo management system performance includes a measure of expected echo return loss enhancement provided by the echo management system.
5. The method of claim 1 , wherein the one or more types of audio processing modifications are based at least in part on acoustic models of inter-device and intra-device echoes.
6. The method described in claim 1, wherein the mutual audibility matrix represents the energy of the echo path between the audio devices.
7. The method of claim 1 , wherein the one or more types of audio processing modifications are based at least in part on an estimated location of the person.
8. The method of claim 7 , wherein the estimated location of the person is based at least in part on output signals from multiple microphones in the audio environment.
9. The method of claim 7 , wherein the one or more types of audio processing modifications include modifying a rendering process to warp a rendering of an audio signal away from the estimated location of the person.
10. The method of claim 1 , wherein the one or more types of audio processing modifications are based at least in part on listening intent.
11. The method of claim 10 , wherein the listening objective includes at least one of a spatial component or a frequency component.
12. The method of claim 1 , wherein the one or more types of audio processing modifications are based at least in part on one or more constraints.
13. The method of claim 12 , wherein the one or more constraints are based on a perceptual model.
14. The method of claim 12 , wherein the one or more constraints include one or more of an audio content energy preservation, an audio spatiality preservation, an audio energy vector, or a regularization constraint.
15. 10. The method of claim 1, further comprising updating at least one of an acoustic model of the audio environment or a model of echo management system performance after applying the one or more types of audio processing modifications.
16. The method of claim 1 , wherein determining the one or more types of audio processing modifications is based on optimizing a cost function.
17. The method of claim 1 , wherein the one or more types of audio processing modifications include spectral modifications.
18. 18. The method of claim 17, wherein the spectral modification comprises reducing the level of audio data in a frequency band from 500 Hz to 3 KHz.
19. The method of claim 1 , wherein the current utterance comprises a wake word utterance.
20. An apparatus configured to perform a method according to any one of claims 1 to 19.
21. A system configured to perform a method according to any one of claims 1 to 19.
22. 20. One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of any one of claims 1 to 19.
Citation Information
Patent Citations
Multichannel acoustic coupling evaluator
JP1998257585A
Method and system for in-hall loudspeaking
JP1999055784A
Acoustic echo cancellation control for distributed audio devices
WO2021021857A1