Acoustic echo cancellation control for distributed audio devices

By coordinating audio equipment systems and adjusting audio processing variations, the problem of low signal-to-echo ratio in multi-device environments was resolved, improving full-duplex audio capabilities and user experience.

CN114207715BActive Publication Date: 2025-08-12DOLBY LABORATORIES LICENSING CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080055689.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-21
Filing Date
2020-07-29
Publication Date
2025-08-12
Estimated Expiration
2040-07-29

AI Technical Summary

Technical Problem

Existing audio devices face challenges in full-duplex audio capabilities, especially when multiple audio devices are present simultaneously, making it difficult to effectively improve the signal-to-echo ratio and impacting the user experience.

Method used

By coordinating the audio equipment system, selective adjustments can be made to audio processing, including reducing the amplifier's reproduction level, altering the rendering process, inserting spectral correction, and dynamic range compression, to improve the signal-to-echo ratio at the microphone.

Benefits of technology

It improves the signal-to-echo ratio at the microphone in multi-audio device environments, enhances full-duplex audio capabilities, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114207715B_ABST
    Figure CN114207715B_ABST
Patent Text Reader

Abstract

An audio processing method may involve receiving an output signal from each of a plurality of microphones in an audio environment, the output signal corresponding to current utterances of a person, and determining one or more aspects of contextual information about the person based on the output signal, including an estimated current proximity of the person to one or more microphone locations. The method may involve selecting two or more loudspeaker-equipped audio devices based at least in part on the one or more aspects of the contextual information, determining one or more types of audio processing changes to apply to audio data rendered to the loudspeaker feeds of the audio devices, and causing the one or more types of audio processing changes to be applied. In some examples, the audio processing changes have the effect of increasing a speech-to-echo ratio at one or more microphones.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 705,897, filed on July 21, 2020, U.S. Provisional Patent Application No. 62 / 705,410, filed on June 25, 2020, U.S. Provisional Patent Application No. 62 / 971,421, filed on February 7, 2020, U.S. Provisional Patent Application No. 62 / 950,004, filed on December 18, 2019, U.S. Provisional Patent Application No. 62 / 880,122, filed on July 30, 2019, U.S. Provisional Patent Application No. 62 / 880,113, filed on July 30, 2019, European Patent Application No. 19212391.7, filed on November 29, 2019, and Spanish Patent Application No. P201930702, filed on July 30, 2019, all of which are incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure relates to systems and methods for coordinating (orchestrating) and implementing audio devices (eg, smart audio devices) and controlling the rendering of audio by the audio devices. Background Art

[0004] Audio devices, including but not limited to smart audio devices, have been widely deployed and are becoming a common feature of many homes. Although existing systems and methods for controlling audio devices provide benefits, improved systems and methods would still be desirable.

[0005] Symbols and terminology

[0006] Throughout this disclosure, including in the claims, "speaker" and "loudspeaker" are used synonymously to refer to any sound-emitting transducer (or set of transducers) driven by a single speaker feed. A typical headphone includes two speakers. A speaker can be implemented to include multiple transducers (e.g., a woofer and a tweeter), which can be driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) can undergo different processing in different circuit branches coupled to different transducers.

[0007] Throughout this disclosure, including in the claims, expressions referring to “operating on” a signal or data (e.g., filtering, scaling, transforming, or applying a gain to a signal or data) are used broadly to refer to operating directly on a signal or data or on a processed version of a signal or data (e.g., a version of a signal that has been initially filtered or preprocessed before being operated on).

[0008] Throughout this disclosure, including in the claims, the expression "system" is used in a broad sense to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to a plurality of inputs, where the subsystem generates M inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.

[0009] Throughout this disclosure, including in the claims, the term "processor" is used in a broad sense to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, video, or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors that are programmed and / or otherwise configured to perform pipeline processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0010] Throughout this disclosure, including in the claims, the terms "couple" or "coupled" are used to refer to either a direct or indirect connection. Thus, if a first device couples to a second device, that connection may be through a direct connection or through an indirect connection via other devices and connections.

[0011] As used herein, a "smart device" is an electronic device that can operate interactively and / or autonomously to some extent, and is typically configured to communicate with one or more other devices (or networks) via various wireless protocols such as Bluetooth, Zigbee, near-field communication, Wi-Fi, Light Fidelity (Li-Fi), 3G, 4G, 5G, etc. Several notable types of smart devices are smart phones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smart watches, smart bracelets, smart key chains, and smart audio devices. The term "smart device" may also refer to a device that exhibits certain properties of ubiquitous computing, such as artificial intelligence.

[0012] The expression "smart audio device" is used herein to refer to a smart device, which can be a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of a virtual assistant's functionality). A single-purpose audio device is a device (e.g., a television (TV) or a mobile phone) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera) and is largely or primarily designed to perform a single purpose. For example, while a TV can typically play (and is considered capable of playing) audio from program material, in most cases, modern TVs run some kind of operating system on which applications (including applications for watching TV) run natively. Similarly, the audio input and output in a mobile phone can do many things, but these are all served by applications running on the phone. In this sense, a single-purpose audio device with speaker(s) and microphone(s) is typically configured to run local applications and / or services to directly use the speaker(s) and microphone(s). Some single-purpose audio devices can be configured to be grouped together to enable audio playback in certain zones or user-configured areas.

[0013] A common type of multi-purpose audio device is an audio device that implements at least some aspects of a virtual assistant's functionality, although other aspects of the virtual assistant's functionality may be implemented by one or more other devices, such as one or more servers, with which the multi-purpose audio device is configured to communicate. Such a multi-purpose audio device may be referred to herein as a "virtual assistant." A virtual assistant is a device (e.g., a smart speaker or voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera). In some examples, a virtual assistant may provide the ability to use multiple devices (other than the virtual assistant) for applications that are cloud-enabled in some sense or that are otherwise not fully implemented in or on the virtual assistant itself. In other words, at least some aspects of the virtual assistant's functionality (e.g., speech recognition functionality) may be implemented (at least in part) by one or more servers or other devices, with which the virtual assistant may communicate via a network (e.g., the Internet). Virtual assistants may sometimes work together, for example, in a discrete and conditionally defined manner. For example, two or more virtual assistants may work together in the sense that one of them (e.g., the virtual assistant that is most confident of having heard the wake word) responds to the wake word. In some implementations, connected virtual assistants may form a constellation, which may be managed by a master application, which may be (or implement) the virtual assistant.

[0014] As used herein, a "wake-up word" is used broadly to refer to any sound (e.g., a human-spoken word or other sound) that a smart audio device is configured to wake up in response to detecting ("hearing") the sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "waking up" means that the device enters a state of waiting (in other words, listening) for a voice command. In some instances, what may be referred to herein as a "wake-up word" may include more than one word, for example, a phrase.

[0015] As used herein, the expression "wake-up word detector" refers to a device that is configured (or includes software that instructs to configure the device) to continuously search for an alignment between real-time sound (e.g., speech) features and a trained model. Typically, a wake-up word event is triggered whenever the wake-up word detector determines that the probability of detecting a wake-up word exceeds a predefined threshold. For example, the threshold may be a predetermined threshold that is adjusted to give a reasonable compromise between a false acceptance rate and a false rejection rate. Following a wake-up word event, the device may enter a state (which may be referred to as an "awake" state or an "attention" state) in which the device listens for commands and passes received commands to a larger, more computationally intensive recognizer.

[0016] As used herein, the expression "microphone position" refers to the position of one or more microphones. In some examples, a single microphone position may correspond to a microphone array residing in a single audio device. For example, a microphone position may be a single position corresponding to an entire audio device including one or more microphones. In some such examples, a microphone position may be a single position corresponding to the centroid of a microphone array of a single audio device. However, in some instances, a microphone position may be the position of a single microphone. In some such examples, an audio device may only have a single microphone. Summary of the Invention

[0017] Some disclosed embodiments provide a method for managing the listener or "user" experience to improve a key criterion for successful full-duplex at one or more audio devices. This criterion is called the signal-to-echo ratio (SER), also referred to herein as the speech-to-echo ratio, and can be defined as the ratio between the voice (or other desired) signal captured from the environment (e.g., a room) by one or more microphones and the echo from output program content, interactive content, etc. presented on the audio device including the one or more microphones. Many audio devices designed for audio environments may have built-in loudspeakers and microphones while providing other functionality. However, other audio devices in an audio environment may have one or more loudspeakers but no microphone(s), or one or more microphones but no loudspeaker(s). In certain use cases or scenarios, some embodiments intentionally avoid using (or do not primarily use) the loudspeaker(s) closest to the user. Alternatively or additionally, some embodiments may cause one or more other types of audio processing changes to the audio data rendered by one or more loudspeakers of the audio environment in order to increase the SER at one or more microphones of the environment.

[0018] Some embodiments are configured to implement a system including coordinated (orchestrated) audio devices, which in some embodiments may include smart audio devices. According to some such embodiments, two or more smart audio devices are (or are configured to implement) wake-up word detectors. Thus, in such examples, multiple microphones (e.g., asynchronous microphones) are available. In some instances, each microphone may be included in at least one of the smart audio devices, or configured to communicate with at least one of the smart audio devices. For example, at least some of the microphones may be discrete microphones (e.g., in home appliances) that are not included in any smart audio device but are configured to communicate with at least one of the smart audio devices (so that their output can be captured by at least one of the smart audio devices). In some embodiments, each wake-up word detector (or each smart audio device that includes a wake-up word detector) or another subsystem of the system (e.g., a classifier) is configured to estimate the zone in which a person is located by applying a classifier driven by multiple acoustic features from at least some microphones (e.g., asynchronous microphones). In some embodiments, the goal may not be to estimate the exact location of a person, but to form a robust estimate of a discrete zone that includes the person's current location.

[0019] In some embodiments, a person (also referred to herein as a "user"), a smart audio device, and a microphone are in an audio environment (e.g., the user's home, car, or place of business) where sound can propagate from the user to the microphone, and the audio environment can include predetermined zones. According to some examples, the environment can include at least the following zones: a food preparation area; a dining area; an open area of a living space; a TV area of a living space (including a TV sofa); and the like. During system operation, it is assumed that the user is physically located in one of the zones (the "user's zone") at any given time, and the user's zone may change from time to time.

[0020] In some examples, the microphones can be asynchronous (e.g., digitally sampled using different sampling clocks) and randomly positioned (or at least not in predetermined locations, symmetrical arrangements, on a grid, etc.). In some instances, the user's zone can be estimated by a data-driven approach driven at least in part by a plurality of high-level features derived from at least one of the wake-up word detectors. In some examples, these features (e.g., wake-up word confidence and reception level) can consume very little bandwidth and can be transmitted (e.g., asynchronously) to the device implementing the classifier with very little network load.

[0021] Aspects of some embodiments relate to implementing smart audio devices and / or coordinating smart audio devices.

[0022] Aspects of some disclosed embodiments include a system configured (e.g., programmed) to perform one or more disclosed methods or steps thereof, and a tangible, non-transitory computer-readable medium (e.g., a disk or other tangible storage medium) that implements non-transitory storage of data, the tangible, non-transitory computer-readable medium storing code for performing one or more disclosed methods or steps thereof (e.g., code executable to perform one or more disclosed methods or steps thereof). For example, some disclosed embodiments may be or include a programmable general-purpose processor, a digital signal processor, or a microprocessor that is programmed and / or otherwise configured with software or firmware to perform any of a variety of operations on data, including one or more disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system that includes an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more disclosed methods (or steps thereof) in response to data asserted thereto.

[0023] In some embodiments, a control system can be configured to implement one or more methods disclosed herein, such as one or more audio session management methods. Some such methods involve (e.g., by a control system) receiving an output signal from each of a plurality of microphones in an audio environment. In some examples, each of the plurality of microphones resides in a microphone position in the audio environment. In some instances, the output signal includes a signal corresponding to a person's current speech. According to some examples, the output signal includes a signal corresponding to non-speech audio data such as noise and / or echo.

[0024] Some such methods involve determining one or more aspects of contextual information related to the person based on the output signal (e.g., by a control system). In some examples, the contextual information includes an estimated current location of the person and / or an estimated current proximity of the person to one or more microphone locations. Some such methods involve selecting two or more audio devices of an audio environment based at least in part on one or more aspects of the contextual information. In some embodiments, each of the two or more audio devices includes at least one loudspeaker.

[0025] Some such methods involve determining (e.g., by a control system) one or more types of audio processing changes to be applied to audio data rendered to loudspeaker feeds of the two or more audio devices. In some examples, the audio processing changes have the effect of increasing a speech-to-echo ratio at one or more microphones. Some such methods involve causing the one or more types of audio processing changes to be applied.

[0026] According to some embodiments, the one or more types of audio processing changes can reduce the loudspeaker reproduction level of the loudspeakers of the two or more audio devices. In some embodiments, at least one of the audio processing changes of the first audio device can be different from the audio processing change of the second audio device. In some examples, selecting the two or more audio devices of the audio environment (e.g., by a control system) can involve selecting N loudspeaker-equipped audio devices of the audio environment, where N is an integer greater than 2.

[0027] In some embodiments, selecting two or more audio devices for the audio environment can be based at least in part on an estimated current location of the person relative to at least one of a microphone location or a location of a loudspeaker-equipped audio device. According to some such embodiments, the method can involve determining a nearest loudspeaker-equipped audio device that is closest to the person's estimated current location or to a microphone location that is closest to the person's estimated current location. In some such examples, the two or more audio devices can include the nearest loudspeaker-equipped audio device.

[0028] In some examples, the one or more types of audio processing changes involve altering a rendering process to distort the rendering of the audio signal away from the estimated current location of the person. In some embodiments, the one or more types of audio processing changes may involve spectral modification. According to some such embodiments, the spectral modification may involve reducing the level of audio data in a frequency band between 500 Hz and 3 kHz.

[0029] In some implementations, the one or more types of audio processing changes may involve inserting at least one gap into at least one selected frequency band of the audio playback signal. In some examples, the one or more types of audio processing changes may involve dynamic range compression.

[0030] In accordance with some embodiments, selecting the two or more audio devices can be based at least in part on a signal-to-echo ratio estimate for one or more microphone locations. For example, selecting the two or more audio devices can be based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold. In some instances, determining the one or more types of audio processing changes can be based on an optimization of a cost function that is based at least in part on the signal-to-echo ratio estimate. For example, the cost function can be based at least in part on rendering performance. In some embodiments, selecting the two or more audio devices can be based at least in part on a proximity estimate.

[0031] In some examples, the method can involve determining (e.g., by a control system) a plurality of current acoustic features from the output signal of each microphone and applying a classifier to the plurality of current acoustic features. According to some embodiments, applying the classifier can involve applying a model trained on previously determined acoustic features, the acoustic features being derived from a plurality of previous utterances made by the person in a plurality of user zones in an environment.

[0032] In some such examples, determining one or more aspects of contextual information related to the person can involve determining an estimate of a user zone in which the person is currently located based at least in part on output from the classifier. According to some embodiments, the estimate of the user zone can be determined without reference to the geometric positions of the plurality of microphones. In some instances, the current utterance and the previous utterance can be or can include a wake word utterance.

[0033] According to some embodiments, the one or more microphones may reside in multiple audio devices in the audio environment. However, in other instances, the one or more microphones may reside in a single audio device in the audio environment. In some examples, at least one of the one or more microphone locations may correspond to multiple microphones of a single audio device. Some disclosed methods may involve selecting at least one microphone based on one or more aspects of the contextual information.

[0034] At least some aspects of the present disclosure can be implemented by methods such as audio session management methods. As noted elsewhere herein, in some instances, the methods can be implemented, at least in part, by control systems such as those disclosed herein. Some such methods involve receiving an output signal from each of a plurality of microphones in an audio environment. In some examples, each of the plurality of microphones resides in a microphone location in the audio environment. In some instances, the output signal includes a signal corresponding to a person's current speech. According to some examples, the output signal includes a signal corresponding to non-speech audio data such as noise and / or echo.

[0035] Some such methods involve determining one or more aspects of contextual information related to the person based on the output signal. In some examples, the contextual information includes an estimated current location of the person and / or an estimated current proximity of the person to one or more microphone locations. Some such methods involve selecting two or more audio devices of an audio environment based at least in part on one or more aspects of the contextual information. In some embodiments, each of the two or more audio devices includes at least one loudspeaker.

[0036] Some such methods involve determining one or more types of audio processing changes to be applied to audio data rendered to loudspeaker feed signals of the two or more audio devices. In some examples, the audio processing changes have the effect of increasing a speech-to-echo ratio at one or more microphones. Some such methods involve causing the one or more types of audio processing changes to be applied.

[0037] According to some embodiments, the one or more types of audio processing changes can cause a reduction in the loudspeaker reproduction level of the loudspeakers of the two or more audio devices. In some embodiments, at least one of the audio processing changes of the first audio device can be different from the audio processing change of the second audio device. In some examples, selecting the two or more audio devices of the audio environment can involve selecting N loudspeaker-equipped audio devices of the audio environment, where N is an integer greater than 2.

[0038] In some embodiments, selecting two or more audio devices for the audio environment can be based at least in part on an estimated current location of the person relative to at least one of a microphone location or a location of a loudspeaker-equipped audio device. According to some such embodiments, the method can involve determining a nearest loudspeaker-equipped audio device that is closest to the person's estimated current location or to a microphone location that is closest to the person's estimated current location. In some such examples, the two or more audio devices can include the nearest loudspeaker-equipped audio device.

[0039] In some examples, the one or more types of audio processing changes involve altering a rendering process to distort the rendering of the audio signal away from the estimated current location of the person. In some embodiments, the one or more types of audio processing changes may involve spectral modification. According to some such embodiments, the spectral modification may involve reducing the level of audio data in a frequency band between 500 Hz and 3 kHz.

[0040] In some implementations, the one or more types of audio processing changes may involve inserting at least one gap into at least one selected frequency band of the audio playback signal. In some examples, the one or more types of audio processing changes may involve dynamic range compression.

[0041] In accordance with some embodiments, selecting the two or more audio devices can be based at least in part on a signal-to-echo ratio estimate for one or more microphone locations. For example, selecting the two or more audio devices can be based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold. In some instances, determining the one or more types of audio processing changes can be based on an optimization of a cost function that is based at least in part on the signal-to-echo ratio estimate. For example, the cost function can be based at least in part on rendering performance. In some embodiments, selecting the two or more audio devices can be based at least in part on a proximity estimate.

[0042] In some examples, the method can involve determining a plurality of current acoustic features from the output signal of each microphone and applying a classifier to the plurality of current acoustic features. According to some embodiments, applying the classifier can involve applying a model trained on previously determined acoustic features, the acoustic features being derived from a plurality of previous utterances made by the person in a plurality of user zones in an environment.

[0043] In some such examples, determining one or more aspects of contextual information related to the person can involve determining an estimate of a user zone in which the person is currently located based at least in part on output from the classifier. According to some embodiments, the estimate of the user zone can be determined without reference to the geometric positions of the plurality of microphones. In some instances, the current utterance and the previous utterance can be or can include a wake word utterance.

[0044] According to some embodiments, the one or more microphones may reside in multiple audio devices in the audio environment. However, in other instances, the one or more microphones may reside in a single audio device in the audio environment. In some examples, at least one of the one or more microphone locations may correspond to multiple microphones of a single audio device. Some disclosed methods may involve selecting at least one microphone based on one or more aspects of the contextual information.

[0045] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as the memory devices described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, and the like. Thus, some innovative aspects of the subject matter described in this disclosure may be implemented in non-transitory media having software stored thereon.

[0046] For example, the software may include instructions for controlling one or more devices to perform a method involving receiving an output signal from each of a plurality of microphones in an audio environment. In some examples, each of the plurality of microphones resides in a microphone location in the audio environment. In some instances, the output signal includes a signal corresponding to a person's current speech. According to some examples, the output signal includes a signal corresponding to non-speech audio data such as noise and / or echo.

[0047] Some such methods involve determining one or more aspects of contextual information related to the person based on the output signal. In some examples, the contextual information includes an estimated current location of the person and / or an estimated current proximity of the person to one or more microphone locations. Some such methods involve selecting two or more audio devices of an audio environment based at least in part on one or more aspects of the contextual information. In some embodiments, each of the two or more audio devices includes at least one loudspeaker.

[0048] Some such methods involve determining one or more types of audio processing changes to be applied to audio data rendered to loudspeaker feed signals of the two or more audio devices. In some examples, the audio processing changes have the effect of increasing a speech-to-echo ratio at one or more microphones. Some such methods involve causing the one or more types of audio processing changes to be applied.

[0049] According to some embodiments, the one or more types of audio processing changes can cause a reduction in the loudspeaker reproduction level of the loudspeakers of the two or more audio devices. In some embodiments, at least one of the audio processing changes of the first audio device can be different from the audio processing change of the second audio device. In some examples, selecting the two or more audio devices of the audio environment can involve selecting N loudspeaker-equipped audio devices of the audio environment, where N is an integer greater than 2.

[0050] In some embodiments, selecting two or more audio devices for the audio environment can be based at least in part on an estimated current location of the person relative to at least one of a microphone location or a location of a loudspeaker-equipped audio device. According to some such embodiments, the method can involve determining a nearest loudspeaker-equipped audio device that is closest to the person's estimated current location or to a microphone location that is closest to the person's estimated current location. In some such examples, the two or more audio devices can include the nearest loudspeaker-equipped audio device.

[0051] In some examples, the one or more types of audio processing changes involve altering a rendering process to distort the rendering of the audio signal away from the estimated current location of the person. In some embodiments, the one or more types of audio processing changes may involve spectral modification. According to some such embodiments, the spectral modification may involve reducing the level of audio data in a frequency band between 500 Hz and 3 kHz.

[0052] In some implementations, the one or more types of audio processing changes may involve inserting at least one gap into at least one selected frequency band of the audio playback signal. In some examples, the one or more types of audio processing changes may involve dynamic range compression.

[0053] In accordance with some embodiments, selecting the two or more audio devices can be based at least in part on a signal-to-echo ratio estimate for one or more microphone locations. For example, selecting the two or more audio devices can be based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold. In some instances, determining the one or more types of audio processing changes can be based on an optimization of a cost function that is based at least in part on the signal-to-echo ratio estimate. For example, the cost function can be based at least in part on rendering performance. In some embodiments, selecting the two or more audio devices can be based at least in part on a proximity estimate.

[0054] In some examples, the method can involve determining a plurality of current acoustic features from the output signal of each microphone and applying a classifier to the plurality of current acoustic features. According to some embodiments, applying the classifier can involve applying a model trained on previously determined acoustic features, the acoustic features being derived from a plurality of previous utterances made by the person in a plurality of user zones in an environment.

[0055] In some such examples, determining one or more aspects of contextual information related to the person can involve determining an estimate of a user zone in which the person is currently located based at least in part on output from the classifier. According to some embodiments, the estimate of the user zone can be determined without reference to the geometric positions of the plurality of microphones. In some instances, the current utterance and the previous utterance can be or can include a wake word utterance.

[0056] According to some embodiments, the one or more microphones may reside in multiple audio devices in the audio environment. However, in other instances, the one or more microphones may reside in a single audio device in the audio environment. In some examples, at least one of the one or more microphone locations may correspond to multiple microphones of a single audio device. Some disclosed methods may involve selecting at least one microphone based on one or more aspects of the contextual information.

[0057] The details of one or more embodiments of the subject matter described in this specification are set forth in the following drawings and description. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Note that the relative dimensions of the following figures may not be drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1A Represents an audio environment according to an example.

[0059] Figure 1B Another example of an audio environment is shown.

[0060] Figure 2A is a block diagram illustrating an example of components of an apparatus capable of implementing various aspects of the present disclosure.

[0061] Figure 2B is a flow chart including blocks of an audio session management method according to some implementations.

[0062] Figure 3A is a block diagram of a system configured to implement separate rendering control and listening or capturing logic across multiple devices.

[0063] Figure 3B is a block diagram of a system according to another disclosed embodiment.

[0064] Figure 3C is a block diagram of an embodiment configured to implement an energy balancing network according to one example.

[0065] Figure 4 is a diagram illustrating an example of audio processing that may increase the speech-to-echo ratio at one or more microphones of an audio environment.

[0066] Figure 5 is a diagram illustrating another type of audio processing that may increase the speech-to-echo ratio at one or more microphones of an audio environment.

[0067] Figure 6 Another type of audio processing that may increase the speech-to-echo ratio at one or more microphones of an audio environment is illustrated.

[0068] Figure 7 is a diagram illustrating another type of audio processing that may increase the speech-to-echo ratio at one or more microphones of an audio environment.

[0069] Figure 8 is a diagram of an example in which the audio device to be turned down may not be the audio device closest to the person who is speaking.

[0070] Figure 9 The diagram illustrates a situation where a device with very high SER is very close to the user.

[0071] Figure 10 This is an overview of what can be done by Figure 2A The apparatus shown in the figure is a flowchart of an example of a method performed by the apparatus.

[0072] Figure 11 is a block diagram of elements configured to implement one example of an embodiment of a zone classifier.

[0073] Figure 12 This is an overview of what can be done by Figure 2AA flowchart of an example of a method performed by an apparatus such as apparatus 200.

[0074] Figure 13 This is an overview of what can be done by Figure 2A Flowchart of another example of a method performed by apparatus 200 or the like.

[0075] Figure 14 This is an overview of what can be done by Figure 2A Flowchart of another example of a method performed by apparatus 200 or the like.

[0076] Figure 15 and Figure 16 is a diagram illustrating a set of example speaker activations and object rendering positions.

[0077] Figure 17 This is an overview of what can be done by Figure 2A The apparatus or system shown in the flowchart is an example of a method performed by the apparatus or system.

[0078] Figure 18 is a diagram of speaker activation in an example embodiment.

[0079] Figure 19 is a diagram of object rendering locations in an example embodiment.

[0080] Figure 20 is a diagram of speaker activation in an example embodiment.

[0081] Figure 21 is a diagram of object rendering locations in an example embodiment.

[0082] Figure 22 is a diagram of speaker activation in an example embodiment.

[0083] Figure 23 is a diagram of object rendering locations in an example embodiment. DETAILED DESCRIPTION

[0084] Currently, designers typically view audio devices as a single interface point for audio, which can be a mix of entertainment, communication, and information services. Using audio for notifications and voice control offers the advantage of avoiding visual or physical distractions. The ever-expanding device landscape is fragmented, with more systems competing for our ears.

[0085] In all forms of interactive audio, the problem of improving full-duplex audio capabilities remains a challenge. When there is audio output in a room that is unrelated to the transmission or information-based capture in the room, it is desirable to remove this audio from the captured signal (e.g., through echo cancellation and / or echo suppression). Some disclosed embodiments provide methods and management for improving the user experience of signal-to-echo ratio (SER), which is a key criterion for successful full-duplex at one or more devices.

[0086] Such an embodiment is expected to be useful in situations where there is more than one audio device within acoustic range of a user, so that each audio device will be able to present appropriately loud audio program material at the user for the desired entertainment, communication, or information service. The value of such an embodiment is expected to be particularly high when there are three or more audio devices in similar proximity to the user.

[0087] Rendering applications is sometimes the primary function of an audio device, and therefore it is sometimes desirable to use as many audio output devices as possible. If the audio device is closer to the user, the audio device may be more advantageous in terms of the ability to accurately locate the sound or transmit a specific audio signal to the user and to image it. However, if these audio devices include one or more microphones, the audio device may also be more suitable for picking up the user's voice. When considering the challenge of the signal-to-echo ratio, it is seen that if a device closer to the user or moving toward the device is implemented in simplex (input only) mode, the signal-to-echo ratio will be significantly improved.

[0088] In various disclosed embodiments, the audio device may have a built-in speaker and microphone while providing other functions (e.g., Figure 1A Some disclosed embodiments implement the concept of intentionally not primarily using the loudspeaker(s) closest to the user in certain situations.

[0089] Consider that in the disintermediation between connected operating systems or applications (e.g., cloud-based applications), many different types of devices (that enable audio input, output, and / or real-time interaction) may be included. Examples of such devices include wearable devices, home audio devices, mobile devices, automobiles and mobile computing devices, and smart speakers. Smart speakers may include network-connected speakers and microphones for cloud-based services. Other examples of such devices may incorporate speakers and / or microphones, including lamps, clocks, televisions, personal assistant devices, refrigerators, and trash cans. Some embodiments are particularly relevant to situations where there is a common platform for multiple audio devices to orchestrate an audio environment via an orchestration device, such as a smart home hub or another device configured for audio session management, which may be referred to herein as an audio session manager. Some such implementations may involve commands between the audio session manager and a locally implemented software application, the language of the commands not being device-specific, but rather involving the orchestration device routing audio content to and from people and places specified by the software application. Some embodiments implement methods for dynamically managing rendering, including, for example, pushing sound away from the nearest device and maintaining constraints for spatial imaging, and / or methods for locating a user in a zone, and / or methods for mapping and positioning devices relative to each other and the user.

[0090] Typically, a system comprising multiple smart audio devices needs to indicate when it hears a "wake word" (as defined above) from a user and is paying attention (in other words, listening) to a command from the user.

[0091] Figure 1A 1 represents an audio environment according to an example. Some disclosed embodiments may be particularly useful in scenarios where there are many audio devices (e.g., as disclosed herein) in an environment (e.g., a living or work space) that can transmit sound and capture audio. Figure 1A system.

[0092] Figure 1A is a diagram of an audio environment (living space) including a system comprising a set of smart audio devices (devices 1.1) for audio interaction, speakers (1.3) for audio output, and controllable lights (1.2). As with other disclosed embodiments, Figure 1AThe types, numbers, and arrangements of the elements in the examples are merely examples. Other embodiments may provide more, fewer, and / or different elements. In some instances, one or more microphones 1.5 may be part of or associated with one of the devices 1.1, the light 1.2, or the speaker 1.3. Alternatively or in addition, one or more microphones 1.5 may be attached to another part of the environment, for example, to a wall, a ceiling, a piece of furniture, a household appliance, or another device of the environment. In the examples, each device 1.1 includes (and / or is coupled to) at least one microphone 1.5. Although Figure 1A Not shown, but some audio environments may include one or more cameras. According to some disclosed embodiments, one or more devices of the audio environment (e.g., devices configured for audio session management, such as one or more devices 1.1, devices implementing an audio session manager, a smart home hub, etc.) may be able to estimate where (e.g., in which zone of a living space) a user (1.4) issuing a wake word, command, etc. is located. Figure 1A One or more devices of the system shown in FIG. 3 (e.g., device 1.1 thereof) may be configured to implement various disclosed embodiments. Using various methods, information may be collectively obtained from the devices of FIG. 3 to provide an estimate of the location of the user who uttered the wake word. According to some disclosed methods, the location of the user who uttered the wake word may be obtained from the device. Figure 1A The microphones 1.5 collectively obtain information and provide the information to a device implementing a classifier (e.g., a device configured for audio session management), which is configured to provide an estimate of the location of the user who spoke the wake-up word.

[0093] In living spaces (e.g. Figure 1A In some examples, a person's "context" may include or correspond to a user zone or an estimate of a user zone in which the person is currently located. Figure 1A In this example, the user area includes:

[0094] 1. Kitchen sink and food preparation area (in the upper left area of the living space);

[0095] 2. Refrigerator door (to the right of the sink and food preparation area);

[0096] 3. Dining area (in the lower left area of the living space);

[0097] 4. Open area of living space (to the right of the sink and food preparation area and dining area);

[0098] 5. TV sofa (on the right side of the open area);

[0099] 6. TV itself;

[0100] 7. Table; and

[0101] 8. Door area or entryway (in the upper right area of the living space). Other audio environments may include more, fewer, and / or other types of user zones, such as one or more bedroom zones, garage zones, patio or deck zones, etc.

[0102] According to some embodiments, a system that estimates where a sound (e.g., a wake word or other attention-getting signal) occurs or originates (e.g., determines an uncertain estimate of where a sound occurs or originates) can have some determined confidence (or multiple hypotheses) in the estimate. For example, if a person happens to be near a boundary between user zones of an audio environment, the uncertain estimate of the person's location can include a determined confidence that the person is in each zone. In some conventional implementations of voice interfaces, the voice assistant's voice is required to be uttered from only one location at a time, which forces a single selection of a single location (e.g., Figure 1A However, based on simple imaginary role playing, it is clear that (in such conventional implementations) the likelihood that the selected location of the source of the assistant's speech (i.e., the location of the speaker included in or coupled to the assistant) is the focal point or a natural return response expressing concern is likely to be low.

[0103] Figure 1B Another example of an audio environment is shown. Figure 1B Another audio environment is depicted, comprising a user 101 speaking direct speech 102, and a system comprising a set of smart audio devices 103 and 105, speakers for audio output, and a microphone. The system can be configured according to some disclosed embodiments. The speech spoken by user 101 (sometimes referred to herein as a speaker) can be recognized as a wake-up word by one or more elements of the system.

[0104] More specifically, Figure 1B The system components include:

[0105] 102: direct local speech (generated by user 101);

[0106] 103: Voice assistant device (coupled to one or more speakers). Device 103 is located closer to user 101 than device 105, and therefore device 103 is sometimes referred to as the "near" device and device 105 is referred to as the "far" device;

[0107] 104: multiple microphones in the near device 103 (or coupled to the near device);

[0108] 105: Voice assistant device (coupled to one or more speakers);

[0109] 106: Multiple microphones in the remote device 105 (or coupled to the remote device);

[0110] 107: Household appliances (such as lamps); and

[0111] 108: Multiple microphones in (or coupled to) home appliance 107. In some examples, each of microphones 108 can be configured to communicate with a device configured to implement a classifier, which in some cases can be at least one of devices 103 or 105. In some embodiments, the device configured to implement the classifier can also be a device configured for audio session management, such as a device configured to implement CHASM or a smart home hub.

[0112] Figure 1B The system may also include at least one classifier (e.g., Figure 11 For example, device 103 (or device 105) may include a classifier. Alternatively or additionally, the classifier may be implemented by another device that may be configured to communicate with devices 103 and / or 105. In some examples, the classifier may be implemented by another local device (e.g., a device within environment 109), while in other examples, the classifier may be implemented by a remote device (e.g., a server) located outside of environment 109.

[0113] According to some embodiments, at least two devices (e.g., Figure 1A 1.1. Figure 1B In some embodiments, two devices 103 and 105 (e.g., devices 103 and 105) work together in some manner (e.g., under the control of an orchestration device such as a device configured for audio session management) to deliver sound, because the audio can be controlled jointly across the devices. For example, two devices 103 and 105 can play sound individually or jointly. In a simple case, devices 103 and 105 act as a joint pair for each rendered portion of the audio (e.g., without loss of generality, a stereo signal where one renders substantially L and the other renders substantially R).

[0114] The home appliance 107 (or another device) may include one microphone 108 that is closest to the user 101 and does not include any loudspeakers, in which case there may be a situation where, for that particular audio environment and that particular position of the user 101, there may already be a preferred signal-to-echo ratio or speech-to-echo ratio (SER) that cannot be improved by changing the audio processing of the audio reproduced by the speaker(s) of the device 105 and / or 107. In some embodiments, no such microphone is present.

[0115] Some disclosed embodiments provide detectable and significant SER performance impacts. Some implementations provide such advantages without having to implement various aspects of zone locations and / or dynamically variable rendering. However, some embodiments implement audio processing changes that involve rendering by repelling or "distorting" sound objects (or audio objects) away from the device. In some instances, the reason for distorting audio objects from a specific audio device, location, etc. may be to improve the signal-to-echo ratio at a specific microphone used to capture human speech. Such distortion may involve, but may not be limited to, turning down the playback level of one, two, three or more nearby audio devices. In some cases, the audio processing changes used to improve SER may be notified by zone detection techniques so that one, two or more nearby audio devices (e.g., turned down) to which the audio processing changes are implemented are those devices closest to the user, closest to the specific microphone that will be used to capture the user's speech, and / or closest to the sounds of interest.

[0116] Various aspects of some embodiments involve context, decision, and audio processing changes, which may be referred to herein as "rendering changes." In some examples, these aspects are:

[0117] Context (e.g., location and / or time). In some examples, both location and time are part of the context, and each can be obtained or determined in various ways;

[0118] Decisions (which may involve thresholds or continuous modulation of the change(s). This component may be simple or complex, depending on the particular embodiment. In some embodiments, decisions may be made on a continuous basis, for example, based on feedback. In some instances, the decisions may create system stability, for example, benign feedback stability as described below; and

[0119] Rendering (the nature of the audio processing change(s). Although referred to herein as "rendering," the audio processing change(s) may or may not involve rendering change(s), depending on the particular implementation. In some implementations, there are several options for audio processing changes, ranging from implementations where the audio processing changes are barely perceptible to implementations where the rendering of the audio processing changes is severe and noticeable.

[0120] In some examples, "context" can involve information about both location and intent. For example, context information can include at least a rough idea of the user's location, such as an estimate of the user zone corresponding to the user's current location. Context information can correspond to audio object locations, such as the location of an audio object corresponding to the user's wake word utterance. In some examples, context information can include information about the timing and likelihood of the object or person uttering a sound.

[0121] Examples of scenarios include, but are not limited to, the following:

[0122] A. Know where the possible locations are. This can be based on

[0123] i) weak or low probability detection (e.g., detecting sounds that may be of interest but may or may not be clear enough to act on);

[0124] ii) specific activation (e.g., the wake word is spoken and clearly detected);

[0125] iii) habits and patterns (e.g., based on pattern recognition, e.g., certain locations (such as a sofa near a TV) may be associated with one or more people watching video material on a TV and listening to related audio while sitting on the sofa);

[0126] iv) and / or integration of some other form of proximity sensing based on other means (such as one or more infrared (IR) sensors, cameras, capacitive sensors, radio frequency (RF) sensors, thermal sensors, pressure sensors (e.g., in or on furniture in the audio environment), wearable beacons, etc.); and

[0127] B. Knowing or estimating the likelihood that a sound a person might want to hear has, for example, improved detectability. This may include some or all of the following:

[0128] i) Events based on certain audio detection (such as wake-up word detection);

[0129] ii) events or contexts based on known activities or sequences of events, such as a pause in the display of video content, a space for interaction in scripted Automatic Speech Recognition (ASR)-style interactive content, or a change in activity and / or conversational dynamics in a full-duplex communication activity (e.g., a pause in one or more participants in a conference call);

[0130] iii) additional sensory input from other modalities;

[0131] iv) The choice to continually improve listening in some way – either by improving preparation or improving listening.

[0132] The key difference between A (knowing where the likely location is) and B (knowing or estimating the likelihood of a desired sound, e.g., with improved detectability) is that A may involve specific location information or knowledge, but not necessarily knowing whether there is anything else to listen to, while B may be more focused on specific timing or event information without necessarily knowing exactly where to listen. Some aspects of A and B can certainly overlap, for example, weak or full detection of a wake word will have information about both location and timing.

[0133] For some use cases, it may be important that the "context" involves information about both the location (e.g., the location of the person and / or nearby microphones) and the timing of desired listening. This contextual information can drive one or more associated decisions and one or more possible audio processing changes (e.g., one or more possible rendering changes). Thus, various embodiments allow for many possibilities based on various types of information that can be used to form a context.

[0134] Next, the "decision" aspect is described. For example, this aspect may involve determining one, two, three, or more output devices whose associated audio processing is to be changed. A simple way to make such a decision is:

[0135] Given information from a context (e.g., a location and / or event (or a belief that the location is significant or important in some sense)), in some examples, the audio session manager can determine or estimate the distance from the location to some or all audio devices in the audio environment. In some implementations, the audio session manager can also create a set of activation potentials for each loudspeaker (or group of loudspeakers) of some or all audio devices in the audio environment. According to some such examples, the set of activation potentials can be determined as [f_1, f_2, ..., f_n] and, without loss of generality, lie in the range [0..1]. In another example, the result of the decision can describe a target speech-to-echo ratio improvement [s_1, s_2, ..., s_n] for each device for the "rendering" aspect. In further examples, both the activation potential and the speech-to-echo ratio improvement can be generated by the "decision" aspect.

[0136] In some embodiments, the activation potential assigned to the "rendering" aspect should ensure the degree of improvement in SER at the desired microphone position. In some such examples, the maximum value of f_n can indicate that the rendered audio is actively ducked or distorted, or when the value s_n is provided, the audio is limited and ducked to achieve a speech-to-echo ratio of s_n. In some embodiments, intermediate values of f_n close to 0.5 indicate that only moderate rendering changes are required, and indicate that distorting the audio source to these locations may be appropriate. Furthermore, in some embodiments, low values of f_n can be considered not to be critical for attenuation. In some such embodiments, f_n values at or below a threshold level may not be asserted. According to some examples, f_n values at or below a threshold level may correspond to locations that distort the rendering of the audio content. In some instances, according to some processes described later, the loudspeakers corresponding to f_n values at or below the threshold level may even be boosted at the playback level.

[0137] According to some embodiments, the aforementioned method (or one of the alternative methods described below) can be used to create control parameters for each selected audio processing variation for all selected audio devices, e.g., for each device of the audio environment, for one or more devices of the audio environment, for two or more devices of the audio environment, for three or more devices of the audio environment, etc. The selection of audio processing variations can vary depending on the particular embodiment. For example, the determination can involve determining:

[0138] - a group of two or more loudspeakers for which audio processing is changed; and

[0139] - varying the degree of audio processing for the set of two or more loudspeakers. In some examples, the degree of variation can be determined within the context of a designed or defined range and can be based at least in part on the capabilities of one or more loudspeakers in the loudspeaker set. In some instances, the capabilities of each loudspeaker can include frequency response, playback level limits, and / or parameters of one or more loudspeaker dynamics processing algorithms.

[0140] For example, a design choice can be that the best option in a particular situation is to turn down the loudspeakers. In some such examples, the maximum and / or minimum extent of audio processing changes can be determined, for example, the extent to which any loudspeaker will be turned down is limited to a specific threshold, for example, 15dB, 20dB, 25dB, etc. In some such embodiments, the decision can be based on heuristics or logic for selecting one, two, three, or more loudspeakers, and based on the confidence level of the activity of interest, the loudspeaker locations, etc. The decision can be to avoid the audio reproduced by one, two, three, or more loudspeakers by an amount within a range of minimum and maximum values (e.g., between 0 and 20dB). In some instances, the decision method (or system element) can create an activation potential set for each audio device equipped with a loudspeaker.

[0141] In a simple example, the decision process can be as simple as determining that all but one audio device have render activation variation 0, and determining that the one audio device has render activation variation 1. In some examples, the design of the audio processing variation(s) (e.g., ducking) and the extent of the audio processing variation(s) (e.g., time constant, etc.) can be independent of the decision logic. This approach creates a simple and effective design.

[0142] However, alternative embodiments may involve selecting two or more loudspeaker-equipped audio devices and changing the audio processing of at least two, at least three (and in some instances, all) of the two or more loudspeaker-equipped audio devices. In some such examples, at least one of the audio processing changes of the first audio device (e.g., a reduction in playback level) may be different from the audio processing changes of the second audio device. In some examples, the differences between the audio processing changes may be based at least in part on the estimated current position of the person or the microphone position relative to the position of each audio device. According to some such embodiments, the audio processing changes may involve applying different speaker activations at different loudspeaker positions as part of changing the rendering process to distort the rendering of the audio signal away from the estimated current position of the person of interest. In some examples, the differences between the audio processing changes may be based at least in part on loudspeaker capabilities. For example, if the audio processing change involves reducing the audio level in the bass range, such a change may be more actively applied to an audio device that includes one or more loudspeakers capable of high volume reproduction in the bass range.

[0143] Next, more details are described regarding audio processing changes, which may be referred to herein as "rendering changes." This disclosure may sometimes refer to this aspect as "turning down the nearest" (e.g., reducing the volume at which the audio content to be played by the nearest one, two, three, or more speakers is rendered), although (as described elsewhere herein) more generally, in many embodiments, what may be affected is one or more changes in audio processing that are intended to improve the overall estimate, measurement, and / or standard of the signal-to-echo ratio to capture or perceive the desired audio emitter (e.g., the person speaking the wake word). In some cases, the audio processing change (e.g., "turning down" the volume of the rendered audio content) is some continuous parameter of the amount of influence or can be adjusted by the parameter. For example, in the case of turning down the loudspeaker, some embodiments may be able to apply an adjustable (e.g., continuously adjustable) attenuation (dB). In some such examples, the adjustable attenuation may have a first range for just noticeable changes (e.g., 0dB-3dB) and a second range for SER that is particularly effective but may be noticeable to the listener (e.g., 0dB-20dB).

[0144] In some embodiments implementing the described scheme (context, decision, and rendering or rendering change), there may not be a specific hard boundary for "nearest" (e.g., for the loudspeaker or device that is "nearest" to the user, or another individual or system element), and without loss of generality, a rendering change may be or include changing (e.g., continuously changing) one or more of:

[0145] A. Changing the output to reduce a mode of audio output from one or more audio devices, wherein the change(s) in audio output may involve one or more of the following:

[0146] i) reduce the overall level of the audio equipment output (turn down one or more speakers, turn them off);

[0147] ii) shaping the spectrum of the output of the one or more loudspeakers, for example using a substantially linear equalization (EQ) filter designed to produce an output that is different from the audio spectrum desired to be detected. In some examples, if the output spectrum is shaped to detect human speech, the filter may reduce frequencies in the range of approximately 500-3 kHz (e.g., plus or minus 5% or 10% at each end of the frequency range), or shape the loudness to emphasize low and high frequencies, leaving room in the middle frequency band (e.g., in the range of approximately 500-3 kHz);

[0148] iii) changing the ceiling or peak of the output to reduce peak levels and / or reduce distortion products that might otherwise degrade the performance of any echo cancellation as part of the overall system creating the implemented SER for audio detection, e.g., a time domain dynamic range compressor or a multi-band frequency dependent compressor. Such audio signal modification can effectively reduce the amplitude of the audio signal and can help limit the excursion of the loudspeaker;

[0149] iv) spatially manipulating the audio in a manner that tends to reduce the energy or coupling of the output of one or more loudspeakers to one or more microphones at which the system (e.g., the audio processing manager) achieves a higher SER, e.g., as in the "warping" example described herein;

[0150] v) using temporal time slicing or adjustments to create 'gaps' sufficient to obtain segments of audio or periods of sparse time-frequency lower output, such as the gap insertion examples described below; and / or

[0151] vi) altering the audio in some combination of the above ways; and / or

[0152] B. Conserving energy and / or creating continuity across a specific or broad set of listening positions, including, for example, one or more of the following:

[0153] i) In some examples, energy removed from one loudspeaker can be compensated by providing additional energy in or to another loudspeaker. In some instances, the overall loudness remains unchanged, or substantially unchanged. This is not an essential feature, but can be an effective way to allow more rigorous changes to the audio processing of the 'nearest' device or group of devices without losing content. However, continuity and / or conservation of energy may be particularly relevant when dealing with complex audio outputs and audio scenarios; and / or

[0154] ii) Time constants of activation, in particular changes to audio processing may be applied somewhat faster (e.g., 100ms-200ms) than they return to normal (e.g., 1000ms-10000ms), such that the change(s) in audio processing appear intentional (if noticeable at all), but the subsequent return from the change(s) may appear unrelated to any actual event or change (from the user's perspective) and, in some instances, may be so slow as to be barely noticeable.

[0155] Additional examples of how scenarios and decisions may be formulated and determined are now presented.

[0156] Example A:

[0157] (Context) As an example, context information can be mathematically expressed as follows:

[0158] H(a,b), the approximate physical distance between devices a and b in meters:

[0159]

[0160]

[0161] Where D represents the group of all devices in the system. The estimated SERS at each device can be expressed as follows:

[0162]

[0163] Determine H and S:

[0164] H is a property of the physical location of the device and can therefore be determined or estimated by:

[0165] (1) Direct indication by the user, for example, by using a smartphone or tablet device to mark or indicate the approximate location of the device on a floor plan or similar graphical representation of the environment. Such digital interfaces are already commonplace for managing the configuration, grouping, naming, purpose, and identity of smart home devices. For example, such direct indication may be provided by the Amazon Alexa smartphone app, the Sonos S2 controller app, or similar applications.

[0166] (2) Using the measured signal strength (sometimes called received signal strength indication or RSSI) of common wireless communication technologies (such as Bluetooth, Wi-Fi, ZigBee, etc.) to solve the basic trilateration problem to produce an estimate of the physical distance between devices, for example, as disclosed in J. Yang and Y. Chen, "Indoor Localization Using Improved RSS-Based Latency Methods," GLOBECOM 2009-2009 IEEE Global Telecommunications Conference, Honolulu, Hawaii, 2009, pp. 1-6, doi:10.1109 / GLOCOM.2009.5425237, and / or as disclosed in Mardeni, R. and Othman, Shaifull and Nizam, (2010) "Node Positioning in ZigBee Network Using Trilateration Method Based on the Received Signal Strength Indicator" Strength Indicator (RSSI) [Node localization in ZigBee networks based on trilateration of received signal strength indicator (RSSI)]” 46, which are incorporated herein by reference.

[0167] S(a) is an estimate of the speech-to-echo ratio at device a. By definition, the speech-to-echo ratio in dB is given by:

[0168]

[0169] In the above expression, represents an estimate of the speech energy in dB, and represents an estimate of the residual echo energy after echo cancellation in dB. Various methods for estimating these quantities are disclosed herein, for example:

[0170] (1) The speech energy and residual echo energy may be estimated by an offline measurement process performed on a particular device, taking into account the acoustic coupling between the device's microphone and speaker and the performance of the onboard echo cancellation circuitry. In some such examples, the average speech energy level "AvgSpeech" may be determined by the average level of human speech measured by the device at a nominal distance. For example, the device may record the speech of a small number of people standing 1 meter away from the microphone-equipped device during production, and the energies may be averaged to produce AvgSpeech. According to some such examples, the average residual echo energy level "AvgEcho" may be estimated by playing music content from the device during production and running the onboard echo cancellation circuitry to produce an echo residual signal. Averaging the energy of the echo residual signal for a small sample of the music content may be used to estimate AvgEcho. When the device is not playing audio, AvgEcho may instead be set to a nominal low value, such as -96.0 dB. In some such embodiments, the speech energy and residual echo energy may be expressed as follows:

[0171]

[0172]

[0173] (2) According to some examples, the average speech energy can be determined by taking the energy of the microphone signal corresponding to the user's speech as determined by a voice activity detector (VAD). In some such examples, when the VAD does not indicate speech, the average residual echo energy can be estimated by the energy of the microphone signal. If x represents the pulse code modulation (PCM) samples of the microphone of device a at some sampling rate, and V represents a VAD flag that takes the value of 1.0 for samples corresponding to voice activity and 0.0 otherwise, then the speech energy and the residual echo energy can be expressed as follows:

[0174]

[0175]

[0176] (3) In addition to the previous approach, in some embodiments, the energy in the microphone can be treated as a random variable and modeled separately based on the VAD determination. The statistical models Sp and E of the speech and echo energies can be estimated separately using any number of statistical modeling techniques. The average values of both speech and echo in dB used to approximate S(a) can then be derived from Sp and E, respectively. Common methods for achieving this can be found in the field of statistical signal processing, for example:

[0177] Assume a Gaussian distribution of energy and compute biased second-order statistics and

[0178] Construct a histogram of energy values for the discrete bins to produce the underlying multimodal distribution, after which an expectation-maximization (EM) parameter estimation step is applied to the mixture model (e.g., a Gaussian mixture model), using the maximum mean value belonging to any subdistribution in the mixture.

[0179] (Decision) As described elsewhere herein, in various disclosed embodiments, a decision aspect determines which devices receive audio processing modifications, such as rendering modifications, and in some embodiments, an indication of how much modification is required for a desired SER improvement. Some such embodiments may be configured to improve the SER at the device with the best initial SER value, e.g., as determined by finding the maximum value of S across all devices in a set D. Other embodiments may be configured to opportunistically improve the SER at devices that are regularly addressed by a user, as determined based on historical usage patterns. Other embodiments may be configured to attempt to improve the SER at multiple microphone locations, e.g., to select multiple devices for purposes of the following discussion.

[0180] Once one or more microphone positions are determined, in some such embodiments, the expected SER improvement (SERI) may be determined as follows:

[0181] SERI=S(m)-TargetSER[dB]

[0182] In the foregoing expressions, m represents the device / microphone position being improved, and TargetSER represents a threshold value, which can be set by the application in use. For example, a wake-up word detection algorithm can tolerate a lower operating SER than a large vocabulary speech recognizer. Typical values for TargetSER can be around -6dB to 12dB. As mentioned, if S(m) is unknown or not easy to estimate in some embodiments, a preset value may be sufficient based on offline measurements of speech and echo recorded in a typical echoic room or environment. Some embodiments may determine the device for which the audio processing (e.g., rendering) is to be modified by specifying f_n as being in the range of 0 to 1. Other embodiments may involve specifying the degree to which the audio processing (e.g., rendering) should be modified in decibels of speech-echo ratio improvement, s_n, possibly calculated according to the following:

[0183] s n =SERI*f n

[0184] Some embodiments may compute f_n directly from the device geometry, for example, as follows:

[0185]

[0186] In the foregoing expressions, m represents the index of the device to be selected for maximum audio processing (eg, rendering) modification, as described above. Other embodiments may involve other choices of easing or smoothing functions for device geometry.

[0187] Example B (reference user area):

[0188] In some embodiments, the context and decision aspects of the present disclosure will be made in the context of one or more user zones. As described in detail later in this document, a set of acoustic features Can be used to estimate a group of region labels C k The posterior probability p(C k |W(j)), for k = {1...K}, for K different user zones in the environment. The association of each audio device with each user zone can be provided by the user themselves as part of the training process described in this document, or alternatively by means of an application (e.g., the Alexa smartphone app or the Sonos S2 controller smartphone app). For example, some embodiments may associate the jth device with a zone label C k The association of the user area is expressed as z(C k , n)∈[0,1]. In some embodiments, z(C k ,n) and the posterior probability p(C k |W(j)) can be considered as context information. Some embodiments may instead consider the acoustic features W(j) themselves as part of the context. In other embodiments, these quantities (z(C k ,n), posterior probability p(C k One or more quantities in |W(j)) and the acoustic feature W(j) itself) and / or a combination of these quantities may be part of the context information.

[0189] Aspects of the decision making of various embodiments may use quantities associated with one or more user zones in device selection. Where both z and p are available, an example decision may be made as follows:

[0190]

[0191] According to such an embodiment, the device with the highest association with the user zone most likely to contain the user will apply the most audio processing (e.g., rendering) changes. In some examples, δ can be a positive number in the range of [0.5, 4.0]. According to some such examples, δ can be used to spatially control the range of rendering changes. In such an embodiment, if δ is selected to be 0.5, more devices will receive a larger rendering change, while a value of 4.0 will limit the rendering change to only the devices closest to the most likely user zone.

[0192] The inventors also envision another class of embodiments in which acoustic features W(j) are used directly in the decision making. For example, if the wake word confidence score associated with utterance j is w n (j), then the device selection can be performed according to the following expression:

[0193]

[0194] In the foregoing expression, δ has the same interpretation as in the previous example, and further has the utility of compensating for the typical wake word confidence distribution that may occur for a particular wake word system. If most devices tend to report high wake word confidence, then δ may be chosen to be a relatively higher number, such as 3.0, to increase the spatial specificity of the rendering variation application. If the confidence of the wake word drops rapidly as the user is positioned further away from the device, then δ may be chosen to be a relatively low number, such as 1.0 or even 0.5, to include more devices in the rendering variation application. The reader should understand that in some alternative embodiments, formulas similar to the above formulas for acoustic features, such as an estimate of the speech level at the device's microphone, and / or the direct to reverb ratio of the user's speech may be substituted for the wake word confidence.

[0195] Figure 2A is a block diagram illustrating an example of components of an apparatus or system capable of implementing various aspects of the present disclosure. As with the other figures provided herein, Figure 2A The types and quantities of elements shown in the figures are provided as examples only. Other embodiments may include more, fewer, and / or different types and quantities of elements. According to some examples, apparatus 200 may be or may include a device configured to perform at least some of the methods disclosed herein. In some embodiments, apparatus 200 may be or may include a smart speaker, a laptop computer, a cellular phone, a tablet device, a smart home hub, or another device configured to perform at least some of the methods disclosed herein. In some embodiments, apparatus 200 may be configured to implement an audio session manager. In some such embodiments, apparatus 200 may be or may include a server.

[0196] In this example, the device 200 includes an interface system 205 and a control system 210. In some embodiments, the interface system 205 can be configured to communicate with one or more devices that are executing or configured to execute a software application. Such software applications may sometimes be referred to as "applications" or simply "apps" in this document. In some embodiments, the interface system 205 can be configured to exchange control information and associated data related to the application. In some embodiments, the interface system 205 can be configured to communicate with one or more other devices of an audio environment. In some examples, the audio environment can be a home audio environment. In some embodiments, the interface system 205 can be configured to exchange control information and data associated with audio devices of the audio environment. In some examples, the control information and associated data can be related to one or more applications with which the device 200 is configured to communicate.

[0197] In some embodiments, the interface system 205 can be configured to receive audio data. The audio data can include audio signals that are arranged to be reproduced by at least some speakers of the audio environment. The audio data can include one or more audio signals and associated spatial data. For example, the spatial data can include channel data and / or spatial metadata. The interface system 205 can be configured to provide rendered audio signals to at least some of the speakers in the environment. In some embodiments, the interface system 205 can be configured to receive input from one or more microphones in the environment.

[0198] The interface system 205 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more universal serial bus (USB) interfaces). According to some embodiments, the interface system 205 may include one or more wireless interfaces. The interface system 205 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 205 may include a control system 210 and a memory system (e.g., Figure 2A However, in some examples, the control system 210 may include a memory system.

[0199] The control system 210 may include, for example, a general purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0200] In some embodiments, the control system 210 can reside in more than one device. For example, a portion of the control system 210 can reside in a device within one of the environments described herein, and another portion of the control system 210 can reside in a device outside the environment, such as a server, a mobile device (e.g., a smart phone or tablet), etc. In other examples, a portion of the control system 210 can reside in a device within one of the environments described herein, and another portion of the control system 210 can reside in one or more other devices of the environment. For example, the control system functionality can be distributed across multiple smart audio devices of the environment, or can be shared by an orchestration device (such as a device that may be referred to as an audio session manager or smart home hub herein) and one or more other devices of the environment. In some such examples, the interface system 205 can also reside in more than one device.

[0201] In some implementations, the control system 210 can be configured to at least partially perform the methods disclosed herein. According to some examples, the control system 210 can be configured to implement an audio session management method that, in some instances, can involve determining one or more types of audio processing changes to be applied to audio data of loudspeaker feeds rendered to two or more audio devices in an audio environment. According to some implementations, the audio processing changes can have the effect of increasing a speech-to-echo ratio at one or more microphones in the audio environment.

[0202] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as the memory devices described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may reside, for example, on a computer. Figure 2A . Thus, various innovative aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media having software stored thereon. For example, the software may include instructions for controlling at least one device to implement the audio session management method. In some examples, the software may include instructions for controlling one or more audio devices of an audio environment to obtain, process, and / or provide audio data. In some examples, the software may include instructions for determining one or more types of audio processing changes to be applied to audio data of loudspeaker feed signals rendered to two or more audio devices of an audio environment. According to some embodiments, the audio processing changes may have the effect of increasing the speech-to-echo ratio at one or more microphones in the audio environment. For example, the software may be provided by Figure 2AThe control system 210 and other control system components are executed.

[0203] In some examples, apparatus 200 may include Figure 2A Optional microphone system 220 is shown in FIG. Optional microphone system 220 may include one or more microphones. In some embodiments, one or more microphones may be part of or associated with another device (e.g., a speaker of a speaker system, a smart audio device, etc.). In some examples, apparatus 200 may not include microphone system 220. However, in some such embodiments, apparatus 200 may still be configured to receive microphone data from one or more microphones in the audio environment via interface system 210.

[0204] According to some embodiments, the apparatus 200 may include Figure 2A Optional loudspeaker system 225 is shown in . Optional loudspeaker system 225 may include one or more loudspeakers. Loudspeakers may sometimes be referred to herein as "speakers." In some examples, at least some of the loudspeakers of optional loudspeaker system 225 may be positioned arbitrarily. For example, at least some of the loudspeakers of optional loudspeaker system 225 may be placed in positions that do not correspond to any standard-specified loudspeaker layout, such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, Hamasaki 22.2, etc. In some such examples, at least some of the loudspeakers of optional loudspeaker system 225 may be placed in positions that are convenient for the space (e.g., where there is room to accommodate the loudspeakers), but not in any standard-specified loudspeaker layout. In some examples, device 200 may not include optional loudspeaker system 225.

[0205] In some embodiments, the apparatus 200 may include Figure 2A Optional sensor system 230 is shown in . Optional sensor system 230 may include one or more cameras, touch sensors, gesture sensors, motion detectors, and the like. According to some embodiments, optional sensor system 230 may include one or more cameras. In some embodiments, the camera may be a stand-alone camera. In some examples, one or more cameras of optional sensor system 230 may reside in a smart audio device, which may be a single-purpose audio device or a virtual assistant. In some such examples, one or more cameras of optional sensor system 230 may reside in a TV, a mobile phone, or a smart speaker. In some examples, device 200 may not include sensor system 230. However, in some such embodiments, device 200 may still be configured to receive sensor data from one or more sensors in the audio environment via interface system 210.

[0206] In some embodiments, the apparatus 200 may include Figure 2A Optional display system 235 is shown in FIG. Optional display system 235 may include one or more displays, such as one or more light emitting diode (LED) displays. In some instances, optional display system 235 may include one or more organic light emitting diode (OLED) displays. In some examples where device 200 includes display system 235, sensor system 230 may include a touch sensor system and / or a gesture sensor system proximate to one or more displays of display system 235. According to some such embodiments, control system 210 may be configured to control display system 235 to present one or more graphical user interfaces (GUIs).

[0207] According to some examples, the apparatus 200 may be or may include a smart audio device. In some such embodiments, the apparatus 200 may be or may (at least in part) implement a wake-up word detector. For example, the apparatus 200 may be or may (at least in part) implement a virtual assistant.

[0208] Figure 2B is a flow diagram including blocks of an audio session management method according to some embodiments. As with other methods described herein, the blocks of method 250 do not have to be performed in the order indicated. In some embodiments, one or more blocks of method 250 may be performed simultaneously. In addition, some embodiments of method 250 may include more or fewer blocks than shown and / or described. The blocks of method 250 may be performed by one or more devices, which may be (or may include) a control system, such as Figure 2A The control system 210 shown in and described above, or one of the other disclosed control system examples. According to some implementations, the blocks of the method 250 may be performed at least in part by a device implementing what may be referred to herein as an audio session manager.

[0209] According to this example, block 255 involves receiving an output signal from each of a plurality of microphones in the audio environment. In this example, each of the plurality of microphones resides in a microphone position in the audio environment, and the output signal includes a signal corresponding to the person's current utterance. In some instances, the current utterance may be a wake-up word utterance. However, the output signal may also include a signal corresponding to times when the person is not speaking. For example, such a signal may be used to establish a baseline level of echo, noise, etc.

[0210] In this example, block 260 involves determining one or more aspects of contextual information related to the person based on the output signal. In this embodiment, the contextual information includes an estimated current location of the person and / or an estimated current proximity of the person to one or more microphone locations. As described above, the expression "microphone location" as used herein indicates the location of one or more microphones. In some examples, a single microphone location may correspond to a microphone array residing in a single audio device. For example, a microphone location may be a single location corresponding to an entire audio device including one or more microphones. In some such examples, the microphone location may be a single location corresponding to the center of mass of the microphone array of the single audio device. However, in some instances, the microphone location may be the location of a single microphone. In some such examples, the audio device may have only a single microphone.

[0211] In some examples, determining contextual information may involve making an estimate of a user zone in which a person is currently located. Some such examples may involve determining a plurality of current acoustic features from the output signal of each microphone and applying a classifier to the plurality of current acoustic features. Applying the classifier may, for example, involve applying a model trained on previously determined acoustic features, the previously determined acoustic features being derived from a plurality of previous utterances made by the person in a plurality of user zones in an environment. In some such examples, determining one or more aspects of contextual information related to the person may involve determining an estimate of the user zone in which the person is currently located based at least in part on an output from the classifier. In some such examples, the estimate of the user zone may be determined without reference to the geometric positions of the plurality of microphones. According to some examples, the current utterance and the previous utterance may be or may include a wake word utterance.

[0212] According to this embodiment, block 265 involves selecting two or more audio devices of the audio environment based at least in part on one or more aspects of the contextual information, the two or more audio devices each including at least one loudspeaker. In some examples, selecting the two or more audio devices of the audio environment can involve selecting N loudspeaker-equipped audio devices of the audio environment, where N is an integer greater than 2. In some instances, selecting the two or more audio devices of the audio environment or selecting N loudspeaker-equipped audio devices of the audio environment can involve selecting all loudspeaker-equipped audio devices of the audio environment.

[0213] In some examples, selecting two or more audio devices for an audio environment can be based at least in part on an estimated current location of a person relative to a microphone position and / or a location of a loudspeaker-equipped audio device. Some such examples can involve determining the closest loudspeaker-equipped audio device to the person's estimated current location or to a microphone position closest to the person's estimated current location. In some such examples, the two or more audio devices can include the closest loudspeaker-equipped audio device.

[0214] According to some embodiments, selecting the two or more audio devices may be based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold.

[0215] According to this example, block 270 involves determining one or more types of audio processing changes to be applied to audio data rendered to microphone feeds of two or more audio devices. In this embodiment, the audio processing changes have the effect of increasing the speech-to-echo ratio at one or more microphones. In some examples, one or more microphones may reside in multiple audio devices in an audio environment. However, according to some embodiments, one or more microphones may reside in a single audio device in an audio environment. In some examples, the audio processing change(s) may reduce the microphone reproduction level of the microphones of the two or more audio devices.

[0216] According to some examples, at least one of the audio processing changes of the first audio device can be different from the audio processing changes of the second audio device. For example, the audio processing change(s) can result in a first reduction in the loudspeaker reproduction level of a first loudspeaker of the first audio device, and can result in a second reduction in the loudspeaker reproduction level of a second loudspeaker of the second audio device. In some such examples, the reduction in loudspeaker reproduction level can be relatively greater for an audio device that is closer to the person's estimated current position (or the microphone position closest to the person's estimated current position).

[0217] However, the inventors contemplate various types of audio processing changes that may be made in some instances. According to some embodiments, one or more types of audio processing changes may involve altering the rendering process to distort the rendering of the audio signal away from the person's estimated current location (or away from the microphone position closest to the person's estimated current location).

[0218] In some embodiments, the one or more types of audio processing changes may involve spectral modification. For example, spectral modification may involve reducing the level of audio data in a frequency band between 500 Hz and 3 kHz. In other examples, spectral modification may involve reducing the level of audio data in frequency bands having higher maximum frequencies and / or lower minimum frequencies. According to some embodiments, one or more types of audio processing changes may involve inserting at least one gap into at least one selected frequency band of the audio playback signal.

[0219] In some implementations, determining one or more types of audio processing changes can be based on optimization of a cost function based at least in part on a signal-to-echo ratio estimate.In some instances, the cost function can be based at least in part on rendering performance.

[0220] According to this example, block 275 involves causing one or more types of audio processing changes to be applied. In some instances, block 275 may involve applying one or more types of audio processing changes by one or more devices controlling audio processing in the audio environment. In other instances, block 275 may involve causing one or more types of audio processing changes to be applied by one or more other devices of the audio environment (e.g., via commands or control signals from an audio session manager).

[0221] Some embodiments of method 250 may involve selecting at least one microphone based on one or more aspects of contextual information. In some such embodiments, method 250 may involve selecting at least one microphone based on an estimated current proximity of a person to one or more microphone locations. Some embodiments of method 250 may involve selecting at least one microphone based on an estimate of a user zone. In some such embodiments, method 250 may involve implementing, at least in part, a virtual assistant functionality based on microphone signals received from the selected microphone(s). In some such embodiments, method 250 may involve providing a teleconferencing functionality based at least in part on microphone signals received from the selected microphone(s).

[0222] Some embodiments provide a system (including two or more devices, e.g., smart audio devices) configured to perform rendering and mapping, and configured to use software or other manifestations of logic (e.g., including system components that implement the logic) to change audio processing (e.g., turn down one, two, or more nearest speakers). The logic may implement a supervisor, such as a device configured to implement an audio session manager, which in some examples may run separately from the system components configured for rendering.

[0223] Figure 3Ais a block diagram of a system configured to implement separate rendering control and listening or capturing logic across multiple devices. As with other disclosed diagrams, Figures 3A to 3C The number, type, and arrangement of the elements shown are merely examples. Other embodiments may include more, fewer, and / or different types of elements. For example, other embodiments may include more than three audio devices, different types of audio devices, etc.

[0224] Depending on the specific example, Figures 3A to 3C The modules shown in FIG. 1 and other modules shown and described in this disclosure may be implemented via hardware, software, firmware, etc. In some embodiments, one or more of the disclosed modules (in some instances, referred to herein as "elements") may be implemented via a processor as described above. Figure 2A In some such examples, one or more of the disclosed modules may be implemented according to software executed by one or more such control systems.

[0225] Figure 3A The components include:

[0226] Audio devices 302, 303, and 304 (which may be smart audio devices in some examples). According to this example, each of the audio devices 302, 303, and 304 includes at least one loudspeaker and at least one microphone;

[0227] - Element 300 represents a form of content, including audio data, to be played on one or more of audio devices 302, 303, and 304. Content 300 may be linear or interactive content, depending on the particular implementation;

[0228] - Module 301 is configured for audio processing, including but not limited to rendering according to rendering logic. For example, in some embodiments, module 301 may be simply configured to copy the audio (e.g., mono or stereo) of content 300 equally to all three audio devices 302, 303, and 304. In some alternative implementations, one or more of audio devices 302, 303, and 304 may be configured to perform audio processing functions, including but not limited to rendering functions;

[0229] - Element 305 represents a signal assigned to audio devices 302, 303, and 304. In some examples, signal 305 may be or include a speaker feed signal. As described above, in some embodiments, the functionality of module 301 may be implemented via one or more of audio devices 302, 303, and 304, in which case signal 305 may be local to one or more of audio devices 302, 303, and 304. However, the signal may be local to one or more of the audio devices 302, 303, and 304. Figure 3A is shown as a set of speaker feed signals, because some embodiments (e.g., as described below with reference to Figure 4 ) implements a simple final interception or post-processing of the signal 305;

[0230] Element 306 represents the raw microphone signals captured by the microphones of audio devices 302 , 303 and 304 .

[0231] Module 307 is configured to implement microphone signal processing logic, and in some examples, microphone signal capture logic. Because each of audio devices 302, 303, and 304 in this example has one or more microphones, the captured raw signal 306 is processed by module 307. In some embodiments, such as herein, module 307 can be configured to implement echo cancellation and / or echo detection functionality.

[0232] - Element 308 represents the local and / or global echo reference signals supplied by module 301 to module 307. According to this example, module 307 is configured to implement echo cancellation and / or echo detection functionality based on the local and / or global echo reference signals 308. In some embodiments, microphone capture processing and / or processing of raw microphone signals may be distributed on each of the audio devices 302, 303, and 304 along with local echo cancellation and / or detection logic. The specific implementation of the capture and capture processing is not important for calculating and understanding the impact of any changes to the rendering on the overall SER and the concept of the efficacy of the capture processing and logic;

[0233] - Module 309 is the system element that performs the overall mixing or combining of the captured audio (e.g., for the purpose of causing the desired audio to be perceived as emanating from a specific single or widespread location). In some embodiments, module 307 may also provide the mixing functionality of element 309; and

[0234] Module 310 is a system element that implements the final aspect of processing the detected audio in order to make some kind of determination regarding what was said or whether an activity of interest occurred in the audio environment. For example, module 310 can provide automatic speech recognition (ASR) functionality, background noise level and / or type sensing functionality, e.g., context about what people are doing in the audio environment, what the overall noise level in the audio environment is, etc. In some embodiments, some or all of the functionality of module 310 can be implemented outside of the audio environment where audio devices 302, 303, and 304 reside, e.g., in one or more devices (e.g., one or more servers) of a cloud-based service provider.

[0235] Figure 3B is a block diagram of a system according to another disclosed embodiment. In this example, Figure 3B The system shown in FIG. 1 includes Figure 3A components of the system and expands Figure 3A A system may include functionality according to some disclosed embodiments. Figure 3B The system includes elements for implementing context, decision, and rendering action aspects, such as those used in operating a distributed audio system. According to some examples, feedback to the elements for implementing context, decision, and rendering action aspects can cause a confidence level to increase when there is activity (e.g., detected speech), or can cause a sense of activity to be confidently reduced (likelihood of activity is low), and thus return audio processing to its initial state.

[0236] Figure 3B The components include the following:

[0237] - Module 351 is a system element that represents (and implements) contextual steps, for example, for obtaining an indication of where it may be desirable to better detect audio (e.g., to increase the speech-to-echo ratio at one or more microphones), and the likelihood or perception of wanting to listen (e.g., the likelihood that speech (such as a wake word or command) can be captured by one or more microphones). In this example, modules 351 and 353 are implemented via a control system, which in this instance is Figure 2A In some embodiments, blocks 301 and 307 may also be implemented by a control system, which in some instances may be the control system 210. In some embodiments, blocks 356, 357, and 358 may also be implemented by a control system, which in some instances may be the control system 210;

[0238] - Element 352 represents a feedback path to module 351. In this example, feedback 352 is provided by module 310. In some embodiments, feedback 352 may correspond to results from captured audio processing (such as audio processing by ASR) of microphone signals that may be relevant to the determined context—for example, a weak or early detection of a wake word or a sense of some low detection of speech activity may be used to begin to increase confidence or sense of the context in which improved listening is desired (e.g., to increase the speech-to-echo ratio at one or more microphones);

[0239] - Module 353 is the system element in which (or by which) decisions are made regarding which audio device to change audio processing for and by how much. Depending on the particular implementation, module 353 may or may not use specific audio device information, such as the type and / or capabilities of the audio device (e.g., loudspeaker capabilities, echo suppression capabilities, etc.), the possible orientation of the audio device, etc. As described in some examples below, the decision-making process(es) of module 353 may be very different for headphone devices compared to smart speakers or other loudspeakers;

[0240] - Element 354 is the output of module 353, which in this example is a set of control functions. Figure 3B , represented as f_n values, to individual rendering blocks via a control path 355, which may also be referred to as a signal path 355. The set of control functions may be distributed (e.g., via wireless transmission) such that the signal path 355 is local to the audio environment. In this example, the control functions are provided to modules 356, 357, and 358; and

[0241] Modules 356, 357, and 358 are system elements configured to effect changes in audio processing, which may include, but are not limited to, output rendering (the rendering aspect of some embodiments). In this example, modules 356, 357, and 358 are controlled by the control function of output 354 (in this example, the f_n value) when activated. In some implementations, the functionality of modules 356, 357, and 358 can be implemented via block 301.

[0242] exist Figure 3BIn embodiments and other implementations of the present invention, a virtuous cycle of feedback can occur. If the output 352 of element 310 (in some instances, automatic speech recognition or ASR can be implemented) detects speech, even if it is weak (e.g., with low confidence), then according to some examples, the context element 351 can estimate the location based on which microphone(s) in the audio environment captures the sound (e.g., which microphone(s) has the maximum energy excluding the echo). According to some such examples, decision box 353 can select one, two, three or more loudspeakers of the audio environment and activate a small value (e.g., f_n=0.25) related to the rendered change. With an overall 20dB avoidance, this value will then perform a volume reduction of approximately 5dB at the selected (multiple) devices, which can be noticeable to ordinary human listeners. When combined with time constant and / or event detection and with other loudspeakers in the audio environment that reproduce similar content, the reduction in level(s) may be less noticeable. In one example, it may be that the audio device 303 (the audio device closest to the person 311 who is speaking) is turned down. In other examples, both audio devices 302 and 303 can be turned down, in some instances by different amounts (e.g., depending on the estimated proximity to person 311). In other examples, audio devices 302, 303, and 304 can all be turned down, in some instances by different amounts. Due to the reduced playback level of one or more loudspeakers of one or more of audio devices 302, 303, and 304, the speech-to-echo ratio can be increased at one or more microphones near person 311 (e.g., one or more microphones of audio device 303). Thus, if person 311 continues to speak (e.g., repeats the wake word or issues a command), the system can now better “hear” person 311. In some such embodiments, during the next time interval (e.g., during the next few seconds), and in some instances in a continuous manner, the system (e.g., an audio session manager implemented at least in part via blocks 351 and 353) can quickly tend to turn down the volume of one or more loudspeakers near person 311, for example, by selecting f_2=1.

[0243] Figure 3C is a block diagram of an embodiment configured to implement an energy balancing network according to one example. Figure 3C is a block diagram of a system comprising Figure 3B components of the system and will Figure 3B The system is expanded to include an element (eg, element 371 ) for implementing energy compensation (eg, 'turn other devices up a bit').

[0244] In some examples, configured for Figure 3C system (or Figure 3CAn audio session management device (audio session manager) of a system) can evaluate the segment energy lost at a listener (311) due to the effects of audio processing (e.g., reducing the level of one or more selected loudspeakers (e.g., speakers of an audio device receiving a control signal, where f_n>0) applied to increase the speech-to-echo ratio at one or more microphones). The audio session manager can then apply level boosting and / or some other form of energy balancing to other speakers of the audio environment to compensate for the SER-based audio processing changes.

[0245] Often, when rendering somewhat related content and there are related or spectrally similar audio components reproduced by multiple loudspeakers in the audio environment (a simple example is monophonic), then not much energy balancing may be necessary. For example, if there are three loudspeakers in the audio environment, with a distance range ratio of 1 to 2, with 1 being the closest, then if the loudspeakers reproduce the same content, turning down the closest loudspeaker by 6dB will only have a 2dB-3dB impact. And turning off the closest loudspeaker may only have an overall impact of 3dB-4dB on the sound at the listener.

[0246] In more complex cases (e.g., intervening gaps or spatial turns), energy conservation and perceived continuity may, in some examples, take the form of a more multi-factored energy balance.

[0247] exist Figure 3C In some examples, the element(s) used to implement context may simply be the audio level of a weak detection of the wake word (reciprocity of proximity). In other words, one example of determining context may be based on detecting the level of any wake word utterance by the detected echo. This approach may or may not involve actually determining the speech-to-echo ratio, depending on the particular implementation. However, in some examples, simply detecting and evaluating the level of the wake word utterance detected at each of a plurality of microphone positions may provide a sufficient level of context.

[0248] Implemented by system components to implement scenarios (e.g., Figure 3C Examples of the method of (in a system) may include, but are not limited to, the following:

[0249] - When a partial wake word is detected, proximity to the microphone-equipped audio device can be inferred from the wake word confidence. The timing of the wake word utterance can also be inferred from the wake word confidence; and

[0250] In addition to echo cancellation and suppression applied to the raw microphone signal, some audio activity may also be detected. Some embodiments may use a set of energy levels and classifications to determine the likelihood that the audio activity is voice activity (voice activity detection). This process may determine a confidence level or likelihood of voice activity. The location of the voice may be based on the probability of the best microphone for similar interaction situations. For example, a device implementing an audio session manager may know in advance that a microphone-equipped audio device is closest to the user, such as a desktop device located at or near the user's common location, rather than a wall-mounted device that is not near the user's common location.

[0251] System components used to implement decisions (e.g., Figure 3C An example embodiment of a system) is an element configured for determining a confidence value regarding voice activity and for determining which is the nearest microphone-equipped audio device.

[0252] exist Figure 3C In the system (and other embodiments), the amount of audio processing change(s) applied to increase the SER at a location can be a function of the distance and confidence level of the speech activity.

[0253] For example, in Figure 3C Examples of methods for implementing rendering in a system include:

[0254] Just turn down dB; and / or

[0255] Speech band equalization (EQ) (e.g., as described below with reference to Figure 4 described); and / or

[0256] Rendering changes in temporal modulation (as referenced Figure 5 described); and / or

[0257] Use temporal time slicing or adjustments to create (e.g., insert audio content) "gaps" or periods with sparse time-frequency lower output sufficient to obtain audio segments of interest. Figure 9 Some examples are described.

[0258] Figure 4 is a diagram illustrating an example of audio processing that may increase the speech-to-echo ratio at one or more microphones of an audio environment. Figure 4 The figure below provides an example of spectrum correction. Figure 4In some embodiments, spectral modification involves reducing the level of frequencies known to correspond to speech, which in these examples are frequencies in the range of approximately 200 Hz and 10 kHz (e.g., within 5% or 10% of the upper and / or lower frequencies of the range). Other examples can involve reducing the level of frequencies in a different frequency band, for example, between approximately 500 Hz and 3 kHz (e.g., within 5% or 10% of the upper and / or lower frequencies of the range). In some embodiments, frequencies outside of this range can be reproduced at a higher level to at least partially compensate for the reduction in loudness caused by the spectral modification.

[0259] Figure 4 The elements include:

[0260] 601: represents the curve of flat EQ;

[0261] 602: A curve showing partial attenuation of the indicated frequency range. This partial attenuation may be relatively unnoticeable but may have a useful effect on speech detection; and

[0262] 603: Curve showing significantly greater attenuation of the indicated frequency range. Spectral modification such as that represented by curve 603 can have a significant impact on hearing speech. In some instances, aggressive spectral modification such as that represented by curve 603 can provide an alternative to significantly reducing the levels of all frequencies.

[0263] In some examples, the audio session manager may cause audio processing changes corresponding to time-varying spectral modifications, such as the sequence represented by curves 601 , 602 , and 603 .

[0264] According to some examples, one or more spectral modifications may be used in the context of other audio processing changes (e.g., rendering changes) to "warp" the reproduced audio away from a certain location, such as an office, a bedroom, a sleeping baby, etc. The spectral modification(s) used in conjunction with such warping may, for example, reduce the level in the bass frequency range, e.g., in the range of 20 Hz-250 Hz.

[0265] Figure 5 is a diagram illustrating another type of audio processing that can increase the speech-to-echo ratio at one or more microphones of an audio environment. In this example, the vertical axis represents the "f" value ranging from 0 to 1, and the horizontal axis represents time in seconds. Figure 5 is a graph of the activation of a rendering effect over time (indicated by curve 701). In some examples, one or more of modules 356, 357, or 358 may implement Figure 5. According to this example, the asymmetry of the time constant (indicated by curve 701) indicates that the system modulates to the controlled value (f_n) in a short time (e.g., 100 ms to 1 second), but returns to zero from value f_n (also identified as value 703) much more slowly (e.g., 10 seconds or more). In some examples, the time interval between 2 seconds and N seconds can be multiple seconds, for example, in the range of 4 seconds to 10 seconds.

[0266] Figure 5 Also shown is a second activation curve 702 having a staircase shape, with a maximum value in this example equal to f_n. According to this embodiment, the staircase rise corresponds to a sudden change in the level of the content itself (eg speech onset or syllable rate).

[0267] As described above, in some embodiments, temporal time slicing or frequency adjustments can create (e.g., by inserting gaps into the audio content) sufficient to obtain "gaps" in the audio segment of interest or periods with sparse time-frequency output (e.g., increasing or decreasing the degree of "gaps" in the audio content and its perception).

[0268] Figure 6 Another type of audio processing that may increase the speech-to-echo ratio at one or more microphones of an audio environment is illustrated. Figure 6 is an example of a spectrogram of a modified audio playback signal into which an imposed gap has been inserted according to one example. More specifically, to generate Figure 6 , the imposed gaps G1, G2 and G3 are inserted into the frequency band of the playback signal, thereby generating a modified audio playback signal. Figure 6 In the spectrogram shown in , positioning along the horizontal axis indicates time, and positioning along the vertical axis indicates the frequency of the content of the modified audio playback signal at a certain moment.

[0269] The density of points in each small area (each such area centered around a point having vertical and horizontal coordinates) indicates the energy of the content of the modified audio playback signal at the corresponding frequency and time (areas with greater density indicate content with greater energy, and areas with less density indicate content with lower energy). Therefore, gap G1 appears at a time (i.e., in a time interval) earlier than the time when gap G2 or G3 appears (or the time interval in which gap G2 or G3 appears), and gap G1 has been inserted in a frequency band higher than the frequency band in which gap G2 or G3 has been inserted.

[0270] Introducing an imposed gap into the playback signal is distinct from simplex device operation, in which the device pauses the playback stream of the content (e.g., in order to better hear the user and the user's environment). According to some disclosed embodiments, introducing an imposed gap into the playback signal can be optimized to significantly reduce (or eliminate) the perceptibility of artifacts caused by the introduced gap during playback, preferably such that the imposed gap has no or only minimal perceptible effect on the user, but such that the output signal of a microphone in the playback environment indicates the imposed gap (e.g., so that a ubiquitous listening method can be implemented with the gap). By using an imposed gap introduced according to some disclosed embodiments, a ubiquitous listening system can monitor for non-playback sounds (e.g., sounds indicative of background activity and / or noise in the playback environment) even without using an acoustic echo canceller.

[0271] According to some examples, gaps may be inserted in the temporal spectral output from a single channel, which may create a sense of sparsity, i.e., an improved listening ability to "hear through the gaps."

[0272] Figure 7 is a diagram illustrating another type of audio processing that may increase the speech-to-echo ratio at one or more microphones of an audio environment. In this embodiment, the audio processing change involves dynamic range compression.

[0273] This example involves a transition between two extremes of dynamic range limiting. In one case, represented by curve 801, the audio session manager causes no dynamic range control to be applied, while in another case, represented by curve 802, the audio session manager causes a relatively aggressive limiter to be applied. The limiter corresponding to curve 802 can reduce the peak value of the audio output by 10 dB or more. According to some examples, the compression ratio can not exceed 3:1. In some embodiments, curve 802 (or another dynamic range compression curve) can include an inflection point at or about -20 dB from the peak output value of the device (e.g., within + / - 1 dB, within + / - 2 dB, within + / - 3 dB, etc.).

[0274] Next, another example of an embodiment of a system element for implementing rendering is described (the audio processing changes have the effect of increasing the speech-to-echo ratio at one or more microphones, e.g., Figure 3B or Figure 3C In this embodiment, energy balancing is performed. As described above, in a simple example, the audio session manager can estimate the segmental energy of the audio lost at the listener's location or zone due to the impact of other audio processing changes used to increase the SER at one or more microphones of the audio environment. The audio session manager can then add enhancements to other speakers to compensate for this energy lost at the listener's location or zone.

[0275] Often, when rendering somewhat related content and there are related or spectrally similar components across multiple devices (a simple example is mono), then there may not be much to do at all. For example, if there are three loudspeakers with a distance range ratio of 1 to 2, with 1 being closest, then turning down the closest loudspeaker by 6dB (if the loudspeakers are reproducing the same content) will only have a 2dB to 3dB impact. And turning down the closest loudspeaker will probably only have an overall impact of 3dB to 4dB on the sound at the listener's position.

[0276] Aspects of additional embodiments are described next.

[0277] 1. Define the 'nearest' second-order factor

[0278] As the following two examples will illustrate, the measure of "closeness" or "nearest" may not be a simple measure of distance, but may be a scalar ranking involving the estimated speech-to-echo ratio. If the audio devices of the audio environment are not identical, each loudspeaker-equipped audio device may have a different coupling from its loudspeaker(s) to its own microphone(s), which has a strong impact on the echo level in the ratio. Similarly, audio devices may have different microphone arrangements that are relatively more or relatively less suitable for listening, e.g., for detecting sounds from a particular direction, for detecting sounds in or from a particular location in the audio environment. Thus, in some embodiments, the calculation (decision) may take into account more than just proximity and reciprocity of listening.

[0279] Figure 8 8 is a diagram of an example in which the audio device to be turned down may not be the audio device closest to the person who is speaking. In this example, audio device 802 is relatively closer to the person 100 who is speaking than audio device 805. According to some examples, in a situation like Figure 8 In the case shown in , the audio session manager can consider the different baseline SER and audio device characteristics and turn off the device(s) that have the best cost / benefit ratio of reducing the impact of the output on the audio presentation, compared to the benefit of turning down the output to better capture the speech of person 101.

[0280] Figure 8Shown is an example where a greater functionality metric of "nearest" can have complexity and practicality. In this instance, there is a person 101 who makes a sound (speech 102), an audio session manager is configured to capture the sound, and two audio devices 802 and 805, both of which have loudspeakers (806 and 804) and microphones (803 and 807). Given that microphone 803 is so close to loudspeaker 804 on audio device 802, which is closer to person 101, there may not be an amount of lowering the loudspeaker of the device that would produce a feasible SER. In this example, microphone 807 on audio device 805 is configured for beamforming (producing a more favorable SER on average), and therefore lowering the loudspeaker of audio device 805 will have a smaller impact than lowering the loudspeaker of audio device 802. In some such examples, the best decision is to lower loudspeaker 806.

[0281] Will refer to Figure 9 Let's describe another example. In this case, consider the most significant difference in baseline SER that might occur between two devices: one is a headset, and the other is a smart speaker.

[0282] Figure 9 The diagram shows a situation where a device with very high SER is very close to the user. Figure 9 In Figure 1, user 101 is wearing headphones 902 and speaking (emitting sound 102 in a manner that is captured by both microphone 903 on headphones 902 and by the microphone of smart speaker device 904). In this case, smart speaker device 904 may also emit some sound to match the headphones (e.g., for near / far rendering of immersive sound). While headphones 902 are clearly the output device closest to user 101, there is virtually no echo path from the headphones to the nearest microphone 903. Therefore, the SER of this device will be very high, and turning it down would have a significant impact, as the headphones are responsible for almost all of the rendering to the listener. In this case, turning down smart speaker 904 would be more beneficial, albeit only slightly and not to the detriment of the overall rendering change (other nearby listeners hear the sound), and may not dictate an actual action—in terms of turning down the speaker or otherwise changing the audio processing parameters, which could improve the SER of the user's voice pickup in a way that better alters the audio provided in the audio environment—in a sense, this is already very practical due to the inherent device SER in the headphones.

[0283] Regarding devices over a certain size with multiple speakers and distributed microphones, in some cases, a single audio device with many speakers and many microphones can be treated as a series of independent devices that are just rigidly connected. In this case, decisions about tuning down can be applied to individual speakers. Therefore, in some embodiments, the audio session manager can treat this type of audio device as a set of individual microphones and loudspeakers, while in other examples, the audio session manager can treat this type of audio device as a single device with a composite speaker and microphone array. Similarly, it can be seen that there is a duality between treating the speakers on a single device as separate devices and the idea that one rendering method in a single multi-loudspeaker audio device is spatial steering, which necessarily brings different changes to the output of the loudspeakers on the single audio device.

[0284] Regarding the secondary effect of the nearest audio device(s) avoiding the spatial imaging sensitivity from audio devices close to the moving listener, in many cases, it may not make sense to reproduce a particular audio object or rendered material output from the nearest loudspeaker(s) even if the loudspeaker is close to the moving listener. This is simply related to the loudness of the direct audio path being proportional to 1 / r. 2 This is related to the fact that r varies directly, where r is the distance the sound travels, and as a loudspeaker becomes closer to any listener (r->0), the level of the sound reproduced by that loudspeaker becomes less stable with respect to the overall mix.

[0285] In some such instances, it may be advantageous to implement the following embodiments, where (for example):

[0286] - A context is some general listening area (e.g., a couch near a TV) where being able to hear the audio of a program someone is watching on TV is considered always useful;

[0287] - Decision: For a device with speakers on a coffee table in the general listening area (e.g., near the sofa), set f_n=1; and

[0288] -Rendering: The device is turned off and energy is used for rendering elsewhere.

[0289] The effect of this audio processing change is to allow the person on the couch to hear better. If the coffee table is to one side of the couch, this approach will avoid the sensitivity of the listener being close to the audio device. In some instances, while the audio device may be ideally positioned for, say, surround channels, the fact that there may be a level difference of 20dB or more between the couch and the speakers means that unless the exact location of the listener / speaker is known, it may be a good idea to turn down or even turn off the nearest device.

[0290] Figure 10This is an overview of what can be done by Figure 2A 1000. As with other methods described herein, the blocks of method 1000 need not be performed in the order shown. Moreover, this method may include more or fewer blocks than shown and / or described. In this embodiment, method 1000 involves estimating a user's position in an environment.

[0291] In this example, block 1005 involves receiving an output signal from each of a plurality of microphones in the environment. In this instance, each of the plurality of microphones resides in a microphone location of the environment. According to this example, the output signal corresponds to the current utterance of the user. In some examples, the current utterance may be or may include a wake word utterance. For example, block 1005 may, for example, involve a control system (e.g., Figure 2A The control system 120) is connected to the interface system (such as Figure 2A The interface system 205) receives an output signal from each of a plurality of microphones in the environment.

[0292] In some examples, at least some of the microphones in the environment can provide output signals that are asynchronous with respect to output signals provided by one or more other microphones. For example, a first microphone among the plurality of microphones can sample audio data according to a first sampling clock, and a second microphone among the plurality of microphones can sample audio data according to a second sampling clock. In some instances, at least one of the microphones in the environment can be included in, or configured to communicate with, a smart audio device.

[0293] According to this example, block 1010 involves determining a plurality of current acoustic features from the output signals of each microphone. In this example, the "current acoustic features" are the acoustic features derived from the "current utterance" of block 1005. In some embodiments, block 1010 may involve receiving the plurality of current acoustic features from one or more other devices. For example, block 1010 may involve receiving at least some of the plurality of current acoustic features from one or more wake-up word detectors implemented by one or more other devices. Alternatively or additionally, in some embodiments, block 1010 may involve determining the plurality of current acoustic features from the output signals.

[0294] Whether the acoustic features are determined by a single device or multiple devices, the acoustic features can be determined asynchronously. If the acoustic features are determined by multiple devices, the acoustic features will generally be determined asynchronously unless the devices are configured to coordinate the process of determining the acoustic features. If the acoustic features are determined by a single device, in some embodiments, the acoustic features can still be determined asynchronously because the single device can receive the output signal of each microphone at different times. In some examples, the acoustic features can be determined asynchronously because at least some of the microphones in the environment can provide output signals that are asynchronous with respect to the output signals provided by one or more other microphones.

[0295] In some examples, the acoustic features may include a wake word confidence metric, a wake word duration metric, and / or at least one received level metric. The received level metric may indicate a received level of sound detected by the microphone and may correspond to a level of an output signal of the microphone.

[0296] Alternatively or additionally, the acoustic signature may include one or more of the following:

[0297] The average state entropy (purity) of each wake word state along the 1-best (Viterbi) alignment with the acoustic model.

[0298] CTC loss (Connectionist Temporal Classification Loss) for the acoustic model of the wake-up word detector.

[0299] In addition to the wake-up word confidence, the wake-up word detector can also be trained to provide an estimate of the speaker's distance from the microphone and / or an RT60 estimate. The distance estimate and / or RT60 estimate can be acoustic features.

[0300] Instead of or in addition to the wideband received level / power at the microphone, the acoustic feature may be the received level in a plurality of logarithmic / Mel / Bark spaced frequency bands. The frequency bands may vary depending on the particular implementation (e.g., 2 bands, 5 bands, 20 bands, 50 bands, 1 octave bands, or 1 / 3 octave bands).

[0301] • A cepstral representation of the spectral information in the previous point, computed by performing a DCT (Discrete Cosine Transform) on the logarithm of the band power.

[0302] • Band power in a frequency band weighted for human speech. For example, the acoustic features may be based only on a specific frequency band (e.g., 400 Hz to 1.5 kHz). In this example, higher and lower frequencies may be ignored.

[0303] • Voice activity detector confidence per band or bin.

[0304] • The acoustic signature may be based at least in part on long term noise estimation in order to ignore microphones with poor signal to noise ratio.

[0305] Kurtosis as a measure of speech "peakiness." Kurtosis can be an indicator of long reverberation tails.

[0306] Estimated wake word onset. We would expect onset and duration to be equal across all microphones within a frame or so. Outliers can provide clues to unreliable estimates. This assumes some degree of synchronization—not necessarily with the samples—but, for example, with frames within a few tens of milliseconds.

[0307] According to this example, block 1015 involves applying a classifier to the plurality of current acoustic features. In some such examples, applying the classifier can involve applying a model trained on previously determined acoustic features derived from a plurality of prior utterances made by the user in a plurality of user zones in the environment. Various examples are provided herein.

[0308] In some examples, the user zones may include a sink area, a food preparation area, a refrigerator area, a dining area, a sofa area, a TV area, a bedroom area, and / or a doorway area. According to some examples, one or more user zones may be predetermined user zones. In some such examples, one or more predetermined user zones may have been selected by the user during the training process.

[0309] In some embodiments, applying the classifier may involve applying a Gaussian mixture model trained on previous utterances. According to some such embodiments, applying the classifier may involve applying a Gaussian mixture model trained on one or more of normalized wake word confidence, normalized average reception level, or maximum reception level of previous utterances. However, in alternative embodiments, applying the classifier may be based on a different model, such as one of the other models disclosed herein. In some instances, the model may be trained using training data labeled with user zones. However, in some examples, applying the classifier involves applying a model trained using unlabeled training data, where the training data does not have user zones labeled.

[0310] In some examples, the previous utterance may have been or may include a wake-up word utterance. According to some such examples, the previous utterance and the current utterance may have been utterances of the same wake-up word.

[0311] In this example, block 1020 involves determining an estimate of the user zone in which the user is currently located based at least in part on the output from the classifier. In some such examples, the estimate can be determined without reference to the geometric positions of the plurality of microphones. For example, the estimate can be determined without reference to the coordinates of the individual microphones. In some examples, the estimate can be determined without estimating the geometric position of the user.

[0312] Some embodiments of method 1000 may involve selecting at least one speaker based on the estimated user zone. Some such embodiments may involve controlling the at least one selected speaker to provide sound to the estimated user zone. Alternatively or additionally, some embodiments of method 1000 may involve selecting at least one microphone based on the estimated user zone. Some such embodiments may involve providing a signal output by the at least one selected microphone to the smart audio device.

[0313] Figure 11 is a block diagram of components configured to implement one example of an embodiment of a zone classifier. According to this example, system 1100 includes distributed in an environment (e.g., Figure 1A or Figure 1B 1 . In this example, the system 1100 includes a plurality of loudspeakers 1104 in at least a portion of an environment (as illustrated in FIG). In this example, the system 1100 includes a multi-channel loudspeaker renderer 1101. According to this embodiment, the output of the multi-channel loudspeaker renderer 1101 is used as both a loudspeaker drive signal (for driving the loudspeaker feeds of the loudspeakers 1104) and an echo reference. In this embodiment, the echo reference is provided to the echo management subsystem 1103 via a plurality of loudspeaker reference channels 1102, the echo reference including at least some of the loudspeaker feed signals output from the renderer 1101.

[0314] In this embodiment, system 1100 includes multiple echo management subsystems 1103. According to this example, echo management subsystems 1103 are configured to implement one or more echo suppression processes and / or one or more echo cancellation processes. In this example, each echo management subsystem 1103 provides a corresponding echo management output 1103A to one of the wake word detectors 1106. Echo management output 1103A has attenuated echo relative to the input of the corresponding echo management subsystem in echo management subsystems 1103.

[0315] According to this embodiment, the system 1100 includes a system distributed in an environment (e.g., Figure 1A or Figure 1BN microphones 1105 (N is an integer) are provided in at least a portion of an environment (as shown in FIG). The microphones may include array microphones and / or point microphones. For example, one or more smart audio devices located in the environment may include a microphone array. In this example, the output of the microphones 1105 is provided as input to the echo management subsystem 1103. Depending on the embodiment, each echo management subsystem 1103 captures the output of a separate microphone 1105 or a separate group or subset of microphones 1105.

[0316] In this example, the system 1100 includes multiple wake-up word detectors 1106. According to this example, each wake-up word detector 1106 receives audio output from one of the echo management subsystems 1103 and outputs multiple acoustic features 1106A. The acoustic features 1106A output from each echo management subsystem 1103 may include, but are not limited to: a measure of wake-up word confidence, wake-up word duration, and reception level. Although the three arrows depicting three acoustic features 1106A are shown as being output from each echo management subsystem 1103, more or fewer acoustic features 1106A may be output in alternative embodiments. Furthermore, although the three arrows collide with the classifier 1107 along a more or less vertical line, this does not indicate that the classifier 1107 must receive acoustic features 1106A from all wake-up word detectors 1106 at the same time. As described elsewhere herein, in some instances, the acoustic features 1106A may be determined and / or provided to the classifier asynchronously.

[0317] According to this embodiment, system 1100 includes a zone classifier 1107, which may also be referred to as classifier 1107. In this example, the classifier receives a plurality of features 1106A from a plurality of wake-up word detectors 1106 for a plurality (e.g., all) of microphones 1105 in the environment. According to this example, an output 1108 of zone classifier 1107 corresponds to an estimate of the user zone in which the user is currently located. According to some such examples, output 1108 may correspond to one or more posterior probabilities. Based on Bayesian statistics, the estimate of the user zone in which the user is currently located may be or may correspond to a maximum a posteriori probability.

[0318] Next, example implementations of a classifier are described, which in some examples may be used with Figure 11 The area classifier 1107 corresponds to x. i (n) is the i-th microphone signal at discrete time n, i = {1...N} (i.e., microphone signal x i (n) is the output of N microphones 1105). In the echo management subsystem 1103, N signals x i (n) is processed to generate a 'clean' microphone signal e i(n), where i = {1...N}, each of the microphone signals at discrete time n. In this example, Figure 11 The clean signal called 1103A i (n) is fed to the wake-up word detector 1106. Here, each wake-up word detector 1106 generates Figure 11 The eigenvector w is called 1106A i (j), where j = {1...J} is the index corresponding to the jth wake-up word utterance. In this example, the classifier 1107 aggregates the feature set as input.

[0319] According to some embodiments, a set of region labels C for k = {1...K] k The user zones K may correspond to the number of different user zones in the environment. For example, user zones may include a sofa zone, a kitchen zone, a reading chair zone, and so on. Some examples may define more than one zone within a kitchen or other room. For example, a kitchen zone may include a sink zone, a food preparation zone, a refrigerator zone, and a dining zone. Similarly, a living room zone may include a sofa zone, a TV zone, a reading chair zone, one or more doorway zones, and so on. Zone labels for these zones may be user-selectable, for example, during a training phase.

[0320] In some embodiments, the classifier 1107 estimates the posterior probability p(C k |W(j)). Probability p(C k |W(j)) indicates the user in each area C k The probability in (for the jth utterance and the kth region, for each region C k , and each utterance), and is an example of the output 1108 of the classifier 1107.

[0321] According to some examples, training data can be collected (e.g., for each user zone) by prompting the user to select or define a zone (e.g., a sofa zone). The training process can involve prompting the user to utter a training utterance, such as a wake-up word, near the selected or defined zone. In the sofa zone example, the training process can involve prompting the user to utter a training utterance at the center and extreme edges of the sofa. The training process can involve prompting the user to repeat the training utterance several times at each location within the user zone. The user can then be prompted to move to another user zone and continue until all designated user zones are covered.

[0322] Figure 12 This is an overview of what can be done by Figure 2A1200 . As with other methods described herein, the blocks of method 1200 do not necessarily need to be executed in the order indicated. Moreover, this method may include more or fewer blocks than those shown and / or described. In this embodiment, method 1200 involves training a classifier for estimating a user's location in an environment.

[0323] In this example, block 1205 involves prompting the user to utter at least one training utterance at each of a plurality of locations within a first user zone of the environment. In some examples, the training utterance(s) may be one or more instances of a wake word utterance. According to some embodiments, the first user zone may be any user zone selected and / or defined by the user. In some examples, the control system may create a corresponding zone tag (e.g., the zone tag C described above). k ) and the region label can be associated with the training data obtained for the first user region.

[0324] An automated prompt system can be used to collect this training data. As described above, the interface system 205 of the device 200 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. For example, during the training process, the device 200 may provide the following prompts to the user on the screen of the display system or have the user hear the prompts announced via one or more speakers:

[0325] "Move to the couch."

[0326] "Say the wake word ten times while moving your head."

[0327] "Move to a spot between the sofa and the reading chair and say the wake-up word ten times."

[0328] "Stand in the kitchen as if you're cooking and say the wake word ten times."

[0329] In this example, block 1210 involves receiving a first output signal from each of a plurality of microphones in an environment. In some examples, block 1210 may involve receiving the first output signal from all active microphones in the environment, while in other examples, block 1210 may involve receiving the first output signal from a subset of all active microphones in the environment. In some examples, at least some of the microphones in the environment may provide output signals that are asynchronous with respect to output signals provided by one or more other microphones. For example, a first microphone in the plurality of microphones may sample audio data according to a first sampling clock, and a second microphone in the plurality of microphones may sample audio data according to a second sampling clock.

[0330] In this example, each of the plurality of microphones resides in a microphone location of the environment. In this example, the first output signal corresponds to an instance of a detected training utterance received from the first user zone. Because block 1205 involves prompting the user to utter at least one training utterance at each of a plurality of locations within the first user zone of the environment, in this example, the term "first output signal" refers to the set of all output signals corresponding to the training utterances for the first user zone. In other examples, the term "first output signal" may refer to a subset of all output signals corresponding to the training utterances for the first user zone.

[0331] According to this example, block 1215 involves determining one or more first acoustic features from each first output signal. In some examples, the first acoustic features may include a wake word confidence metric and / or a received level metric. For example, the first acoustic features may include a normalized wake word confidence metric, an indication of a normalized average received level, and / or an indication of a maximum received level.

[0332] As described above, because block 1205 involves prompting the user to utter at least one training utterance at each of a plurality of locations within a first user zone of the environment, the term "first output signal" in this example refers to the set of all output signals corresponding to the training utterances for the first user zone. Thus, in this example, the term "first acoustic features" refers to the set of acoustic features derived from the set of all output signals corresponding to the training utterances for the first user zone. Thus, in this example, the first set of acoustic features is at least as large as the first set of output signals. For example, if two acoustic features are determined from each output signal, the first set of acoustic features will be twice as large as the first set of output signals.

[0333] In this example, block 1220 involves training a classifier model to establish a correlation between the first user zone and the first acoustic feature. For example, the classifier model can be any of those disclosed herein. According to this embodiment, the classifier model is trained without reference to the geometric positions of the plurality of microphones. In other words, in this example, during the training process, the classifier model is not provided with data regarding the geometric positions of the plurality of microphones (e.g., microphone coordinate data).

[0334] Figure 13 This is an overview of what can be done by Figure 2A200 or the like. As with other methods described herein, the blocks of method 1300 do not have to be executed in the order indicated. For example, in some embodiments, at least a portion of the acoustic feature determination process of block 1325 may be performed before block 1315 or block 1320. Moreover, this method may include more or fewer blocks than those shown and / or described. In this embodiment, method 1300 involves training a classifier for estimating a user's location in an environment. Method 1300 provides an example of extending method 1200 to multiple user zones of an environment.

[0335] In this example, block 1305 involves prompting the user to utter at least one training utterance at a location within the user zone of the environment. In some instances, block 1305 may be referred to above, except that block 1305 is related to a single location within the user zone. Figure 12 In some examples, the training utterance(s) may be one or more instances of a wake word utterance. According to some embodiments, a user zone may be any user zone selected and / or defined by a user. In some examples, the control system may create a corresponding zone label (e.g., the zone label C described above). k ) and the region label can be associated with training data obtained for the user region.

[0336] According to this example, block 1310 is substantially as described above with reference to Figure 12 1210 . However, in this example, the process of block 1310 is generalized to any user zone, not necessarily the first user zone for which training data is acquired. Thus, the output signal received in block 1310 is “an output signal from each of a plurality of microphones in the environment, each of the plurality of microphones residing at a microphone location in the environment, the output signal corresponding to an instance of a detected training utterance received from the user zone.” In this example, the term “output signal” refers to the set of all output signals corresponding to one or more training utterances in the location of the user zone. In other examples, the term “output signal” may refer to a subset of all output signals corresponding to one or more training utterances in the location of the user zone.

[0337] According to this example, block 1315 involves determining whether sufficient training data has been acquired for the current user zone. In some such examples, block 1315 may involve determining whether output signals corresponding to a threshold number of training utterances have been acquired for the current user zone. Alternatively or additionally, block 1315 may involve determining whether output signals corresponding to training utterances in a threshold number of locations within the current user zone have been acquired. If not, in this example, method 1300 returns to block 1305 and prompts the user to utter at least one additional utterance at a location within the same user zone.

[0338] However, if it is determined in block 1315 that sufficient training data has been acquired for the current user zone, then in this example, the process continues to block 1320. According to this example, block 1320 involves determining whether to acquire training data for additional user zones. According to some examples, block 1320 may involve determining whether training data has been acquired for each user zone that the user has previously identified. In other examples, block 1320 may involve determining whether training data has been acquired for a minimum number of user zones. The user may have selected a minimum number. In other examples, the minimum number may be a recommended minimum number per environment, a recommended minimum number per room of an environment, etc.

[0339] If it is determined in block 1320 that training data should be obtained for an additional user zone, then in this example, the process continues to block 1322, which involves prompting the user to move to another user zone of the environment. In some examples, the next user zone may be selectable by the user. According to this example, after the prompt at block 1322, the process continues to block 1305. In some such examples, the prompt at block 1322 may be followed by a prompt for the user to confirm that the user has arrived at the new user zone. According to some such examples, the user may be asked to confirm that the user has arrived at the new user zone before providing the prompt at block 1305.

[0340] If it is determined in block 1320 that training data should not be obtained for additional user zones, then in this example, the process continues to block 1325. In this example, method 1300 involves obtaining training data for K user zones. In this embodiment, block 1325 involves determining first to G-th acoustic features from first to H-th output signals corresponding to each of the first to K-th user zones for which training data has been obtained. In this example, the term "first output signal" refers to the set of all output signals corresponding to the training utterances for the first user zone, and the term "H-th output signal" refers to the set of all output signals corresponding to the training utterances for the K-th user zone. Similarly, the term "first acoustic feature" refers to the set of acoustic features determined from the first output signal, and the term "G-th acoustic feature" refers to the set of acoustic features determined from the H-th output signal.

[0341] According to these examples, block 1330 involves training a classifier model to establish correlations between the first to Kth user regions and the first to Kth acoustic features, respectively. For example, the classifier model can be any of the classifier models disclosed herein.

[0342] In the foregoing example, the user zone is labeled (eg, according to the zone label C described above). k ). However, the model can be trained on labeled or unlabeled user regions, depending on the specific implementation. In the labeled case, each training utterance can be paired with a label corresponding to a user region, for example, as follows:

[0343]

[0344] Training a classifier model may involve determining the best fit for the labeled training data. Without loss of generality, appropriate classification methods for the classifier model may include:

[0345] Bayesian classifiers, e.g., with each class distribution described by a multivariate normal distribution, a full-covariance Gaussian mixture model, or a diagonal-covariance Gaussian mixture model;

[0346] Vector quantization;

[0347] Nearest neighbor (k-means);

[0348] A neural network with a SoftMax output layer, where one output corresponds to each class;

[0349] Support Vector Machine (SVM); and / or

[0350] Boosting techniques such as Gradient Boosting Machines (GBM)

[0351] In one example implementation of the unlabeled case, the data can be automatically split into K clusters, where K may also be unknown. For example, the unlabeled automatic splitting can be performed by using classical clustering techniques, such as the k-means algorithm or Gaussian mixture modeling.

[0352] To improve robustness, regularization can be applied to classifier model training, and model parameters can be updated over time as new utterances are uttered.

[0353] Further aspects of the embodiments are described next.

[0354] Example acoustic feature sets (e.g., Figure 11The acoustic features 1106A) may include the likelihood of the wake word confidence, the average reception level over the estimated duration of the most confident wake word, and the maximum reception level over the duration of the most confident wake word. For each wake word utterance, the features may be normalized relative to the maximum value of the features. The training data may be labeled and a full covariance Gaussian mixture model (GMM) may be trained to maximize the expected value of the training labels. The estimated region may be the class that maximizes the posterior probability.

[0355] The above description of some embodiments discusses learning acoustic zone models from a set of training data collected during a collection process of prompts. In the described model, training time (or configuration mode) and runtime (or regular mode) can be viewed as two different modes in which the microphone system can be placed. An extension of this approach is online learning, where some or all of the acoustic zone models are learned or adapted online (e.g., at runtime or in regular mode). In other words, even after a classifier is applied during "runtime" to estimate the user zone in which the user is currently located (e.g., based on Figure 10 1000), in some embodiments, the process of training the classifier can continue.

[0356] Figure 14 This is an overview of what can be done by Figure 2A 1400 is a flowchart of another example of a method performed by a device such as device 200. As with other methods described herein, the blocks of method 1400 do not have to be executed in the order indicated. Moreover, this method may include more or fewer blocks than those shown and / or described. In this embodiment, method 1400 involves continuous training of a classifier during the "runtime" process of estimating the user's position in the environment. Method 1400 is an example of what is referred to herein as an online learning model.

[0357] In this example, block 1405 of method 1400 corresponds to blocks 1005 through 1020 of method 1000. Here, block 1405 involves providing an estimate of the user zone in which the user is currently located based, at least in part, on the output from the classifier. Depending on the embodiment, block 1410 involves obtaining implicit or explicit feedback regarding the estimate of block 1405. In block 1415, the classifier is updated based on the feedback received in block 1405. For example, block 1415 may involve one or more reinforcement learning methods. As indicated by the dashed arrow from block 1415 to block 1405, in some embodiments, method 1400 may involve returning to block 1405. For example, method 1400 may involve providing a future estimate of the user zone in which the user will be located at that future time based on applying the updated model.

[0358] Explicit techniques for obtaining feedback can include:

[0359] Using a voice user interface (UI) to ask the user whether the prediction is correct. (For example, a voice may be provided to the user indicating: “I think you are on the couch, please say ‘true’ or ‘false’”).

[0360] Use voice UI to notify users that incorrect predictions can be corrected at any time. (For example, a voice could be provided to the user indicating something like, “I can now predict where you are when you speak to me. If I mispredict, just say something like, ‘Amanda, I’m not on the couch. I’m in the reading chair.’”)

[0361] Use voice UI to inform users that correct predictions can be rewarded at any time. (For example, a voice could be provided to the user indicating something like, “I’m now able to predict where you are when you speak to me. If I’m right, you can help improve my prediction further by saying something like, ‘Amanda, yes. I’m on the couch.’”)

[0362] • Include physical buttons or other UI elements that the user can operate to provide feedback (e.g., thumbs up and / or thumbs down buttons on a physical device or in a smartphone app).

[0363] The goal of predicting the user zone in which the user is located can be to inform microphone selection or adaptive beamforming schemes that attempt to more effectively pick up sound from the user's acoustic zone, for example, to better recognize commands following a wake word.

[0364] In this scenario, implicit techniques for obtaining feedback on the quality of the region prediction may include:

[0365] Penalize predictions that result in incorrect recognition of a command following a wake word. An agent that can indicate a misrecognition can include the user shortening the voice assistant's response to the command, for example, by issuing a counter-command such as "Amanda, stop!"

[0366] Penalize predictions that result in low confidence that the speech recognizer has successfully recognized the command. Many automatic speech recognition systems have the ability to return a confidence level along with the result that can be used for this purpose;

[0367] Penalizing the second-pass wake word detector from retrospectively detecting the wake word prediction with high confidence; and / or

[0368] Strengthen predictions to identify the wake word and / or correctly recognize the user’s command with high confidence.

[0369] The following is an example of a second-pass wake-word detector that fails to retrospectively detect the wake-word with high confidence. Assume that after obtaining an output signal corresponding to the current utterance from a microphone in the environment and determining an acoustic feature based on the output signal (e.g., via multiple first-pass wake-word detectors configured to communicate with the microphones), the acoustic feature is provided to a classifier. In other words, assume that the acoustic feature corresponds to the detected wake-word utterance. Assume further that the classifier determines that the person uttering the current utterance is most likely in zone 3, which corresponds to the reading chair in this example. For example, there may be a specific microphone or learned combination of microphones that is known to be best suited for listening to a person's speech when the person is in zone 3, for example, to send the person's speech to a cloud-based virtual assistant service for voice command recognition.

[0370] Further assume that after determining which microphone(s) will be used for speech recognition, but before the person's speech is actually sent to the virtual assistant service, a second-pass wake-word detector operates on the microphone signals corresponding to the speech detected by the selected microphone(s) for zone 3 to be submitted for command recognition. If the second-pass wake-word detector disagrees with your multiple first-pass wake-word detectors about which wake-word was actually uttered, it may be because the classifier incorrectly predicted zone 3. Therefore, the classifier should be penalized.

[0371] Techniques for a posteriori updating of the region mapping model after one or more wake words are spoken may include:

[0372] Maximum a posteriori (MAP) adaptation of Gaussian mixture models (GMM) or nearest neighbor models; and / or

[0373] Reinforcement learning of neural networks, e.g., by associating appropriate “one-hot” (in case of correct prediction) or “one-cold” (in case of incorrect prediction) ground truth labels with the SoftMax output and applying online backpropagation to determine new network weights.

[0374] In this context, some examples of MAP adaptation can involve adjusting the mean in the GMM each time the wake word is spoken. In this way, the mean can become more similar to the acoustic features observed when subsequent wake words are spoken. Alternatively or additionally, such examples can involve adjusting the variance / covariance or mixing weight information in the GMM each time the wake word is spoken.

[0375] For example, the MAP adaptation scheme can be as follows:

[0376] μ i,new =μ i,old *α+x*(1-α)

[0377] In the above equation, μ i,old Denotes the mean of the i-th Gaussian in the mixture, α denotes a parameter that controls how aggressively MAP adaptation should occur (α may be in the range [0.9, 0.999]), and x denotes the feature vector of the new wake word utterance. Index "i" will correspond to the element of the mixture that returns the highest prior probability of containing the speaker location at the wake word time.

[0378] Alternatively, each mixture element can be adjusted based on its prior probability of containing the wake word, for example, as follows:

[0379] M i,new =μ i,old *β i *x(l-β i )

[0380] In the above equation, β i =α*(1-P(i)), where P(i) represents the prior probability that observation x is due to mixture element i.

[0381] In a reinforcement learning example, there may be three user zones. Suppose that for a particular wake-up word, the model predicts three user zones with probabilities of [0.2, 0.1, 0.7]. If a second source of information (e.g., a second-pass wake-up word detector) confirms that the third zone is correct, the ground truth label might be [0, 0, 1] (a "one-hot"). The posterior update of the zone mapping model can involve backpropagating the error through the neural network, which effectively means that if the same input is shown again, the neural network will predict zone 3 more strongly. Conversely, if the second source of information shows that zone 3 is an incorrect prediction, then in one example, the ground truth label might be [0.5, 0.5, 0.0]. If the same input is shown in the future, backpropagating the error through the neural network will make the model less likely to predict zone 3.

[0382] Flexible rendering allows spatial audio to be rendered on any number of arbitrarily placed speakers. Given the widespread deployment of audio devices, including but not limited to smart audio devices (e.g., smart speakers) in the home, there is a need for flexible rendering techniques that allow consumer products to perform flexible rendering of audio and playback of such rendered audio.

[0383] Several techniques have been developed to implement flexible rendering. They frame the rendering problem as one of minimizing a cost function consisting of two terms: the first modeling the desired spatial impression the renderer is trying to achieve, and the second assigning the cost of activating a speaker. To date, this second term has focused on creating sparse solutions, where only speakers very close to the desired spatial location of the audio being rendered are activated.

[0384] Playing back spatial audio in a consumer environment is typically associated with a specified number of loudspeakers placed in specified locations: for example, 5.1 and 7.1 surround sound. In these cases, the content is written specifically for the associated loudspeakers and encoded as discrete channels, one for each loudspeaker (e.g., Dolby Digital or Dolby Digital Plus, etc.). More recently, immersive, object-based spatial audio formats (Dolby Atmos) have been introduced that break this association between content and specific loudspeaker locations. Instead, content can be described as a collection of individual audio objects, each with metadata that may change over time, describing the desired perceived position of the audio object in three-dimensional space. At playback time, the content is converted into loudspeaker feeds by a renderer that adapts to the number and position of the loudspeakers in the playback system. However, many such renderers still restrict the position of a set of loudspeakers to one of a set of prescribed layouts (e.g., 3.1.2, 5.1.2, 7.1.4, 9.1.6 with Dolby Atmos, etc.).

[0385] Beyond this limited rendering, methods have been developed that allow object-based audio to be flexibly rendered on a virtually arbitrary number of loudspeakers placed in arbitrary locations. These methods require the renderer to know the number and physical location of the loudspeakers in the listening space. To make such systems practical for the average consumer, automated methods for locating the loudspeakers are desired. One such method relies on the use of multiple microphones, which may be co-located with the loudspeakers. By playing the audio signal through the loudspeakers and recording it with the microphones, the distance between each loudspeaker and the microphone is estimated. From these distances, the positions of both the loudspeaker and the microphone are then derived.

[0386] Coinciding with the introduction of object-based spatial audio in the consumer space has been the rapid adoption of so-called “smart speakers” such as the Amazon Echo line of products. The immense popularity of these devices can be attributed to the simplicity and convenience they offer through wireless connectivity and an integrated voice interface (e.g., Amazon’s Alexa), but the sonic capabilities of these devices are generally limited, particularly when it comes to spatial audio. In most cases, these devices are restricted to mono or stereo playback. However, combining the flexible rendering and automatic position technologies described above with multiple orchestrated smart speakers can produce systems with very sophisticated spatial playback capabilities that remain very simple for consumers to set up. Consumers can place as many or as few speakers as needed, wherever convenient, without having to run speaker wires thanks to the wireless connectivity, and built-in microphones can be used to automatically position the speakers for the associated flexible renderer.

[0387] Conventional flexible rendering algorithms aim to achieve as closely as possible a specific desired perceptual spatial impression. In an orchestrated smart speaker system, sometimes, maintaining that spatial impression may not be the most important or desired goal. For example, if a person is simultaneously trying to speak to an integrated voice assistant, it may be desirable to temporarily change the spatial rendering in a way that reduces the relative playback levels of speakers near certain microphones in order to increase the signal-to-noise ratio and / or signal-to-echo ratio (SER) of the microphone signals including the detected speech. Some embodiments described herein may be implemented as modifications to existing flexible rendering methods to allow such dynamic modification of the spatial rendering, for example, for the purpose of achieving one or more additional goals.

[0388] Existing flexible rendering techniques include Centroid Amplitude Panning (CMAP) and Flexible Virtualization (FV). At a high level, these two techniques render a set of one or more audio signals, each audio signal having an associated desired perceptual spatial position, for playback on a set of two or more loudspeakers, where the relative activation of the set of loudspeakers is a function of a model of the perceptual spatial position of the audio signal played back by the loudspeakers and the proximity of the desired perceptual spatial position of the audio signal to the loudspeaker positions. The model ensures that the listener hears the audio signal near its expected spatial position, and the proximity term controls which loudspeakers are used to achieve this spatial impression. Specifically, the proximity term favors activating loudspeakers that are close to the desired perceptual spatial position of the audio signal. For both CMAP and FV, this functional relationship can be conveniently derived from a cost function written as the sum of two terms, one for the spatial aspect and one for proximity:

[0389]

[0390] Here, gather represents the positions of a set of M loudspeakers, denotes the desired perceptual spatial location of the audio signal, and g denotes the M-dimensional vector of loudspeaker activations. For CMAP, each activation in the vector represents a gain for each loudspeaker, while for FV, each activation represents a filter (in the second case, g can equivalently be viewed as a vector of complex values at a particular frequency, and different g are computed across multiple frequencies to form the filter). The optimal vector of activations is found by minimizing the cost function across the activations:

[0391]

[0392] Under certain definitions of the cost function, it is difficult to control the absolute level of the optimal activation resulting from the above minimization, although g opt The relative levels between the components are appropriate. To solve this problem, you can perform g optSubsequent normalization of in order to control the absolute level of activation. For example, one may wish to normalize the vector to have unit length, which conforms to the commonly used constant power shift rule:

[0393]

[0394] The exact behavior of the flexible rendering algorithm depends on the C of the cost function. spatial and C proximity For CMAP, C spatial is obtained from a model that places the perceptual spatial location of an audio signal played from a set of loudspeakers at the position indicated by its associated activation gain g i The centroids of the positions of these loudspeakers weighted (elements of vector g):

[0395]

[0396] Equation 3 is then manipulated into a spatial cost representing the squared error between the desired audio position and the desired audio position produced by the activated loudspeakers:

[0397]

[0398] For FV, the spatial term of the cost function is defined differently. The goal is to produce the same spatial coordinates at the left and right ears as the audio object positions. The corresponding binaural response b. Conceptually, b is a 2×1 vector of filters (one filter for each ear), but it is more convenient to think of it as a 2×1 vector of complex values at a specific frequency. Continuing this representation at a specific frequency, the desired binaural response can be obtained from a set of HRTFs indexed by the object position:

[0399]

[0400] Meanwhile, the 2×1 binaural response e produced by the loudspeaker at the listener’s ears is modeled as a 2×M acoustic transfer matrix H multiplied by an M×1 vector g of complex loudspeaker activation values:

[0401] e=Hg(6)

[0402] The acoustic transmission matrix H is based on the set of loudspeaker positions Finally, the spatial component of the cost function is defined as the desired binaural response (Equation 5) versus the desired binaural response produced by the loudspeakers (Equation 5).

[0403] The square error between

[0404]

[0405] Conveniently, the spatial terms of the cost functions for CMAP and FV defined in both Equations 4 and 7 can be rearranged as matrix quadratic functions as a function of the speaker activation g:

[0406]

[0407] Where A is an M×M square matrix, B is a 1×M vector, and C is a scalar. The rank of the matrix A is 2, and therefore when M>2, there are infinitely many speaker activations g for which the spatial error term is zero. The second term C of the cost function is introduced pr o ximity This uncertainty is removed and a specific solution is produced that has perceptually beneficial properties compared to other possible solutions. For both CMAP and FV, C pr o ximity is constructed so that the position Far from the desired audio signal location The activation of speakers located near the desired position is penalized more than the activation of speakers located closer to the desired position. This construction produces an optimal set of sparse speaker activations, where only speakers close to the position of the desired audio signal are significantly activated, and in fact leads to a spatial reproduction of the audio signal that is perceptually more robust to listener movement around the speaker set.

[0408] To this end, the second term C of the cost function pr o ximity can be defined as the distance-weighted sum of the squared absolute values of the speaker activations. This is concisely expressed in matrix form as:

[0409]

[0410] where D is the diagonal matrix of distance penalties between the desired audio position and each speaker:

[0411]

[0412] The distance penalty function can take many forms, but the following is a useful parameterization:

[0413]

[0414] in, is the Euclidean distance between the desired audio position and the loudspeaker position, and α and β are adjustable parameters. The parameter α indicates the global strength of the penalty; d0 corresponds to the spatial extent of the distance penalty (loudspeakers at approximately d0 distance or further will be penalized), and β accounts for the abruptness of the onset of the penalty at distance d0.

[0415] Combining the two terms of the cost functions defined in Equations 8 and 9a yields the overall cost function:

[0416] C(g)=g * Ag+Bg+C+g * Dg=g * (A+D)g+Bg+C (10)

[0417] Setting the derivative of this cost function with respect to g to zero and solving for g yields the optimal speaker activation solution:

[0418]

[0419] Typically, the optimal solution in Equation 11 may produce speaker activations with negative values. For CMAP construction of flexible renderers, such negative activations may be undesirable, and therefore Equation (11) can be minimized while keeping all activations positive.

[0420] Figure 15 and Figure 16 is a diagram illustrating a set of example speaker activation and object rendering positions. In these examples, the speaker activation and object rendering positions correspond to speaker positions of 4, 64, 165, -87, and -4 degrees. Figure 15 Speaker activations 1505a, 1510a, 1515a, 1520a, and 1525a are shown, comprising the optimal solution to Equation 11 for these particular speaker locations. Figure 16 Individual speaker locations are plotted as points 1605, 1610, 1615, 1620, and 1625, corresponding to speaker activations 1505a, 1510a, 1515a, 1520a, and 1525a, respectively. Figure 16 The ideal object positions (in other words, the positions where audio objects are to be rendered) for a number of possible object angles are also shown as points 1630a, and the corresponding actual rendering positions for these objects are shown as points 1635a, connected to the ideal object positions by dashed lines 1640a.

[0421] One class of embodiments relates to methods for rendering audio for playback by at least one (e.g., all or some) of a plurality of coordinated (orchestrated) smart audio devices. For example, a group of smart audio devices present in a user's home (in a system) can be orchestrated to handle a variety of simultaneous use cases, including flexible rendering (depending on the embodiment) of audio for playback by all or some of the smart audio devices (i.e., by all or some of the smart audio device's (multiple) speakers). Many interactions with the system are contemplated that require dynamic corrections to the rendering. Such corrections can, but are not necessarily focused on spatial fidelity.

[0422] Some embodiments are methods for rendering audio for playback by at least one (e.g., all or some) of a group of smart audio devices (or for playback by at least one (e.g., all or some) of another group of speakers). Rendering may include minimization of a cost function, where the cost function includes at least one dynamic speaker activation term. Examples of such dynamic speaker activation terms include (but are not limited to):

[0423] The proximity of the loudspeaker to one or more listeners;

[0424] The proximity of the loudspeaker to attractive or repulsive forces;

[0425] The audibility of the loudspeaker with respect to some location (e.g. the listener's position or a baby's room);

[0426] The speaker's capabilities (e.g., frequency response and distortion);

[0427] Synchronization of speakers with respect to other speakers;

[0428] Wake-up word performance; and

[0429] Echo canceller performance.

[0430] The dynamic speaker activation item(s) may enable at least one of a variety of behaviors, including distorting the spatial rendering of the audio away from a particular smart audio device so that the microphone of the particular smart audio device can better hear the speaker or so that a secondary audio stream can be better heard from the speaker(s) of the smart audio device.

[0431] Some embodiments implement rendering for playback through the speaker(s) of a coordinated (orchestrated) plurality of smart audio devices. Other embodiments implement rendering for playback through the speaker(s) of another set of speakers.

[0432] Pairing a flexible rendering approach (implemented in accordance with some embodiments) with a group of wireless smart speakers (or other smart audio devices) can result in a very capable and easy-to-use spatial audio rendering system. When considering interactions with such a system, it is clearly desirable to dynamically modify the spatial rendering so that it is optimized for other goals that may arise during use of the system. To achieve this goal, one class of embodiments enhances an existing flexible rendering algorithm (in which speaker activation is a function of the previously disclosed spatial terms and proximity terms) with one or more additional dynamically configurable functions that depend on one or more properties of the audio signal being rendered, the speaker group, and / or other external inputs. According to some embodiments, the cost function for the existing flexible rendering given in Equation 1 is augmented with these one or more additional dependencies according to the following equation

[0433]

[0434] In Equation 12, the term represents an additional cost item, and represents a set of one or more properties of an audio signal being rendered (e.g., an object-based audio program), One or more properties representing a set of speakers to which the audio is being rendered, and Represents one or more additional external inputs. Each item Returns the cost as a function of the activation g associated with a combination of one or more properties of the audio signal, the loudspeaker, and / or the external input, typically represented by the set It should be understood that the set At least from or An element of any one of .

[0435] Examples include, but are not limited to:

[0436] The desired perceived spatial location of the audio signal;

[0437] the level of the audio signal (which may vary over time); and / or

[0438] The frequency spectrum of the audio signal (possibly varying over time).

[0439] Examples include, but are not limited to:

[0440] The location of the loudspeakers in the listening space;

[0441] The frequency response of the loudspeaker;

[0442] ·Playback level limit for loudspeakers;

[0443] Parameters of the speaker's internal dynamics processing algorithms, such as limiter gain;

[0444] A measurement or estimate of the acoustic transmission from each loudspeaker to the other loudspeakers;

[0445] Measurement of the performance of echo cancellers on loudspeakers; and / or

[0446] • The relative synchronization of the loudspeakers with respect to each other.

[0447] Examples include, but are not limited to:

[0448] The location of one or more listeners or speakers in the playback space;

[0449] A measurement or estimate of the acoustic transmission from each loudspeaker to the listening position;

[0450] The measurement or estimation of the acoustic transmission from a speaker to a set of loudspeakers;

[0451] The location of some other landmarks in the replay space; and / or

[0452] A measurement or estimate of the acoustic transmission from each loudspeaker to some other landmark in the playback space;

[0453] Using the new cost function defined in Equation 12, the optimal activation group can be found by minimization with respect to g as previously specified in Equations 2a and 2b and possible post-normalization.

[0454] Figure 17 This is an overview of what can be done by Figure 2A 1700. As with other methods described herein, the blocks of method 1700 do not have to be performed in the order shown. Moreover, this method may include more or fewer blocks than shown and / or described. The blocks of method 1700 may be performed by one or more devices, which may be (or may include) a control system, such as Figure 2A The control system 210 is shown in FIG.

[0455] In this embodiment, block 1705 involves receiving audio data by the control system and via the interface system. In this example, the audio data includes one or more audio signals and associated spatial data. According to this embodiment, the spatial data indicates an expected perceptual spatial position corresponding to the audio signal. In some instances, the expected perceptual spatial position can be explicit, for example, as indicated by positional metadata such as Dolby Atmos positional metadata. In other instances, the expected perceptual spatial position can be implicit, for example, the expected perceptual spatial position can be an assumed position associated with channels according to Dolby 5.1, Dolby 7.1, or other channel-based audio format. In some examples, block 1705 involves a rendering module of the control system receiving the audio data via the interface system.

[0456] According to this example, block 1710 involves rendering, by a control system, audio data for reproduction via a set of loudspeakers in an environment to produce a rendered audio signal. In this example, rendering each of the one or more audio signals included in the audio data involves determining the relative activation of the set of loudspeakers in the environment by optimizing a cost function. According to this example, the cost is a function of a model of the perceived spatial location of the audio signal when played back through the set of loudspeakers in the environment. In this example, the cost is also a function of a measure of the proximity of the expected perceived spatial location of the audio signal to the location of each loudspeaker in the set of loudspeakers. In this embodiment, the cost is also a function of one or more additional dynamically configurable functions. In this example, the dynamically configurable functionality is based on one or more of: proximity of the loudspeaker to one or more listeners; proximity of the loudspeaker to a location of attraction, where attraction is a factor that favors activation of relatively taller loudspeakers that are closer to the location of attraction; proximity of the loudspeaker to a location of repulsion, where repulsion is a factor that favors activation of relatively taller loudspeakers that are closer to the location of repulsion; the capabilities of each loudspeaker relative to other loudspeakers in the environment; synchronization of the loudspeakers with respect to other loudspeakers; wake word performance; or echo canceller performance.

[0457] In this example, block 1715 involves providing the rendered audio signal to at least some of the set of loudspeakers of the environment via the interface system.

[0458] According to some examples, the model of perceptual spatial position can generate binaural responses corresponding to the positions of the audio objects at the left and right ears of the listener. Alternatively or additionally, the model of perceptual spatial position can place the perceptual spatial position of an audio signal played from a set of loudspeakers at the centroid of the positions of the set of loudspeakers weighted by the loudspeakers' associated activation gains.

[0459] In some examples, the one or more additional dynamically configurable functions may be based at least in part on a level of the one or more audio signals. In some examples, the one or more additional dynamically configurable functions may be based at least in part on a frequency spectrum of the one or more audio signals.

[0460] Some examples of method 1700 involve receiving loudspeaker layout information.In some examples, one or more additional dynamically configurable functions can be based at least in part on the location of each loudspeaker in the environment.

[0461] Some examples of method 1700 involve receiving loudspeaker specification information. In some examples, one or more additional dynamically configurable functions can be based at least in part on the capabilities of each loudspeaker, which can include one or more of: frequency response, playback level limits, or parameters of one or more loudspeaker dynamics processing algorithms.

[0462] According to some examples, one or more additional dynamically configurable functions may be based at least in part on a measurement or estimate of acoustic transmission from each loudspeaker to the other loudspeakers. Alternatively or additionally, one or more additional dynamically configurable functions may be based at least in part on a listener or speaker position of one or more people in the environment. Alternatively or additionally, one or more additional dynamically configurable functions may be based at least in part on a measurement or estimate of acoustic transmission from each loudspeaker to the listener or speaker position. The estimate of acoustic transmission may, for example, be based at least in part on walls, furniture, or other objects that may be located between each loudspeaker and the listener or speaker position.

[0463] Alternatively or additionally, one or more additional dynamically configurable functionalities may be based at least in part on the object locations of one or more non-microphone objects or landmarks in the environment. In some such embodiments, the one or more additional dynamically configurable functionalities may be based at least in part on measurements or estimates of acoustic transmission from each microphone to the object locations or landmark locations.

[0464] Flexible rendering can be implemented to achieve many new and useful behaviors by employing one or more appropriately defined additional cost terms. All of the example behaviors listed below are designed to penalize certain loudspeakers under certain conditions that are considered undesirable. The end result is that these loudspeakers are less activated in the spatial rendering of a set of audio signals. In many of these cases, one might consider simply turning down the undesirable loudspeakers, without resorting to any corrections to the spatial rendering, but such a strategy could significantly degrade the overall balance of the audio content. For example, certain components of the mix could become completely inaudible. On the other hand, for the disclosed embodiments, integrating these penalties into the core optimization of the rendering allows the rendering to adapt and use the remaining less penalized loudspeakers to perform the best possible spatial rendering. This is a more elegant, more adaptable, and more efficient solution.

[0465] Example use cases include, but are not limited to:

[0466] Provides a more balanced spatial presentation around the listening area

[0467] o It has been found that spatial audio is best presented across speakers that are approximately the same distance from the intended listening area. Costs can be structured so that loudspeakers that are significantly closer or further than the average distance from the listening area are penalized, thereby reducing the activation of those loudspeakers;

[0468] Move audio away from or towards the listener or speaker

[0469] o If a user of the system is attempting to speak to the system or an intelligent voice assistant associated with the system, it may be beneficial to create a cost that penalizes microphones that are closer to the speaker. In this way, these microphones are activated less frequently, allowing their associated microphones to better hear the speaker;

[0470] o To provide a more intimate experience for a single listener, i.e., to minimize the playback level for others in the listening space, speakers located far from the listener's position may be heavily penalized so that only the speakers closest to the listener are most prominently activated;

[0471] Move audio away from or towards a landmark, area, or region

[0472] o Certain locations near the listening space may be considered sensitive, such as a baby's room, crib, office, reading area, study area, etc. In this case, a cost may be constructed that penalizes the use of speakers close to that location, zone, or area;

[0473] o Alternatively, for the same situation above (or a similar situation), the speaker system may already generate a measure of the acoustic transmission from each speaker into the baby's room, particularly when one of the speakers (with an attached or associated microphone) resides in the baby's room. In this case, rather than using the physical proximity of the speakers to the baby's room, a cost may be constructed that penalizes the use of speakers with high measured acoustic transmission into the room; and / or

[0474] Optimal use of the speaker's capabilities

[0475] o The capabilities of different loudspeakers can vary significantly. For example, one popular smart speaker may contain only a single 1.6" full-range driver with limited low-frequency capabilities. On the other hand, another smart speaker may contain a more capable 3" woofer. These capabilities are typically reflected in the frequency response of the speaker, and as such, a set of responses associated with a speaker may be utilized in the cost term. At certain frequencies, a speaker that is less powerful relative to other speakers (as measured by its frequency response) may be penalized and therefore activated to a lesser extent. In some embodiments, such frequency response values may be stored with the smart speaker and then reported to the computational unit responsible for optimizing flexible rendering;

[0476] o Many speakers contain more than one driver, each responsible for playing a different frequency range. For example, a popular smart speaker is a two-way design, containing a woofer for lower frequencies and a tweeter for higher frequencies. Typically, such speakers contain a crossover circuit for dividing the full-range playback audio signal into the appropriate frequency ranges and sending it to the corresponding drivers. Alternatively, such speakers can provide flexible renderer playback access to each individual driver, as well as information about the capabilities (such as frequency response) of each individual driver. By applying the cost terms as described above, in some examples, the flexible renderer can automatically establish a crossover between two drivers based on their relative capabilities at different frequencies;

[0477] o The example use of frequency response described above focuses on the inherent capabilities of the loudspeaker, but may not accurately reflect the capabilities of the loudspeaker when placed in the listening environment. In some cases, the frequency response of a loudspeaker as measured at the intended listening position can be obtained through some calibration procedure. Such measurements can be used instead of pre-calculated responses to better optimize the use of the loudspeaker. For example, a loudspeaker may be inherently very capable at a particular frequency, but due to its placement (e.g., behind a wall or a piece of furniture) may produce a very limited response at the intended listening position. A measurement that captures this response and feeds into an appropriate cost term can prevent significant activation of such a loudspeaker;

[0478] o Frequency response is only one aspect of a loudspeaker's playback capabilities. Many smaller loudspeakers begin to distort and then reach their excursion limits as the playback level increases, especially for lower frequencies. To reduce this distortion, many loudspeakers implement dynamic processing that limits the playback level to below certain limiting thresholds that can vary with frequency. In the event that a loudspeaker approaches or is at these thresholds while other loudspeakers participating in the flexible rendering are not, it makes sense to reduce the signal level in the limiting loudspeaker and transfer that energy to the other, less burdened loudspeakers. According to some embodiments, this behavior can be automatically implemented by appropriately configuring the associated cost terms. Such cost terms may involve one or more of the following:

[0479] Monitor the global playback volume relative to the limit thresholds of the loudspeakers. For example, loudspeakers with volume levels close to their limit thresholds may be penalized more.

[0480] Monitoring dynamic signal levels (which may also vary with frequency) relative to a loudspeaker's limiting threshold (which may also vary with frequency). For example, a loudspeaker with a monitored signal level close to its limiting threshold may be penalized more;

[0481] Directly monitoring parameters of the loudspeaker's dynamics processing, such as limiting gain. In some such examples, parameters indicating more limiting loudspeakers may be penalized more; and / or

[0482] Monitor the actual instantaneous voltage, current, and power delivered by the amplifier to the loudspeaker to determine if the loudspeaker is operating within its linear range. For example, a loudspeaker operating less linearly may be penalized more.

[0483] o Smart speakers with integrated microphones and interactive voice assistants typically employ some type of echo cancellation to reduce the level of the audio signal played back by the speaker as picked up by the recording microphone. The greater the reduction, the greater the chance that the speaker can hear and understand the speaker in the space. If the echo canceller residual is consistently high, this can indicate that the speaker is being driven into a non-linear region where prediction of the echo path becomes challenging. In this case, it can make sense to divert signal energy away from the speaker, and as such, a cost term that accounts for echo canceller performance can be beneficial. Such a cost term can assign a high cost to speakers whose associated echo cancellers perform poorly;

[0484] o In order to achieve predictable imaging when rendering spatial audio on multiple speakers, it is generally necessary to reasonably synchronize the playback on a set of loudspeakers across time. For wired loudspeakers, this is a given, but for a large number of wireless loudspeakers, synchronization can be challenging and the end result variable. In this case, it may be possible for each loudspeaker to report its relative degree of synchronization with the target, and this degree can then be fed into a synchronization cost term. In some such examples, loudspeakers with a lower degree of synchronization may be penalized more and therefore excluded from rendering. Additionally, certain types of audio signals may not require tight synchronization, for example, components of an audio mix that is intended to be diffuse or non-directional. In some embodiments, the components can be marked as such with metadata, and the synchronization cost term can be modified so that the penalty is reduced.

[0485] Additional examples of embodiments are described below. Similar to the proximity cost defined in Equations 9a and 9b, each new cost function term It may also be convenient to express it as a weighted sum of the squared absolute values of the loudspeaker activations, for example as follows:

[0486]

[0487] Among them, W j is the weight A diagonal matrix describing the cost associated with activating speaker i for item j:

[0488]

[0489] Combining Equations 13a and b with the matrix quadratic versions of the CMAP and FV cost functions given in Equation 10 yields a potentially beneficial implementation of the generalized extended cost function (of some embodiments) given in Equation 12:

[0490] C(g)=g * Ag+Bg+C+g * Dg+∑ j g * W j g=g * (A+D+∑ j W j )g+Bg+C (14)

[0491] With this definition of the new cost function term, the overall cost function is still matrix quadratic, and the optimal set of activations gopt can be found by differentiating Equation 14 to produce

[0492]

[0493] Consider each of the weight terms wij as a given continuous penalty value for each of the loudspeakers A function of is useful. In one example embodiment, the penalty value is the distance from the object (to be rendered) to the loudspeaker under consideration. In another example embodiment, the penalty value indicates that a given loudspeaker cannot reproduce some frequencies. Based on the penalty value, the weight term wij can be parameterized as:

[0494]

[0495] Among them, α j represents the pre-factor (which takes into account the global strength of the weight term), where τ j represents a penalty threshold (around or above which the weight term becomes significant), and where fj(x) represents a monotonically increasing function. For example, with The weight term has the following form:

[0496]

[0497] Among them, α j , β j , τ j are adjustable parameters that indicate the global strength of the penalty, the abruptness of the onset of the penalty, and the degree of the penalty, respectively. Care should be taken in setting these adjustable values so that the cost term C j Relative to any other additional cost items and C spatial and C proximityThe relative influence of α is appropriate to achieve the desired outcome. For example, as a rule of thumb, if one wants a particular penalty to clearly dominate over other penalties, then one should set its intensity α to j A setting of approximately ten times the next largest penalty strength might be appropriate.

[0498] If all loudspeakers are penalized, it is often convenient to subtract the minimum penalty from all weight terms in post-processing so that at least one of the loudspeakers is not penalized:

[0499] w ij →w′ ij =w ij -min i (w ij ) (18)

[0500] As described above, many possible use cases can be implemented using the new cost function terms described herein (and similar new cost function terms employed according to other embodiments). Next, more specific details are described using the following three examples: moving audio toward a listener or speaker, moving audio away from a listener or speaker, and moving audio away from a landmark.

[0501] In a first example, what will be referred to herein as "attraction" is used to pull the audio toward a location, which in some examples may be the location of the listener or speaker, a landmark, furniture, etc. The location may be referred to herein as an "attraction location" or "attractor location." As used herein, "attraction" is a factor that favors relatively higher loudspeaker activation closer to the attraction location. According to this example, the weight w ij Using the form of Equation 17, the continuous penalty value p ij The distance between the i-th speaker and the fixed attractor The distance is given, and the threshold τ j is given by the maximum of these distances for all loudspeakers:

[0502]

[0503]

[0504] To illustrate the use case of "pulling" the audio towards the listener or speaker, α is specifically set to j =20,β j =3, and Set to the vector corresponding to the listener / speaker position of 180 degrees (bottom center of the plot). j , β j and These values of are examples only. In some embodiments, αj can be in the range of 1 to 100 and β j Can be in the range of 1 to 25. Figure 18 is a diagram of speaker activation in an example embodiment. In this example, Figure 18 Speaker activations 1505b, 1510b, 1515b, 1520b, and 1525b are shown, including Figure 15 and Figure 16 The optimal solution to the cost function for the same loudspeaker positions in , plus the solution given by w ij Indicates attraction. Figure 19 is a diagram of object rendering locations in an example embodiment. In this example, Figure 19 The corresponding ideal object positions 1630b for a number of possible object angles and the corresponding actual rendered positions 1635b for those objects are shown, connected to the ideal object positions 1630b by dashed lines 1640b. The actual rendered positions 1635b are oriented toward a fixed position. The tilted orientation of illustrates the influence of the attractor weights on the optimal solution of the cost function.

[0505] In the second and third examples, a "repulsion force" is used to "push" the audio away from a certain location, which may be a person's location (e.g., a listener's location, a speaker's location, etc.) or other location, such as a landmark location, a furniture location, etc. In some examples, a repulsion force may be used to push the audio away from an area or zone of the listening environment, such as an office area, a reading area, a bed or bedroom area (e.g., a crib or bedroom), etc. According to some such examples, a particular location may be used as a representative of a zone or area. For example, the location representing the crib may be the estimated location of the baby's head, the estimated location of the sound source corresponding to the baby, etc. The location may be referred to herein as a "repulsion force location" or a "repulsion location." As used herein, a "repulsion force" is a factor that favors the activation of relatively lower loudspeakers that are closer to the repulsion force location. According to this example, relative to a fixed repulsion location Define p ij and τ j , similar to the attraction in Equation 19:

[0506]

[0507]

[0508] To illustrate the use case of pushing audio away from the listener or speaker, in one example, α can be specifically set j =5,β j =2, and Set to the vector corresponding to the listener / speaker position of 180 degrees (at the bottom center of the plot). j, β j and These values of are examples only. As mentioned above, in some examples, α j can be in the range of 1 to 100 and β j Can be in the range of 1 to 25. Figure 20 is a diagram of speaker activation in an example embodiment. According to this example, Figure 20 Speaker activations 1505c, 1510c, 1515c, 1520c, and 1525c are shown, which include the optimal solution to the cost function for the same speaker positions as the previous figure, plus the value given by w ij Indicates the repulsive force. Figure 21 is a diagram of object rendering locations in an example embodiment. In this example, Figure 21 The ideal object positions 1630c and the corresponding actual rendered positions 1635c for those objects are shown, connected to the ideal object position 1630c by dashed lines 1640c. The actual rendered positions 1635c are off the fixed position The tilted orientation of illustrates the influence of the repeller weights on the optimal solution of the cost function.

[0509] A third example use case is to "push" the audio away from an acoustically sensitive landmark, such as a door to a sleeping baby's room. Similar to the last example, Set to the vector corresponding to a door position of 180 degrees (bottom center of the plot). To achieve a stronger repulsion and tilt the soundstage fully to the front of the main listening space, set α j =20,β j =5. Figure 22 is a diagram of speaker activation in an example embodiment. Again, in this example, Figure 22 Speaker activations 1505d, 1510d, 1515d, 1520d, and 1525d are shown and comprise the optimal solution for the same set of speaker locations, plus a stronger repulsive force. Figure 23 is a diagram of object rendering locations in an example embodiment. And again, in this example, Figure 23 Shown are ideal object positions 1630d for a number of possible object angles and the corresponding actual rendered positions 1635d for those objects, connected to the ideal object position 1630d by dashed lines 1640d. The tilted orientation of the actual rendered position 1635d illustrates the effect of stronger repeller weights on the optimal solution to the cost function.

[0510] exist Figure 2BIn a further example of method 250, the use case responds to the selection of two or more audio devices in the audio environment (block 265) and applies a "repulsive" force to the audio (block 275). Following the previous example, in some embodiments, the selection of two or more audio devices can take the form of a unitless parameter value f_n that controls the degree to which the audio processing change occurs. Many combinations are possible. In a simple example, the weight corresponding to the repulsive force can be directly selected as Penalizes the device as selected by the "Decision" aspect.

[0511] In addition to the previous examples of determining weights, in some implementations, weights may be determined as follows:

[0512]

[0513] In the above equation, α j , β j , τ j denote adjustable parameters that indicate the global strength of the penalty, the abruptness of the onset of the penalty, and the degree of the penalty, respectively, as described above with reference to Equation 17. Thus, the above equation can be understood as a combination of multiple penalty terms generated from multiple simultaneous use cases. For example, the term p can be used as described in the previous example ij and τ j Push the audio "away" from the sensitive landmark while still being "pushed away" from the microphone location where it is desired to use the term f as determined by the decision aspect i To improve SER.

[0514] The previous example also introduced s_n, directly expressed as speech-echo ratio improvement in decibels. Some embodiments may involve selecting values of α and β (the strength of the penalty and the suddenness of the onset of the penalty, respectively) in some parts based on the s_n value in dB, and the previously indicated ij The formula can be used to ij and β ij Alternative α j and β j For example, a value of s_i = -20 dB would correspond to a high cost of activating the i-th speaker. In some such embodiments, α ij Can be set as the ratio cost function C spatial and C proximity Typical values of other terms in are many times higher. For example, Determining a new value for α, the equation for a value of s_i = -20 dB will result in α ij The value of is 10 times the value usually used in the cost function. In some examples, β ij Correction will be set to 0.5 < βij The range < 1.0 may be a suitable correction based on large negative values of s_i to "push" the audio out of a larger area around the ith speaker. For example, the value of s_i can be mapped to β according to ij :

[0515]

[0516] In this example, for s_i=-20.0 dB, β ij will be 0.8333.

[0517] Aspects of the example embodiments include the following Enumerated Example Embodiments ("EEEs"):

[0518] EEE1. A method (or system) for improving a signal-to-echo ratio to detect a voice command from a user, whereby

[0519] a. There are multiple devices used to create output audio program material

[0520] b. There is a known set of distances or ordered relationships between the device and the listener

[0521] c. The system selectively reduces the volume of the device closest to the user

[0522] EEE2. A method or system as described in EEE1, wherein the detection of signals includes signals from any noise-emitting objects or desired audio monitoring points that have a known distance relationship with the group of devices.

[0523] EEE3. A method or system as described in EEE1 or EEE2, wherein the ranking of devices includes consideration of distance to a nominal source distance and a signal-to-echo ratio of the devices.

[0524] EEE4. A method or system as described in any of EEE1 to EEE3, wherein the ranking takes into account the generalized proximity of the device to the user and the approximate reciprocity of the generalized proximity to estimate the most effective signal-to-echo ratio improvement and rank the devices in this sense. Various aspects of some disclosed embodiments include a system or device configured (e.g., programmed) to perform one or more disclosed methods, and a tangible computer-readable medium (e.g., a disk) storing code for implementing one or more disclosed methods or steps thereof. For example, the system can be or include a programmable general-purpose processor, a digital signal processor, or a microprocessor, which is programmed and / or otherwise configured with software or firmware to perform any of a variety of operations on data, including one or more disclosed methods or steps thereof. Such a general-purpose processor can be or include a computer system, which includes an input device, a memory, and a processing subsystem, which is programmed (and / or otherwise configured) to perform one or more disclosed methods (or steps thereof) in response to data asserted thereto.

[0525] Some disclosed embodiments are implemented as configurable (e.g., programmable) digital signal processors (DSPs), which are configured (e.g., programmed and otherwise configured) to perform the required processing on (multiple) audio signals, including the execution of one or more disclosed methods. Alternatively, some embodiments (or their elements) can be implemented as general-purpose processors (e.g., personal computers (PCs) or other computer systems or microprocessors, which may include input devices and memory), which are programmed with software or firmware and / or otherwise configured to perform any of the various operations including one or more disclosed methods or their steps. Alternatively, the elements of some disclosed embodiments are implemented as general-purpose processors or DSPs configured (e.g., programmed) to perform one or more disclosed methods or their steps, and the system further includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to perform one or more disclosed methods or their steps will typically be coupled to an input device (e.g., a mouse and / or keyboard), a memory, and a display device.

[0526] Another aspect of some disclosed embodiments is a computer-readable medium (e.g., a disk or other tangible storage medium) storing code for performing any embodiment of one or more disclosed methods or steps thereof (e.g., a codec executable to perform any embodiment of one or more disclosed methods or steps thereof).

[0527] Although specific embodiments and applications have been described herein, it will be apparent to those skilled in the art that many changes may be made to the embodiments and applications described herein without departing from the scope of the material described and claimed herein. It should be understood that although certain implementations have been shown and described, the present disclosure is not limited to the specific embodiments described and shown or the specific methods described.

Claims

1. A method for managing an audio session, comprising: receiving an output signal from each of a plurality of microphones in an audio environment, each of the plurality of microphones residing in a microphone location of the audio environment, the output signal comprising a signal corresponding to a current utterance of a person; determining one or more aspects of contextual information about the person based on the output signal, the contextual information comprising at least one of an estimated current location of the person or an estimated current proximity of the person to one or more microphone locations; determining a nearest loudspeaker-equipped audio device that is closest to a microphone position that is closest to the estimated current position of the person; selecting two or more audio devices of the audio environment based at least in part on the one or more aspects of the contextual information, the two or more audio devices each comprising at least one loudspeaker, and wherein the two or more audio devices include the nearest loudspeaker-equipped audio device; determining one or more types of audio processing changes to apply to audio data rendered to loudspeaker feeds of the two or more audio devices, the audio processing changes having the effect of increasing a speech-to-echo ratio at the microphone closest to the estimated current location of the person, wherein the echo comprises at least some of the audio output by the two or more audio devices, and wherein at least one of the audio processing changes for the proximate audio device is different from the audio processing change for a second audio device of the at least two audio devices, and wherein the one or more types of audio processing changes cause a reduction in loudspeaker reproduction level of the proximate audio device; and The one or more types of audio processing changes are caused to be applied.

2. The method according to claim 1, wherein The one or more types of audio processing changes involve spectral modifications.

3. The method according to any one of claims 1 to 2, wherein The one or more types of audio processing changes cause a decrease in loudspeaker reproduction levels of the loudspeakers of the two or more audio devices.

4. The method according to any one of claims 1 to 2, wherein Selecting two or more audio devices of the audio environment includes selecting N loudspeaker-equipped audio devices of the audio environment, N being an integer greater than two.

5. The method according to any one of claims 1 to 2, wherein Selecting the two or more audio devices of the audio environment is based at least in part on an estimated current position of the person relative to at least one of a microphone location or a loudspeaker-equipped audio device location.

6. The method according to any one of claims 1 to 2, wherein The one or more types of audio processing changes involve altering a rendering process to distort the rendering of an audio signal away from the estimated current position of the person.

7. The method of claim 2, wherein: The spectral modification involves reducing the level of audio data in the frequency band between 500 Hz and 3 KHz.

8. The method according to any one of claims 1 to 2, wherein The one or more types of audio processing changes involve inserting at least one gap into at least one selected frequency band of the audio playback signal.

9. The method according to any one of claims 1 to 2, wherein The one or more types of audio processing changes involve dynamic range compression.

10. The method according to any one of claims 1 to 2, wherein Selecting the two or more audio devices is based at least in part on a signal-to-echo ratio estimate for one or more microphone locations.

11. The method according to claim 10, wherein: Selecting the two or more audio devices is based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold.

12. The method of claim 10, wherein: Determining the one or more types of audio processing changes is based on an optimization of a cost function, the optimization being based at least in part on the signal-to-echo ratio estimate.

13. The method of claim 12, wherein: The cost function is based at least in part on rendering performance.

14. The method according to any one of claims 1 to 2, wherein Selecting the two or more audio devices is based at least in part on a proximity estimate.

15. The method of any one of claims 1 to 2, further comprising: determining a plurality of current acoustic features from the output signal of each microphone; applying a classifier to the plurality of current acoustic features, wherein applying the classifier involves applying a model trained on previously determined acoustic features derived from a plurality of previous utterances made by the person in a plurality of user zones in the environment; and Wherein one or more aspects of determining contextual information associated with the person involve determining an estimate of a user zone in which the person is currently located based at least in part on output from the classifier.

16. The method of claim 15, wherein: The estimate of the user zone is determined without reference to geometric positions of the plurality of microphones.

17. The method of claim 15, wherein: The current utterance and the previous utterance include a wake-up word utterance.

18. The method of any one of claims 1 to 2, further comprising selecting at least one microphone based on the one or more aspects of the contextual information.

19. The method according to any one of claims 1 to 2, wherein The one or more microphones reside in a plurality of audio devices of the audio environment.

20. The method according to any one of claims 1 to 2, wherein The one or more microphones reside in a single audio device of the audio environment.

21. The method according to any one of claims 1 to 2, wherein At least one microphone location of the one or more microphone locations corresponds to multiple microphones of a single audio device.

22. An apparatus configured to perform the method of any one of claims 1 to 21.

23. A system configured to perform the method of any one of claims 1 to 21.

24. One or more non-transitory media having software stored thereon, the software comprising instructions for controlling one or more devices to perform the method of any one of claims 1 to 21.

25. A computer program product comprising a computer program which, when executed by one or more processors, causes the method of any one of claims 1 to 21 to be performed.

Citation Information

Patent Citations

  • Information processing system and recording medium

    EP2874411A1

  • Speech recognition models based on location indicia

    US20140039888A1

  • Multi-channel echo cancellation and noise suppression

    US20140328490A1