Acoustic echo cancellation control for distributed audio devices

The system coordinates audio devices to enhance speech-to-echo ratio by applying audio processing modifications, addressing full-duplex performance challenges in multi-device environments through dynamic audio management and optimization.

JP7746513B2Active Publication Date: 2025-09-30DOLBY LABORATORIES LICENSING CORP +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024214353
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-21
Filing Date
2024-12-09
Publication Date
2025-09-30
Estimated Expiration
2040-07-29

AI Technical Summary

Technical Problem

Existing audio devices struggle with improving full-duplex audio performance, particularly in environments with multiple devices, leading to challenges in maintaining a high signal-to-echo ratio (SER) for effective communication.

Method used

A system that coordinates multiple audio devices to apply audio processing modifications, such as reducing loudspeaker playback levels and modifying audio signals to increase the speech-to-echo ratio, based on the estimated location of the user relative to microphone and loudspeaker positions, using classifiers and cost functions to optimize audio rendering.

Benefits of technology

Enhances the speech-to-echo ratio by dynamically managing audio output to improve full-duplex communication effectiveness in multi-device environments, reducing echo and optimizing audio delivery based on user location and device proximity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007746513000093
    Figure 0007746513000093
  • Figure 0007746513000094
    Figure 0007746513000094
  • Figure 0007746513000095
    Figure 0007746513000095
Patent Text Reader

Abstract

To provide a system and a method for coordinating and implementing multiple audio devices and for controlling the rendering of audio sounds by the audio devices.SOLUTION: An audio session management method includes receiving output signals from a plurality of microphones in an audio environment, determining one or more aspects of contextual information on the person on the basis of the output signals, selecting two or more loudspeaker-embedded audio devices based at least in part on the one or more aspects of the contextual information, determining one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for the audio devices, and applying the one or more audio processing modifications, the audio processing modifications have the effect of increasing a speech-to-echo ratio at one or more microphones.SELECTED DRAWING: Figure 2B
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application is a continuation of U.S. Provisional Patent Application No. 62 / 705,897, filed July 21, 2020; U.S. Provisional Patent Application No. 62 / 705,410, filed June 25, 2020; U.S. Provisional Patent Application No. 62 / 971,421, filed February 7, 2020; U.S. Provisional Patent Application No. 62 / 950,004, filed December 18, 2019; U.S. Provisional Patent Application No. 62 / 950,004, filed July 30, 2019. This application claims priority to Provisional Patent Application No. 62 / 880,122, U.S. Provisional Patent Application No. 62 / 880,113, filed July 30, 2019, European Patent Application No. 19212391.7, filed November 29, 2019, and Spanish Patent Application No. P201930702, filed July 30, 2019, the disclosures of each of which are incorporated herein by reference in their entirety.

[0002] The present application relates to systems and methods for orchestrating and implementing multiple audio devices (e.g., smart audio devices) and controlling the rendering of audio sounds by the audio devices. [Background technology]

[0003] Audio devices (including, but not limited to, smart audio devices) are widely used and are becoming commonplace elements in many homes. While existing systems and methods for controlling audio devices offer benefits, improved systems and methods are desired.

[0004] [Notation and Naming] Throughout this disclosure, including the claims, "speaker" and "loudspeaker" are used interchangeably to refer to any acoustically emitting transducer (or set of transducers) driven by a single speaker feed. A typical headphone set includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter) driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may receive different processing in different circuit branches connected to different transducers.

[0005] Throughout this disclosure, including the claims, the expression "performing" an operation on a signal or data (e.g., filtering, scaling, transforming, or applying a gain to a signal or data) is used broadly to mean performing the operation directly on the signal or data, or performing the operation on a processed version of the signal or data (e.g., a pre-filtered or pre-processed version of the signal before undergoing the operation).

[0006] Throughout this disclosure, including the claims, the term "system" is used broadly to mean a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where M of the inputs are generated by the subsystem and the other XM inputs are received from external sources) may also be referred to as a decoder system.

[0007] Throughout this disclosure, including the claims, the term "processor" refers to a processor that processes data (e.g., The term "processor" is used broadly to mean a system or device that is programmable or otherwise configurable (e.g., by software or firmware) to perform operations on audio or other sound data, for example, audio, or video or other image data. Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors that are programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0008] Throughout this disclosure, including the claims, the term "couples" or "connected" is used. The term "coupled" is intended to mean either a direct or an indirect connection. Thus, when a first device is connected to a second device, this connection may be achieved by a direct connection or by an indirect connection via other devices and connections.

[0009] As used herein, a "smart device" generally refers to an electronic device that is capable of operating interactively and / or autonomously to some degree and is configured to communicate with one or more other devices (or networks) via various wireless protocols, such as Bluetooth, Zigbee, near field communication, Wi-Fi, light fidelity (Li-Fi), 3G, 4G, 5G, etc. Some representative examples of smart devices include smartphones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smart watches, smart bands, smart keychains, and smart audio devices. The term "smart device" may also refer to devices that exhibit some characteristics of ubiquitous computing, such as artificial intelligence.

[0010] The term "smart audio device" is used herein to refer to a smart device that is either a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of virtual assistant functionality). A single-purpose audio device is a device (e.g., a smart speaker, television (TV), or mobile phone) that includes or is connected to at least one microphone (optionally also including or connected to at least one speaker and / or camera) and is generally or primarily designed to achieve a single purpose. For example, while a TV can typically play (or is considered to be able to play) audio from program material, modern TVs in most cases run some kind of operating system on which multiple applications run locally, including applications for watching television. Similarly, the audio inputs and outputs of a mobile phone can do many things, but these are provided by applications running on the phone. In this sense, a single-purpose audio device having speaker(s) and microphone(s) is often configured to run local applications and / or services to directly use the speaker(s) and microphone(s). There are also single-purpose audio devices that are configured to be grouped together to provide audio playback over a zone or user-defined area.

[0011] One general type of multipurpose audio device is an audio device that implements at least some aspects of virtual assistant functionality, although other aspects of the virtual assistant functionality may be implemented by one or more other devices, such as one or more servers with which the multipurpose audio device is configured to communicate. In the present specification, such a multipurpose audio device may be referred to as a "virtual assistant." A virtual assistant is a device (e.g., a smart speaker or a voice-assistant-embedded device) that includes or is connected to at least one microphone (and, optionally, includes or is connected to at least one speaker and / or at least one camera). In some examples, a virtual assistant provides the ability to utilize multiple devices (separate from the virtual assistant) for applications that are cloud-enabled in some sense or that are not fully implemented in the virtual assistant itself. In other words, at least some aspects of the virtual assistant functionality, such as voice recognition functionality, may be implemented (at least in part) by one or more servers or other devices with which the virtual assistant may communicate via a network such as the Internet. Multiple virtual assistants may cooperate, for example, in a discrete and conditionally defined manner. For example, two or more virtual assistants may cooperate in the sense that one of them (i.e., the one that is most confident that it heard the wake word) responds to the wake word. In some embodiments, multiple connected devices may form a kind of collection managed by a single main application, which may be (or may implement) a virtual assistant.

[0012] As used herein, the term "wake word" is used broadly to mean any sound (e.g., a word uttered by a human being, or any other sound). A smart audio device is configured to wake up in response to detecting ("hearing") a sound (using at least one microphone included in or connected to the smart audio device, or at least one other microphone). In this context, "Awake" in this context means that the device is listening for sound commands (i.e., listening In some cases, what may be referred to herein as a "wake word" may include multiple words, e.g., a phrase.

[0013] As used herein, the term "wake word detector" refers to a device (or software including instructions for configuring a device) configured to continuously search for a match between real-time sound (e.g., speech) features and a trained model. Typically, a wake word event is triggered whenever the wake word detector determines that the probability of detecting the wake word exceeds a predefined threshold. For example, the threshold may be a predetermined threshold adjusted to provide a reasonable compromise between false acceptance and false rejection rates. After a wake word event, the device enters a state where it is attentive to commands (sometimes referred to as an "awakened" or "attentiveness" state), and in this state, it listens for received commands. It can be passed to a larger, more computationally intensive recognizer.

[0014] As used herein, the term "microphone location" refers to the location of one or more microphones. In some examples, a microphone location may correspond to a microphone array consisting of multiple microphones provided in an audio device. For example, a microphone location may be a location corresponding to an entire audio device including one or more microphones. In some such examples, a microphone location may be a location corresponding to the center of gravity of the microphone array of an audio device. However, in some examples, a microphone location may be the location of a single microphone. In some such examples, the audio device may have only one microphone. Summary of the Invention

[0015] Some disclosed embodiments provide a listener or "user" experience that improves key criteria for successful full duplex communication in one or more audio devices. This criterion provides an approach to managing the experience. This is known as the Speech to Echo ratio (SER), and is sometimes referred to as the Speech to Echo Ratio. It may be defined as the ratio of a voice signal (or other desired signal) captured from an environment (e.g., a room) via one or more microphones to the echo presented at an audio device equipped with one or more microphones from output program content, interactive content, etc. It is conceivable that many audio devices in an audio environment incorporate both loudspeakers and microphones while also engaging in other functions. However, other audio devices in the audio environment may have one or more loudspeakers but no microphones, or one or more microphones but no loudspeakers. Some embodiments intentionally avoid using (or do not primarily use) the loudspeaker(s) closest to the user in a given use case or scenario. Alternatively or additionally, some embodiments may apply one or more other types of audio processing modifications to audio data rendered by one or more loudspeakers of the audio environment to increase the SER at one or more microphones of the audio environment.

[0016] Some embodiments are configured to implement a system including coordinated audio devices. In some implementations, the audio devices may include smart audio devices. According to some such implementations, two or more of the smart audio devices are wake word detectors (or are configured to implement a wake word detector). Thus, in such examples, multiple microphones (e.g., asynchronous microphones) are available. In some examples, each microphone may be included in at least one of the smart audio devices or configured to communicate with at least one of the smart audio devices. For example, at least some of the microphones may be independent microphones (e.g., in a consumer electronics device) that are not included in any of the smart audio devices but are configured to communicate with at least one of the smart audio devices (such that the microphone output can be captured by at least one of the smart audio devices). In some embodiments, each wake word detector (or each smart audio device that includes a wake word detector), or another subsystem of the system (e.g., a classifier), is configured to estimate a person zone by applying a classifier driven by multiple acoustic features obtained from at least some of the microphones (e.g., asynchronous microphones). In some implementations, the goal may not be to estimate the person's exact location, but to form a robust estimate of a discrete zone that includes the person's current location.

[0017] In some embodiments, a person (sometimes referred to herein as a "user"), a smart audio device, and a microphone are present in an audio environment (e.g., the user's residence, automobile, or workplace). Within this audio environment, sound may propagate from the user to the microphone. The audio environment may include multiple predefined zones. According to some examples, the environment may include at least the following zones: a cooking area, a dining area, an open area of ​​the living space, a TV area of ​​the living space (including a TV sofa), etc. During operation of the system, it is assumed that the user is physically present in one of the zones (the user's zone) at any given time, and that the user's zone may vary over time.

[0018] In some examples, the microphones may be asynchronous (e.g., digitally sampled using different sampling clocks) and randomly positioned (or at least not positioned at predetermined locations, not positioned symmetrically, not positioned in a grid, etc.). In some examples, the user's zone may be estimated via a data-driven approach driven, at least in part, by multiple high-level features obtained from at least one of the wake word detectors. These features (e.g., wake word confidence and receive level) may, in some examples, use little bandwidth and may be transmitted (e.g., asynchronously) to devices implementing the classifier with very little network load. [Problem to be solved by the invention]

[0019] Aspects of some embodiments relate to implementing and / or coordinating smart audio devices. [Means for solving the problem]

[0020] Aspects of some disclosed embodiments include a system configured (e.g., programmed) to perform one or more of the disclosed methods or steps thereof, and a tangible, non-transitory, computer-readable medium (e.g., a disk or other tangible storage medium) implementing non-transitory data storage having stored thereon code for performing one or more of the disclosed methods or steps thereof (e.g., code executable to perform one or more of the disclosed methods or steps thereof). For example, some disclosed embodiments may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of various operations on data comprising one or more of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system including an input device, memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more of the disclosed methods (or steps thereof) in response to asserted data.

[0021] In some embodiments, a control system may be configured to implement one or more methods disclosed herein, such as one or more audio session management methods. Some such methods include receiving (e.g., by a control system) an output signal from each of a plurality of microphones in an audio environment. In some examples, each of the plurality of microphones is present at a microphone position in the audio environment. In some examples, the output signal includes a signal corresponding to a person's current vocalization. According to some examples, the output signal includes a signal corresponding to non-speech audio data, such as noise and / or echo.

[0022] Some such methods include determining (e.g., by a control system) one or more aspects of contextual information about the person based on the output signal. In some examples, the contextual information includes an estimated current location of the person and / or an estimated current proximity of the person to one or more microphone locations. Some such methods include selecting two or more audio devices of the audio environment based at least in part on the one or more aspects of the contextual information. In some embodiments, each of the two or more audio devices includes at least one loudspeaker.

[0023] Some such methods include determining (e.g., by a control system) one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for the two or more audio devices. In some examples, the audio processing modifications have the effect of increasing the speech-to-echo ratio at one or more microphones. Some such methods include applying the one or more audio processing modifications.

[0024] According to some embodiments, the one or more audio processing modifications may result in a reduction in loudspeaker playback levels of the loudspeakers of the two or more audio devices. In some embodiments, at least one of the audio processing modifications for a first audio device may be different from the audio processing modifications for a second audio device. In some examples, selecting two or more audio devices of the audio environment (e.g., by a control system) may include selecting N loudspeaker-equipped audio devices of the audio environment, where N is an integer greater than 2.

[0025] In some embodiments, selecting the two or more audio devices of the audio environment may be based at least in part on an estimated current location of the person relative to at least one of a microphone location and a loudspeaker-equipped audio device location. According to some such embodiments, the method may include determining a nearest loudspeaker-equipped audio device closest to the estimated current location of the person or a nearest loudspeaker-equipped audio device closest to the microphone location closest to the estimated current location of the person. In some such examples, the two or more audio devices may include the nearest loudspeaker-equipped audio device.

[0026] In some examples, the one or more audio processing modifications include modifying a rendering process to warp a rendering of the audio signal away from the estimated current location of the person. In some embodiments, the one or more audio processing modifications may include spectral modification. According to some such embodiments, the spectral modification may include reducing the level of audio data in a frequency band between 500 Hz and 3 KHz.

[0027] In some embodiments, the one or more audio processing modifications may include inserting at least one gap in at least one selected frequency band of the audio playback signal. In some examples, the one or more audio processing modifications may include dynamic range compression.

[0028] According to some embodiments, selecting the two or more audio devices may be based at least in part on a signal-to-echo ratio estimate for one or more microphone positions. For example, selecting the two or more audio devices may be based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold. In some examples, determining the one or more audio processing modifications may be based on optimizing a cost function based at least in part on the signal-to-echo ratio estimate. For example, the cost function may be based at least in part on rendering performance. In some embodiments, selecting the two or more audio devices may be based at least in part on a proximity estimate.

[0029] In some examples, the method may include determining (e.g., by a control system) a plurality of current acoustic features from the output signal of each microphone, and applying a classifier to the plurality of current acoustic features. According to some embodiments, applying the classifier may include applying a model trained on previously determined acoustic features obtained from a plurality of past utterances made by the person in a plurality of user zones in the environment.

[0030] In some such examples, determining one or more aspects of context information about the person may include determining an estimate of a user zone in which the person is currently located based at least in part on output from the classifier. According to some embodiments, the estimate of the user zone may be determined without reference to geometric locations of the plurality of microphones. In some examples, the current utterance and the past utterance may be or may include an utterance of a wake word.

[0031] According to some embodiments, the one or more microphones may be located within multiple audio devices of the audio environment. However, in other examples, the one or more microphones may be located within a single audio device of the audio environment. In some examples, at least one of the one or more microphone locations may correspond to multiple microphones of a single audio device. Some disclosed methods may include selecting at least one microphone depending on the one or more aspects of the context information.

[0032] At least some aspects of the present disclosure may be implemented by a method, such as an audio session management method. As indicated elsewhere herein, in some examples, the method may be implemented, at least in part, by a control method as disclosed herein. Some such methods include receiving an output signal from each of a plurality of microphones in an audio environment. In some examples, each of the plurality of microphones is present at a microphone position in the audio environment. In some examples, the output signal includes a signal corresponding to a person's current vocalization. According to some examples, the output signal includes a signal corresponding to non-speech audio data, such as noise and / or echo.

[0033] Some such methods include determining one or more aspects of contextual information about the person based on the output signal. In some examples, the contextual information includes an estimated current location of the person and / or an estimated current proximity of the person to one or more microphone locations. Some such methods include selecting two or more audio devices of the audio environment based at least in part on the one or more aspects of the contextual information. In some embodiments, each of the two or more audio devices includes at least one loudspeaker.

[0034] Some such methods include determining one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for the two or more audio devices. In some examples, the audio processing modifications have the effect of increasing the speech-to-echo ratio at one or more microphones. Some such methods include applying the one or more audio processing modifications.

[0035] According to some embodiments, the one or more audio processing modifications are In some embodiments, at least one of the audio processing changes for a first audio device may be different from the audio processing changes for a second audio device. In some examples, selecting two or more audio devices of the audio environment may include selecting N loudspeaker-integrated audio devices of the audio environment, where N is an integer greater than 2.

[0036] In some embodiments, selecting the two or more audio devices of the audio environment may be based at least in part on an estimated current location of the person relative to at least one of a microphone location and a loudspeaker-equipped audio device location. According to some such embodiments, the method may include determining a nearest loudspeaker-equipped audio device closest to the estimated current location of the person or a nearest loudspeaker-equipped audio device closest to the microphone location closest to the estimated current location of the person. In some such examples, the two or more audio devices may include the nearest loudspeaker-equipped audio device.

[0037] In some examples, the one or more audio processing modifications include modifying a rendering process to warp a rendering of the audio signal away from the estimated current location of the person. In some embodiments, the one or more audio processing modifications may include spectral modification. According to some such embodiments, the spectral modification may include reducing the level of audio data in a frequency band between 500 Hz and 3 KHz.

[0038] In some embodiments, the one or more audio processing modifications may include inserting at least one gap in at least one selected frequency band of the audio playback signal. In some examples, the one or more audio processing modifications may include dynamic range compression.

[0039] According to some embodiments, selecting the two or more audio devices may be based at least in part on a signal-to-echo ratio estimate for one or more microphone positions. For example, selecting the two or more audio devices may be based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold. In some examples, determining the one or more audio processing modifications may be based on optimizing a cost function based at least in part on the signal-to-echo ratio estimate. For example, the cost function may be based at least in part on rendering performance. In some embodiments, selecting the two or more audio devices may be based at least in part on a proximity estimate.

[0040] In some examples, the method may include determining a plurality of current acoustic features from the output signal of each microphone and applying a classifier to the plurality of current acoustic features. According to some embodiments, applying the classifier may include applying a model trained on previously determined acoustic features obtained from a plurality of past utterances made by the person in a plurality of user zones in the environment.

[0041] In some such examples, determining one or more aspects of context information about the person may include determining an estimate of a user zone in which the person is currently located based at least in part on output from the classifier. According to an embodiment, the estimate of the user zone may be determined without reference to the geometric locations of the microphones. In some examples, the current utterance and the past utterance may be or may include an utterance of a wake word.

[0042] According to some embodiments, the one or more microphones may be located within multiple audio devices of the audio environment. However, in other examples, the one or more microphones may be located within a single audio device of the audio environment. In some examples, at least one of the one or more microphone locations may correspond to multiple microphones of a single audio device. Some disclosed methods may include selecting at least one microphone depending on the one or more aspects of the context information.

[0043] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including, but not limited to, random access memory (RAM) devices and read-only memory (ROM) devices. Thus, some innovative aspects of the subject matter described in this disclosure may be implemented in software-stored non-transitory media.

[0044] For example, the software may include instructions for controlling one or more devices to perform a method including receiving output signals from each of a plurality of microphones in an audio environment. In some examples, each of the plurality of microphones is present at a microphone position in the audio environment. In some examples, the output signals include signals corresponding to a person's current vocalization. According to some examples, the output signals include signals corresponding to non-speech audio data, such as noise and / or echo.

[0045] Some such methods include determining one or more aspects of contextual information about the person based on the output signal. In some examples, the contextual information includes an estimated current location of the person and / or an estimated current proximity of the person to one or more microphone locations. Some such methods include selecting two or more audio devices of the audio environment based at least in part on the one or more aspects of the contextual information. In some embodiments, each of the two or more audio devices includes at least one loudspeaker.

[0046] Some such methods include determining one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for the two or more audio devices. In some examples, the audio processing modifications have the effect of increasing the speech-to-echo ratio at one or more microphones. Some such methods include applying the one or more audio processing modifications.

[0047] According to some embodiments, the one or more audio processing modifications may result in a reduction in loudspeaker playback levels of the loudspeakers of the two or more audio devices. In some embodiments, at least one of the audio processing modifications for a first audio device may be different from an audio processing modification for a second audio device. In some examples, selecting two or more audio devices of the audio environment may include selecting two or more audio devices of the audio environment that have N loudspeakers. The method may include selecting an audio device, where N is an integer greater than 2.

[0048] In some embodiments, selecting the two or more audio devices of the audio environment may be based at least in part on an estimated current location of the person relative to at least one of a microphone location and a loudspeaker-equipped audio device location. According to some such embodiments, the method may include determining a nearest loudspeaker-equipped audio device closest to the estimated current location of the person or a nearest loudspeaker-equipped audio device closest to the microphone location closest to the estimated current location of the person. In some such examples, the two or more audio devices may include the nearest loudspeaker-equipped audio device.

[0049] In some examples, the one or more audio processing modifications include modifying a rendering process to warp a rendering of the audio signal away from the estimated current location of the person. In some embodiments, the one or more audio processing modifications may include spectral modification. According to some such embodiments, the spectral modification may include reducing the level of audio data in a frequency band between 500 Hz and 3 KHz.

[0050] In some embodiments, the one or more audio processing modifications may include inserting at least one gap in at least one selected frequency band of the audio playback signal. In some examples, the one or more audio processing modifications may include dynamic range compression.

[0051] According to some embodiments, selecting the two or more audio devices may be based at least in part on a signal-to-echo ratio estimate for one or more microphone positions. For example, selecting the two or more audio devices may be based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold. In some examples, determining the one or more audio processing modifications may be based on optimizing a cost function based at least in part on the signal-to-echo ratio estimate. For example, the cost function may be based at least in part on rendering performance. In some embodiments, selecting the two or more audio devices may be based at least in part on a proximity estimate.

[0052] In some examples, the method may include determining a plurality of current acoustic features from the output signal of each microphone and applying a classifier to the plurality of current acoustic features. According to some embodiments, applying the classifier may include applying a model trained on previously determined acoustic features obtained from a plurality of past utterances made by the person in a plurality of user zones in the environment.

[0053] In some such examples, determining one or more aspects of context information about the person may include determining an estimate of a user zone in which the person is currently located based at least in part on output from the classifier. According to some embodiments, the estimate of the user zone may be determined without reference to geometric locations of the plurality of microphones. In some examples, the current utterance and the past utterance may be or may include an utterance of a wake word.

[0054] According to some embodiments, the one or more microphones may be located within multiple audio devices of the audio environment. However, in other examples, the one or more microphones may be located within a single audio device of the audio environment. In some examples, at least one of the one or more microphone locations may correspond to multiple microphones of a single audio device. Some disclosed methods may include selecting at least one microphone depending on the one or more aspects of the context information.

[0055] The details of one or more embodiments of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the following description, drawings, and claims. It should be noted that the relative dimensions of the figures below may not be drawn to scale. [Brief explanation of the drawings]

[0056] [Figure 1A] FIG. 1A illustrates an audio environment according to an example. [Figure 1B] FIG. 1B shows another example of an audio environment. [Figure 2A] FIG. 2A is a block diagram illustrating example components of an apparatus capable of implementing various aspects of the present disclosure. [Figure 2B] FIG. 2B is a flow diagram including blocks of an audio session management method according to some embodiments. [Figure 3A] FIG. 3A is a block diagram of a system configured to implement separate rendering control and listening or capture logic across multiple devices. [Figure 3B] FIG. 3B is a block diagram of a system according to another disclosed embodiment. [Figure 3C] FIG. 3C is a block diagram of an embodiment configured to implement an energy balancing network, according to an example. [Figure 4] FIG. 4 is a graph illustrating an example of audio processing that may increase the speech-to-echo ratio in one or more microphones of an audio environment. [Figure 5] FIG. 5 is a graph illustrating another type of audio processing that may increase the speech-to-echo ratio at one or more microphones in an audio environment. [Figure 6] FIG. 6 illustrates another type of audio processing that may increase the speech-to-echo ratio at one or more microphones in an audio environment. [Figure 7] FIG. 7 is a graph illustrating another type of audio processing that may increase the speech-to-echo ratio at one or more microphones in an audio environment. [Figure 8] FIG. 8 is an illustration of an example where the audio device to be turned down may not be the audio device closest to the person speaking. [Figure 9] FIG. 9 shows a situation where a device with a very high SER is very close to the user. [Figure 10] FIG. 10 is a flow chart outlining an example of a method that may be performed by an apparatus such as that shown in FIG. 2A. [Figure 11] FIG. 11 is a block diagram of elements of an example embodiment configured to implement a zone classifier. [Figure 12] FIG. 12 is a flow chart outlining an example of a method that may be performed by an apparatus such as apparatus 200 of FIG. 2A. [Figure 13] FIG. 13 is a flow chart outlining another example of a method that may be performed by an apparatus such as apparatus 200 of FIG. 2A. [Figure 14] FIG. 14 is a flow chart outlining another example of a method that may be performed by an apparatus such as apparatus 200 of FIG. 2A. [Figure 15] FIG. 15 is a diagram showing an example set of speaker activation potentials and object rendering positions. [Figure 16] FIG. 16 is a diagram showing an example set of speaker activation potentials and object rendering positions. [Figure 17] FIG. 17 is a flow chart outlining an example of a method that may be performed by an apparatus or system such as that shown in FIG. 2A. [Figure 18] FIG. 18 is a graph of speaker activation potential in an example embodiment. [Figure 19] FIG. 19 is a graph of object rendering positions in an example embodiment. [Figure 20] FIG. 20 is a graph of speaker activation potential in an example embodiment. [Figure 21] FIG. 21 is a graph of object rendering positions in an example embodiment. [Figure 22] FIG. 22 is a graph of speaker activation potential in an example embodiment. [Figure 23] FIG. 23 is a graph of object rendering positions in an example embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0057] Currently, designers generally think of audio devices as a single interface point for audio voice, which may be a mix of entertainment, communication, and information services. Using audio voice for notifications and voice control has the advantage of avoiding visual or physical interruptions. The expanding device landscape is fragmented, with more systems competing for our pair of ears.

[0058] The challenge of improving full-duplex audio performance in all forms of interactive audio speech remains challenging. When audio output is present in a room that is not relevant to transmission or information capture in that room, it is desirable to remove this audio output from the captured signal (e.g., by echo cancellation and / or echo suppression). Some disclosed embodiments provide a user experience approach and management that improves the signal-to-echo ratio (SER), a key criterion for successful full-duplex communication in one or more devices.

[0059] Such embodiments may be useful in situations where there are multiple audio devices within hearing range of a user, each capable of providing audio program material at an appropriate volume at the user's location for a desired entertainment, communication, or information service. Such embodiments may be particularly valuable when there are three or more audio devices in similar proximity to the user.

[0060] Because rendering applications may be the primary function of audio devices, it may be desirable to use as many audio output devices as possible. If some audio devices are closer to the user, these audio devices may be more advantageous in terms of their ability to accurately position audio or deliver specific audio signaling and imaging to the user. However, if these audio devices contain one or more microphones, they may be preferable for picking up the user's voice. When considered together with the challenges of signal-to-echo ratio, it can be seen that using devices that are closer to the user in simplex (input-only) mode or moving the user closer to these devices dramatically improves the signal-to-echo ratio.

[0061] In various disclosed embodiments, an audio device may have both a built-in speaker and microphone while providing other functionality (e.g., the functionality shown in FIG. 1A ). Some disclosed embodiments implement the concept of intentionally not primarily using the loudspeaker(s) closest to the user in some situations.

[0062] In a connected operating system or disintermediation between applications (e.g., cloud-based applications), It is understood that audio devices may include many different types of devices (devices that provide input, output, and / or real-time interaction). Examples of such devices include wearable home audio devices, mobile devices, automated mobile computing devices, and smart speakers. Smart speakers may include network-connected speakers and microphones for cloud-based services. Other examples of such devices incorporate speakers and / or microphones and include lights, clocks, televisions, personal digital assistants, refrigerators, and trash cans. Some embodiments are particularly relevant to situations where a common platform exists for orchestrating multiple audio devices in an audio environment via an orchestrating device such as a smart home hub or other device configured to perform audio session management (sometimes referred to herein as an “audio session manager”). Some such implementations may utilize a device-non-specific language where the orchestrating device routes audio content among multiple users and / or locations specified by a software application. In some embodiments, the audio session manager may include commands between the audio session manager and a locally implemented software application. Some embodiments implement methods for dynamically managing rendering (e.g., including constraints to maintain spatial imaging by directing audio away from the nearest device), and / or for locating a user within a zone, and / or for mapping and locating devices relative to each other and to the user.

[0063] Typically, a system that includes multiple smart audio devices will need to indicate when it hears the "wake word" (as defined above) from the user and is paying attention to commands from the user (in other words, listening for commands from the user).

[0064] 1A illustrates an example audio environment. Some disclosed embodiments may be particularly useful in scenarios where there are multiple audio devices within any environment (e.g., a living space or workplace) that can transmit sound and capture audio, for example, as disclosed herein. The system of FIG. 1A may be configured in accordance with various disclosed embodiments.

[0065] FIG. 1A is a diagram of an audio environment (living space) including a set of smart audio devices (devices 1.1) for audio interaction, speakers (1.3) for audio output, and controllable lights (1.2). As with other disclosed embodiments, the types, numbers, and arrangements of elements in FIG. 1A are exemplary only. Other embodiments may provide more, fewer, and / or different elements. In some examples, one or more microphones of microphones 1.5 may be part of or associated with one of devices 1.1, lights 1.2, and speakers 1.3. Alternatively or additionally, one or more microphones of microphones 1.5 may be mounted in another part of the environment (e.g., a wall, ceiling, furniture, appliances, or another device in the environment). In one example, each of devices 1.1 includes (and / or is connected to) at least one microphone 1.5. Although not shown in FIG. 1A, some audio environments may include one or more cameras. According to some disclosed embodiments, one or more devices in the audio environment (e.g., one or more devices in device 1.1) may be connected to an audio session manager. A device (e.g., a device configured for audio session management, a device implementing an audio session manager, a smart home hub, etc.) may be able to estimate the location (e.g., in which zone of a living space) of a user (1.4) who uttered a wake word, command, etc. One or more devices (e.g., device 1.1) of the system shown in FIG. 1A may be configured to implement various disclosed embodiments. Various methods can be used to collectively obtain information from the devices of FIG. 3 to provide a location estimate of a user who uttered a wake word. According to some disclosed methods, information is collectively obtained from microphone 1.5 of FIG. 1A and provided to a device (e.g., a device configured for audio session management) implementing a classifier configured to provide a location estimate of a user who uttered a wake word.

[0066] Within a living space (e.g., the living space of FIG. 1A ), there exists a set of typical activity zones in which a person performs a task or activity or crosses a threshold. These areas (referred to herein as “user zones”) may, in some examples, be defined by the user without specifying coordinates or other indicators of a geometric location. According to some examples, a person’s “context” may include or correspond to the user zone in which the user is currently located or an estimate of that user zone. In FIG. 1A , the user zones include the following zones: 1. Kitchen sink and cooking area (in the upper left area of ​​the living space); 2. Refrigerator door (to the right of the sink and cooking area); 3. Dining area (in the lower left area of ​​the living space); 4. Open area of ​​living space (to the right of the sink and cooking area as well as the dining area); 5. TV sofa (right side of open area); 6. Television itself; 7. Table; 8. Door area or passageway (within the upper right area of ​​the living space). Other audio environments may include more user zones, fewer user zones, and / or other types of user zones, such as one or more bedroom zones, garage zones, patio or deck zones, etc.

[0067] According to some embodiments, a system that estimates (e.g., determines an uncertain estimate of) where a sound (e.g., a wake word or other attention-seeking signal) originated (or occurred) may have a certain confidence in the estimate (or may have multiple hypotheses). For example, if a person happens to be near the boundary of multiple user zones in an audio environment, the uncertain estimate of the person's location may include a certain confidence that the person is within each of these zones. In some conventional implementations of voice interfaces, the voice assistant's voice emanates from only one location at a time, requiring a single choice for one location (e.g., one of the eight speaker locations (1.1 and 1.3) in FIG. 1A ). However, based on simple hypothetical role-play, it is clear that (in such conventional implementations) the selected location of the assistant's voice source (i.e., the location of a speaker included in or connected to the voice assistant) may be unlikely to be a focus point for a natural response to express attention.

[0068] 1B illustrates another example of an audio environment. The audio environment illustrated in FIG. 1B includes a user 101 uttering direct speech 102 and a system including a pair of smart audio devices 103 and 105, a speaker for audio output, and a microphone. This system may be configured according to some disclosed embodiments. An utterance uttered by the user 101 (sometimes referred to herein as the "speaker") may be recognized as a wake word by one or more elements of the system.

[0069] More specifically, the elements of the system of FIG. 1B include: 102: Direct, local voice (generated by user 101). 103: A voice assistant device (connected to one or more loudspeakers). Device 103 is located closer to user 101 than device 105. Device 103 is sometimes referred to as the "proximal" device, and device 105 is sometimes referred to as the "distal" device. 104: Multiple microphones in (or connected to) the proximal device 103. 105: A voice assistant device (connected to one or more loudspeakers). 106: Multiple microphones in (or connected to) the distal device 105. 107: Home appliances (e.g., lamps). 108: Multiple microphones within (or connected to) home appliance 107. In some examples, each of microphones 108 may be configured to communicate with a device configured to implement a classifier (in some examples, at least one of devices 103 or 105). In some embodiments, the device configured to implement the classifier may also be a device configured for audio session management, such as a device configured to implement CHASM or a smart home hub.

[0070] The system of FIG. 1B may also include at least one classifier (e.g., classifier 1107 of FIG. 11 , described below). For example, device 103 (or device 105) may include the classifier. Alternatively or additionally, the classifier may be implemented by another device that may be configured to communicate with device 103 and / or device 105. In some examples, the classifier may be implemented by another local device (e.g., a device within environment 109). Conversely, in other examples, the classifier may be implemented by a remote device (e.g., a server) located outside environment 109.

[0071] According to some implementations, at least two devices (e.g., device 1.1 in FIG. 1A , devices 103 and 105 in FIG. 1B , etc.) cooperate in any manner (e.g., under the control of an orchestration device, such as a device configured for audio session management) to deliver audio such that the audio may be jointly controlled between the devices. For example, the two devices 103 and 105 may play audio individually or jointly. In one simple case, devices 103 and 105 operate as a joint pair, each rendering a portion of the audio (e.g., without loss of generality, one rendering substantially the left side and the other rendering substantially the right side of a stereo signal).

[0072] A home appliance 107 (or another device) may include one microphone 108 that is closest to the user 101 and does not have a loudspeaker. Consider then a situation where, for this particular audio environment and the particular location of this user 101, there is already a favorable signal-to-echo ratio or speech-to-echo ratio (SER) that cannot be improved by changing the audio processing on the audio played by the speaker(s) of the device 105 and / or home appliance 107. In some embodiments, no such microphone exists.

[0073] Some disclosed embodiments have a detectable and significant SER performance impact Some implementations provide such benefits without implementing aspects of zone location and / or dynamically variable rendering. However, some embodiments implement audio processing modifications that include repelling or “warping” sound objects (or audio objects) away from a device. A reason for warping audio objects away from a particular audio device, a particular location, etc., is, in some examples, to improve the signal-to-echo ratio at a particular microphone used to capture human speech. Such warping may include, but is not limited to, lowering the playback level of one, two, or more nearby audio devices. In some cases, audio processing modifications to improve SER may be signaled by zone detection techniques such that the one, two, or more nearby audio devices for which audio processing modifications are implemented (e.g., lowering playback levels) are the audio devices closest to the user, the audio devices closest to a particular microphone used to capture the user's speech, and / or the audio devices closest to a voice of interest.

[0074] Aspects of some embodiments include context, decisions, and audio processing modifications (referred to herein as "rendering modifications"). In some examples, these aspects are as follows: CONTEXT (such as location and / or time). In some cases , location and time are part of the context, and each can be provided or determined in a variety of ways. DECISION (which may include continuous adjustment of thresholds or changes). This component may be simple or complex, depending on the particular embodiment. In some embodiments, decisions may be made continuously, for example, in response to feedback. In some instances, decisions may result in system stability, such as virtuous feedback stability, as described below. RENDER (the essence of audio processing modification). Although the term "rendering" is used in this specification, audio processing modification may also include rendering modification. In some embodiments, there are multiple options for audio processing modifications, ranging from implementing barely perceptible audio processing modifications to rendering audio processing modifications that are severe and obvious.

[0075] In some examples, "context" may include information about location and intent. For example, context information may include at least a rough knowledge of the user's location, such as an estimate of the user's zone that matches the user's current location. Context information may also correspond to audio object locations (e.g., audio object locations that match the user's utterance of a wake word). In some examples, context information may include information about the timing and likelihood that an object or individual made a sound. Examples of context include, but are not limited to, the following: A. Knowing where the likely location is. This is based on the following: i) Weak or low probability detection (eg, detection of a sound that may be of interest but may not be clear enough to act upon). ii) Specific activation (e.g., a spoken and clearly detected wake word). iii) Habits and patterns (e.g., based on pattern recognition, such that a given location, such as a sofa near a television, is associated with one or more people sitting on the sofa watching video material on the television and listening to associated audio). iv) Integration of other forms of proximity sensing based on and / or other modalities (e.g., one or more infrared sensors, cameras, capacitive sensors, radio frequency (RF) sensors, heat sensors, pressure sensors, wearable beacons, etc. (e.g., located in or on furniture in the audio environment)). B. Knowing or estimating the likelihood of a sound that a person wants to hear, for example with improved detection. This may include some or all of the following: i) Events based on the detection of any audio sound, such as wake word detection. ii) Events or contexts based on a known activity or series of events (e.g., a pause in the display of video content, a space for interaction in scripted automatic speech recognition (ASR) type interactive content, or a change in activity and / or a change in the interaction dynamics of a full-duplex communication activity (e.g., a pause by one or more participants in a video conference). iii) Additional sensory input from other modalities iv) The option to frequently improve listening in any manner – increased readiness or improved listening.

[0076] To illustrate the key difference between A (knowing where the likely locations are) and B (knowing or estimating the likelihood of the audio the user wants to hear, e.g., with improved detection), A includes specific location information or knowledge without necessarily knowing whether there is more to hear, whereas B focuses more on specific timing or event information without necessarily knowing exactly where to listen. Of course, there may be overlap in some aspects of A and B. For example, weak or full detection of the wake word has information about both location and timing.

[0077] For some use cases, what is important is that the "context" includes information about both the location of the listening interest (e.g., the location of the person and / or the nearest microphone) and timing. This contextual information drives one or more associated decisions and one or more possible audio processing changes (e.g., one or more possible rendering changes). Thus, various embodiments consider many possibilities based on various information that may be used to form the context.

[0078] Next, we will discuss the "Determine" aspect. This aspect may involve, for example, determining one, two, or more output devices for which associated audio processing will be modified. One simple way to form such a determination is as follows:

[0079] Given information from the context (e.g., location and / or events (or confidence that there is something important about that location in some sense)), in some examples, the audio session manager may determine or estimate the distance from that location to some or all audio devices in the audio environment. In some embodiments, the audio session manager may also generate a set of activation potentials for each loudspeaker (or set of loudspeakers) for some or all audio devices in the audio environment. According to some such examples, the set of activation potentials may be determined as [f_1, f_2, ..., f_n], which, without loss of generality, may be in the range [0..1]. In another example, the result of the determination may describe a target speech-to-echo ratio improvement value [s_1, s_2, ..., s_n] for each device in the "rendering" aspect. In a further example, both the activation potentials and the speech-to-echo ratio improvement value may be generated by the "determining" aspect.

[0080] In some embodiments, the activation potential conveys the degree to which the "rendering" aspect ensures improved SER at the desired microphone location. In some such examples, a maximum value of f_n indicates aggressive ducking or warping of the rendered audio, or, given a value s_n, indicates that the audio is limited and ducked to achieve a speech-to-echo ratio of s_n. Intermediate values ​​of f_n closer to 0.5 may indicate that only moderate rendering changes are required, and that warping the audio source toward these locations is appropriate in some embodiments. Furthermore, in some implementations, low values ​​of f_n may be deemed insignificant to attenuate. In some such implementations, f_n values ​​below a threshold level may not be asserted. According to some examples, f_n values ​​below a threshold level may correspond to a location to which the rendering of the audio content is warped. In some examples, loudspeakers corresponding to f_n values ​​below a threshold level may have their playback level increased according to a process described below.

[0081] According to some implementations, the above method (or one of the alternative methods described below) may be used to generate control parameters for each selected audio processing modification for all selected audio devices, e.g., for each device in the audio environment, for one or more devices in the audio environment, for two or more devices in the audio environment, for three or more devices in the audio environment, etc. The selection of audio processing modifications may vary depending on the particular implementation. For example, a decision may be made based on: - a set of two or more loudspeakers for which audio processing is to be modified; - determining an extent to which to modify the audio processing for the set of two or more loudspeakers. The extent of modification may, in some examples, be determined in the context of a designed or determined range, which may be based at least in part on the capabilities of one or more loudspeakers included in the set of loudspeakers. In some examples, the capabilities of each loudspeaker may include frequency response, playback level limits, and / or parameters of one or more loudspeaker dynamics processing algorithms.

[0082] For example, a design choice may determine that the best option in a particular situation is to reduce the volume of a loudspeaker. In some such examples, a maximum and / or minimum range of audio processing modifications may be determined. For example, the range for reducing the volume of any loudspeaker may be limited to a particular threshold, such as 15 dB, 20 dB, 25 dB, etc. In some such implementations, the decision may be based on heuristics or logic for selecting one, two, or more loudspeakers and may be based on factors such as the confidence of the activity of interest and the loudspeaker locations. The decision may be to reduce (duck) the volume of audio played by one, two, or more loudspeakers by any amount within a minimum and maximum range (e.g., 0 dB to 20 dB). The decision method (or system element) may, in some examples, generate a set of activation potentials for each loudspeaker-embedded audio device.

[0083] In one simple example, the decision process may be as simple as determining that all audio devices except one have a rendering-triggered change value of 0, and determining that one audio device has a rendering-triggered change value of 1. The design of the audio processing change (e.g., volume reduction (ducking)) and the range of the audio processing change (e.g., time constant) may be independent of the decision logic in some examples. This approach results in a simple and effective design.

[0084] However, other embodiments may include selecting two or more audio devices with built-in loudspeakers and modifying the audio processing for at least two, at least three (and, in some examples, all) of the two or more audio devices with built-in loudspeakers. In some such examples, at least one of the audio processing modifications (e.g., reducing the playback level) for a first audio device may be different from the audio processing modification for a second audio device. The difference between these audio processing modifications may, in some examples, be based at least in part on the person's estimated current location or microphone position relative to the location of each audio device. According to some such embodiments, the audio processing modification may include, as part of modifying the rendering process, applying different speaker activation potentials at different loudspeaker locations to warp the rendering of the audio signal away from the person of interest's estimated current location. The difference between these audio processing modifications may, in some examples, be based at least in part on the capabilities of the loudspeakers. For example, if an audio processing change includes reducing the level of audio sounds in the bass range, such a change may be applied more aggressively to an audio device that includes one or more loudspeakers capable of playing loud sounds in the bass range.

[0085] Next, further details regarding the audio processing modification aspect (sometimes referred to herein as the "rendering modification" aspect) are provided. This disclosure defines this aspect as "turn nearest down" (e.g., one, two, or three speakers). This may be referred to as "reducing the volume at which audio content played by the most proximal loudspeaker is rendered," but more generally (as illustrated elsewhere herein) many embodiments may implement one or more modifications to audio processing aimed at improving the overall estimate, measure, and / or criterion of signal-to-echo ratio with respect to the ability to capture or detect a desired audio source (e.g., the speaker of the wake word). In some cases, the audio processing modification (e.g., "lowering" the volume of rendered audio content) is (or can be) adjusted by some continuous parameter of the resulting volume. For example, in the context of lowering the volume of a loudspeaker, some embodiments may be able to apply an adjustable (e.g., continuously adjustable) amount of attenuation (dB). In some such examples, the adjustable amount of attenuation may have a first range (e.g., 0-3 dB) for a barely noticeable change and a second range (e.g., 0-20 dB) that provides a particularly effective improvement in SER but is absolutely noticeable to a listener.

[0086] The above schema (CONTEXT, DECISION, and RENDERING) In some embodiments that implement RENDER or RENDERING CHANGE, there may not be a particular hard boundary of "closest" (e.g., for a loudspeaker or device that is "closest" to a user or another individual or system element), and without loss of generality, a rendering change may be implemented as follows: The method may or may include varying (e.g., continuously varying) one or more of A and B.

[0087] A. A mode of modifying output to reduce audio output from one or more audio devices, where the modification of audio output may include one or more of the following i) to vi): i) Reducing the overall level of the audio device output (reducing the volume of one or more loudspeakers, muting one or more loudspeakers). ii) For example, by using a nearly linear equalization (EQ) filter designed to produce an output that is different from the spectrum of the audio sound we want to detect. Shaping the spectrum of the loudspeaker output. In some examples, if the output spectrum is being shaped to detect human voices, the filter may reduce frequencies in the range of approximately 500 Hz to 3 kHz (e.g., ±5% or ±10% at each end of this frequency range), or may emphasize the low and high frequency bands and emphasize the mid-band (e.g., approximately 500 The loudness can be shaped so that space remains in the frequency range (Hz to 3 kHz). iii) Modifying the output ceiling or peak to either reduce the peak level and / or reduce distortion products that may additionally degrade the performance of any echo cancellation that is part of the overall system (e.g., a time-domain dynamic range compressor or a multi-band frequency-dependent compressor) that produces the achieved SER for audio detection. Such audio signal modifications may effectively reduce the amplitude of the audio signal and may contribute to limiting the loudspeaker excursion. iv) Spatially steering the audio in a way that tends to reduce energy, or connecting the output of one or more loudspeakers to one or more microphones for which the system (e.g., audio processing manager) achieves a higher SER, as in the "warping" example described herein. v) Using temporal time slicing or time adjustment to create "gaps" or periods of lower output sparse time frequencies sufficient to obtain glimpses of audio, similar to the gap insertion example described below. vi) Modifying the audio voice in any combination of the above ways.

[0088] B. Conserving energy and / or creating continuity at a particular listening position or a broad set of listening positions, including, for example, one or more of the following i) and ii): i) In some instances, energy removed from one loudspeaker can be compensated for by providing additional energy to another loudspeaker. In some instances, the overall loudness remains the same, or substantially the same. While this is not a required feature, it can be an effective means of allowing more drastic changes to be made to the audio processing of the "nearest" device or set of nearest-neighbor devices without losing content. However, continuity and / or energy conservation can be particularly relevant when dealing with complex audio outputs and audio scenes. ii) Activation time constant. In particular, the audio processing change may be applied a little earlier (e.g., 100-200 ms) than the return to normal state (e.g., 1000-10000 ms), so that the audio processing change, if perceptible, appears intentional, but then the return from the changed state to the normal state may not appear to be related to any actual event or change (from the user's perspective) and may in some instances be so slow as to be barely perceptible.

[0089] We now describe further examples of how contexts and decisions can be formulated and determined.

[0090] Embodiment A (CONTEXT) For example, context information can be expressed mathematically as follows: can be formulated as follows: H(a,b), the approximate physical distance between device a and device b (meters):

number

number

[0091] Determine H and S: H is a property of the physical location of the device, which can be determined or estimated by (1) and (2) below. (1) Direct user instruction, such as using a smartphone or tablet device to mark or indicate the approximate location of a device on a floor plan or similar graphical representation of the environment. Such digital interfaces are already commonplace in managing the configuration, grouping, names, purposes, and identities of smart home devices. For example, such direct instruction may be provided via an Amazon Alexa smartphone application, a Sonos S2 controller application, or a similar application. (2) For example, as disclosed in J. Yang and Y. Chen, "Indoor Localization Using Improved RSS-Based Lateration Methods," GLOBECOM 2009 - 2009 IEEE Global Telecommunications Conference, Honolulu, HI, 2009, pp. 1-6, doi: 10.1109 / GLOCOM.2009.5425237 and / or Mardeni, R. & Othman, Shaifull & Nizam, (2010) "Node Positioning in ZigBee Network Using Trilateration Method Based on the Received Signal Strength Indicator (RSSI)" 46 (both of which are incorporated herein by reference), a basic trilateration problem is solved using the measured signal strength (sometimes referred to as Received Signal Strength Indicator or RSSI) of common wireless communication technologies such as Bluetooth, Wi-Fi, and ZigBee to generate an estimate of the physical distance between devices.

[0092] S(a) is an estimate of the speech-to-echo ratio at device a. By definition, the speech-to-echo ratio (dB) is given by:

number

[0093] In the above formula, JPEG0007746513000004.jpg75 is the estimated speech energy (dB), JPEG0007746513000005.jpg74 is an estimate of the residual echo energy (in dB) after echo cancellation. Various methods for estimating these quantities are disclosed herein, including, for example:

[0094] (1) Speech energy and residual echo energy may be estimated by an offline measurement process performed on a particular device, taking into account the acoustic connection between the device's microphone and speaker and the performance of its on-board echo cancellation circuitry. In some such examples, the average speech energy level "AvgSpeech" may be determined by the average level of human speech measured by the device at a nominal distance. For example, speech from a small number of people standing one meter away from a microphone-equipped device may be recorded by the device during production, and the energy may be averaged to generate AvgSpeech. According to some such examples, the average residual echo energy level "AvgEcho" may be estimated by playing music content from the device during production and running the on-board echo cancellation circuitry to generate an echo residual signal. AvgEcho may be estimated by averaging the energy of the echo residual signal for a small sample of music content. When the device is not playing audio, AvgEcho may be set to a nominal low value (e.g., -96.0 dB). In some such implementations, speech energy and residual echo energy may be expressed as follows:

number

[0095] (2) According to some examples, the average speech energy may be determined by obtaining the energy of a microphone signal corresponding to a user's speech as determined by a voice activity detector (VAD). In some such examples, the average residual echo energy may be estimated by the energy of the microphone signal when the VAD does not indicate speech. If x is the pulse code modulation (PCM) samples of device a's microphone at a sampling rate, and V is a VAD flag that takes the value 1.0 for samples corresponding to a speech activity and the value 0.0 otherwise, the speech energy and residual echo energy may be expressed as follows:

number

[0096] (3) Further to the above method, in some embodiments, the energy in the microphone may be treated as a random variable and modeled separately based on the VAD decision. Statistical models S and E of the speech energy and echo energy, respectively, may be estimated using any number of statistical modeling techniques. Average values ​​(dB) for both speech and echo to approximate S(a) may be derived from S and E, respectively. Common methods for achieving this exist in the field of statistical signal processing. For example, Assuming a Gaussian distribution of energy and biased second-order statistics JPEG0007746513000008.jpg72 and Calculate JPEG0007746513000009.jpg73. A histogram of energy values ​​consisting of discrete bins is created to obtain a possibly multimodal distribution, where, after applying an expectation-maximization (EM) parameter estimation step for a mixture model (e.g., a Gaussian mixture model), the largest mean value belonging to any of the sub-distributions in the mixture model is calculated. JPEG0007746513000010.jpg723

[0097] (DECISION) As described elsewhere herein, in various disclosed embodiments, the determining aspect determines which devices have received audio processing modifications, such as rendering modifications, and in some embodiments, which devices have received an indication of how much modification is required for the desired SER improvement. Some such embodiments may be configured to improve the SER on devices with the best initial SER values, determined, for example, by finding the maximum value of S across all devices included in set D. Other embodiments may be configured to opportunistically improve the SER on devices that are regularly spoken to by the user, determined based on historical usage patterns. Other embodiments may , may be configured to attempt to improve the SER at multiple microphone positions. For example, multiple devices are selected for the following discussion.

[0098] Once one or more microphone positions have been determined, in some such implementations, a desired SER improvement (SERI) may be determined as follows:

number

[0099] In the above equation, m indicates the device / microphone position to be improved, and TargetSER is a threshold value that may be set by the application in use. For example, a wake word detection algorithm may tolerate a lower operating SER than a large vocabulary speech recognizer. Typical values ​​for TargetSER may be on the order of -6 dB to 12 dB. As previously mentioned, in some embodiments, if S(m) is not known or cannot be easily estimated, any preset value based on offline measurements of speech and echo recorded in a typical reverberant room or setting may suffice. Some embodiments may determine which device attempts to modify its audio processing (e.g., rendering) by specifying f_n, which ranges from 0 to 1. Other embodiments may include specifying the degree to which the audio processing (e.g., rendering) should be modified in units of a speech-to-echo ratio improvement value (in decibels), s_n, where s_n may be calculated as follows:

number

[0100] Some embodiments may calculate f_n directly from the device geometry, for example:

number

[0101] In the above equation, m is the index of the device selected for the largest modification of audio processing (e.g., rendering). Other implementations may include other options for relaxing or smoothing the function over device geometry.

[0102] Embodiment B (User Zone Reference) In some embodiments, the contextual and decisional aspects of the present disclosure may be generated in a context where one or more user zones exist, such as a set of acoustic features, as described in more detail later in this specification. Using JPEG0007746513000014.jpg738, the posterior probability JPEG0007746513000015.jpg717(C k is a set of zone labels, JPEG0007746513000016.jpg718, and there are K distinct user zones in the environment). Associating each audio device with each user zone may be accomplished by the user themselves as part of a training process described herein, or via an application such as, for example, the Alexa smartphone app or the Sonos S2 Controller smartphone app. For example, some implementations may associate the jth device with a zone label C k to associate a user zone with the In some embodiments, the image may be expressed as JPEG0007746513000017.jpg723. JPEG0007746513000018.jpg712 and posterior probability JPEG0007746513000019.jpg717 may be considered as context information. Some embodiments may instead consider the acoustic features W(j) themselves as part of the context. In other embodiments, these quantities ( JPEG0007746513000020.jpg712, posterior probability JPEG0007746513000021.jpg717, and the acoustic feature W(j) itself), and / or the quantity thereof may be part of the context information.

[0103] The decision aspect of various embodiments may use quantities related to one or more user zones in selecting a device. If both z and p are available, an example decision may be made as follows:

number

[0104] We also consider another type of embodiment where the acoustic features W(j) are used directly in the decision aspect. For example, we define the wake word confidence score associated with utterance j as w(j). n If (j), then the selection of the device can be made according to the following formula:

number

[0105] FIG. 2A is a block diagram illustrating example components of a device or system capable of implementing various aspects of the present disclosure. As with other figures herein, the types and numbers of elements shown in FIG. 2A are exemplary only. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 200 may be configured to implement various aspects of the present disclosure. The device 200 may be or include a device configured to perform at least some of the methods described herein. In some embodiments, the device 200 may be or include a smart speaker, a laptop computer, a mobile phone, a tablet device, a smart home hub, or another device configured to perform at least some of the methods disclosed herein. In some embodiments, the device 200 may be configured to implement an audio session manager. In some such embodiments, the device 200 may be or include a server.

[0106] In this example, device 200 includes interface system 205 and control system 210. Interface system 205, in some embodiments, may be configured to communicate with one or more devices running (or configured to run) software applications. Such software applications may also be referred to as “applications” or simply “apps.” Interface system 205, in some embodiments, may be configured to exchange control information and associated data related to the applications. Interface system 205, in some embodiments, may be configured to communicate with one or more other devices in an audio environment. The audio environment may, in some examples, be a home audio environment. Interface system 205, in some embodiments, may be configured to exchange control information and associated data with audio devices in the audio environment. The control information and associated data, in some examples, relate to one or more applications with which device 200 is configured to communicate.

[0107] The interface system 205, in some embodiments, may be configured to receive audio data. The audio data may include audio signals intended to be played by at least some speakers of the audio environment. The audio data may include one or more audio signals and associated spatial data. The spatial data may include, for example, channel data and / or spatial metadata. The interface system 205 may be configured to provide rendered audio signals to at least some loudspeakers of a set of loudspeakers of the audio environment. The interface system 205, in some embodiments, may be configured to receive input from one or more microphones in the environment.

[0108] The interface system 205 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 205 may include one or more wireless interfaces. The interface system 205 may include one or more devices that implement a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 205 may include one or more interfaces between the control system 210 and a memory system (such as optional memory system 215 shown in FIG. 2A ). However, in some examples, the control system 210 may include a memory system.

[0109] The control system 210 may be implemented, for example, by a single-chip or multi-chip general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device. The circuit may include discrete gate or transistor logic and / or discrete hardware components.

[0110] In some embodiments, control system 210 may be provided in multiple devices. For example, a portion of control system 210 may be provided in a device present in one of the environments illustrated herein, and another portion of control system 210 may be provided in a device present outside of the environment, such as a server or a mobile device (e.g., a smartphone or tablet computer). In other examples, a portion of control system 210 may be provided in a device present in one of the environments illustrated herein, and another portion of control system 210 may be provided in one or more other devices in the environment. For example, the functionality of the control system may be distributed among multiple smart audio devices in the environment, or may be shared by an orchestration device (such as what is referred to herein as an “audio session manager” or “smart home hub”) and one or more other devices in the environment. Interface system 205 may also be provided in multiple devices in some such examples.

[0111] In some embodiments, the control system 210 may be configured to at least partially perform the methods disclosed herein. According to some examples, the control system 210 may be configured to implement an audio session management method. This audio session management method may, in some examples, include determining one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for two or more audio devices in the audio environment. According to some embodiments, the audio processing modifications may have the effect of increasing the speech-to-echo ratio at one or more microphones in the audio environment.

[0112] Some or all of the methods disclosed herein may be performed by one or more devices in response to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices (including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc.) as described herein. The one or more non-transitory media may be provided, for example, within any of memory system 215 and / or control system 210 shown in FIG. 2A . Thus, various innovative aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media storing software. The software may include, for example, instructions for controlling at least one device to implement an audio session management method. The software may, in some examples, include instructions for controlling one or more audio devices of an audio environment to acquire, process, and / or provide audio data. The software may, in some examples, include determining one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for two or more audio devices of the audio environment. According to some implementations, the audio processing modifications may have the effect of increasing the speech-to-echo ratio at one or more microphones in the audio environment. The software may be executable by one or more components of a control system, such as, for example, control system 210 of FIG. 2A.

[0113] In some examples, device 200 may include optional microphone system 220 shown in FIG. 2A. Optional microphone system 220 may include one or more microphones. In some embodiments, one or more of the microphones may be part of or associated with another device, such as a speaker of a speaker system or a smart audio device. In some examples, device 200 may not include microphone system 220. However, In some such embodiments, device 200 may still be configured to receive microphone data for one or more microphones in the audio environment via interface system 210.

[0114] According to some embodiments, device 200 may include any of loudspeaker systems 225 shown in FIG. 2A . Any of loudspeaker systems 225 may include one or more loudspeakers. A loudspeaker may also be referred to herein as a “speaker.” In some examples, at least some of the loudspeakers of any of loudspeaker systems 225 may be arranged in any location. For example, at least some of the speakers of any of loudspeaker systems 225 may be arranged in locations that do not correspond to any standard, prescribed speaker layout, such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, Hamasaki 22.2, etc. In some such examples, at least some of the loudspeakers of any of loudspeaker systems 225 may be arranged in locations that are convenient for the space (e.g., where there is space to accommodate the loudspeakers) and may not be arranged in any standard, prescribed speaker layout. In some examples, the device 200 may not include any loudspeaker system 225 .

[0115] In some embodiments, device 200 may include any of sensor systems 230 shown in FIG. 2A . Optional sensor system 230 may include one or more cameras, touch sensors, gesture sensors, motion detectors, etc. According to some embodiments, optional sensor system 230 may include one or more cameras. In some embodiments, the camera may be a freestanding camera. In some examples, one or more cameras of optional sensor system 230 may be provided within a smart audio device, where the smart audio device may be a single-purpose audio device or a virtual assistant. In some such examples, one or more cameras of optional sensor system 230 may be provided within a television, a mobile phone, or a smart speaker. In some examples, device 200 may not include sensor system 230. However, in some such embodiments, device 200 may still be configured to receive sensor data for one or more sensors in the audio environment via interface system 210.

[0116] In some embodiments, device 200 may include optional display system 235 shown in FIG. 2A . Optional display system 235 may include one or more display devices, such as one or more light-emitting diode (LED) displays. In some examples, optional display system 235 may include one or more organic light-emitting diode (OLED) displays. In some examples in which device 200 includes display system 235, sensor system 230 may include a touch sensor system and / or a gesture sensor system proximate one or more display devices of display system 235. According to some such embodiments, control system 210 may be configured to control display system 235 to present one or more graphical user interfaces (GUIs).

[0117] According to some examples, device 200 may be or may include a smart audio device. In some such implementations, device 200 may be or may (at least partially) implement a wake word detector. For example, device 200 may be or may (at least partially) implement a virtual assistant.

[0118] 2B is a flow diagram including blocks of an audio session management method according to some embodiments. The blocks of method 250, as well as other methods described herein, do not necessarily have to be performed in the order shown. In some embodiments, the blocks of method 250 One or more of the blocks may be performed simultaneously. Additionally, some implementations of method 250 may include more or fewer blocks than those shown and / or described. The blocks of method 250 may be performed by one or more devices, which may be (or may include) a control system, such as control system 210 described above with reference to FIG. 2A or one of the other disclosed example control systems. According to some implementations, the blocks of method 250 may be performed, at least in part, by a device implementing what is referred to herein as an audio session manager.

[0119] According to this example, block 255 includes receiving an output signal from each of a plurality of microphones in the audio environment. In this example, each of the plurality of microphones is located at a microphone position in the audio environment, and the output signal includes a signal corresponding to a person's current utterance. In some examples, the current utterance may be an utterance of a wake word. However, the output signal may also include a signal corresponding to a time when the person is not speaking. Such a signal may be used, for example, to establish a baseline level of echo, noise, etc.

[0120] In this example, block 260 includes determining one or more aspects of context information about the person based on the output signal. In this embodiment, the context information includes the person's estimated current location and / or the person's estimated current proximity to one or more microphone locations. As noted above, the phrase "microphone location" as used herein refers to the location of one or more microphones. In some examples, a microphone location may correspond to a microphone array consisting of multiple microphones in an audio device. For example, a microphone location may be a location corresponding to an entire audio device including one or more microphones. In some such examples, a microphone location may be a location corresponding to the center of gravity of the microphone array of an audio device. However, in some examples, a microphone location may be the location of a single microphone. In some such examples, an audio device may have only one microphone.

[0121] In some examples, determining the context information may include generating an estimate of a user zone in which the person is currently located. Some such examples may include determining a plurality of current acoustic features from the output signals of each microphone and applying a classifier to the plurality of current acoustic features. Applying the classifier may include, for example, applying a model trained on previously determined acoustic features obtained from a plurality of past utterances made by the person within a plurality of user zones in the environment. In some such examples, determining one or more aspects of the context information about the person may include determining an estimate of a user zone in which the person is currently located based at least in part on output from the classifier. In some such examples, the estimate of the user zone may be determined without reference to the geometric locations of the plurality of microphones. According to some examples, the current utterance and the past utterance may be or may include an utterance of a wake word.

[0122] According to this embodiment, block 265 includes selecting two or more audio devices of the audio environment based at least in part on one or more aspects of the context information, each of the two or more audio devices including at least one loudspeaker. In some examples, selecting two or more audio devices of the audio environment may include selecting N loudspeaker-equipped audio devices of the audio environment, where N is an integer greater than 2. In some examples, selecting two or more audio devices of the audio environment may include selecting N loudspeaker-equipped audio devices of the audio environment, where N is an integer greater than 2. Selecting two or more audio devices of the audio environment, or selecting N audio devices with built-in loudspeakers of the audio environment may include selecting all audio devices with built-in loudspeakers of the audio environment.

[0123] In some examples, selecting two or more audio devices of the audio environment may be based at least in part on the person's estimated current location relative to the microphone location and / or the loudspeaker-equipped audio device location. Some such examples may include determining the nearest loudspeaker-equipped audio device closest to the person's estimated current location, or determining the nearest loudspeaker-equipped audio device closest to the microphone location closest to the person's estimated current location. In some such examples, the two or more audio devices may include the nearest loudspeaker-equipped audio device.

[0124] According to some implementations, selecting two or more audio devices may be based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold.

[0125] According to this example, block 270 includes determining one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for two or more audio devices. In this embodiment, the audio processing modifications have the effect of increasing the speech-to-echo ratio at one or more microphones. In some examples, the one or more microphones may be located within multiple audio devices of the audio environment. However, according to some implementations, the one or more microphones may be located within a single audio device of the audio environment. In some examples, the audio processing modifications may result in a reduction in loudspeaker playback levels for the loudspeakers of the two or more audio devices.

[0126] According to some examples, at least one of the audio processing changes for a first audio device may be different from the audio processing changes for a second audio device. For example, the audio processing change(s) may cause a first reduction in the loudspeaker playback level of a first loudspeaker of the first audio device and a second reduction in the loudspeaker playback level of a second loudspeaker of the second audio device. In some such examples, the reduction in loudspeaker playback level may be greater relative to an audio device that is closer to the person's estimated current location (or to a microphone location closest to the person's estimated current location).

[0127] However, the inventors contemplate many types of audio processing modifications that may be made in some examples. According to some implementations, one or more types of audio processing modifications may include modifying the rendering process to warp the rendering of the audio signal away from the person's estimated current location (or away from the microphone location closest to the person's estimated current location).

[0128] In some embodiments, the one or more audio processing modifications may include spectral modifications. For example, the spectral modifications may include reducing the level of audio data in a frequency band between 500 Hz and 3 KHz. In other examples, the spectral modifications may include reducing the level of audio data in a frequency band having a higher maximum frequency and / or a lower minimum frequency. According to some embodiments, the one or more audio processing modifications may include inserting at least one gap in at least one selected frequency band of the audio playback signal. stomach.

[0129] In some embodiments, determining one or more audio processing modifications may be based on optimizing a cost function that is based at least in part on the signal-to-echo ratio estimate. In some examples, the cost function may be based at least in part on rendering performance.

[0130] According to this example, block 275 may include applying one or more audio processing changes. In some examples, block 275 may include applying the one or more audio processing changes by one or more devices controlling audio processing within the audio environment. In other examples, block 275 may include applying the one or more audio processing changes by one or more other devices in the audio environment (e.g., via commands or control signals from an audio session manager).

[0131] Some implementations of method 250 may include selecting at least one microphone depending on one or more aspects of the context information. In some such implementations, method 250 may include selecting at least one microphone depending on an estimated current proximity of a person to one or more microphone locations. Some implementations of method 250 may include selecting at least one microphone depending on an estimate of a user zone. According to some such implementations, method 250 may include implementing a virtual assistant function at least in part depending on microphone signals received from the selected microphone(s). In some such implementations, method 250 may include providing a videoconferencing function based at least in part on microphone signals received from the selected microphone(s).

[0132] Some embodiments are configured to implement the rendering and mapping and to use software or other logic manifestations (e.g., the system implementing the logic) to modify audio processing (e.g., to reduce the volume of one, two, or more of the nearest loudspeakers). The system may include a system (including two or more devices, e.g., smart audio devices) configured to implement an audio session manager (including a system element). The logic may implement a supervisor, such as a device configured to implement an audio session manager. The supervisor may, in some examples, execute separately from the system element configured for rendering.

[0133] Figure 3A is a block diagram of a system configured to implement separate rendering control and listening or capture logic across multiple devices. As with the other disclosed figures, the number, types, and arrangement of elements shown in Figures 3A, 3B, and 3C are exemplary only. Other implementations may include more elements, fewer elements, and / or different types of elements. For example, other implementations may include four or more audio devices, different types of audio devices, etc.

[0134] The modules shown in Figures 3A, 3B, and 3C, as well as other modules shown and described in this disclosure, may be implemented via hardware, software, firmware, etc., depending on the particular example. In some embodiments, one or more of the disclosed modules (which may also be referred to in some examples as "elements") may be implemented via a control system, such as control system 210 described with reference to Figure 2A. In some such examples, one or more of the disclosed modules may be implemented via a control system, such as control system 210 described with reference to Figure 2A. The module may be implemented in accordance with software executed by one or more such control systems.

[0135] The elements of FIG. 3A include: Audio devices 302, 303, and 304 (which may in some examples be smart audio devices). According to this example, each of audio devices 302, 303, and 304 includes at least one loudspeaker and at least one microphone.

[0136] Element 300 represents a form of content, including audio data, that is played across one or more of audio devices 302, 303, and 304. Content 300 may be linear or interactive content, depending on the particular implementation.

[0137] Module 301 is configured for audio processing, including, but not limited to, rendering according to rendering logic. For example, in some embodiments, module 301 may be configured to simply replicate the audio (e.g., mono or stereo) of content 300 equally across all three audio devices 302, 303, and 304. In some other embodiments, one or more of audio devices 302, 303, and 304 may be configured to implement audio processing functionality, including, but not limited to, rendering functionality.

[0138] Element 305 represents a signal distributed to audio devices 302, 303, and 304. In some examples, signal 305 may be or include a speaker feed signal. As noted above, in some implementations, the functionality of module 301 may be implemented via one or more of audio devices 302, 303, and 304, in which case signal 305 may be limited to one or more of audio devices 302, 303, and 304. However, FIG. 3A illustrates them as a set of speaker feed signals because some embodiments (e.g., the embodiments described below with reference to FIG. 4) implement simple final interception or post-processing of signal 305.

[0139] - Element 306 represents the raw microphone signal captured by the microphones of audio devices 302, 303 and 304.

[0140] Module 307 is configured to implement microphone signal processing logic and, in some examples, microphone signal capture logic. In this example, since audio devices 302, 303, and 304 each have one or more microphones, captured raw signal 306 is processed by module 307. In some embodiments, as here, module 307 may be configured to implement echo cancellation and / or echo detection functionality.

[0141] Element 308 represents a local echo reference signal and / or a global echo reference signal provided by modules 301 to 307. According to this example, module 307 is configured to implement an echo cancellation and / or echo detection function in response to the local echo reference signal and / or the global echo reference signal 308. In some embodiments, microphone capture processing and / or processing of raw microphone signals is performed in each of audio devices 302, 303 and 304. The particular implementation of the capture and capture processing is not important to the idea of ​​calculating and understanding the impact that any changes made to the rendering have on the overall SER and the effectiveness of the capture processing and logic.

[0142] Module 309 is a system element that implements the overall mixing or combining of captured audio sounds (e.g., to make desired audio sounds perceived as emanating from a specific single location or a range of locations). In some embodiments, module 307 may also provide the mixing functionality of element 309.

[0143] Module 310 is the system element that implements the final aspect of processing detected audio sounds to make a determination about what was said or whether an activity of interest occurred in the audio environment. Module 310 may provide, for example, automatic speech recognition (ASR) functionality, background noise level and / or type detection functionality, for example, context regarding what people are doing in the audio environment, the overall noise level in the audio environment, etc. In some embodiments, some or all of the functionality of module 310 may be implemented outside the audio environment in which audio devices 302, 303, and 304 are located (e.g., in one or more devices (e.g., one or more servers) of a cloud-based service provider).

[0144] 3B is a block diagram of a system according to another disclosed embodiment. In this example, the system shown in FIG. 3B includes elements of the system of FIG. 3A and extends the system of FIG. 3A to include functionality according to some disclosed embodiments. The system of FIG. 3B includes components that implement the context, decision, and rendering action aspects that apply to an operational distributed audio system. Some examples include CONTEXT, DECISION, and Feedback to elements implementing the RENDERING ACTION aspect can be an increase in confidence when there is activity (e.g., detected speech) or a decrease in confidence in the sense of activity (e.g., activity detection). This may result in either: a low likelihood of an activity; or the ability to reset audio processing to an initial state.

[0145] The elements of FIG. 3B include: - Module 351 shows (and implements) the steps of the CONTEXT. The modules 351 and 353 are system elements that, for example, obtain an indication of where better detection of audio sounds (e.g., increasing the speech-to-echo ratio in one or more microphones) may be desired, and the likelihood or feeling that we want to hear (e.g., the likelihood that speech, such as a wake word or command, will be captured by one or more microphones). In this example, modules 351 and 353 are implemented via a control system (in this example, control system 210 of FIG. 2A ). In some embodiments, blocks 301 and 307 may also be implemented by the control system (which may be control system 210 in some examples). According to some embodiments, blocks 356, 357, and 358 may also be implemented by the control system (which may be control system 210 in some examples).

[0146] Element 352 indicates a feedback path to module 351. In this example, feedback 352 is provided by module 310. In some embodiments, feedback 352 is provided by a microprocessor that may be relevant to determining the context. The system may respond to the results of audio processing (such as audio processing for ASR) resulting from the capture of audio signals. For example, a sense of weak or early detection of a wake word or low detection of speech activity may be used to begin increasing the confidence or sense of a context requiring improved listening (e.g., increasing the speech-to-echo ratio in one or more microphones).

[0147] Module 353 is a system element in which (or by which) decisions are made regarding which audio devices' audio processing to modify and by how much to modify the audio processing. Module 353 may or may not use specific audio device information, such as the type and / or capabilities of the audio device (e.g., loudspeaker capabilities, echo suppression capabilities, etc.) and the likely orientation of the audio device, depending on the particular implementation. As will be described later in some examples, the decision-making process of module 353 may be significantly different for headphone devices compared to smart speakers or other loudspeakers.

[0148] Element 354 is the output of module 353, which is output to the individual rendering blocks via control path 355 (sometimes referred to as signal path 355). In this example, the output of module 353 is a set of control functions, shown as f_n values ​​in FIG. 3B. This set of control functions may be communicated (e.g., via wireless transmission) so that signal path 355 is limited to the audio environment. In this example, the control functions are provided to modules 356, 357, and 358.

[0149] Modules 356, 357, and 358 are system elements configured to modify audio processing, which may include, but is not limited to, output rendering (the RENDER aspect of some embodiments). In this example, modules 356, 357, and 358 are activated by a control function (the f_n value in this example) of output 354. In some embodiments, the functionality of modules 356, 357, and 358 may be implemented via block 301.

[0150] In the embodiment of FIG. 3B and other embodiments, a virtuous cycle of feedback can occur. If the output 352 of element 310 (which may, in some examples, implement automatic speech recognition (ASR)) detects speech according to some examples, even if it is a weak detection (e.g., with low confidence), the CONTEXT element 351 can determine where in the audio environment the utterance is located. The location may be estimated based on which microphone(s) captured the sound (e.g., which microphone(s) had the most non-echo energy). According to some such examples, DECISION block 353 may select one, two, or more loudspeakers in the audio environment and activate a small value for rendering modification (e.g., f_n=0.25). For an overall volume reduction (ducking) of 20 dB, this value results in a volume decrease of approximately 5 dB at the selected device(s), perceptible to average human hearing. When combined with time constants and / or event detection, and if other loudspeakers in the audio environment are playing similar content, the level reduction may be less perceptible. In one example, it may be audio device 303 (the audio device closest to person 311) that reduces the volume. In other examples, the volume of both audio devices 302 and 303 may be reduced, in some examples by different amounts (e.g., depending on the estimated proximity to person 311). In other examples, the volume of all audio devices 302, 303, and 304 may be reduced, in some examples by different amounts. As a result of reducing the level of playback by one or more of audio devices 302, 303, and 304, the speech-to-echo ratio may be reduced by one or more microphones (e.g., The volume of one or more loudspeakers near person 311 may be increased (e.g., at one or more microphones of audio device 303). Thus, if person 311 continues to speak (e.g., continues to repeat the wake word or continues to issue commands), the system may better "hear" person 311. In some such embodiments, during the next period of time (e.g., the next few seconds), in some examples in a continuous manner, the system (e.g., an audio session manager implemented at least in part via blocks 351 and 353) may quickly switch to turning off the volume of one or more loudspeakers near person 311, for example, by selecting f_2=1.

[0151] Figure 3C is a block diagram of an embodiment configured to implement an energy balancing network according to an example, which includes elements of the system of Figure 3B and extends the system of Figure 3B to include elements (e.g., element 371) that implement energy compensation (e.g., "turn up the volume on other devices a little").

[0152] In some examples, a device configured for audio session management (audio session manager) of the system of FIG. 3C (or a system like that of FIG. 3C) may evaluate banded energy at the listener (311) that is lost as a result of audio processing (e.g., reducing the level of one or more selected loudspeakers (e.g., loudspeakers of an audio device that receive a control signal with f_n>0)) applied to increase the speech-to-echo ratio at one or more microphones. The audio session manager may then apply level increases and / or other forms of energy balancing to other speakers in the audio environment to compensate for the SER-type audio processing changes.

[0153] When multiple loudspeakers in an audio environment are rendering slightly related content and audio components that are correlated or have similar spectra are being reproduced (a simple example is mono playback), less energy balancing may be necessary. For example, if an audio environment contains three loudspeakers spaced apart by a factor of 1 to 2 (1 being the closest), and the loudspeakers are reproducing the same content, reducing the volume of the nearest loudspeaker by 6 dB will only have an impact of 2 to 3 dB. Turning off the nearest loudspeaker will only have an overall impact of 3 to 4 dB on the sound at the listener's position.

[0154] In more complex situations (e.g., gap insertion or spatial steering), in some instances, the form of energy preservation and perceptual continuity may be a more multi-dimensional energy balance.

[0155] In Figure 3C, the element(s) that implement CONTEXT are: In some examples, it may be the audio level (reciprocity of proximity) of the weak wake word detection. In other words, determining the context. An example of determining the level of a wake word utterance detected via detected echo may be based on the level of any wake word utterance detected via detected echo. Such a method may or may not actually include determining a speech-to-echo ratio, depending on the particular implementation. However, in some examples, simply detecting and evaluating the level of the wake word utterance detected at each of multiple microphone locations may be sufficient to determine the context. Various levels can be provided.

[0156] A system that implements a context (e.g., in the system of Figure 3C) Some examples of methods implemented by elements include, but are not limited to:

[0157] Upon detecting a portion of the wake word, proximity to an audio device with a microphone can be inferred from the wake word confidence. The timing of the wake word utterance can also be inferred from the wake word confidence.

[0158] - In addition to echo cancellation and echo suppression applied to the raw microphone signal, some audio activity is detected. Some embodiments use a set of energy levels and classifications to determine how likely the audio activity is voice activity (voice activity detection). This process may determine a confidence or likelihood of voice activity. The voice location may be based on the best microphone possibilities for similar situations of interaction. For example, a device implementing an audio session manager may know in advance that an audio device with one microphone is closest to the user, such as a device on a table at or near the user's frequent location, rather than a wall-mounted device that is not near the user's frequent location.

[0159] An example embodiment of a system element that implements DECISION (e.g., in the system of FIG. 3C) is an element configured to determine a confidence value for voice activity and to determine which device is the nearest audio device with an integrated microphone.

[0160] In the system of FIG. 3C (and other embodiments), the amount of audio processing modification(s) to apply to increase the SER at any location may be a function of distance and confidence regarding voice activity.

[0161] How to implement RENDERING (e.g., in the system of Figure 3C) Some examples include the following: Only reduce dB and / or Speech band equalization (EQ) (e.g., as described below with reference to FIG. 4 ) and / or Time modulation of the rendering changes (described with reference to FIG. 5), and and / or Using temporal time slicing or time adjustment to generate (e.g., insert into the audio content) "gaps" or periods of sparse time-frequency lower output sufficient to capture glimpses of the audio speech of interest. Some examples are described below with reference to FIG. 9.

[0162] FIG. 4 is a graph illustrating an example of audio processing that may increase the speech-to-echo ratio at one or more microphones in an audio environment. The graph in FIG. 4 provides an example of spectral modification. In FIG. 4, the spectral modification includes reducing the level of frequencies known to correspond to speech (in these examples, frequencies within a range of approximately 200 Hz to 10 KHz (e.g., within 5% or 10% of the high and / or low frequencies of this range)). Another example may include reducing the level of frequencies within a different frequency band (e.g., between approximately 500 Hz and 3 KHz (e.g., within 5% or 10% of the high and / or low frequencies of this range)). In some embodiments, frequencies outside this range may be reproduced at a higher level to at least partially compensate for the reduction in loudness caused by the spectral modification.

[0163] The elements of Figure 4 include: 601: Curve showing flat EQ. 602: Curve showing partial attenuation in the indicated frequency range. Such partial attenuation may be relatively imperceptible, but may nevertheless have a useful impact on speech detection. 603: Curve showing significantly greater attenuation of the indicated frequency range. Spectral modification such as that shown by curve 603 can have a significant impact on speech intelligibility. In some instances, aggressive spectral modification such as that shown by curve 603 may provide the option of significantly reducing the level of all frequencies.

[0164] In some examples, the audio session manager may make audio processing changes that correspond to time-varying spectral modifications, such as the sequences shown by curves 601, 602, and 603.

[0165] According to some examples, one or more spectral modifications may be used in the context of other audio processing modifications, such as in the context of rendering modifications that result in "warping" the reproduced audio sound away from the location of an office, a bedroom, a sleeping baby, etc. The spectral modification(s) used in conjunction with such warping may, for example, reduce levels within the bass frequency range (e.g., 20-250 Hz).

[0166] FIG. 5 is a graph illustrating another type of audio processing that may increase the speech-to-echo ratio at one or more microphones in an audio environment. In this example, the vertical axis represents "f" values ​​ranging from 0 to 1, and the horizontal axis represents time (seconds). FIG. 5 illustrates a trajectory (shown by curve 701) versus time of activation of a rendering effect. In some examples, one or more of modules 356, 357, or 358 may implement the type of audio processing shown in FIG. 5. According to this example, the asymmetry in the time constants (shown by curve 701) indicates that the system adjusts to a controlled value (f_n) in a short period of time (e.g., 100 ms to 1 second), but relaxes from the value f_n (value 703) to zero over a significant period of time (e.g., 10 seconds or more). In some examples, the interval between 2 seconds and N seconds may be multiple seconds (e.g., in the range of 4 to 10 seconds).

[0167] 5 also shows a second activation curve 702, which in this example is stepped, with a maximum value equal to f_n. According to this embodiment, the rising steps correspond to abrupt changes in the level of the content itself (e.g., voice onset or syllable rate).

[0168] As mentioned above, in some embodiments, temporal time slicing or frequency adjustment may be used to create "gaps" or periods of sparse time-frequency output (e.g., by inserting gaps in the audio content) sufficient to capture glimpses of audio speech of interest (e.g., expanding or reducing the extent of "gappiness" of the audio content and its perception).

[0169] FIG. 6 illustrates another type of audio processing that may increase the speech-to-echo ratio at one or more microphones in an audio environment. FIG. 6 illustrates a forced gap according to one example. This is an example of a spectrogram of a modified audio playback signal with gaps inserted. More specifically, to generate the spectrogram of Figure 6, forced gaps G1, G2, and G3 were inserted into the frequency band of the playback signal to generate a modified audio playback signal. In the spectrogram shown in Figure 6, position along the horizontal axis indicates time, and position along the vertical axis indicates the frequency of the content of the modified audio playback signal at any given time.

[0170] The density of dots in each sub-region (each sub-region is centered at a point having vertical and horizontal coordinates) indicates the energy of the content of the modified audio playback signal at the corresponding frequency and time (higher density regions indicate content with more energy, and lower density regions indicate content with less energy). Thus, gap G1 exists at an earlier time (or period) than gaps G2 or G3, and gap G1 is inserted in a higher frequency band than the frequency bands in which gaps G2 or G3 are inserted.

[0171] Introducing a forced gap into a playback signal is distinct from simplex device operation, in which the device pauses the playback stream of content (e.g., to better hear the user and the user's environment). Introducing a forced gap into a playback signal in accordance with some disclosed embodiments may be optimized to significantly reduce (or eliminate) the likelihood of perceiving artifacts resulting from the introduced gap during playback, preferably such that the forced gap has zero or minimal impact on the user, but the microphone output signal in the playback environment indicates the forced gap (e.g., the gap may be utilized to implement a pervasive listening method). Using a forced gap introduced in accordance with some disclosed embodiments, a pervasive listening system may monitor non-playback audio (e.g., audio indicative of background activity and / or background noise in the playback environment) without using an acoustic echo canceller.

[0172] According to some examples, multiple gaps may be inserted in the time-spectral output from a single channel, which may result in an enhanced listening ability with a looser sense of "hearing through the gaps."

[0173] 7 is a graph illustrating another type of audio processing that may increase the speech-to-echo ratio at one or more microphones of an audio environment. In this embodiment, the audio processing modification includes dynamic range compression.

[0174] This example includes a transition between two extremes that limit the dynamic range. In one case, shown by curve 801, the audio session manager applies no dynamic range control, while in another case, shown by curve 802, the audio session manager applies a relatively aggressive limiter. A limiter corresponding to curve 802 may reduce the peaks of the audio output by 10 dB or more. In some examples, the compression ratio is only 3:1. In some embodiments, curve 802 (or another dynamic range compression curve) may include a knee at (or about) -20 dB (e.g., within + / - 1 dB, within + / - 2 dB, within + / - 3 dB, etc.) of the device's peak output.

[0175] Next, rendering (e.g., in the system of FIG. 3B or FIG. 3C) Another example of an embodiment of a system element that implements an SER (audio processing modification that has the effect of increasing the speech-to-echo ratio at one or more microphones) is described. In this embodiment, energy balancing is performed. As mentioned above, in one simple example, the audio session manager may evaluate the banded energy of the audio speech at a listener's position or zone that has been lost as a result of other audio processing modifications to increase the SER at one or more microphones in the audio environment. Then, The audio session manager may add a boost to other speakers to make up for the energy lost at this listener's position or zone.

[0176] Rendering content that is somewhat related and is correlated or similar If components with this spectrum are present on multiple devices (as is often the case with mono playback as a simple example), then there may not be much to do. For example, if you have three loudspeakers spaced apart by a factor of 1 to 2 (with 1 being the closest), then reducing the volume of the nearest loudspeaker by 6 dB will only have an impact of 2 to 3 dB (assuming the loudspeakers are playing the same content), and turning off the nearest loudspeaker will probably only have an overall impact of 3 to 4 dB on the sound at the listener's position.

[0177] Additional embodiment aspects will now be described.

[0178] 1. Quadratic Factors in the Definition of "NEAREST" As the following two examples show, the measure of "proximity" or "closest" need not be a simple measure of distance, but may be a scalar ranking that includes an estimated speech-to-echo ratio. When multiple audio devices in an audio environment are not identical, each loudspeaker-integrated audio device may have a different connection from its loudspeaker(s) to its own microphone(s), significantly affecting the echo level in the speech-to-echo ratio. These audio devices may also have different microphone placements that make them relatively more or less suitable for listening (e.g., for detecting sounds from a particular direction, or for detecting sounds at or from a particular location in the audio environment). Thus, in some embodiments, the calculation (DECISION) may consider proximity and reciprocity of hearing as factors other than their respective microphone placements.

[0179] 8 is a diagram of an example where the audio device to reduce the volume of may not be the audio device closest to the person speaking. In this example, audio device 802 is closer to person 100 than audio device 805. According to some examples, in a situation such as that shown in FIG. 8, the audio session manager may consider different baseline SER and audio device characteristics and reduce the volume of the device(s) with the best cost / benefit ratio of the benefit of reducing the output power to better capture the speech of person 101 versus the impact of reducing the output power on the audio presentation.

[0180] Figure 8 shows that complexity and utility can exist in a more functional measure of "nearest." An example is shown. In this example, person 101 is making a sound (utterance 102), and an audio session manager is configured to capture this sound. Two audio devices 802 and 805 are also provided, both of which have loudspeakers (806 and 804) and microphones (803 and 807). Considering that microphone 803 is very close to loudspeaker 804 of audio device 802, which is closer to person 101, lowering the volume of this device's loudspeaker may not result in suitable SER. In this example, microphone 807 of audio device 805 is configured to perform beamforming (which generally results in more favorable SER), and therefore lowering the volume of the loudspeaker of audio device 805 may have less impact than lowering the volume of the loudspeaker of audio device 802. In some such examples, the optimal decision may be to lower the volume of loudspeaker 806.

[0181] Another example is described with reference to Figure 9. Here we consider the largest possible difference in baseline SER between two devices, one a pair of headphones and the other a smart speaker.

[0182] Figure 9 illustrates a situation where a device with a very high SER is very close to a user. In Figure 9, a user 101 is wearing headphones 902 and speaking a voice 102, which is captured by both a microphone 903 of the headphones 902 and a microphone of a smart speaker device 904. In this case, the smart speaker device 904 may also generate any voice that suits the headphones (e.g., near / far rendering for immersive sound). While the headphones 902 are indeed the closest output device to the user 101, there is almost no echo path from the headphones to the nearest microphone 903, and therefore the SER of this device is very high, which has a huge impact when the volume of the device is turned down, as the headphones provide almost all of the rendering effect to the listener. In this case, lowering the volume of the smart speaker 904 may be more beneficial, even though the actual action may not be determined, as it is only partial and counter to the overall change in rendering (other listeners nearby are hearing the audio). This is because lowering the speaker volume or otherwise changing the audio processing parameters may improve the SER of the user pickup, altering the audio sound provided in the audio environment for the better. In a sense, the inherent device SER in headphones is already functional enough.

[0183] For devices with multiple speakers and distributed microphones that exceed a given size, under some conditions, a single audio device with multiple speakers and microphones can be considered as a constellation of separate devices that happen to be tightly connected. In this case, volume-reduction decisions may be applied to individual speakers. Thus, in some implementations, an audio session manager may consider this type of audio device as a collection of independent microphones and loudspeakers, whereas in other examples, an audio session manager may consider this type of audio device as a single device with a composite speaker and microphone array. It can also be appreciated that there is a duality between treating speakers in a single device as separate devices and the idea that spatial steering is one approach to rendering in a single audio device with multiple loudspeakers (which necessarily results in differential changes to the output of the loudspeakers in a single audio device).

[0184] Regarding the secondary effect of the nearest audio device(s) avoiding spatial imaging sensitivities from audio devices close to a moving listener, in many cases it may not make sense to play a particular audio object or rendering material from the nearest loudspeaker(s), even if there are loudspeakers close to the moving listener. This is simply related to the fact that the loudness of the direct audio sound path varies directly (1 / r 2 (r is the distance the sound travels). And as a loudspeaker gets closer to a given listener (r->0), the level of the sound being reproduced by this loudspeaker becomes less consistent with the overall sound mix.

[0185] In some such instances, it may be advantageous to implement (for example) the following embodiments: - CONTEXT is the audio information about the program someone is watching on TV. A common listening area (e.g., a sofa near a television) where it is assumed that being able to hear voices is always useful. - DECISION: In a typical listening area (e.g., near a sofa) For a device with a speaker placed on a coffee table, set f_n=1. - RENDERING: Turn off the device and let the energy go somewhere else. be rendered.

[0186] The impact of this audio processing change is better listening for the person sitting on the sofa. If a coffee table is at one end of the sofa, this method can eliminate the listener's proximity sensitivity to this audio device. In some instances, while this audio device is in an ideal location, for example for surround channels, the fact that there may be a 20 dB level difference across the sofa to this speaker means that if the exact location of the listener / speaker is unknown, it may be a good idea to turn down the volume or turn off the closest device.

[0187] 10 is a flow diagram outlining an example of a method that may be performed by an apparatus such as that shown in FIG. 2A. The blocks of method 1000, as well as other methods disclosed herein, do not necessarily have to be performed in the order shown. Moreover, such methods may include more or fewer blocks than those shown and / or described. In this embodiment, method 1000 includes estimating a user's position within an environment.

[0188] In this example, block 1005 includes receiving an output signal from each of a plurality of microphones in the environment. In this example, each of the plurality of microphones is disposed at a microphone location in the environment. According to this example, the output signal corresponds to a current utterance of a user. In some examples, the current utterance may be or may include an utterance of a wake word. Block 1005 may include, for example, a control system (such as control system 120 of FIG. 2A ) receiving the output signal from each of the plurality of microphones in the environment via an interface system (such as interface system 205 of FIG. 2A ).

[0189] In some examples, at least some of the microphones in the environment may provide output signals that are asynchronous with respect to output signals provided by one or more other microphones. For example, a first microphone of the plurality of microphones may sample audio data based on a first sampling clock and a second microphone of the plurality of microphones may sample audio data based on a second sampling clock. In some examples, at least one of the microphones in the environment may be included in or configured to communicate with a smart audio device.

[0190] According to this example, block 1010 includes determining a plurality of current acoustic features from the output signal of each microphone. In this example, the "current acoustic features" are acoustic features obtained from the "current utterance" of block 1005. In some embodiments, block 1010 may include receiving the plurality of current acoustic features from one or more other devices. For example, block 1010 may include receiving at least some of the plurality of current acoustic features from one or more wake word detectors implemented by the one or more other devices. Alternatively or additionally, in some embodiments, block 1010 may include determining the plurality of current acoustic features from the output signal.

[0191] Whether the acoustic features are determined by a single device or multiple devices, the acoustic features may be determined asynchronously. When the acoustic features are determined by multiple devices, the devices may coordinate the process of determining the acoustic features. Unless configured otherwise, the acoustic features may generally be determined asynchronously. If the acoustic features are determined by a single device, in some embodiments, the acoustic features may still be determined asynchronously because the single device may receive the output signal of each microphone at different times. In some examples, the acoustic features may be determined asynchronously because at least some of the microphones in the environment may provide output signals that are asynchronous with respect to the output signals provided by one or more other microphones.

[0192] In some examples, the acoustic features may include a wakeword confidence metric, a wakeword duration metric, and / or at least one received level metric. The Bell index indicates the received level of the sound detected by the microphone and may correspond to the level of the microphone's output signal.

[0193] Alternatively or additionally, the acoustic features may include one or more of the following: Average state entropy (purity) for each wake word state along the 1-best (Viterbi) ordering for the acoustic model. CTC-loss (Connectionist Temporal Classification Loss) for the acoustic model of the wake word detector. The wake word detector may be trained to provide, in addition to the wake word confidence, an estimate of the speaker's distance from the microphone and / or an RT60 estimate, which may be an acoustic feature. Instead of or in addition to the wideband receive level / power at the microphone, the acoustic feature may be the receive level at a number of log / mel / bark-spaced frequency bands, which may vary depending on the particular implementation (e.g., 2 frequency bands, 5 frequency bands, 20 frequency bands, 50 frequency bands, 1 octave frequency band, or 1 / 3 octave frequency band). A cepstral representation of the spectral information at a point in time in the past, calculated by taking the DCT (Discrete Cosine Transform) of the logarithm of the band power. Band power in a frequency band weighted for human speech. For example, the acoustic features may be based only on a specific frequency band (e.g., 400 Hz to 1.5 kHz). In this example, higher and lower frequencies may be ignored. Voice activity detector confidence, per band or per bin. The acoustic signature may be based at least in part on the long-term noise estimate to ignore microphones with insufficient signal-to-noise ratio. Kurtosis as a measure of the "peakiness" of speech. can be an indicator of smearing due to long reverb tails. Estimated wake word start time. Onset and duration are expected to be equal within a frame, or across all microphones. Outliers may be a clue to an unreliable estimate. This assumes a certain level of synchrony, not necessarily across samples, but across frames of, say, tens of milliseconds.

[0194] According to this example, block 1015 includes applying a classifier to the plurality of current acoustic features. In some such examples, applying the classifier may include applying a model trained on previously determined acoustic features obtained from a plurality of past utterances made by the user in a plurality of user zones in the environment. Various examples are described herein.

[0195] In some examples, the user zones may include a sink area, a cooking area, a refrigerator area, These user zones may include a dining area, a sofa area, a television area, a sleeping area, and / or an entryway area. According to some examples, one or more of these user zones may be predetermined user zones. In some such examples, the one or more predetermined user zones are selectable by a user during the training process.

[0196] In some embodiments, applying the classifier may include applying a Gaussian mixture model trained on past utterances. According to some such embodiments, applying the classifier may include applying a Gaussian mixture model trained on one or more of a normalized wake word confidence, a normalized average received level, or a maximum received level of the past utterances. However, in other embodiments, applying the classifier may be based on a different model, such as one of the other models disclosed herein. In some examples, the model may be trained using training data labeled with user zones. However, in some examples, applying the classifier includes applying a model trained with unlabeled training data that is not labeled with user zones.

[0197] In some examples, the past utterance may be or may include an utterance of the wake word. According to some such examples, the past utterance and the current utterance may be an utterance of the same wake word.

[0198] In this example, block 1020 includes determining an estimate of a user zone in which the user is currently located based at least in part on the output from the classifier. In some such examples, this estimate may be determined without reference to the geometric locations of the microphones. For example, this estimate may be determined without reference to the coordinates of individual microphones. In some examples, this estimate may be determined without estimating the geometric location of the user.

[0199] Some implementations of method 1000 may include selecting at least one speaker responsive to the estimated user zone. Some such implementations may include controlling the at least one selected speaker to provide sound to the estimated user zone. Alternatively or additionally, some implementations of method 1000 may include selecting at least one microphone responsive to the estimated user zone. Some such implementations may include providing a signal output by the at least one selected microphone to a smart audio device.

[0200] 11 is a block diagram of elements of an example embodiment configured to implement a zone classifier. According to this example, a system 1100 includes multiple loudspeakers 1104 distributed throughout at least a portion of an environment (e.g., an environment such as that shown in FIG. 1A or FIG. 1B). In this example, the system 1100 includes a multi-channel loudspeaker renderer 1101. According to this implementation, the output of the multi-channel loudspeaker renderer 1101 serves as both a loudspeaker drive signal (a speaker feed signal that drives the speaker 1104) and an echo reference signal. In this implementation, the echo reference signal is provided to an echo management subsystem 1103 via multiple loudspeaker reference channels 1102. Here, the echo reference signal includes at least some of the speaker feed signals output from the renderer 1101.

[0201] In this embodiment, the system 1100 includes multiple echo management subsystems 1103. According to this example, the echo management subsystems 1103 are configured to implement one or more echo suppression processes and / or one or more echo cancellation processes. In this example, each of the echo management subsystems 1103 provides a corresponding echo management output 1103A to one of the wake word detectors 1106. The echo management output 1103A has an attenuated echo compared to the input to the associated one of the echo management subsystems 1103.

[0202] According to this embodiment, the system 1100 includes N microphones 1105 (N is an integer) distributed throughout at least a portion of an environment (e.g., the environment shown in FIG. 1A or FIG. 1B). These microphones may include array microphones and / or spot microphones. For example, one or more smart audio devices disposed within the environment may include an array of microphones. In this example, the outputs of the microphones 1105 are provided as inputs to the echo management subsystems 1103. According to this embodiment, each of the echo management subsystems 1103 captures the output of an individual microphone 1105 or an individual group or subset of the microphones 1105.

[0203] In this example, the system 1100 includes multiple wake word detectors 1106. According to this example, each of the wake word detectors 1106 receives audio output from one of the echo management subsystems 1103 and outputs multiple acoustic features 1106A. The acoustic features 1106A output from each echo management subsystem 1103 may include, but are not limited to, measures of wake word confidence, wake word duration, and receive level. Although three arrows representing three acoustic features 1106A are shown as being output from each echo management subsystem 1103, in other embodiments, a greater or lesser number of acoustic features 1106A may be output. Furthermore, although the three arrows point along substantially vertical lines to the classifier 1107, this does not indicate that the classifier 1107 necessarily receives acoustic features 1106A from all of the wake word detectors 1106 simultaneously. As noted elsewhere herein, the acoustic features 1106A may, in some instances, be determined asynchronously and / or provided to the classifier asynchronously.

[0204] According to this embodiment, system 1100 includes a zone classifier 1107 (sometimes referred to as classifier 1107). In this example, the classifier receives multiple features 1106A from multiple wake word detectors 1106 for multiple microphones 1105 (e.g., all microphones 1105) in the environment. According to this example, output 1108 of zone classifier 1107 corresponds to an estimate of a user zone in which the user is currently located. According to some such examples, output 1108 may correspond to one or more posterior probabilities. The estimate of a user zone in which the user is currently located may be or may correspond to a maximum posterior probability based on Bayesian statistics.

[0205] Next, an example implementation of a classifier is described, which in some instances may correspond to the zone classifier 1107 of FIG. iLet (n) be the i-th (i={1…N}) microphone signal at discrete time n (i.e., microphone signal x i (n) are the outputs of the N microphones 1105. In the echo management subsystem 1103, N signals x i (n), we obtain a "clean" microphone signal e at each discrete time n. i (n) is generated (i={1…N}) In this example, the clean signal e shown at 1103A in FIG. i (n) are fed to wake word detectors 1106, where each wake word detector 1106 receives a vector of features w, denoted as 1106A in FIG. i (j), where j={1...J} is the index corresponding to the jth wake word utterance. In this example, the classifier 1107 generates a total set of features Input is JPEG0007746513000024.jpg738.

[0206] According to some embodiments, a set of zone labels C k (k={1...K}) may correspond to a number (K) of different user zones within the environment. For example, the user zones may include a sofa zone, a kitchen zone, a reading chair zone, etc. Some examples may define multiple zones within a kitchen or other room. For example, a kitchen area may include a sink zone, a cooking zone, a refrigerator zone, and a dining zone. Similarly, a living room area may include a sofa zone, a TV zone, a reading chair zone, one or more entryway zones, etc. Zone labels for these zones may be selectable by the user, for example, during a training period.

[0207] In some embodiments, the classifier 1107 determines the posterior probability of the set of features W(j), for example, by using a Bayesian classifier. Estimate JPEG0007746513000025.jpg717. Probability JPEG0007746513000026.jpg717 is the jth utterance and kth zone, k For each of the above and each of the utterances, the user is in zone C. k , which shows the probability that each of the following is present.

[0208] According to some examples, training data may be collected (e.g., for each user zone) by prompting a user to select or define a zone (e.g., a sofa zone). The training process may include prompting the user to make training utterances (e.g., uttering a wake word) in the vicinity of the selected or defined zone. In the sofa zone example, the training process may include prompting the user to make training utterances at the center and at both ends of the sofa. The training process may include prompting the user to repeat the training utterance multiple times at each location within the user zone. The user may then be prompted to move to another user zone and continue making training utterances until all designated user zones are covered.

[0209] 12 is a flow diagram outlining an example of a method that may be performed by an apparatus such as apparatus 200 of FIG. 2A. The blocks of method 1200, as well as other methods disclosed herein, do not necessarily have to be performed in the order shown. Moreover, such methods may include more or fewer blocks than shown and / or described. In this embodiment, method 1200 includes training a classifier to estimate a user's location within an environment.

[0210] In this example, block 1205 includes prompting a user to perform at least one training utterance at each of a plurality of locations within a first user zone of the environment. The training utterance may, in some examples, be one or more instances of a wake word utterance. According to some implementations, the first user zone may be any user zone selected and / or defined by the user. In some examples, the control system may generate a corresponding zone label (e.g., the previously described zone label C k , a corresponding instance of one of the first user zones may be generated, and a zone label may be associated with the training data obtained for the first user zone.

[0211] An automated facilitation system may be used to collect these training data. The interface system 205 of the device 200 may include one or more devices that implement a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. For example, the device 200 may display the following prompting messages to the user on the screen of the display system or through one or more speakers during the training process: "Move to the couch" "Shake your head from side to side and say the wake word 10 times" "Move to a position halfway between the sofa and the reading chair and say the wake word 10 times." "Stand in the kitchen as if you were cooking and say the wake word 10 times."

[0212] In this example, block 1210 includes receiving a first output signal from each of a plurality of microphones in the environment. In some examples, block 1210 may include receiving a first output signal from all of the active microphones in the environment, whereas in other examples, block 1210 may include receiving a first output signal from a subset including all of the active microphones in the environment. In some examples, at least some of the microphones in the environment may provide output signals that are asynchronous with respect to output signals provided by one or more other microphones. For example, a first microphone of the plurality of microphones may sample audio data based on a first sampling clock, and a second microphone of the plurality of microphones may sample audio data based on a second sampling clock.

[0213] In this example, each of the plurality of microphones is disposed at a microphone location in the environment. In this example, the first output signal corresponds to a detected instance of a training utterance received from a first user zone. Because block 1205 includes prompting a user to perform at least one training utterance at each of a plurality of locations within the first user zone of the environment, in this example, the term "first output signal" refers to the set of all output signals corresponding to the training utterances for the first user zone. In other examples, the term "first output signal" may refer to a subset of all output signals corresponding to the training utterances for the first user zone.

[0214] According to this example, block 1215 includes determining one or more first acoustic features from each of the first output signals. In some examples, the first acoustic features may include a wake word confidence index and / or a received level index. For example, the first acoustic features may include a normalized wake word confidence index, a normalized average received level index, and / or a maximum received level index.

[0215] As described above, because block 1205 includes prompting the user to make at least one training utterance at each of a plurality of locations within a first user zone of the environment, in this example, the term "first output signal" refers to the set of all output signals corresponding to the training utterances for the first user zone. Thus, in this example, the term "first acoustic feature" refers to the set of acoustic features obtained from the set of all output signals corresponding to the training utterances for the first user zone. Thus, in this example, the set of first acoustic features is at least as large as the set of first output signals. For example, if two acoustic features are determined from each of the output signals, the set of first acoustic features will be twice as large as the set of first output signals.

[0216] In this example, block 1220 trains a classifier model to classify the first user zone. and the first acoustic feature. The classifier model may be, for example, any of the classifier models disclosed herein. According to this embodiment, the classifier model is trained without reference to the geometric positions of the multiple microphones. In other words, in this example, data regarding the geometric positions of the multiple microphones (e.g., microphone coordinate data) is not provided to the classifier model during the training process.

[0217] FIG. 13 is a flow diagram outlining another example method that may be performed by an apparatus such as apparatus 200 of FIG. 2A. The blocks of method 1300, as well as other methods disclosed herein, do not necessarily have to be performed in the order shown. For example, in some embodiments, at least a portion of the acoustic feature determination process of block 1325 may be performed prior to block 1315 or block 1320. Furthermore, such methods may include more or fewer blocks than those shown and / or described. In this embodiment, method 1300 includes training a classifier to estimate a user's location within an environment. Method 1300 provides an example of deploying method 1200 to multiple user zones in an environment.

[0218] In this example, block 1305 includes prompting the user to make at least one training utterance at a location within a user zone of the environment. In some examples, block 1305 may be performed in the manner described above with reference to block 1205 of FIG. 12, except that block 1305 relates to a single location within the user zone. The training utterance may, in some examples, be one or more instances of a wake word utterance. According to some implementations, the user zone may be any user zone selected and / or defined by the user. In some examples, the control system may generate a corresponding zone label (e.g., the previously described zone label C k , corresponding instances of one of the zones) and associate the zone labels with the training data obtained for the user zones.

[0219] According to this example, block 1310 is performed substantially as described above with reference to block 1210 of FIG. 12. However, in this example, the process of block 1310 is generalized to any user zone, not necessarily the first user zone for which training data was acquired. Thus, the output signals received from block 1310 may be expressed as "output signals from each of a plurality of microphones in the environment, each of the plurality of microphones being located at a microphone location in the environment, the output signals corresponding to detected instances of training utterances received from the user zone." In this example, the term "output signals" refers to the set of all output signals corresponding to one or more training utterances at a location in the user zone. In other examples, the term "output signals" refers to a subset of all output signals corresponding to one or more training utterances at a location in the user zone.

[0220] According to this example, block 1315 includes determining whether sufficient training data has been acquired for the current user zone. In some such examples, block 1315 may include determining whether output signals corresponding to a threshold number of training utterances have been acquired for the current user zone. Alternatively or additionally, block 1315 may include determining whether output signals corresponding to training utterances at a threshold number of locations within the current user zone have been acquired. If determined not, in this example, method 1300 returns to block 1305 and prompts the user to make at least one additional utterance at a location within the same user zone.

[0221] However, if at block 1315 it is determined that sufficient training data has been obtained for the current user zone, then in this example the process continues to block 1320. According to the method described above, block 1320 determines whether to obtain training data for additional user zones. In some examples, block 1320 may include determining whether training data has been obtained for each user zone previously identified by the user. In other examples, block 1320 may include determining whether training data has been obtained for a minimum number of user zones. The minimum number may be selected by the user. In other examples, the minimum number may be a recommended minimum number per environment, a recommended minimum number per room within the environment, or the like.

[0222] If, at block 1320, it is determined that training data should be acquired for additional user zones, in this example, the process continues to block 1322. Block 1322 includes prompting the user to move to another user zone of the environment. In some examples, the next user zone may be selectable by the user. According to this example, the process continues to block 1305 after the prompting step of block 1322. In some such examples, the user may be prompted to confirm that the user has reached the new user zone after the prompting step of block 1322. According to some such examples, the user may be asked to confirm that the user has reached the new user zone before the prompting step of block 1305.

[0223] If, at block 1320, it is determined that training data should not be acquired for additional user zones, in this example, the process continues to block 1325. In this example, method 1300 includes obtaining training data for K user zones. In this embodiment, block 1325 includes determining 1st through Gth acoustic features from 1st through Hth output signals corresponding to the 1st through Kth user zones for which training data was obtained. In this example, the term "first output signal" refers to the set of all output signals corresponding to training utterances for the first user zone. Also, the term "Hth output signal" refers to the set of all output signals corresponding to training utterances for the Kth user zone. Similarly, the term "first output signal" refers to the set of acoustic features determined from the first output signal, and the term "Gth acoustic feature" refers to the set of acoustic features determined from the Hth output signal.

[0224] According to these examples, block 1330 includes training a classifier model to form correlations between the first through Kth user zones and the first through Kth acoustic features, respectively. The classifier model may be, for example, any of the classifier models disclosed herein.

[0225] In the above example, the user zone is (e.g., the previously explained zone label C k However, the model may be trained according to labeled or unlabeled user zones, depending on the particular implementation. If labeled, each training utterance may be paired with a label corresponding to the user zone, e.g.,

number

[0226] Training a classifier model may involve determining the best fit to the labeled training data. Without loss of generality, the classification approach appropriate for the classifier model may be determined. The approach may include the following: Bayesian classifiers, e.g., where the per-class distribution is a multivariate positive Bayesian classifiers described by normal distributions, full-covariance Gaussian mixture models, or diagonal-covariance Gaussian mixture models. Vector quantization, · Nearest neighbor (k-means), A neural network with a SoftMax output layer, with one output for each class. Support Vector Machines (SVM), and / or Boosting techniques, such as gradient boosting machines (GBM).

[0227] In one example of implementing the unlabeled case, the data may be automatically divided into K clusters (K may be unknown). The unlabeled automatic division may be performed, for example, by using classical clustering techniques (e.g., k-means algorithm or Gaussian mixture modeling).

[0228] To improve robustness, regularization may be applied to the training of the classifier model, and the model parameters may be updated over time as new utterances are made.

[0229] Further aspects of the embodiment will now be described.

[0230] An example set of acoustic features (e.g., acoustic features 1106A in FIG. 11) may include a likelihood of wake word confidence, an average received level for the estimated length of the most confident wake word, and a maximum received level for the estimated length of the most confident wake word. The features may be normalized to their maximum value for each wake word utterance. The training data may be labeled, and a full-covariance mixture Gaussian model (GMM) may be trained to maximize the expectation of the training labels. The estimated zone may be the class that maximizes the posterior probability.

[0231] The above description of some embodiments discussed learning an acoustic zone model from a set of training data collected during an accelerated collection process. In that model, training time (or configuration mode) and runtime (or regular mode) can be thought of as two different modes in which a microphone system can be deployed. A development to this scheme is online learning, in which some or all of the acoustic zone model is learned or adapted online (e.g., at runtime or in regular mode). In other words, even after applying a classifier in a "runtime" process to generate an estimate of the user zone in which the user is currently located (e.g., according to method 1000 of FIG. 10), in some embodiments, the process of training the classifier may continue.

[0232] 14 is a flow diagram outlining another example method that may be performed by a device such as device 200 of FIG. 2A. The blocks of method 1400, as well as other methods disclosed herein, do not necessarily have to be performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described. In this embodiment, method 1400 includes continuous training of a classifier during the "runtime" process of estimating a user's position within an environment. Method 1400 is an example of what is referred to herein as an "online learning mode."

[0233] In this example, block 1405 of method 1400 corresponds to blocks 1005-1020 of method 1000, where block 1405 includes providing an estimate of a user zone in which the user is currently located based at least in part on output from the classifier. According to this embodiment, block 1410 includes obtaining implicit or explicit feedback regarding the estimate of block 1405. At block 1415, the classifier is updated according to the feedback received at block 1405. Block 1415 may include, for example, one or more reinforcement learning methods. As suggested by the dotted arrow extending from block 1415 to block 1405, in some embodiments, method 1400 may include returning to block 1405. For example, method 1400 may include providing a future estimate of a user zone in which the user will be located at a future time based on applying the updated model.

[0234] Explicit techniques for obtaining feedback may include: Use a voice user interface (UI) to ask the user if their prediction was correct or incorrect. For example, you might provide the user with a voice that says: "I think you're sitting on the couch. Please answer 'true' or 'false'." Inform the user that they can always use the voice UI to correct incorrect predictions. (For example, you might provide the user with a voice that says: "Talk to me and I can predict where you are. If my prediction is wrong, respond with something like, 'Amanda, I'm not sitting on the couch. I'm sitting in the reading chair.'") Always use a voice UI to let users know that they can be rewarded for correct predictions. (For example, you might provide a voice that says to the user: "Talk to me and I can predict where you are. If my prediction is correct, you can respond with something like, 'Amanda, that's right. I'm sitting on the couch.' This will help improve my predictions.") Contains physical buttons or other UI elements that the user can interact with to provide feedback (e.g., thumbs-up and / or thumbs-down buttons on the physical device or within a smartphone app).

[0235] The purpose of predicting the user zone in which the user will be located may be to inform a microphone selection scheme or an adaptive beamforming scheme that attempts to more effectively pick up sound from the user's acoustic zone, for example, to better recognize commands following a wake word. In such a scenario, implicit techniques for obtaining feedback on the quality of the zone prediction may include: Penalize predictions that result in misrecognition of the command following the wake word. Proxies that may indicate misrecognition may include the user interrupting the voice assistant's response to the command by issuing a cancel command-like command, such as "Amanda, stop!" Penalize predictions that result in low confidence that the speech recognizer correctly recognized the command. Many automatic speech recognition systems have the ability to return a confidence level with their results, which can be used for this purpose; penalizing predictions that result in the second pass wake word detector failing to retroactively detect the wake word with high confidence; and / or Enhance predictions that result in high-confidence recognition of wake words and / or correct recognition of user commands.

[0236] Described below is an example where a second-pass wake word detector fails to retroactively detect the wake word with high confidence after obtaining output signals corresponding to the current utterance from microphones in the environment and (e.g., using a microphone to communicate with the microphone). Assume that after determining acoustic features based on the output signal (via multiple first-pass wake word detectors configured in a manner similar to the above), the acoustic features are provided to a classifier. In other words, the acoustic features are assumed to correspond to the detected wake word utterance. Furthermore, it is assumed that the person making the current utterance is most likely located in Zone 3 (which in this example corresponds to the reading chair). , Suppose the classifier determines, for example, that there may be a particular microphone or combination of microphones that are known to be optimal for listening to the voices of people in Zone 3 to be sent to a cloud-based virtual assistant for voice command recognition.

[0237] Further, assume that after determining which microphone(s) to use for voice recognition, but before the person's speech is actually sent to the virtual assistant service, a second-path wake word detector operates on the microphone signals corresponding to the speech detected by the microphone(s) selected for Zone 3 that you intend to send for command recognition. If the second-path wake word detector disagrees with the first-path wake word detectors regarding the actual utterance of the wake word, it is likely because the classifier predicted the zone incorrectly. Therefore, the classifier must be penalized.

[0238] Techniques for post-updating (post-updating) the zone mapping model after one or more wake words have been spoken may include the following. Maximum a posteriori (MAP) fitting of Gaussian mixture models (GMM) or nearest neighbor models, and / or For example, reinforcement learning of neural networks, e.g., Reinforcement learning is achieved by associating "one-hot" (for accurate predictions) or "one-cold" (for inaccurate predictions) ground truth labels with SoftMax outputs and applying online backpropagation to determine the weights of the new network.

[0239] Some examples of MAP adaptation in this context may include adjusting the mean within the GMM each time a wake word is spoken, so that the mean more closely matches the acoustic features observed when a subsequent wake word is spoken. Alternatively or additionally, such examples may include adjusting the variance / covariance or mixture weight information within the GMM each time a wake word is spoken.

[0240] For example, a MAP adaptation scheme may be as follows: μ i,new =μ i,old *α+x*(1-α)

[0241] In the above formula, μ i,old denotes the mean of the i-th Gaussian in the mixture, α denotes a parameter that controls how aggressively MAP adaptation should occur (α can be in the range [0.9, 0.999]), and x denotes the feature vector of the new wake word utterance. The index "i" corresponds to the mixture element that returns the highest prior probability of containing the speaker's position at the wake word time.

[0242] Alternatively, each of the mixing elements may be adjusted according to a priori probability of including the wake word, for example, as follows: M i,new =μ i,old *β i *x(1-β i )

[0243] In the above formula, β i=α*(1-P(i)), where P(i) denotes the prior probability that observation x is attributed to mixture element i.

[0244] In one example of reinforcement learning, three user zones may be provided. Suppose for a particular wake word, the model predicts that the probabilities for the three user zones are [0.2, 0.1, 0.7]. If a second source of information (e.g., a second-pass wake word detector) confirms that the third zone was correct, then the correct label will be [0, 0, 1] ( Posterior updates to the zone mapping model may involve backpropagating the error through the neural network, which effectively means that the neural network will more strongly predict zone 3 if the same input is presented again. Conversely, if a second source of information indicates that zone 3 was an incorrect prediction, in one example, the correct label may be [0.5, 0.5, 0.0]. Backpropagating the error through the neural network reduces the likelihood that the model will predict zone 3 if the same input is presented in the future.

[0245] Flexible rendering allows spatial audio to be rendered on any number of arbitrarily positioned speakers. Given the widespread deployment of audio devices, including but not limited to smart audio devices (e.g., smart speakers) in the home, there is a need to provide flexible rendering techniques that enable consumer products to perform flexible rendering of audio and play back the audio so rendered.

[0246] Several techniques have been developed to implement flexible rendering. They treat the rendering problem as one of cost function minimization, where the cost function consists of two terms: the first term models the desired spatial impression the renderer is trying to achieve, and the second term assigns a cost to activating speakers. Currently, this second term focuses on generating a sparse solution, where only speakers that are close to the desired spatial location of the audio sound being rendered are activated.

[0247] Spatial audio playback in consumer environments is typically associated with a specified number of loudspeakers located in specified positions, e.g., 5.1 surround and 7.1 surround. In these cases, content is authored specifically for the associated loudspeakers and encoded as discrete channels, one for each loudspeaker (e.g., Dolby Digital or Dolby Digital Plus). More recently, immersive, object-based spatial audio formats have been introduced (Dolby Atmos) that decouple the content from specific loudspeaker positions. Instead, content may be described as a collection of individual audio objects, each with possibly time-varying metadata describing the desired perceived location of that audio object in three-dimensional space. During playback, the content is converted into loudspeaker feed signals by a renderer that matches the number and positions of the loudspeakers in the playback system. However, many such renderers still constrain the position of a set of loudspeakers to one of a set of prescribed layouts (e.g., 3.1.2, 5.1.2, 7.1.4, 9.1.6, etc. in Dolby Atmos).

[0248] To move beyond such constrained rendering, methods have been developed that allow object-based audio sounds to be flexibly rendered on any number of loudspeakers placed in any positions. These methods require the renderer to know the number and physical locations of the loudspeakers in the listening space. To make such systems practical for the average consumer, an automated method for identifying the loudspeaker locations is desirable. One such method relies on the use of multiple microphones, possibly co-located with the loudspeakers. By playing audio signals through the loudspeakers and recording with the microphones, the distance between each loudspeaker and microphone is estimated. From these distances, the locations of both the loudspeakers and the microphones are derived.

[0249] As object-based spatial audio is introduced into consumer spaces, Amazon There is a rapid adoption of so-called "smart speakers," such as the Echo product line. The immense popularity of these devices can be attributed to the simplicity and convenience offered by wireless connectivity and integrated voice interfaces (e.g., Amazon's Alexa). However, the sonic capabilities of these devices are generally limited, especially for spatial audio. In most cases, these devices are constrained to mono or stereo playback. However, combining the flexible rendering and automatic localization technologies described above with multiple orchestrated smart speakers can result in a system with highly sophisticated spatialization capabilities, while keeping consumer setup extremely simple. Thanks to wireless connectivity, consumers can place as many speakers as they want wherever convenient, without the need for speaker wiring, and the built-in microphones can automatically locate the speakers for the associated flexible renderer.

[0250] Conventional flexible rendering algorithms are designed to achieve a particular desired perceived spatial impression as closely as possible. In a system of orchestrated smart speakers, maintaining this spatial impression may sometimes not be the most important or desired objective. For example, if someone simultaneously attempts to speak to an integrated voice assistant, it may be desirable to temporarily alter the spatial rendering in a manner that reduces the relative playback level at speakers near a given microphone so as to increase the signal-to-noise ratio and / or signal-to-echo ratio (SER) of the microphone signal containing the detected speech. Some embodiments described herein may be implemented as modifications of existing flexible rendering methods to enable dynamic modifications to such spatial rendering, e.g., to achieve one or more additional objectives.

[0251] Existing flexible rendering techniques include Center of Mass Amplitude Panning (CMAP) and Flexible Virtualization (FV). In general terms, both of these techniques involve the creation of a sound image from two or more loudspeakers. A set of one or more audio signals (each audio signal having an associated desired perceived spatial location) is rendered for playback on a set of loudspeakers consisting of a plurality of speakers, where the relative activation of the set of loudspeakers is a function of a model of the perceived spatial location of the audio signals played on the loudspeakers and the proximity of the desired perceived spatial location of the audio signals to the positions of the loudspeakers. The model ensures that listeners near the desired spatial location hear the audio signals, and the proximity term controls which speakers are used to achieve this spatial impression. In particular, the proximity term prefers to activate speakers near the desired perceived spatial location of the audio signals. For both CMAP and FV, this functional relationship is suitably derived from a cost function written as the sum of two terms, one representing spatial aspect and one representing proximity, as follows:

number

number

[0252] In the given definition of the cost function, Although the relative levels between the components of JPEG0007746513000032.jpg77 are appropriate, it is difficult to control the absolute level of the optimum activation potential obtained as a result of the above minimization. Normalization of the vectors may be performed to control the absolute level of the activation potential. For example, it may be desirable to normalize the vectors to have unit length. This is done according to commonly used constant power panning rules.

number

[0253] The exact behavior of the flexible rendering algorithm depends on the two terms in the cost function C spatial and C proximity In the case of CMAP, C spatial defines the perceived spatial position of an audio signal reproduced from a pair of loudspeakers in relation to the associated activation gains g i It is derived from a model that places the loudspeaker at the center of gravity of the weighted positions (elements of vector g).

number

number

number

number

number

number

[0254] For this purpose, the second term of the cost function, C proximity may be defined as the distance-weighted sum of the squared absolute values ​​of the speaker activation potentials, which can be compactly expressed in matrix form as:

number

number

[0255] The overall cost function is obtained by combining the two terms of the cost function defined in equation (8) and equation (9a).

number

[0256] Setting the derivative of this cost function with respect to g to zero and solving for g gives the optimal solution for the speaker activation potential.

number

[0257] In general, the optimal solution in equation (11) may result in negative speaker activation potentials. For the CMAP configuration of the flexible renderer, such negative activation potentials may be undesirable, and therefore equation (11) may be minimized by restricting all activation potentials to be positive.

[0258] Figures 15 and 16 show an example set of speaker activation potentials and object rendering positions. In these examples, the speaker activation potentials and object rendering positions correspond to speaker positions of 4 degrees, 64 degrees, 165 degrees, -87 degrees, and -4 degrees. Figure 15 shows speaker activation potentials 1505a, 1510a, 1515a, 1520a, and 1525a, which contain the optimal solution to equation (11) for the specific speaker positions. Figure 16 plots the individual speaker positions as dots 1605, 1610, 1615, 1620, and 1625, which correspond to speaker activation potentials 1505a, 1510a, 1515a, 1520a, and 1525a, respectively. FIG. 16 also shows the ideal object positions (i.e., the positions where audio objects would be rendered) for a number of possible object angles (dots 1630a), as well as the corresponding actual rendering positions for those objects (dots 1635a connected to the ideal object positions by dotted lines 1640a).

[0259] Certain embodiments include a method for rendering audio for playback by at least one (e.g., all or some) of a plurality of coordinated (orchestrated) smart audio devices. For example, a set of smart audio devices in a user's home (included in a system) may be orchestrated to address a variety of simultaneous use cases. Such cases may include playback by all or some of the smart audio devices (i.e., all or some). This includes flexible rendering of audio sounds (according to an embodiment) for playback by the speaker(s) of a smart audio device. Many interactions with the system are envisioned that require dynamic modifications to the rendering. Such modifications may, but are not necessarily, focused on spatial fidelity.

[0260] Some embodiments are methods of rendering audio sounds for playback by at least one (e.g., all or some) of a set of smart audio devices (or for playback by at least one (e.g., all or some) of another set of speakers). The rendering may include minimizing a cost function, where the cost function includes at least one dynamic speaker activation potential term. Examples of such dynamic speaker activation potential terms include, but are not limited to, the following: · The proximity of the loudspeaker to one or more listeners; The proximity of the speaker to an attractive or repelling force; · The audibility of the loudspeaker for several positions (e.g., the listener's position, or the baby's room); · Speaker capabilities (e.g. frequency response and distortion); · Synchronization of a speaker with other speakers; Wake word performance; and Echo cancellation performance.

[0261] The dynamic speaker activation potential term(s) may enable at least one of a variety of behaviors, including warping the spatial presentation of audio away from a particular smart audio device so that its microphones can better hear the speaker, or so that secondary audio streams can better be heard from the smart audio device's speaker(s).

[0262] Some embodiments implement rendering for playback over the speaker(s) of multiple coordinated (orchestrated) smart audio devices, while other embodiments implement rendering for playback over the speaker(s) of a separate set of speakers.

[0263] Combining the flexible rendering method (as implemented in accordance with some embodiments) with a set of wireless smart speakers (or other smart audio devices) can result in an extremely capable and easy-to-use spatial audio rendering system. When considering interactions with such a system, it becomes apparent that it would be desirable to dynamically modify the spatial rendering to optimize for other objectives that may arise during use of the system. To this end, certain embodiments augment an existing flexible rendering algorithm (in which speaker activation potentials are functions of the spatial and proximity terms already disclosed) with one or more additional, dynamically configurable functions based on one or more characteristics of the audio signal being rendered, the set of speakers, and / or other external inputs. According to some embodiments, the existing flexible rendering cost function given in Equation (1) is augmented with these one or more additional dependencies according to the following equation:

number

[0264] In equation (12), the term JPEG0007746513000052.jpg1031 indicates an additional cost term, where: JPEG0007746513000053.jpg75 describes a set of one or more characteristics of an audio signal being rendered (e.g., of an object-based audio program), JPEG0007746513000054.jpg76 indicates a set of one or more characteristics of the speaker through which the audio is being rendered, JPEG0007746513000055.jpg75 indicates one or more additional external inputs. JPEG0007746513000056.jpg1031 is a set Returns a cost as a function of activation potential g for a combination of one or more characteristics of an audio signal, a speaker, and / or an external input, collectively represented by JPEG0007746513000057.jpg721. JPEG0007746513000058.jpg721 is JPEG0007746513000059.jpg75, JPEG0007746513000060.jpg76, or It should be understood that the image contains at least one element from JPEG0007746513000061.jpg75.

[0265] Below, An example, but not limited to, is JPEG0007746513000062.jpg75. · the desired perceived spatial location of the audio signal; the level of the audio signal (possibly time-varying); and / or The spectrum of the audio signal (possibly time-varying). Below, Take for example (but not limited to) JPEG0007746513000063.jpg76. · The location of the loudspeakers within the listening space; · loudspeaker frequency response; · Upper and lower limits for loudspeaker playback levels; Parameters of dynamics processing algorithms in the loudspeakers (limiter gains, etc.); · Measurements or estimates of the acoustic transmission from each loudspeaker to the other loudspeakers; A measure of the echo cancellation performance of the loudspeaker; and / or Relative synchronization between speakers. Below, An example, but not limited to, is JPEG0007746513000064.jpg75. · The location of one or more listeners or speakers within the playback space; · Measurements or estimates of the sound transmission from each loudspeaker to the listening position; · Measurements or estimates of the acoustic transmission from a talker to a pair of loudspeakers; the location of any other landmarks within the playback space; and / or Measurements or estimates of the acoustic transmission from each loudspeaker to some other landmark within the reproduction space.

[0266] With the new cost function defined in equation (12), an optimal set of actuation potentials can be found through minimization and possible post-regularization for g, as previously specified in equations (2a) and (2b).

[0267] Figure 17 is a flow chart outlining one example of a method that may be performed by an apparatus or system such as that shown in Figure 2A. The blocks of method 1700, as well as other methods disclosed herein, do not necessarily have to be performed in the order shown. Moreover, such methods may include more or fewer blocks than those shown and / or described. The blocks of method 1700 may be performed by one or more devices, which may be (or may include) a control system, such as control system 210 shown in Figure 2A.

[0268] In this embodiment, block 1705 includes receiving audio data by the control system via the interface system. In this example, the audio data includes one or more audio signals and associated spatial data. According to this embodiment, the spatial data indicates a perceived spatial location of a target corresponding to the audio signal. In some examples, the perceived spatial location of the target may be explicit, as indicated by position metadata, such as Dolby Atmos position metadata. In other examples, the perceived spatial location of the target may be implicit. For example, the perceived spatial location of the target may be a hypothetical location associated with a channel based on Dolby 5.1, Dolby 7.1, or other multi-channel audio formats. In some examples, block 1705 includes a rendering module of the control system that receives the audio data via the interface system.

[0269] According to this example, block 1710 includes rendering, by the control system, the audio data for playback through a set of loudspeakers in the environment to generate a rendered audio signal. Rendering each of the above audio signals includes determining the relative activation of a set of loudspeakers in the environment by optimizing a cost function. According to this example, the cost is a function of a model of the perceived spatial location of the audio signal when played on the set of loudspeakers in the environment. In this example, the cost is also a function of a measure of the proximity of the perceived spatial location of the target of the audio signal to the location of each loudspeaker in the set of loudspeakers. In this embodiment, the cost is also a function of one or more additional dynamically configurable functions. In this example, the dynamically configurable functions are based on one or more of the following: the proximity of the loudspeaker to one or more listeners; the proximity of the loudspeaker to an attractive position, where attractive is a factor that favors a relatively higher loudspeaker activation potential closer to an attractive position; or the proximity of the loudspeaker to a repulsive position, where repulsive is a factor that favors a relatively lower loudspeaker activation potential closer to a repulsive position. the performance of each loudspeaker relative to other loudspeakers in the environment; the synchronization of a loudspeaker with other loudspeakers; wake word performance; or echo cancellation performance.

[0270] In this example, block 1715 includes providing the rendered audio signals to at least some of the loudspeakers of the set of loudspeakers in the environment via an interface system.

[0271] According to some examples, the model of perceived spatial position may generate binaural responses corresponding to audio object positions at the left and right ears of a listener. Alternatively or additionally, the model of perceived spatial position may place the perceived spatial position of an audio signal played from a set of loudspeakers at the centroid of the set of loudspeaker positions, weighted by the loudspeakers' associated activation gains.

[0272] In some examples, the one or more additional dynamically configurable functions may be based at least in part on the level of the one or more audio signals. In some examples, the one or more additional dynamically configurable functions may be based at least in part on the spectrum of the one or more audio signals.

[0273] Some examples of the method 1700 include receiving loudspeaker layout information. In some examples, one or more additional dynamically configurable functions may be based at least in part on the position of each of the loudspeakers in the environment.

[0274] Some examples of method 1700 include receiving loudspeaker specification information. In some examples, one or more additional dynamically configurable functions may be based at least in part on the capabilities of each loudspeaker, where the capabilities of each loudspeaker may include one or more of frequency response, upper and lower playback level limits, or one or more loudspeaker dynamics processing algorithms.

[0275] According to some examples, the one or more additional dynamically configurable functions may be based at least in part on measurements or estimates of acoustic transmission from each loudspeaker to the other loudspeakers. Alternatively or additionally, the one or more additional dynamically configurable functions may be based at least in part on positions of listeners or speakers among one or more people in the environment. Alternatively or additionally, the one or more additional dynamically configurable functions may be based at least in part on measurements or estimates of acoustic transmission from each loudspeaker to the listener or speaker's positions. The estimates of acoustic transmission may be based at least in part on walls, furniture, or other objects that may be present between each loudspeaker and the listener or speaker's positions, for example.

[0276] Alternatively or additionally, one or more additional dynamically configurable functions may be added to one or more of the In some such embodiments, one or more additional dynamically configurable functions may be based at least in part on measurements or estimates of acoustic transmission from each loudspeaker to the object or landmark positions.

[0277] By implementing flexible rendering with one or more well-defined additional cost terms, numerous new and useful behaviors may be achieved. All of the example behaviors listed below are cast around penalizing certain loudspeakers under certain conditions that are considered undesirable. The end result is that these loudspeakers are activated relatively less in the spatial rendering of a set of audio signals. In many of these cases, one might consider simply lowering the volume of the undesired loudspeakers, without making any modifications to the spatial rendering. However, such a strategy may significantly disrupt the overall balance of the audio content. For example, certain components of the mixed audio may become completely inaudible. In contrast, the disclosed embodiment integrates these penalizations into the core optimization of the rendering, allowing the rendering to adapt and achieve the best possible spatial rendering using the remaining, less penalized speakers. This is a much more elegant, adaptive, and effective solution.

[0278] Example use cases include, but are not limited to: · Providing a more balanced spatial presentation around the listening area. o It has been found that spatial audio is best presented across loudspeakers that are roughly the same distance from the intended listening area. The costs may be configured such that loudspeakers that are too close or too far away compared to the average distance from the loudspeakers to the listening area are penalized and their activation reduced.

[0279] Move the audio source away from or towards the listener or speaker. o When a user of the system intends to speak to a smart voice assistant of the system (or associated with the system), it may be beneficial to create a cost that penalizes loudspeakers that are closer to the speaker. In this way, these loudspeakers are activated less, allowing their associated microphones to better hear the speaker's utterances. o To provide a more intimate experience for a single listener that minimizes playback levels for others in the listening space, loudspeakers located further away from the listener's position may be heavily penalized so that only the loudspeakers closest to the listener are activated most significantly.

[0280] Move the audio away from or towards a landmark, zone, or area. o Certain locations near the listening space may be considered sensitive (baby room, baby's bed, office, reading area, study area, etc.) In such cases, costs may be configured to penalize the use of speakers near this location, zone, or area. Alternatively, for the same (or similar) case as above, the system of speakers may have generated measurements of the sound transmission from each speaker into the baby's room, especially if one of the speakers (with an attached or associated microphone) is located inside the baby's room. In this case, the physical proximity of the speaker to the baby's room Rather than using a 100 kHz power supply, the cost may be structured to penalize the use of speakers with high measured sound transmission into the baby's room.

[0281] Optimal use of speaker capabilities o The capabilities of different loudspeakers can vary greatly. For example, one popular smart speaker may have only a single 1.6-inch full-range driver with limited low-frequency capabilities, while another smart speaker may have a much more capable 3-inch woofer. These capabilities are generally reflected in the frequency response of the speaker, so the set of frequency responses associated with the speaker may be used in the cost terms. A speaker that is less capable in terms of frequency response than other speakers at a particular frequency may be penalized and therefore activated less frequently. In some embodiments, such frequency response values ​​may be stored with the smart loudspeaker and then communicated to a computing unit responsible for optimizing the flexible rendering.

[0282] Many speakers have multiple drivers, each responsible for reproducing a different frequency band. For example, any given smart speaker is a two-way design with a woofer for low frequencies and a tweeter for high frequencies. Typically, such speakers have crossover circuitry that splits the full-range playback audio signal into appropriate frequency bands and sends them to the respective drivers. Alternatively, such speakers provide the flexible renderer with playback access to each driver, as well as information about the capabilities of each driver (e.g., frequency response). By applying cost terms such as those just described, in some examples, the flexible renderer may automatically form a crossover between two drivers based on their relative capabilities at different frequencies.

[0283] The above-mentioned examples of using frequency responses focus on the inherent capabilities of the speaker, but may not accurately reflect the capabilities of the speaker as placed in the listening environment. In a given case, the measured frequency response of the speaker at the intended listening position may be available through some calibration procedure. Such measurements may be used instead of pre-calculated frequency responses to better optimize speaker usage. For example, a speaker may have high inherent capabilities at certain frequencies, but its placement (e.g., behind a wall or furniture) may severely limit its frequency response at the intended listening position. Capturing this frequency response and feeding it into appropriate cost terms can prevent significant activation of such speakers.

[0284] o Frequency response is only one aspect of a loudspeaker's reproduction capabilities. Many small loudspeakers first begin to distort as the reproduction level increases and then reach an excursion limit, especially for low frequencies. To reduce such distortion, many loudspeakers implement dynamics processing that constrains the reproduction level to a level below some limiting thresholds, which may be variable over frequency. If a speaker is near or above these thresholds, while other speakers participating in flexible rendering are not, it makes sense to reduce the signal level in the limiting speaker and redirect this energy to other, less burdened speakers. Such behavior can be achieved automatically according to some embodiments by appropriately setting the associated cost terms. Such cost terms may include one or more of the following:

[0285] Monitor the overall playback volume with respect to the loudspeaker's limit threshold, for example, penalize loudspeakers whose volume level is closer to their limit threshold more heavily.

[0286] Monitoring dynamic signal levels (possibly varying over frequency) in relation to a loudspeaker limit threshold (possibly varying over frequency). For example, loudspeakers whose monitored signal levels are closer to the limit threshold may be penalized more heavily.

[0287] Directly monitoring loudspeaker dynamics processing parameters, such as limiting gains. In some such instances, loudspeakers for which these parameters exhibit greater limitations may be penalized more.

[0288] Monitor the actual instantaneous voltage, current, and power delivered by the amplifier to the loudspeaker to determine whether the loudspeaker is operating in a linear range. For example, loudspeakers operating with less linearity may be penalized more.

[0289] o Smart speakers with integrated microphones and interactive voice assistants typically use some kind of echo cancellation technology to reduce the level of the audio signal picked up by the recording microphone and being played from the speaker. The greater this reduction, the more likely the speaker is to be able to hear and understand the speech of speakers in the space. If the echo canceller residual is consistently high, this may indicate that the speaker is being driven into a nonlinear region where the echo path is difficult to predict. In such cases, it may make sense to redirect signal energy away from this speaker, and therefore a cost term that takes echo cancellation performance into account may be beneficial. Such a cost term may assign a higher cost to speakers whose associated echo cancellers perform poorly.

[0290] o When rendering spatial audio over multiple loudspeakers, achieving predictable imaging generally requires that the playback over a set of loudspeakers be reasonably synchronized over time. While this is natural for wired loudspeakers, achieving synchronization with a large number of wireless loudspeakers can be difficult and the final result may vary. In such cases, each loudspeaker may be able to report its relative degree of synchronization with the target speaker, and this degree of synchronization may be contributed to a synchronization cost term. In some such instances, loudspeakers that are less synchronized may be penalized more heavily and therefore excluded from rendering. Furthermore, certain audio signals (e.g., components of mixed audio intended for diffuse or omnidirectional playback) may not require strict synchronization. In some embodiments, the components may be tagged as such in metadata, and the synchronization cost term may be modified to reduce the penalty.

[0291] Further examples of embodiments are now described. Similar to the proximity costs defined in equations (9a) and (9b), the new cost function terms It may be convenient to express each of JPEG0007746513000065.jpg1031 as a weighted sum of the squares of the absolute values ​​of the speaker activation potentials, for example:

number

number

[0292] Combining equations (13a) and (13b) with quadratic matrix transformations of the CMAP and FV cost functions given in equation (10) provides a potentially useful implementation of the general augmented cost function (in some embodiments) given in equation (12).

number

[0293] Once the new cost function terms are defined in this way, the overall cost function remains a quadratic matrix, and the optimal set of activation potentials g opt can be found via differentiation of equation (14) as follows:

number

[0294] The weight term w ij , each of which is a given consecutive penalty value for each of the loudspeakers. It is useful to think of the penalty as a function of the object being rendered, JPEG0007746513000071.jpg1039. In one example embodiment, this penalty value is the distance from the object being rendered to the loudspeaker being considered. In another example embodiment, this penalty value represents the inability of a given loudspeaker to reproduce some frequencies. Based on this penalty value, the weighting term w ij can be parameterized as follows:

number

number

[0295] If all loudspeakers are penalized, it is often advantageous to subtract a minimum penalty from all weight terms in post-processing so that at least one of these speakers is not penalized.

number

[0296] As mentioned above, there are many possible use cases that can be realized using the new cost function terms described herein (and similar new cost function terms used according to other embodiments). More specific details will now be provided through three examples: moving audio toward a listener or speaker, moving audio away from a listener or speaker, and moving audio away from a landmark.

[0297] In a first example, an audio signal is pulled toward a location using what is referred to herein as a "gravitational force." This location may be, in some examples, a listener or speaker's location, a landmark location, a piece of furniture, or the like. This location may also be referred to herein as a "gravitational position" or "attractor position." As used herein, a "gravitational force" is a factor that favors a relatively higher loudspeaker activation potential closer to the gravitational position. According to this example, a weight w ij takes the form of equation (17). The continuous penalty value p ij is the fixed attractor position The distance from JPEG0007746513000076.jpg73 to the i-th speaker is given by the threshold τ j is given by the maximum of these distances for all the speakers.

number

number

[0298] To illustrate the use case of "pulling" audio towards the listener or speaker, specifically, α j = 20, β j =3, Set JPEG0007746513000079.jpg73 to the vector corresponding to the listener / speaker position at 180 degrees (bottom center of the plot). α j , β j and These values ​​for JPEG0007746513000080.jpg73 are merely examples. j may be in the range of 1 to 100, and β j may be in the range of 1 to 25. Figure 18 is a graph of speaker activation potentials in an example embodiment. In this example, Figure 18 shows speaker activation potentials 1505b, 1510b, 1515b, 1520b, and 1525b. These are ij16B. FIG. 19 includes an optimal solution to the cost function for the same speaker positions as in FIGS. 15 and 16, with the addition of an attractive force represented by . FIG. 19 is a graph of object rendering positions in an example embodiment. In this example, FIG. 19 shows the corresponding ideal object positions 1630b for a number of possible object angles, and the corresponding actual rendering positions 1635b for those objects. The actual rendering positions 1635b are connected to the ideal object positions 1630b by dotted lines 1640b. The fixed position of the actual rendering positions 1635b The diagonal orientation towards JPEG0007746513000081.jpg73 shows the impact of attractor weighting on the optimal solution to the cost function.

[0299] In the second and third examples, a "repulsive force" is used to "push" audio away from a location. Here, the location may be a person's location (e.g., a listener's location, a speaker's location, etc.) or another location, such as a landmark location, a piece of furniture, etc. In some examples, the repulsive force may be used to push audio away from an area or zone of the listening environment (e.g., an office area, a reading area, a bed or sleeping area (e.g., a baby's bed or sleeping area)). In some such examples, a particular location may be used as a representative of the zone or area. For example, a location representative of a baby's bed may be the estimated location of the baby's head or the location of an estimated sound source corresponding to the baby. This location may also be referred to herein as a "repulsive force location" or "repulsive location." As used herein, a "repulsive force" is a factor that favors a relatively lower loudspeaker activation potential closer to the repulsive force location. In this example, similar to the attractive force in equation (19), a fixed repulsive location For JPEG0007746513000082.jpg73, p ij and τ j is defined as follows:

number

number

[0300] To illustrate the use case of pushing audio away from the listener or speaker, we specifically j = 5, β j =2, Set JPEG0007746513000085.jpg73 to the vector corresponding to the listener / speaker position at 180 degrees (bottom center of the plot). α j , β j and These values ​​for JPEG0007746513000086.jpg73 are for illustrative purposes only. As mentioned above, in some examples, α j may be in the range of 1 to 100, and β j may be in the range of 1 to 25. FIG. 20 is a graph of speaker activation potentials in an example embodiment. According to this example, FIG. 20 shows speaker activation potentials 1505c, 1510c, 1515c, 1520c, and 1525c. These are ij 21 is a graph of object rendering positions in an example embodiment. In this example, FIG. 21 shows ideal object positions 1630c and the corresponding actual rendering positions 1635c for those objects for a number of possible object angles. The actual rendering positions 1635c are connected to the ideal object positions 1630c by dotted lines 1640c. The fixed position of the actual rendering positions 1635c The diagonal orientation away from JPEG0007746513000087.jpg73 shows the impact of the repeller weighting on the optimal solution to the cost function.

[0301] A third use case example is "pushing" audio away from an acoustically sensitive landmark, such as the door leading to a sleeping baby's room. JPEG0007746513000088.jpg73 is set to the vector corresponding to the 180 degree door position (bottom center of the plot). To achieve a stronger repulsion and distort the sound field entirely into the front part of the primary listening space, α j = 20, β j =5. FIG. 22 shows the speaker activation in the embodiment. FIG. 22 is a graph of the potentials. Again, in this example, FIG. 22 shows speaker activation potentials 1505d, 1510d, 1515d, 1520d, and 1525d, which include optimal solutions for the same set of speaker positions with stronger repulsion added. FIG. 23 is a graph of object rendering positions in an example embodiment. Again, in this example, FIG. 23 shows ideal object positions 1630d for a number of possible object angles, and the corresponding actual rendering positions 1635d for those objects. The actual rendering positions 1635d are connected to the ideal object positions 1630d by dotted lines 1640d. The diagonal orientation of the actual rendering positions 1635d indicates the impact of stronger repeller weighting on the optimal solution to the cost function.

[0302] In a further example of method 250 of FIG. 2B, a use case is responsive to selecting two or more audio devices in an audio environment (block 265) and applying a "repulsive" force to audio sounds (block 275). Following the previous example, the selection of two or more audio devices may, in some examples, take the form of a value f_n (a unitless parameter that controls the degree to which audio processing modification occurs). Many combinations are possible. In one simple example, the weight corresponding to the repulsive force is: It may be directly selected as JPEG0007746513000089.jpg718, and the device selected by the "decision" aspect will be penalized.

[0303] Further to the above example of determining weights, in some embodiments, the weights may be determined as follows:

number

[0304] In the above formula, α j , β j , τ j are adjustable parameters that indicate the overall strength of the penalty, the abruptness of the onset of the penalty, and the range of the penalty, respectively, as already explained with reference to equation (17). Thus, the above equation may be understood as a combination of multiple penalty terms, which arise from multiple simultaneous use cases. For example, audio speech can be expressed as the term p ij and the term τ j , which is "pushed away" from the landmarks that need attention, while the term f determined by the decision aspect is i It is also "pushed away" from microphone locations where it is desirable to improve SER using

[0305] The previous example also introduces s_n expressed directly in terms of speech-to-echo improvement (in decibels). Some embodiments may include selecting values ​​of α and β (the strength of the penalty and the abruptness of the onset of the penalty, respectively) based, in part, on the value of s_n (dB), and w ij The formula given above for α j and β j Instead of α, ij and β ij For example, a value of s_i=-20 dB may correspond to a high cost of activating the i-th speaker. In some such examples, α ij is the other term in the cost function, C spatial and C proximity For example, the new value of α can be set to: This can be determined by JPEG0007746513000091.jpg1025. This results in a value of α that is 10 times larger than usual in the cost function for a value of s_i=-20dB. ij It can be β ij , 0.5<β i j <1.0 may be an appropriate modification in some instances based on large negative values ​​of s_i, which "push" the audio away from a significantly larger region around the i-th speaker. For example, the value of s_i may be set to β ij can be mapped to

number

[0306] Aspects of example embodiments include the following enumerated example embodiments (EEE): EEE1. A method (or system) for improving the signal-to-echo ratio for detecting voice commands from a user, comprising: a. Multiple devices are in use to generate the output audio program material. b. There is a known set of distances or ordered relationships from the listener for these devices. c. The system selectively reduces the volume of the device that is closest to the user.

[0307] EEE2. The method or system of EEE1, wherein detecting the signal includes detecting the signal from any noise-producing object or from a desired point of audio monitoring where the distance relationship to a set of devices is known.

[0308] EEE3. The method or system of EEE1 or EEE2 for ordering devices, comprising: This includes considering the signal-to-echo ratio of the

[0309] EEE4. The method or system of any of EEE1-EEE3, wherein the ordering takes into account the generalized proximity of devices to the user and their approximate reciprocity to estimate the most effective signal-to-echo ratio improvement value, and orders the devices in this sense. Aspects of some disclosed embodiments include a system or device configured (e.g., programmed) to perform one or more of the disclosed methods, and a tangible computer-readable medium (e.g., disk) having stored thereon code for implementing one or more of the disclosed methods or steps thereof. For example, the system may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of various operations on data comprising one or more of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system including an input device, memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more of the disclosed methods (or steps thereof) in response to asserted data.

[0310] Some disclosed embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed or otherwise configured) to perform the desired processing on the audio signal(s), including performing one or more of the disclosed methods. Alternatively, some embodiments ( A DSP (or a processor, or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) programmed with software or firmware and / or otherwise configured to perform any of the various operations, including one or more of the disclosed methods or steps thereof. Alternatively, elements of some disclosed embodiments are implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform one or more of the disclosed methods or steps thereof, where the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to perform one or more of the disclosed methods or steps thereof may typically be connected to an input device (e.g., a mouse and / or keyboard), memory, and a display device.

[0311] Another aspect of some disclosed embodiments is a computer-readable medium (e.g., a disk or other tangible storage medium) having stored thereon code (e.g., code executable to perform an embodiment) for performing any embodiment or step of one or more of the disclosed methods.

[0312] While specific embodiments and applications have been described herein, it will be apparent to those skilled in the art that many modifications can be made to the embodiments and applications described herein without departing from the scope of the subject matter described and claimed. While certain implementations have been illustrated and described, it should be understood that the disclosure is not limited to the specific embodiments described and illustrated or to the specific methods described.

Claims

1. 1. A method for managing an audio session, comprising: receiving an output signal from each of a plurality of microphones in an audio environment, each of the plurality of microphones being at a microphone position in the audio environment, the output signal comprising a signal corresponding to a current utterance of a person; determining one or more aspects of context information about the person based on the output signal, the context information including at least one of an estimated current location of the person and an estimated current proximity of the person to one or more microphone locations; selecting two or more audio devices of the audio environment based at least in part on the one or more aspects of the context information, each of the two or more audio devices including at least one loudspeaker; determining one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for the two or more audio devices, the audio processing modifications having the effect of increasing a speech-to-echo ratio at one or more microphones of the plurality of microphones, the one or more audio processing modifications including spectral modifications; applying the one or more audio processing modifications; An audio session management method comprising:

2. 2. The method of claim 1, wherein at least one of the audio processing changes for a first audio device is different from an audio processing change for a second audio device.

3. 2. The method of claim 1, wherein the spectral modification includes reducing the level of audio data in a frequency band between 500 Hz and 3 KHz.

4. 2. The method of claim 1, wherein the one or more audio processing modifications cause a reduction in loudspeaker playback level of at least one loudspeaker of the two or more audio devices.

5. 2. The audio session management method of claim 1, wherein selecting two or more audio devices of the audio environment includes selecting N loudspeaker-equipped audio devices of the audio environment, where N is an integer greater than two.

6. 2. The audio session management method of claim 1, wherein selecting the two or more audio devices of the audio environment is based at least in part on an estimated current location of the person relative to at least one of a microphone location and a loudspeaker-integrated audio device location.

7. 2. The method of claim 1, wherein the one or more audio processing modifications include modifying a rendering process to warp a rendering of an audio signal in a direction away from the estimated current location of the person.

8. 2. The method of claim 1, wherein the one or more audio processing modifications include inserting at least one gap in at least one selected frequency band of the audio playback signal.

9. 2. The method of claim 1, wherein the one or more audio processing modifications include dynamic range compression.

10. The audio session management method of claim 1 , wherein the step of selecting two or more audio devices is based at least in part on signal-to-echo ratio estimates for one or more microphone locations.

11. 11. The audio session management method of claim 10, wherein selecting the two or more audio devices is based at least in part on determining whether the signal-to-echo ratio estimate is less than or equal to a signal-to-echo ratio threshold.

12. 11. The method of claim 10, wherein determining the one or more audio processing modifications is based on optimizing a cost function that is based at least in part on the signal-to-echo ratio estimate.

13. The audio session management method of claim 12 , wherein the cost function is based at least in part on rendering performance.

14. The method of claim 1 , wherein the step of selecting two or more audio devices is based at least in part on a proximity estimate.

15. determining a plurality of current acoustic features from the output signals of each microphone; applying a classifier to the plurality of current acoustic features; applying the classifier includes applying a model trained on previously determined acoustic features obtained from a plurality of previous utterances made by the person in a plurality of user zones within the audio environment; 2. The audio session management method of claim 1, wherein determining one or more aspects of context information about the person comprises determining an estimate of a user zone in which the person is currently located based at least in part on output from the classifier.

16. 16. The method of claim 15, wherein the estimate of the user zone is determined without reference to the geometric locations of the plurality of microphones.

17. 16. The audio session management method of claim 15, wherein the current utterance and the past utterance include utterances of a wake word.

18. The method of claim 1 , further comprising selecting at least one microphone depending on the one or more aspects of the context information.

19. The method of claim 1 , wherein the one or more microphones are located within a plurality of audio devices of the audio environment.

20. The method of claim 1 , wherein the one or more microphones are provided within an audio device of the audio environment.

21. The method of claim 1 , wherein at least one of the one or more microphone locations corresponds to multiple microphones of an audio device.

22. One or more non-transitory media having software stored thereon, said software controlling one or more devices to: receiving an output signal from each of a plurality of microphones in an audio environment, each of the plurality of microphones being at a microphone position in the audio environment, the output signal comprising a signal corresponding to a current utterance of a person; determining one or more aspects of context information about the person based on the output signal, the context information including at least one of an estimated current location of the person and an estimated current proximity of the person to one or more microphone locations; selecting two or more audio devices of the audio environment based at least in part on the one or more aspects of the context information, each of the two or more audio devices including at least one loudspeaker; determining one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for the two or more audio devices, the audio processing modifications having the effect of increasing a speech-to-echo ratio at one or more microphones of the plurality of microphones, the one or more audio processing modifications including spectral modifications; applying the one or more audio processing modifications; One or more non-transitory media containing instructions for performing an audio session management method, including:

23. 23. The one or more non-transitory media of claim 22, wherein at least one of the audio processing modifications for a first audio device is different from an audio processing modification for a second audio device.

24. an interface system; a control system; An apparatus comprising: The control system includes: receiving, via the interface system, an output signal from each of a plurality of microphones in an audio environment, each of the plurality of microphones being at a microphone position in the audio environment, the output signal comprising a signal corresponding to a current utterance of a person; determining one or more aspects of context information about the person based on the output signal, the context information including at least one of an estimated current location of the person and an estimated current proximity of the person to one or more microphone locations; selecting two or more audio devices of the audio environment based at least in part on the one or more aspects of the context information, each of the two or more audio devices including at least one loudspeaker; determining one or more audio processing modifications to apply to audio data being rendered into loudspeaker feed signals for the two or more audio devices, the audio processing modifications having the effect of increasing a speech-to-echo ratio at one or more microphones of the plurality of microphones, the one or more audio processing modifications including spectral modifications; applying the one or more audio processing modifications; 20. An apparatus configured to:

25. 25. The apparatus of claim 24, wherein at least one of the audio processing modifications for a first audio device is different from an audio processing modification for a second audio device.

Citation Information

Patent Citations

  • Method and system for in-hall loudspeaking

    JP1999055784A

  • Loudspeaker system

    JP2006238254A

  • Sound emission / pickup apparatus

    JP2007318274A

  • Sound emitting and collecting apparatus and system

    JP2008294599A