Rendering audio over multiple speakers with multiple activation criteria

The method enhances smart audio device systems by dynamically adjusting speaker activation based on environmental factors, improving audio quality and microphone performance through a cost function optimization approach.

JP2025164836APending Publication Date: 2025-10-30DOLBY LABORATORIES LICENSING CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025136857
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-25
Filing Date
2025-08-20
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing audio rendering systems for smart audio devices lack flexibility and fail to dynamically adjust spatial rendering based on various factors such as speaker proximity, listener position, and microphone performance, limiting their effectiveness in complex environments.

Method used

A method for rendering audio that minimizes a cost function incorporating dynamic speaker activation terms, including proximity to listeners, speaker capabilities, and microphone performance, allowing for dynamic modifications to spatial rendering.

Benefits of technology

Enables flexible and adaptive audio playback that optimizes speaker activation based on environmental factors, enhancing audio quality and microphone performance in smart audio device systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025164836000001_ABST
    Figure 2025164836000001_ABST
Patent Text Reader

Abstract

To provide methods for rendering audio for playback by two or more speakers.SOLUTION: An audio includes one or more audio signals, each with an associated intended perceived spatial position. Relative activation of the speakers may be a cost function of a model of perceived spatial position of the audio signals when played back over the speakers, a measure of proximity of the intended perceived spatial position of the audio signals to positions of the speakers, and one or more additional dynamically configurable functions. The dynamically configurable functions may be based on one or more additional dynamically configurable functions dependent on at least one or more properties of the audio signals, one or more properties of the set of speakers and / or one or more external inputs.SELECTED DRAWING: Figure 3A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 971,421, filed February 7, 2020, U.S. Provisional Patent Application No. 62 / 705,410, filed June 25, 2020, and Spanish Patent Application No. P201930702, filed July 30, 2019, each of which is incorporated herein by reference in its entirety.

[0002] Technical Field SUMMARY This disclosure relates to systems and methods for rendering audio for playback by some or all speakers (e.g., each activated speaker) of a collection of speakers. [Background technology]

[0003] Audio devices, including but not limited to smart audio devices, are widely deployed and are becoming a common feature in many homes. While existing systems and methods for controlling audio devices offer advantages, improved systems and methods would be desirable.

[0004] Notation and Name Throughout this disclosure, including the claims, "speaker" and "loudspeaker" are used interchangeably to refer to any sound-emitting transducer (or collection of transducers) driven by a single speaker feed. A typical set of headphones includes two speakers.

[0005] Throughout this disclosure, including the claims, the phrase performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing the operation directly on the signal or data, or to performing the operation on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing before performing the operation).

[0006] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to an apparatus, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.

[0007] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, video, or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipeline processing on audio or other audio data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0008] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean a direct or indirect connection. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.

[0009] In this paper, we use the term "smart audio device" to refer to a smart device that is either a single-purpose audio device or a virtual assistant (e.g., a connected virtual assistant). A single-purpose audio device is a device (e.g., a television or a mobile phone) that includes or is coupled to at least one microphone and is designed largely or primarily to achieve a single purpose. While televisions are typically capable of (and are considered capable of) playing audio from program material, modern televisions almost always run some kind of operating system on which applications, including television viewing applications, run locally. Similarly, audio input and output on a mobile phone can be numerous, but these are serviced by applications running on the phone. In this sense, single-purpose audio devices with speakers and microphones are often configured to run local applications and / or services for direct use of the speakers and microphones. Some single-purpose audio devices may be configured to be grouped to achieve audio playback in a zone or user-configured area.

[0010] A virtual assistant (e.g., a connected virtual assistant) is a device (e.g., a smart speaker or voice-assistant integrated device) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker), and may provide the ability to utilize multiple devices (different from the virtual assistant) for applications that are cloud-enabled or otherwise not implemented in or on the virtual assistant itself. Virtual assistants sometimes cooperate, e.g., in a very discrete, conditionally defined manner. For example, two or more virtual assistants can cooperate in the sense that one of them, e.g., the virtual assistant that is most confident that it heard the wake word, responds to that word. Connected devices can form a kind of constellation, which may be managed by one main application that may be (or implement) the virtual assistant.

[0011] Here, "wake word" is used broadly to mean any sound (e.g., a word spoken by a human being, or some other sound) that the smart audio device is configured to wake up in response to detecting ("listening") for that sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "awakening" refers to the device entering a state in which it waits for (i.e., listens for) a voice command. In some instances, what may be referred to herein as a "wake word" may include multiple words, e.g., a phrase.

[0012] Here, the term "wake word detector" refers to a device (or software including instructions for configuring a device) configured to continuously search for alignment between real-time audio (e.g., speech) features and a trained model. Typically, a wake word event is triggered whenever the wake word detector determines that the probability that the wake word has been detected exceeds a predetermined threshold. For example, the threshold may be a predetermined threshold adjusted to provide a good compromise between false accept and false reject rates. Following a wake word event, the device may enter a state (which may be referred to as an "awake" or "attention" state) in which it listens for commands and passes received commands to a larger, more computationally intensive recognizer. Summary of the Invention [Means for solving the problem]

[0013] Some embodiments are methods for rendering audio for playback by at least one (e.g., all or a portion) of a set of smart audio devices or for playback by at least one (e.g., all or a portion) of a set of speakers. The rendering may include minimizing a cost function, the cost function including at least one dynamic (e.g., dynamically configurable) speaker activation term. The inclusion of a dynamically configurable term in the activation penalty allows the spatial rendering to be modified in response to a number of considered controls. Examples of dynamic speaker activation terms include, but are not limited to, the following: · The proximity of the speaker to one or more listeners; · The proximity of the speaker to the forces of attraction or repulsion; · The audibility of the speaker with respect to some position (e.g., the listener's position or the baby's room); · Speaker capabilities (frequency response, distortion); ·Synchronization of speakers relative to other speakers; Wake word performance; and / or Echo cancellation performance.

[0014] Minimizing the cost function (including at least one dynamic speaker activation term) may result in deactivation of at least one of the speakers (in the sense that each such speaker does not play the associated audio content) and activation of at least one of the speakers (in the sense that each such speaker plays at least a portion of the rendered audio content). The dynamic speaker activation term may enable at least one of a variety of behaviors, including distorting the spatial presentation of audio away from a particular smart audio device to allow its microphones to better hear a speaker or to allow a secondary audio stream to better be heard from the speaker of the smart audio device.

[0015] Some disclosed implementations may include a system configured (e.g., programmed) to perform any embodiment of the disclosed method, or steps thereof, and a tangible, non-transitory, computer-readable medium embodying non-transitory storage of data (e.g., a disk or other tangible storage medium) that stores code (e.g., executable code to perform) for performing any embodiment of the disclosed method, or steps thereof. For example, embodiments of the disclosed system may be or include a programmable general-purpose processor, digital signal processor, or microprocessor that is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including any embodiment of the disclosed method, or steps thereof. Such a general-purpose processor may be or include a computer system that includes an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform any embodiment of the disclosed method (or steps thereof) in response to data presented thereto.

[0016] At least some aspects of the present disclosure may be implemented via methods such as audio processing methods. In some instances, methods may be implemented, at least in part, by a control system such as those disclosed herein. Some such methods involve receiving audio data by the control system via an interface system. In some examples, the audio data includes one or more audio signals and associated spatial data. According to some examples, the spatial data indicates an intended perceived spatial location corresponding to the audio signal.

[0017] Some such methods involve rendering, by a control system, the audio data for playback through a set of loudspeakers in an environment to generate rendered audio signals. In some examples, rendering each of one or more audio signals included in the audio data involves determining the relative activation of a set of loudspeakers in the environment by optimizing a cost, the cost being a function of: a model of the perceived spatial location of the reproduced audio signal when played through the set of loudspeakers in the environment; a measure of the proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker in the set of loudspeakers; and one or more additional dynamically configurable features.

[0018] According to some examples, the one or more additional dynamically configurable features are based on one or more of the following: proximity of the loudspeaker to one or more listeners; proximity of the loudspeaker to a position of attraction, where attraction is a factor that favors relatively higher activation of loudspeakers closer to the position of attraction; proximity of the loudspeaker to a position of repulsion, where repulsion is a factor that favors relatively lower activation of loudspeakers closer to the position of repulsion; the capabilities of each loudspeaker relative to other loudspeakers in the environment; synchronization of the loudspeaker with respect to other loudspeakers; wake word performance; and / or echo cancellation performance.

[0019] Some such methods involve providing, via an interface system, rendered audio signals to at least some loudspeakers of the set of loudspeakers in the environment, and some such methods involve playing the rendered audio signals by at least some loudspeakers of the set of loudspeakers.

[0020] According to some implementations, the model of perceived spatial position can generate binaural responses corresponding to audio object positions at the listener's left and right ears. In some examples, the model of perceived spatial position can place the perceived spatial position of an audio signal reproduced from a set of loudspeakers at the center of mass of the set of loudspeaker positions, weighted by the loudspeaker's associated activation gains. In some such cases, the model of perceived spatial position can also generate binaural responses corresponding to audio object positions at the listener's left and right ears.

[0021] In some instances, the one or more additional dynamically configurable functions may be based, at least in part, on a level of the one or more audio signals. In some examples, the one or more additional dynamically configurable functions may be based, at least in part, on a spectrum of the one or more audio signals.

[0022] According to some implementations, the one or more additional dynamically configurable capabilities may be based, at least in part, on the position of each loudspeaker in the environment. In some cases, the capabilities of each loudspeaker may include one or more of a frequency response, a playback level limit, or parameters of one or more loudspeaker dynamics processing algorithms. In some examples, the one or more additional dynamically configurable capabilities may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to other loudspeakers.

[0023] According to some examples, the one or more additional dynamically configurable functions may be based, at least in part, on the location of one or more people in the environment. In some such examples, the one or more additional dynamically configurable functions may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the location of the one or more people.

[0024] In some examples, the one or more additional dynamically configurable functions may be based, at least in part, on object positions of one or more non-loudspeaker objects in the environment. In some such examples, the one or more additional dynamically configurable functions may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the object positions.

[0025] In some cases, the one or more additional dynamically configurable features may be based, at least in part, on estimates of acoustic transmission from each speaker to one or more landmarks, regions, or zones of the environment. According to some examples, the intended perceived spatial location may correspond to at least one of channel or location metadata of a channel-based audio format.

[0026] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include one or more memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. Thus, some innovative aspects of the subject matter described in this disclosure may be implemented in non-transitory media having software stored thereon.

[0027] For example, the software may include instructions for controlling one or more devices to perform a method involving receiving, by a control system, audio data via an interface system. In some examples, the audio data includes one or more audio signals and associated spatial data. According to some examples, the spatial data indicates intended perceived spatial locations corresponding to the audio signals.

[0028] Some such methods involve rendering, by a control system, the audio data for playback through a set of loudspeakers in an environment to generate rendered audio signals. In some examples, rendering each of one or more audio signals included in the audio data involves determining the relative activation of a set of loudspeakers in the environment by optimizing a cost, the cost being a function of: a model of the perceived spatial location of the reproduced audio signal when played on the set of loudspeakers in the environment; a measure of the proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker in the set of loudspeakers; and one or more additional dynamically configurable features.

[0029] According to some examples, the one or more additional dynamically configurable features are based on one or more of the following: proximity of the loudspeaker to one or more listeners; proximity of the loudspeaker to a position of attraction, where attraction is a factor that favors relatively higher activation of loudspeakers closer to the position of attraction; proximity of the loudspeaker to a position of repulsion, where repulsion is a factor that favors relatively lower activation of loudspeakers closer to the position of repulsion; the capabilities of each loudspeaker relative to other loudspeakers in the environment; synchronization of the loudspeaker with respect to other loudspeakers; wake word performance; and / or echo cancellation performance.

[0030] Some such methods involve providing, via an interface system, rendered audio signals to at least some loudspeakers of the set of loudspeakers in the environment, and some such methods involve playing the rendered audio signals by at least some loudspeakers of the set of loudspeakers.

[0031] According to some implementations, the model of perceived spatial position can generate binaural responses corresponding to audio object positions at the listener's left and right ears. In some examples, the model of perceived spatial position can center the perceived spatial position of an audio signal reproduced from a set of loudspeakers at the center of mass of the set of loudspeaker positions weighted by the loudspeakers' associated activation gains. In some such examples, the model of perceived spatial position can also generate binaural responses corresponding to audio object positions at the listener's left and right ears.

[0032] In some instances, the one or more additional dynamically configurable functions may be based, at least in part, on a level of the one or more audio signals. In some examples, the one or more additional dynamically configurable functions may be based, at least in part, on a spectrum of the one or more audio signals.

[0033] According to some implementations, the one or more additional dynamically configurable capabilities may be based, at least in part, on the position of each loudspeaker in the environment. In some cases, the capabilities of each loudspeaker may include one or more of a frequency response, a playback level limit, or parameters of one or more loudspeaker dynamics processing algorithms. In some examples, the one or more additional dynamically configurable capabilities may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to other loudspeakers.

[0034] According to some examples, the one or more additional dynamically configurable functions may be based, at least in part, on the location of one or more people in the environment. In some such examples, the one or more additional dynamically configurable functions may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the location of the one or more people.

[0035] In some examples, the one or more additional dynamically configurable functions may be based, at least in part, on object positions of one or more non-loudspeaker objects in the environment. In some such examples, the one or more additional dynamically configurable functions may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the object positions.

[0036] In some cases, the one or more additional dynamically configurable features may be based, at least in part, on estimates of acoustic transmission from each speaker to one or more landmarks, regions, or zones of the environment. According to some examples, the intended perceived spatial location may correspond to at least one of channel or location metadata of a channel-based audio format.

[0037] At least some aspects of the present disclosure may be implemented using an apparatus. For example, one or more apparatuses may be capable of at least partially performing the methods disclosed herein. In some implementations, the apparatus may include an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof.

[0038] In some implementations, a control system may be configured to perform one or more of the methods disclosed herein. Some such methods may involve receiving audio data by the control system via an interface system. In some examples, the audio data includes one or more audio signals and associated spatial data. According to some examples, the spatial data indicates an intended perceived spatial location corresponding to the audio signal.

[0039] Some such methods involve rendering, by a control system, the audio data for playback through a set of loudspeakers in an environment to generate rendered audio signals. In some examples, rendering each of one or more audio signals included in the audio data involves determining the relative activation of a set of loudspeakers in the environment by optimizing a cost, the cost being a function of: a model of the perceived spatial location of the reproduced audio signal when played on the set of loudspeakers in the environment; a measure of the proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker in the set of loudspeakers; and one or more additional dynamically configurable features.

[0040] According to some examples, the one or more additional dynamically configurable features are based on one or more of the following: proximity of the loudspeaker to one or more listeners; proximity of the loudspeaker to a position of attraction, where attraction is a factor that favors relatively higher activation of loudspeakers closer to the position of attraction; proximity of the loudspeaker to a position of repulsion, where repulsion is a factor that favors relatively lower activation of loudspeakers closer to the position of repulsion; the capabilities of each loudspeaker relative to other loudspeakers in the environment; synchronization of the loudspeaker with respect to other loudspeakers; wake word performance; and / or echo cancellation performance.

[0041] Some such methods involve providing, via an interface system, rendered audio signals to at least some loudspeakers of the set of loudspeakers in the environment, and some such methods involve playing the rendered audio signals by at least some loudspeakers of the set of loudspeakers.

[0042] According to some implementations, the model of perceived spatial position can generate binaural responses corresponding to audio object positions at the listener's left and right ears. In some examples, the model of perceived spatial position can center the perceived spatial position of an audio signal reproduced from a set of loudspeakers at the center of mass of the set of loudspeaker positions weighted by the loudspeakers' associated activation gains. In some such examples, the model of perceived spatial position can also generate binaural responses corresponding to audio object positions at the listener's left and right ears.

[0043] In some instances, the one or more additional dynamically configurable functions may be based, at least in part, on a level of the one or more audio signals. In some examples, the one or more additional dynamically configurable functions may be based, at least in part, on a spectrum of the one or more audio signals.

[0044] According to some implementations, the one or more additional dynamically configurable capabilities may be based, at least in part, on the position of each loudspeaker in the environment. In some cases, the capabilities of each loudspeaker may include one or more of a frequency response, a playback level limit, or parameters of one or more loudspeaker dynamics processing algorithms. In some examples, the one or more additional dynamically configurable capabilities may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to other loudspeakers.

[0045] According to some examples, the one or more additional dynamically configurable functions may be based, at least in part, on the location of one or more people in the environment. In some such examples, the one or more additional dynamically configurable functions may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the location of the one or more people.

[0046] In some examples, the one or more additional dynamically configurable functions may be based, at least in part, on object positions of one or more non-loudspeaker objects in the environment. In some such examples, the one or more additional dynamically configurable functions may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the object positions.

[0047] In some cases, the one or more additional dynamically configurable features may be based, at least in part, on estimates of acoustic transmission from each speaker to one or more landmarks, regions, or zones of the environment. According to some examples, the intended perceived spatial location may correspond to at least one of channel or location metadata of a channel-based audio format.

[0048] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. It should be noted that the relative dimensions of the following figures may not be drawn to scale. [Brief explanation of the drawings]

[0049] [Figure 1] FIG. 10 illustrates an exemplary set of speaker activation and object rendering positions. [Figure 2] FIG. 10 illustrates an exemplary set of speaker activation and object rendering positions. [Figure 3A] FIG. 13 is a flow diagram outlining an example of a method that may be performed by an apparatus or system such as that shown in FIG. 11 or FIG. 12. [Figure 3B] 10 is a graph of speaker activation in an example embodiment. [Figure 4] 1 is a graph of object rendering positions in an example embodiment. [Figure 5] 10 is a graph of speaker activation in an example embodiment. [Figure 6] 1 is a graph of object rendering positions in an example embodiment. [Figure 7] 10 is a graph of speaker activation in an example embodiment. [Figure 8]1 is a graph of object rendering positions in an example embodiment. [Figure 9] 10 is a dot graph illustrating speaker activation in an example embodiment. [Figure 10] 10 is a graph of trilinear interpolation between points showing speaker activation according to an example. [Figure 11] 1 is a diagram of an example environment; [Figure 12] FIG. 1 is a block diagram illustrating example components of an apparatus capable of implementing various aspects of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0050] Flexible rendering allows spatial audio to be rendered on any number of arbitrarily positioned speakers. Given the widespread deployment of audio devices in the home, including but not limited to smart audio devices (e.g., smart speakers), there is a need to enable flexible rendering techniques that allow consumer products to perform flexible rendering of audio and playback of the audio so rendered.

[0051] Several techniques have been developed to achieve flexible rendering. They treat the rendering problem as a cost function minimization problem. The cost function consists of two terms: a first term that models the desired spatial impression the renderer is trying to achieve, and a second term that assigns costs to speaker activation. To date, this second term has focused on producing a sparse solution where only speakers close to the desired spatial location of the rendered audio are activated.

[0052] Spatial audio playback in consumer environments has typically been tied to a given number of loudspeakers placed in prescribed locations, such as 5.1 and 7.1 surround sound. In these cases, content is authored specifically for the associated loudspeakers and encoded as discrete channels, one for each loudspeaker (e.g., Dolby Digital or Dolby Digital Plus). More recently, immersive object-based spatial audio formats (Dolby Atmos) have been introduced that break this association between content and specific loudspeaker locations. Instead, content is described as a collection of individual audio objects, each with potentially time-varying metadata that describes the audio object's desired perceived location in three-dimensional space. During playback, content is converted into loudspeaker feeds by a renderer that matches the number and locations of the loudspeakers in the playback system. However, many such renderers constrain the position of a set of loudspeakers to one of a set of prescribed layouts (e.g., 3.1.2, 5.1.2, 7.1.4, 9.1.6, etc. in Dolby Atmos).

[0053] Beyond such constrained rendering, methods have been developed that allow object-based audio to be flexibly rendered on a truly arbitrary number of loudspeakers placed in arbitrary locations. These methods require the renderer to have knowledge of the number and physical locations of the loudspeakers in the listening space. For such systems to be practical for the average consumer, an automated method for locating the loudspeakers would be desirable. One such method relies on the use of multiple microphones, possibly co-located with the loudspeakers. By playing audio signals through the loudspeakers and recording with the microphones, the distance between each loudspeaker and microphone is estimated. From these distances, the locations of both the loudspeaker and the microphone are then estimated.

[0054] The introduction of object-based spatial audio in consumer spaces has coincided with the rapid adoption of so-called “smart speakers,” such as the Amazon Echo line of products. While the immense popularity of these devices is due to their simplicity and convenience offered by wireless connectivity and integrated voice interfaces (e.g., Amazon's Alexa), the acoustic capabilities of these devices have generally been limited, especially with regard to spatial audio. In most cases, these devices are constrained to mono or stereo playback. However, combining the aforementioned flexible rendering and automatic localization techniques with multiple orchestrated smart speakers can provide a system with highly sophisticated spatial playback capabilities that remains extremely simple for consumers to set up. Because of wireless connectivity, consumers can place as many or as few speakers as desired wherever convenient, without the need to run speaker cords, and built-in microphones can be used to automatically localize the speakers for the associated flexible renderers.

[0055] Conventional flexible rendering algorithms are designed to achieve, to the greatest extent possible, a particular desired perceived spatial impression. In an orchestrated smart speaker system, maintaining this spatial impression may sometimes not be the most important or desired objective. For example, if someone is simultaneously attempting to speak to an integrated voice assistant, it may be desirable to temporarily modify the spatial rendering to reduce the relative playback level on speakers near certain microphones in order to increase the signal-to-noise ratio of the recording. Some embodiments described herein may be implemented as modifications to existing flexible rendering methods to allow such dynamic modifications to the spatial rendering, e.g., to achieve one or more additional objectives.

[0056] Existing flexible rendering techniques include Center of Mass Amplitude Panning (CMAP) and Flexible Virtualization (FV). From a high level, both of these techniques render a collection of one or more audio signals, each with an associated desired perceived spatial location, for playback through a collection of two or more speakers, where the relative activation of the speakers in the collection is a function of a model of the perceived spatial location of the audio signals reproduced through the speakers and the proximity of the desired perceived spatial location of the audio signals to the positions of those speakers. The model ensures that the audio signals are heard by the listener near their intended spatial location, and a proximity term controls which speakers are used to achieve this spatial impression. In particular, the proximity term favors the activation of speakers that are close to the desired perceived spatial location of the audio signals.

[0057] For both CMAP and FV, this functional relationship is expressed as a cost function written as a sum of two terms, one for the spatial aspect and one for the proximity:

number

number

number

[0058] In some definitions of the cost function, g opt While the relative levels between the components of are appropriate, it is difficult to control the absolute level of optimal activation that results from the above minimization. To address this issue, a subsequent normalization may be performed so that the absolute level of activation is controlled. For example, it may be desirable to normalize the vector to have unit length, similar to the commonly used constant power panning rule:

number

[0059] The exact behavior of the flexible rendering algorithm depends on the two terms in the cost function C spatial and C proximity For CMAP, C spatialis derived from a model that places the perceived spatial location of an audio signal reproduced from a collection of loudspeakers at the center of mass of those loudspeaker locations weighted by their associated activation gains (elements of vector g):

number

number

[0060] In FV, the spatial terms of the cost function are defined differently. The goal here is to generate a binaural response b corresponding to an audio object location (vector o) at the listener's left and right ear. Conceptually, b is a 2x1 vector of filters (one filter for each ear), but more conveniently it is treated as a 2x1 vector of complex values ​​at a particular frequency. Continuing this representation at a particular frequency, the desired binaural response can be obtained from a set of HRTF indices indexed by the object location:

number

[0061] At the same time, the 2×1 binaural response e produced by the loudspeakers at the listener's ears is modeled as a 2×M acoustic transfer matrix H multiplied by an M×1 vector g of complex speaker activation values:

number

number

number

[0062] Conveniently, the spatial terms of the cost functions for CMAP and FV defined in Equations 4 and 7 can both be reorganized into matrix-quadratic forms as functions of the speaker activation g:

number

number

number

[0063] To this end, the second term of the cost function, C proximitycan be defined as the distance-weighted sum of the squared absolute values ​​of the speaker activations, which can be succinctly expressed in matrix form as follows:

number

number

[0064] The distance penalty function can take many forms, but the following is a useful parameterization:

number

number

[0065] Combining the two terms of the cost function defined in Equation 8 and Equation 9a gives the overall cost function.

number

number

[0066] In general, the optimal solution of Equation (11) may result in speaker activations that are negative in value. For the CMAP construction of a flexible renderer, such negative activations may be undesirable, and thus Equation (11) may be minimized under the condition that all activations remain positive.

[0067] 1 and 2 illustrate an exemplary set of speaker activation and object rendering positions. In these examples, the speaker activations and object rendering positions correspond to speaker positions of 4, 64, 165, -87, and -4 degrees. FIG. 1 shows speaker activations 105a, 110a, 115a, 120a, and 125a that constitute the optimal solution to Equation 11 for these specific speaker positions. FIG. 2 plots the individual speaker positions as dots 205, 211, 215, 220, and 225, corresponding to speaker activations 105a, 110a, 115a, 120a, and 125a, respectively. FIG. 2 also illustrates ideal object positions (i.e., positions where audio objects should be rendered) for a number of possible object angles as dots 230a and the corresponding actual rendering positions for those objects as dots 235a, connected to the ideal object positions by dotted lines 240a.

[0068] One class of embodiments involves a method for rendering audio for playback by at least one (e.g., all or some) of multiple coordinated (orchestrated) smart audio devices. For example, a set of smart audio devices in a user's home may be orchestrated to handle a variety of simultaneous use cases. Such use cases include rendering (according to certain embodiments) audio for playback by all or some of the smart audio devices (i.e., through all or some speakers). Many interactions with the system are contemplated that require dynamic modifications to the rendering. Such modifications may, but need not, focus on spatial fidelity.

[0069] Some embodiments are methods for rendering audio for playback by at least one (e.g., all or a portion) of a set of smart audio devices (or for playback by at least one (e.g., all or a portion) of another set of speakers). The rendering may include minimizing a cost function, the cost function including at least one dynamic speaker activation term. Examples of such dynamic speaker activation terms include, but are not limited to, the following: · The proximity of the speaker to one or more listeners; · The proximity of the speaker to the forces of attraction or repulsion; · The audibility of the speaker with respect to some position (e.g., the listener's position or the baby's room); · Speaker capabilities (frequency response, distortion); ·Synchronization of speakers relative to other speakers; Wake word performance; and Echo cancellation performance.

[0070] The dynamic speaker activation term may enable at least one of a variety of behaviors, including distorting the spatial presentation of audio away from a particular smart audio device to allow its microphone to better hear the speaker or to allow a secondary audio stream to better be heard from the speaker of the smart audio device.

[0071] Some embodiments include: Implement rendering for playback by speakers of multiple orchestrated smart audio devices. Other embodiments implement rendering for playback by speaker(s) of a separate collection of speakers.

[0072] Pairing a flexible rendering method (implemented according to some embodiments) with a collection of wireless smart speakers (or other smart audio devices) can provide an extremely capable and easy-to-use spatial audio rendering system. Considering interactions with such a system, it becomes apparent that dynamic modifications to the spatial rendering may be desirable to optimize for other objectives that may arise during use of the system. To this end, one class of embodiments augments existing flexible rendering algorithms with one or more additional, dynamically configurable features that depend on one or more attributes of the audio signal being rendered, the collection of speakers, and / or other external inputs. According to some embodiments, the existing flexible rendering cost function given in Equation 1 is augmented with these one or more additional dependencies, as follows:

number

[0073] In Equation 12, the term

number

number

number

number

number

number

number

number

[0074]

number

[0075]

number

[0076]

number

[0077] Using the new cost function defined in Equation 12, we can find the optimal set of activations through minimization with respect to g and possible a posteriori regularization, as previously described in Equations 2a and 2b.

[0078] 3A is a flow diagram outlining one example of a method that may be implemented by a device or system such as that shown in FIG. 11 or 12. The blocks of method 300, as with other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than shown and / or described. The blocks of method 300 may be performed by one or more devices that may be (or may include) a control system, such as control system 1210 shown in FIG. 12.

[0079] In this implementation, block 305 involves receiving audio data by a control system via an interface system. In this example, the audio data includes one or more audio signals and associated spatial data. According to this implementation, the spatial data indicates an intended perceived spatial location corresponding to the audio signal. In some cases, the intended perceived spatial location may be explicit, e.g., indicated by positional metadata such as Dolby Atmos positional metadata. In other cases, the intended perceived spatial location may be implicit, e.g., the intended perceived spatial location may be an assumed location associated with a channel according to Dolby 5.1, Dolby 7.1, or other channel-based audio formats. In some examples, block 305 involves a rendering module of the control system receiving the audio data via an interface system.

[0080] According to this example, block 310 involves rendering the audio data by a control system to generate rendered audio signals for playback through a set of loudspeakers in the environment. In this example, rendering each of one or more audio signals included in the audio data involves determining the relative activation of a set of loudspeakers in the environment by optimizing a cost function. According to this example, the cost is a function of a model of the perceived spatial location of the audio signal when played through the set of loudspeakers in the environment. In this example, the cost is also a function of a measure of the proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker in the set of loudspeakers. In this implementation, the cost is also a function of one or more additional dynamically configurable features. In this example, the dynamically configurable features are based on one or more of the following: the proximity of the loudspeaker to one or more listeners; the proximity of the loudspeaker to an attraction position, where attraction is a factor that favors relatively higher activation of loudspeakers closer to the attraction position; the proximity of the loudspeaker to a repulsion position, where repulsion is a factor that favors relatively lower activation of loudspeakers closer to the repulsion position; the capabilities of each loudspeaker relative to other loudspeakers in the environment; the synchronization of the loudspeaker with respect to other loudspeakers; wake word performance; or echo cancellation performance.

[0081] In this example, block 315 involves providing the rendered audio signals to at least some loudspeakers of the set of loudspeakers in the environment via an interface system.

[0082] According to some examples, the model of perceived spatial location can generate binaural responses corresponding to audio object locations at the listener's left and right ears. Alternatively or additionally, the model of perceived spatial location can center the perceived spatial location of an audio signal reproduced from a set of loudspeakers at the center of mass of the positions of said set of loudspeakers weighted by the associated activation gains of the loudspeakers.

[0083] In some examples, the one or more additional dynamically configurable functions may be based, at least in part, on a level of the one or more audio signals. In some instances, the one or more additional dynamically configurable functions may be based, at least in part, on a spectrum of the one or more audio signals.

[0084] Some examples of method 300 involve receiving speaker layout information. In some examples, the one or more additional dynamically configurable features may be based, at least in part, on the location of each loudspeaker in the environment.

[0085] Some examples of method 300 involve receiving loudspeaker specification information. In some examples, the one or more additional dynamically configurable features can be based, at least in part, on the capabilities of each loudspeaker, which can include one or more of frequency response, playback level limits, or parameters of one or more loudspeaker dynamics processing algorithms.

[0086] According to some examples, the one or more additional dynamically configurable functions may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the other loudspeakers. Alternatively or additionally, the one or more additional dynamically configurable functions may be based, at least in part, on the positions of one or more human listeners or speakers in the environment. Alternatively or additionally, the one or more additional dynamically configurable functions may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the listener or speaker positions. The estimates of acoustic transmission may, for example, be based, at least in part, on walls, furniture, or other objects that may be between each loudspeaker and the listener or speaker positions.

[0087] Alternatively or additionally, the one or more additional dynamically configurable functions may be based, at least in part, on object positions of one or more non-loudspeaker objects or landmarks in the environment. In some such implementations, the one or more additional dynamically configurable functions may be based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the object or landmark positions.

[0088] By employing one or more well-defined additional cost terms to achieve flexible rendering, many new and useful behaviors can be achieved. All of the example behaviors listed below are created by penalizing certain loudspeakers under certain conditions that are deemed undesirable. The end result is that these loudspeakers are activated less in the spatial rendering of the set of audio signals. While in many of these cases it might be considered to simply reduce the volume of the undesired loudspeakers without modifying the spatial rendering, such a strategy can significantly degrade the overall balance of the audio content. Certain components of the mix may, for example, become completely inaudible. On the other hand, in the disclosed embodiment, by integrating these penalizations into the core optimization of the rendering, the rendering can adapt and perform the best possible spatial rendering using the remaining speakers with lower penalties. This is a much more elegant, adaptive, and effective solution.

[0089] Example use cases include, but are not limited to:

[0090] Provides a more balanced spatial presentation around the listening area It has been found that spatial audio is best presented through loudspeakers that are approximately the same distance from the intended listening area. Costs may be constructed such that loudspeakers that are significantly closer or farther than the average distance of the loudspeakers to the listening area are penalized, thereby reducing their activation.

[0091] ● Moving the audio away from or towards the listener or speaker If a user of the system is going to speak to the system or to a smart voice assistant associated with the system, it is beneficial to create a cost that penalizes loudspeakers that are closer to the speaker, so that these loudspeakers are activated less, allowing the associated microphones to hear the speaker better. To provide a more private experience for a single listener that minimizes playback levels for other listeners in the listening space, speakers farther from the listener's position may be heavily penalized, so that only the speakers closest to the listener are activated most prominently.

[0092] Move audio away from or closer to a landmark, zone, or area Certain locations in the vicinity of the listening space may be considered sensitive, e.g., baby rooms, cribs, offices, reading areas, study areas, etc. In such cases, a cost may be constructed that penalizes the use of speakers close to this location, zone or area. Alternatively, for the same (or similar) case as above, a system of speakers may generate measurements of sound transmission from each speaker to the baby room, especially if one of the speakers (with an attached or associated microphone) is located within the baby room itself. In this case, rather than using the physical proximity of the speakers to the baby room, a cost may be constructed that penalizes the use of speakers with high measured sound transmission into the baby room. And / or

[0093] ●Optimal use of speaker capabilities The capabilities of different loudspeakers can vary significantly. For example, one popular smart speaker may only contain a single 1.6-inch full-range driver with limited low-frequency capabilities, while another smart speaker may contain a much more capable 3-inch woofer. These capabilities are generally reflected in the frequency response of the speaker, and thus the aggregate response associated with the speaker may be utilized in cost terms. At certain frequencies, speakers that are less capable than other speakers, as measured by their frequency response, may be penalized and therefore activated to a lesser extent. In some implementations, such frequency response values ​​may be stored in the smart loudspeaker and then reported to a computational unit responsible for optimizing the dynamic rendering.

[0094] Many speakers contain multiple drivers, each responsible for reproducing a different frequency range. For example, some popular smart speakers are two-way designs that include a woofer for low frequencies and a tweeter for high frequencies. Typically, such speakers include crossover circuitry to split the full-range playback audio signal into appropriate frequency ranges and route them to the respective drivers. Alternatively, such speakers can provide the Flexible Renderer playback access to each individual driver and provide information about each individual driver's capabilities, such as frequency response. By applying cost terms such as those described above, in some instances, the Flexible Renderer can automatically construct a crossover between two drivers based on their relative capabilities at different frequencies.

[0095] The above use examples of frequency response focus on the inherent capabilities of a speaker, but may not accurately reflect the capabilities of a speaker placed in a listening environment. In some cases, the frequency response of a speaker measured at the intended listening position may be available through some calibration procedure. Such measurements may be used instead of pre-calculated responses to better optimize the use of the speaker. For example, a certain speaker may be inherently very capable at certain frequencies, but due to its placement (e.g., against a wall or behind furniture), may produce a very limited response at the intended listening position. Capturing this response and inputting the measurements into the appropriate cost terms can prevent significant activation of such speakers.

[0096] ○ Frequency response is only one aspect of a loudspeaker's reproduction capabilities. Many small loudspeakers begin to distort as the reproduction level increases, and then reach an excursion limit, especially at low frequencies. To reduce such distortion, many loudspeakers implement dynamics processing that constrains the reproduction level below some limiting thresholds that may be variable across frequency. If a speaker is close to or at these thresholds, and other speakers participating in flexible rendering are not, it makes sense to reduce the signal level of the limiting speaker and direct this energy to other, less burdened speakers. Such behavior can be achieved automatically according to some embodiments by appropriately configuring the associated cost terms. Such cost terms may involve one or more of the following: Monitoring of global playback volume relative to a loudspeaker's limiting threshold. For example, loudspeakers whose volume level is closer to their limiting threshold may be penalized more heavily; Monitoring of dynamic signal levels, possibly varying over frequency, in relation to a loudspeaker limiting threshold, also possibly varying over frequency. For example, loudspeakers whose monitored signal levels are closer to their limiting thresholds may be penalized more heavily; Direct monitoring of loudspeaker dynamics processing parameters, such as limiting gain. In some such instances, loudspeakers whose parameters indicate stronger limiting may be penalized more heavily; and / or Monitoring the actual instantaneous voltage, current, and power being delivered by the amplifier to the loudspeaker to determine if the loudspeaker is operating in its linear range. For example, loudspeakers operating with less linearity may be penalized more.

[0097] Smart speakers with integrated microphones and interactive voice assistants typically use some type of echo cancellation to reduce the level of the audio signal played from the speaker that is picked up by the recording microphone. The greater this reduction, the more likely the speaker is to hear and understand speakers in the space. If the echo canceller residual is consistently high, this may be an indicator that the speaker is being driven into a nonlinear region where the echo path becomes difficult to predict. In such cases, it makes sense to divert signal energy away from that speaker, and thus a cost term that takes echo canceller performance into account may be beneficial. Such a cost term may assign a higher cost to speakers whose associated echo cancellers perform poorly.

[0098] o To achieve predictable imaging when rendering spatial audio with multiple loudspeakers, it is generally necessary for playback on a set of loudspeakers to be reasonably synchronized over time. For wired loudspeakers, this is natural, but with a large number of wireless loudspeakers, synchronization can be difficult and the final result can be variable. In such cases, each loudspeaker may be able to report its relative degree of synchronization with the target, and this degree may be input into the synchronization cost term. In some such instances, loudspeakers with lower degrees of synchronization may be penalized more heavily and thus excluded from rendering. Furthermore, strict synchronization may not be required for certain types of audio signals, for example, components of an audio mix that are intended to be diffuse or non-directional. In some implementations, components may be tagged as such using metadata, and the synchronization cost term may be modified so that the penalty is reduced.

[0099] Next, an example embodiment will be described.

[0100] Similar to the proximity costs defined in Equations 9a and 9b, the terms in the new cost function

number

number

number

number

[0101] Combining Equations 13a and 13b with matrix-quadratic versions of the CMAP and FV cost functions given in Equation 10 yields a potentially useful implementation of the general extended cost function (in some embodiments) given in Equation 12:

number

[0102] With this definition of the new cost function term, the overall cost function remains a matrix quadratic form, and the optimal set of activations g opt can be found through differentiation of Equation 14, resulting in:

number

[0103] Weight term w ij , respectively, for a given consecutive penalty value for each loudspeaker.

number

number

number

number

[0104] If all loudspeakers are penalized, it is often convenient in post-processing to subtract the minimum penalty from all weight terms to ensure that at least one of the speakers is not penalized:

number

[0105] As noted above, there are many possible use cases that can be realized using the new cost function terms described herein (and similar new cost function terms used according to other embodiments). Three examples will now be used to provide more specific details: moving audio toward a listener or speaker, moving audio away from a listener or speaker, and moving audio away from a landmark.

[0106] In a first example, what is referred to herein as a "gravitational force" is used to pull audio toward a location. That location may be, in some examples, a listener or speaker location, a landmark location, a furniture location, etc. This location may be referred to herein as a "gravitational position" or "attractor position." As used herein, "gravitational force" is a factor that favors relatively higher loudspeaker activation in the vicinity closer to the attraction position. According to this example, the weight w ij takes the form of Equation 17, and the continuous penalty value p ij is the fixed attractor position of the i-th speaker

number

number

[0107] To illustrate the use case of "pulling" audio towards a listener or speaker, specifically α j = 20, β j Set =3,

number

number

[0108] In the second and third examples, a "repulsion force" is used to "push" audio away from a location, which may be a person's location (e.g., a listener's location, a speaker's location, etc.) or other location, such as a landmark location, a furniture location, etc. In some examples, a repulsion force may be used to push audio away from an area or zone of an auditory environment, such as an office area, a reading area, a bed or sleeping area (e.g., a crib or bedroom). According to some such examples, a specific location may be used to represent a zone or area. For example, a location representing an infant's bed may be an estimated location of the infant's head, an estimated sound source location corresponding to the infant, etc. This location may be referred to herein as a "repulsion force location" or "repulsion location." As used herein, a "repulsion force" is a factor that promotes relatively lower speaker activation closer to the repulsion force location. According to this example, a fixed repulsion location

number

number

[0109] To illustrate the use case of directing audio away from the listener or speaker, specifically α j = 5, β j Set =2,

number

number

[0110] A third example use case is to "push" audio away from an acoustically sensitive landmark, such as the door to a sleeping baby's room. Similar to the previous example, j to the vector corresponding to the 180 degree door position (bottom center of the plot). To achieve a stronger repulsion force and to fully bias the sound field towards the front of the primary listening space, we set α j = 20, β j =5. FIG. 7 is a graph of speaker activations in an exemplary embodiment. Again, in this example, FIG. 7 shows speaker activations 105d, 110d, 115d, 120d, and 125d that constitute an optimal solution to the same set of speaker positions, but with a stronger repulsion weighting. FIG. 8 is a graph of object rendering positions in an exemplary embodiment. Again, in this example, FIG. 8 shows ideal object positions 230d for a number of possible object angles and the corresponding actual rendering positions 235d for those objects connected to the ideal object positions 230d by dotted lines 240d. The skewed orientation of actual rendering positions 235d illustrates the impact of a stronger repulsion weighting on the optimal solution to the cost function.

[0111] One practical consideration when implementing dynamic cost-flexible rendering (according to some embodiments) is computational complexity. In some cases, solving a unique cost function for each frequency band for each audio object in real time may not be feasible, given that object positions (which may be indicated by metadata for each rendered audio object) can change many times per second. An alternative approach that reduces computational complexity at the expense of memory is to use a lookup table that samples the three-dimensional space of all possible object positions. The sampling does not need to be the same in all dimensions. Figure 9 is a point graph illustrating speaker activations in one exemplary embodiment. In this example, the x and y dimensions are sampled with 15 points, and the z dimension is sampled with 5 points. Other implementations may include more or fewer samples. According to this example, each point represents M speaker activations for a CMAP or FV solution.

[0112] At runtime, to determine the actual activation for each speaker, in some examples, tri-linear interpolation between the last eight speaker activation points may be used. Figure 10 is a graph of tri-linear interpolation between points showing speaker activations according to one example. In this example, the process of successive linear interpolation includes interpolating each pair of points in the top surface to determine first and second interpolated points 1005a and 1005b, interpolating each pair of points in the bottom surface to determine third and fourth interpolated points 1010a and 1010b, interpolating the first and second interpolated points 1005a and 1005b to determine a fifth interpolated point 1015 in the top surface, interpolating the third and fourth interpolated points 1010a and 1010b to determine a sixth interpolated point 1020 in the bottom surface, and interpolating the fifth and sixth interpolated points 1015 and 1020 to determine a seventh interpolated point 1025 between the top and bottom surfaces. While trilinear interpolation is a useful interpolation method, those skilled in the art will understand that trilinear interpolation is only one possible interpolation method that may be used in implementing aspects of the present disclosure, and that other examples may include other interpolation methods.

[0113] In the first example above, where repulsion forces are used to create an acoustic space for a voice assistant, another key concept is the transition from a rendered scene without repulsion forces to one with repulsion forces: both the previous set of speaker activations without repulsion forces and the new set of speaker activations with repulsion forces are computed and interpolated over a period of time to create a smooth transition and give the impression that the sound field is dynamically distorted.

[0114] An example of audio rendering implemented according to an embodiment is an audio rendering method comprising: 1. A method comprising: rendering a set of one or more audio signals, each having an associated desired perceived spatial location, through a set of two or more loudspeakers, wherein the relative activation of the sets of loudspeakers is a function of a model of the perceived spatial location of the audio signals played through those loudspeakers, the proximity of the desired perceived spatial location of the audio object to the position of the loudspeakers, and one or more additional dynamically configurable features that depend on at least one or more attributes of the set of audio signals, one or more attributes of the set of loudspeakers, or one or more external inputs.

[0115] A further example of an embodiment will now be described with reference to FIG.

[0116] Figure 11 is a diagram of an environment according to one example. In this example, the environment is a living space that includes a set of smart audio devices (device 1.1) for audio interaction, speakers (1.3) for audio output, and controllable lights (1.2). In one example, only device 1.1 contains a microphone and therefore knows where the user (1.4) is located when issuing a wake word command. Using various methods, information can be obtained from these devices collectively to provide a location estimate (e.g., a fine-grained location estimate) of the user issuing (e.g., speaking) the wake word.

[0117] In such a living space, there is a collection of natural action zones where people perform tasks, activities, or cross thresholds. These action areas (zones) are where there may be efforts to estimate the user's position (e.g., determining an uncertain position) or the user's context to assist with other aspects of the interface. In the example of Figure 11, the key action areas are: 1. Kitchen sink and cooking area (upper left area of ​​living space); 2. Refrigerator door (to the right of the sink and cooking area); 3. Dining area (lower left area of ​​the living space); 4. Open area of ​​living space (to the right of the sink and cooking and dining areas); 5. TV couch (right of open area); 6. The TV itself; 7. Table; 8. Door area or entrance (top right area of ​​living space).

[0118] In some examples, an area or zone may correspond to all or a portion of a room in the environment. According to some such examples, an area or zone may correspond to all or a portion of a bedroom. In one such example, an area or zone may correspond to all or a portion of a baby's bedroom, for example, an area near the crib.

[0119] Often it will be apparent that there will be a similar number of lights in similar positions to fit the action area, some or all of which may be individually controllable networked agents.

[0120] According to some embodiments, the audio is rendered (e.g., by one of devices 1.1 or other devices in the system of FIG. 11) for playback by one or more of the speakers (and / or one or more speakers of device (1.1)) (according to any embodiment of the disclosed method).

[0121] Many embodiments are technically possible, and it will be clear to one skilled in the art from this disclosure how to implement them. Several embodiments of the disclosed systems and methods are described herein.

[0122] 12 is a block diagram illustrating example components of a device capable of implementing various aspects of the present disclosure. According to some examples, device 1200 may be or include a smart audio device configured to perform at least a portion of the methods disclosed herein. In other implementations, device 1200 may be or include another device configured to perform at least a portion of the methods disclosed herein, such as a laptop computer, a cellular phone, a tablet device, a smart home hub, etc. In some such implementations, device 1200 may be or include a server.

[0123] In this example, device 1200 includes an interface system 1205 and a control system 1210. Interface system 1205, in some implementations, may be configured to receive an audio program stream. The audio program stream may include audio signals scheduled to be played by at least some speakers in the environment. The audio program stream may include spatial data, such as channel data and / or spatial metadata. Interface system 1205, in some implementations, may be configured to receive input from one or more microphones in the environment.

[0124] The interface system 1205 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 1205 may include one or more wireless interfaces. The interface system 1205 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 1205 may include one or more interfaces between the control system 1210 and a memory system, such as the optional memory system 1215 shown in FIG. 12, although the control system 1210 may include the memory system.

[0125] The control system 1210 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0126] In some implementations, the control system 1210 may reside in more than one device. For example, a portion of the control system 1210 may reside in a device within one of the environments described herein, and another portion of the control system 1210 may reside in a device outside the environment, such as a server, a mobile device (e.g., a smartphone or tablet computer), or another portion of the control system 1210 may reside in one or more other devices in the environment. For example, the functionality of the control system may be distributed across multiple smart audio devices in the environment, or may be shared by an orchestration device (e.g., what may be referred to herein as a smart home hub) and one or more other devices in the environment. The interface system 1205 may also reside in more than one device in some such examples.

[0127] In some implementations, control system 1210 may be configured, at least in part, to perform the methods disclosed herein. According to some examples, control system 1210 may be configured to implement a method for rendering audio on multiple speakers with multiple activation criteria.

[0128] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may reside, for example, in optional memory system 1215 and / or control system 1210 shown in FIG. 12. Thus, various innovative aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media storing software. The software may include, for example, instructions for controlling at least one device to process audio data. The software may be executable by one or more components of a control system, such as control system 1210 of FIG. 12.

[0129] In some examples, device 1200 may include an optional microphone system 1220 shown in FIG. 12. Optional microphone system 1220 may include one or more microphones. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker in a speaker system, a smart audio device, etc.

[0130] According to some implementations, device 1200 may include optional loudspeaker system 1225, as shown in FIG. 12 . Optional speaker system 1225 may include one or more speakers. In some examples, at least some of the speakers of optional speaker system 1225 may be positioned arbitrarily. For example, at least some of the speakers of optional speaker system 1225 may be positioned in locations that do not correspond to any standard-defined speaker layout, such as Dolby 5.1, Dolby 7.1, Hamasaki 22.2, etc. In some such examples, at least some of the speakers of optional speaker system 1225 may be positioned in locations that are convenient for the space (e.g., where there is space to accommodate the speakers), but may be in locations that are not in any standard-defined speaker layout.

[0131] According to some such examples, device 1200 may be or include a smart audio device. In some such implementations, device 1200 may be or include a wake word detector. For example, device 1200 may be or include a virtual assistant.

[0132] Some disclosed implementations include systems or devices configured (e.g., programmed) to perform any embodiment of the disclosed methods, and tangible computer-readable media (e.g., disks) storing code for implementing any embodiment of the disclosed methods, or steps thereof. For example, the disclosed systems may be or include a programmable general-purpose processor, digital signal processor, or microprocessor that is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the disclosed methods, or steps thereof. Such a general-purpose processor may be or include a computer system that includes an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform embodiments of the disclosed methods (or steps thereof) in response to data presented thereto.

[0133] Some embodiments of the disclosed system are implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on an audio signal, including performing embodiments of the disclosed methods. Alternatively, embodiments of the disclosed methods (or elements thereof) are implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) that is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations, including embodiments of the disclosed methods. Alternatively, elements of some embodiments are implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform embodiments of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to perform embodiments of the disclosed methods is typically coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device.

[0134] Another aspect of the present disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) having stored thereon code (e.g., executable code) for performing any of the disclosed methods or steps thereof.

[0135] Various features and aspects will be understood from the following enumerated example embodiments (EEE).

[0136] EEE1. A method for rendering audio for playback by at least two speakers of at least one smart audio device of a collection of smart audio devices, the audio being one or more audio signals, each audio signal having an associated desired perceived spatial location, and wherein relative activation of speakers of the collection of speakers is a function of a model of the perceived spatial location of the audio signals played on those speakers, a proximity of the desired perceived spatial location of the audio signals to the position of the speakers, and one or more additional dynamically configurable features that depend on at least one or more attributes of the audio signals, one or more attributes of the collection of speakers, or one or more external inputs.

[0137] EEE2. The additional dynamically configurable features include at least one of: proximity of a speaker to one or more listeners; proximity of a speaker to attractive or repulsive forces; audibility of a speaker relative to some position; capability of a speaker; synchronization of a speaker relative to other speakers; wake word performance; or echo cancellation performance.

[0138] EEE3. The method of claim EEE1 or EEE2, wherein the rendering comprises minimizing a cost function, the cost function comprising at least one dynamic speaker activation term.

[0139] EEE4. A method for rendering audio for playback by at least two speakers of a set of speakers, the audio being one or more audio signals, each audio signal having an associated desired perceived spatial location, and wherein relative activation of speakers of the set of speakers is a function of a model of the perceived spatial location of the audio signals played on those speakers, a proximity of the desired perceived spatial location of the audio signals to the position of the speakers, and one or more additional dynamically configurable features that depend on at least one or more attributes of the audio signals, one or more attributes of the set of speakers, or one or more external inputs.

[0140] EEE5. The additional dynamically configurable features include at least one of: proximity of a speaker to one or more listeners; proximity of a speaker to attractive or repulsive forces; audibility of a speaker relative to some position; capability of a speaker; synchronization of a speaker relative to other speakers; wake word performance; or echo cancellation performance.

[0141] EEE6. The method of claim EEE4 or 5, wherein said rendering comprises minimizing a cost function, said cost function comprising at least one dynamic speaker activation term.

[0142] EEE7. An audio rendering method comprising: rendering a set of one or more audio signals, each having an associated desired perceived spatial location, to a set of two or more loudspeakers, wherein the relative activation of the sets of loudspeakers is a function of a model of the perceived spatial location of the audio signals played on those loudspeakers, the proximity of the desired perceived spatial location of the audio object to the position of the loudspeakers, and one or more additional dynamically configurable features that depend on at least one or more attributes of the set of audio signals, one or more attributes of the set of loudspeakers, or one or more external inputs.

[0143] While particular embodiments and applications are described herein, it will be apparent to those skilled in the art that many variations of the embodiments and applications described herein are possible without departing from the scope described and claimed herein. While certain forms have been shown and described, it should be understood that the scope of the disclosure is not limited to the specific embodiments described and shown or to the specific methods described.

[0144] Several aspects will be described. [Aspect 1] 1. A method of audio processing comprising: receiving, by a control system via an interface system, audio data including one or more audio signals and associated spatial data, the spatial data indicating intended perceived spatial locations corresponding to the audio signals; rendering, by the control system, the audio data for playback through a set of loudspeakers in an environment to generate rendered audio signals, wherein rendering each of the one or more audio signals included in the audio data includes determining a relative activation of a set of loudspeakers in the environment by optimizing a cost, the cost being: a model of the perceived spatial location of the reproduced audio signal when reproduced on the set of loudspeakers in the environment; a measure of the proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker of the set of loudspeakers; and a step in which the step is a function of one or more additional dynamically configurable functions based on one or more of: the proximity of the loudspeaker to one or more listeners; the proximity of the loudspeaker to an attraction position, where attraction is a factor that favors relatively higher activation of loudspeakers closer to the attraction position; the proximity of the loudspeaker to a repulsion position, where repulsion is a factor that favors relatively lower activation of loudspeakers closer to the repulsion position; the performance of each loudspeaker relative to other loudspeakers in the environment; the synchronization of the loudspeaker with respect to other loudspeakers; wake word performance; or echo canceller performance; and providing, via the interface system, the rendered audio signals to at least some loudspeakers of the set of loudspeakers in the environment. Audio processing methods. [Aspect 2] 2. The audio processing method of aspect 1, wherein the model of perceived spatial location generates a binaural response corresponding to audio object location at the listener's left ear and right ear. Aspect 3 2. The audio processing method of claim 1, wherein the model of perceived spatial position places the perceived spatial position of an audio signal reproduced from a set of loudspeakers at the center of mass of the positions of the set of loudspeakers weighted by the associated activation gains of the loudspeakers. Aspect 4 4. The audio processing method of claim 3, wherein the model of perceived spatial position also generates binaural responses corresponding to audio object positions at the listener's left and right ears. Aspect 5 5. The audio processing method of any one of aspects 1 to 4, wherein the one or more additional dynamically configurable functions are based, at least in part, on levels of the one or more audio signals. Aspect 6 6. The audio processing method of any one of aspects 1 to 5, wherein the one or more additional dynamically configurable functions are based, at least in part, on a spectrum of the one or more audio signals. Aspect 7 7. The audio processing method of any one of aspects 1 to 6, wherein the one or more additional dynamically configurable features are based, at least in part, on a position of each loudspeaker in the environment. Aspect 8 8. The audio processing method of any one of aspects 1 to 7, wherein the capabilities of each loudspeaker include one or more of a frequency response, a playback level limit, or parameters of one or more loudspeaker dynamics processing algorithms. Aspect 9 9. The audio processing method of any one of aspects 1 to 8, wherein the one or more additional dynamically configurable features are based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the other loudspeakers. Aspect 10 10. The audio processing method of any one of aspects 1 to 9, wherein the one or more additional dynamically configurable functions are based, at least in part, on the position of one or more people in the environment. Aspect 11 An audio processing method as described in aspect 10, wherein the one or more additional dynamically configurable functions are based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the one or more person positions. Aspect 12 12. The audio processing method of any one of aspects 1 to 11, wherein the one or more additional dynamically configurable functions are based, at least in part, on object positions of one or more non-loudspeaker objects in the environment. Aspect 13 13. The audio processing method of claim 12, wherein the one or more additional dynamically configurable functions are based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the object position. Aspect 14 14. The audio processing method of any one of aspects 1 to 13, wherein the one or more additional dynamically configurable functions are based, at least in part, on an estimation of acoustic transmission from each speaker to one or more landmarks, areas, or zones of the environment. Aspect 15 15. The audio processing method of any one of aspects 1 to 14, wherein the intended perceived spatial position corresponds to at least one of channel or position metadata of a channel-based audio format. Aspect 16 16. A system configured to perform the method of any one of aspects 1 to 15. Aspect 17 One or more non-transitory media storing software including instructions for controlling one or more devices to perform the method of any one of aspects 1 to 15.

Claims

1. 1. A method of audio processing comprising: receiving, by a control system via an interface system, audio data including one or more audio signals and associated spatial data, the spatial data indicating intended perceived spatial locations corresponding to the audio signals; Rendering, by the control system, the audio data for playback through a set of loudspeakers in an environment to generate rendered audio signals, wherein rendering each of the one or more audio signals included in the audio data comprises at least in part adjusting an activation gain for each loudspeaker of the set of loudspeakers in the environment by: a model of the perceived spatial location of the reproduced audio signal when reproduced on the set of loudspeakers in the environment; a measure of the proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker of the set of loudspeakers; and One or more additional dynamically configurable features determining the one or more additional dynamically configurable functions based on: one or more characteristics of the audio signal, one or more characteristics of one or more loudspeakers, or one or more external inputs; and providing, via the interface system, the rendered audio signals to at least some loudspeakers of the set of loudspeakers in the environment. Audio processing methods.

2. 2. The audio processing method of claim 1, wherein the model of perceived spatial location generates a binaural response corresponding to audio object locations at the left and right ears of a listener.

3. 2. The audio processing method of claim 1, wherein the model of perceived spatial position places the perceived spatial position of an audio signal reproduced from a set of loudspeakers at the center of mass of the positions of the set of loudspeakers weighted by the associated activation gains of the loudspeakers.

4. 4. The audio processing method of claim 3, wherein the model of perceived spatial location also generates binaural responses corresponding to audio object locations at the left and right ears of a listener.

5. 10. The audio processing method of claim 1, wherein the one or more additional dynamically configurable features are based, at least in part, on the level of the one or more audio signals.

6. 10. The audio processing method of claim 1, wherein the one or more additional dynamically configurable features are based, at least in part, on a spectrum of the one or more audio signals.

7. 10. The audio processing method of claim 1, wherein the one or more additional dynamically configurable features are based, at least in part, on the position of each loudspeaker in the environment.

8. 2. The audio processing method of claim 1, wherein the one or more characteristics of the one or more loudspeakers include one or more of a frequency response, a playback level limit, or parameters of one or more loudspeaker dynamics processing algorithms.

9. 10. The audio processing method of claim 1, wherein the one or more additional dynamically configurable functions are based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the other loudspeakers.

10. The audio processing method of claim 1 , wherein the one or more additional dynamically configurable features are based, at least in part, on the location of one or more people in the environment.

11. 11. The audio processing method of claim 10, wherein the one or more additional dynamically configurable features are based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the one or more person positions.

12. 10. The audio processing method of claim 1, wherein the one or more additional dynamically configurable features are based, at least in part, on object positions of one or more non-loudspeaker objects in the environment.

13. 13. The audio processing method of claim 12, wherein the one or more additional dynamically configurable features are based, at least in part, on measurements or estimates of acoustic transmission from each loudspeaker to the object position.

14. 2. The audio processing method of claim 1, wherein the one or more additional dynamically configurable features are based, at least in part, on an estimation of acoustic transmission from each speaker to one or more landmarks, areas, or zones of the environment.

15. 2. The audio processing method of claim 1, wherein the intended perceived spatial location corresponds to at least one of channel or location metadata of a channel-based audio format.

16. 2. The audio processing method of claim 1, wherein the rendering comprises optimizing a cost that is a function of: the model of the perceived spatial location of the one or more audio signals when played on the set of loudspeakers in the environment; the indicator of proximity of the intended perceived spatial location of the one or more audio signals to a position of each loudspeaker of the set of loudspeakers; and the one or more additional dynamically configurable features.

17. 2. The audio processing method of claim 1, wherein the one or more external inputs include one or more of: one or more listener or speaker positions in the environment; measurements or estimates of acoustic transmission from each loudspeaker to a listening position; measurements or estimates of acoustic transmission from a speaker to the set of loudspeakers; environmental positions other than speaker, listener or loudspeaker positions in the environment; measurements or estimates of acoustic transmission from each loudspeaker to the environmental positions; or combinations thereof.

18. One or more non-transitory media storing instructions for controlling one or more devices to perform an audio processing method, the audio processing method comprising: receiving, by a control system via an interface system, audio data including one or more audio signals and associated spatial data, the spatial data indicating intended perceived spatial locations corresponding to the audio signals; Rendering, by the control system, the audio data for playback through a set of loudspeakers in an environment to generate rendered audio signals, wherein rendering each of the one or more audio signals included in the audio data comprises at least in part adjusting an activation gain for each loudspeaker of the set of loudspeakers in the environment by: a model of the perceived spatial location of the reproduced audio signal when reproduced on the set of loudspeakers in the environment; a measure of the proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker of the set of loudspeakers; and One or more additional dynamically configurable features determining the one or more additional dynamically configurable functions based on: one or more characteristics of the audio signal, one or more characteristics of one or more loudspeakers, or one or more external inputs; and providing, via the interface system, the rendered audio signals to at least some loudspeakers of the set of loudspeakers in the environment. Non-transient medium.

19. an interface system; Control system and 10. An apparatus comprising: receiving audio data via the interface system, the audio data including one or more audio signals and associated spatial data, the spatial data indicating intended perceived spatial locations corresponding to the audio signals; Rendering the audio data for playback through a set of loudspeakers in an environment to generate rendered audio signals, wherein rendering each of the one or more audio signals included in the audio data comprises at least in part adjusting an activation gain for each loudspeaker of the set of loudspeakers in the environment by: a model of the perceived spatial location of the reproduced audio signal when reproduced on the set of loudspeakers in the environment; a measure of the proximity of the intended perceived spatial location of the audio signal to the location of each loudspeaker of the set of loudspeakers; and determining based on one or more additional dynamically configurable functions, the one or more additional dynamically configurable functions being based on: one or more characteristics of the audio signal, one or more characteristics of one or more loudspeakers, or one or more external inputs; and providing, via the interface system, the rendered audio signals to at least some loudspeakers of the set of loudspeakers in the environment. Device.

20. 20. The apparatus of claim 19, wherein the one or more additional dynamically configurable functions are based, at least in part, on a location of one or more people in the environment.