Managing playback of multiple audio streams on multiple speakers

CN117499852BActive Publication Date: 2026-08-11DOLBY LABORATORIES LICENSING CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-27
Publication Date
2026-08-11

Smart Images

  • Figure CN117499852B_ABST
    Figure CN117499852B_ABST
Patent Text Reader

Abstract

A multi-stream rendering system and method can simultaneously render and play multiple audio program streams on multiple arbitrarily placed loudspeakers. At least one of the program streams can be a spatial mix. The rendering of the spatial mix can be dynamically adjusted based on the simultaneous rendering of one or more additional program streams.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese invention patent application filed on July 27, 2020, with application number 202080067801.3 and invention title "Managing playback of multiple audio streams on multiple speakers".

[0002] Inventors: Alan J. Seefeldt, Joshua B. Lando, Daniel Arteaga, Mark R.P. Thomas, Glenn N. Dickins

[0003] Cross-references to related applications

[0004] This application claims U.S. Provisional Patent Application No. 62 / 992,068, filed March 19, 2020; U.S. Provisional Patent Application No. 62 / 949,998, filed December 18, 2019; European Patent Application No. 19217580.0, filed December 18, 2019; Spanish Patent Application No. P201930702, filed July 30, 2019; U.S. Provisional Patent Application No. 62 / 971,421, filed February 7, 2020; U.S. Provisional Patent Application No. 62 / 705,410, filed June 25, 2020; and U.S. Provisional Patent Application No. 62 / 880,1, filed July 30, 2019. 11. Priority to U.S. Provisional Patent Application No. 62 / 704,754, filed May 27, 2020; U.S. Provisional Patent Application No. 62 / 705,896, filed July 21, 2020; U.S. Provisional Patent Application No. 62 / 880,114, filed July 30, 2019; U.S. Provisional Patent Application No. 62 / 705,351, filed June 23, 2020; U.S. Provisional Patent Application No. 62 / 880,115, filed July 30, 2019; and U.S. Provisional Patent Application No. 62 / 705,143, filed June 12, 2020, each of which is incorporated herein by reference in its entirety. Technical Field

[0005] This disclosure relates to systems and methods for playing back and rendering audio for playback from some or all of a group of speakers (e.g., each active speaker). Background Technology

[0006] Audio devices, including but not limited to smart audio devices, have been widely deployed and are becoming a common feature in many homes. While existing systems and methods for controlling audio devices offer benefits, improvements in systems and methods will continue to be desired.

[0007] Symbols and terms

[0008] Throughout this disclosure, including in the claims, the terms "speaker" and "loudspeaker" are used synonymously to refer to any sound-emitting transducer (or a group of transducers) driven by a single loudspeaker feed. A typical headphone assembly includes two loudspeakers. A loudspeaker can be implemented to include multiple transducers (e.g., a woofer and a tweeter), which can be driven by a single common loudspeaker feed or multiple loudspeaker feeds. In some examples, the loudspeaker feeds can undergo different processing in different circuit branches coupled to different transducers.

[0009] Throughout this disclosure, including in the claims, the expression “operating on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used broadly to refer to directly operating on a signal or data or operating on a processed version of a signal or data (e.g., a signal version that has been pre-filtered or pre-processed before being operated on).

[0010] Throughout this disclosure, including in the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem implementing a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, wherein the subsystem generates M inputs, and the other XM inputs are received from an external source) may also be referred to as a decoder system.

[0011] Throughout this disclosure, including in the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) for performing operations on data (e.g., audio, video, or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing of audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0012] Throughout this disclosure, including in the claims, the terms "couples" or "coupled" are used to refer to direct or indirect connections. Therefore, if a first device is coupled to a second device, the connection can be achieved either through a direct connection or through an indirect connection via other devices and connections.

[0013] As used herein, a “smart device” is an electronic device that can operate interactively and / or autonomously to some extent, typically configured to communicate with one or more other devices (or networks) via various wireless protocols such as Bluetooth, Zigbee, Near Field Communication, Wi-Fi, Li-Fi, 3G, 4G, and 5G. Several notable types of smart devices include smartphones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smartwatches, smart bracelets, smart keychains, and smart audio devices. The term “smart device” can also refer to devices that exhibit certain properties of ubiquitous computing, such as artificial intelligence.

[0014] The term "smart audio device" is used herein to refer to a smart device, which can be a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of a virtual assistant function). A single-purpose audio device is a device that includes or is coupled to at least one microphone (and optionally includes or is coupled to at least one speaker and / or at least one camera) and is largely or primarily designed to perform a single purpose (e.g., a television (TV) or a mobile phone). For example, while a TV can generally play (and is considered capable of playing) audio from program material, in most cases, modern TVs run some kind of operating system on which applications (including TV-watching applications) run natively. Similarly, audio input and output in a mobile phone can do many things, but these are served by applications running on the phone. In this sense, a single-purpose audio device with (multiple) speakers and (multiple) microphones is typically configured to run native applications and / or services to directly use said (multiple) speakers and (multiple) microphones. Some single-purpose audio devices can be configured to be combined to enable audio playback in a specific area or user-configured area.

[0015] A common type of multipurpose audio device is an audio device that implements at least some aspects of virtual assistant functionality, although other aspects of the virtual assistant functionality may be implemented by one or more other devices, such as one or more servers, with the multipurpose audio device configured to communicate with said one or more servers. Such a multipurpose audio device may be referred to herein as a “virtual assistant.” A virtual assistant is a device (e.g., a smart speaker or voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera). In some examples, a virtual assistant may provide the ability to use multiple devices (different from the virtual assistant) for applications that are, in a sense, cloud-enabled or otherwise not fully implemented in or on top of the virtual assistant itself. In other words, at least some aspects of the virtual assistant functionality (e.g., speech recognition) may (at least partially) be implemented by one or more servers or other devices, with the virtual assistant communicating with said one or more servers or other devices via a network (such as the Internet). Virtual assistants may sometimes work together, for example, in a discrete and conditionally defined manner. For example, two or more virtual assistants may work together in the sense that one of them (e.g., the virtual assistant most certain that it has heard the wake word) responds to the wake word. In some implementations, the connected virtual assistants can form a constellation, which can be managed by a main application, which may be (or implement) the virtual assistants.

[0016] In this document, "wake word" is used broadly to refer to any sound (e.g., a human-spoken word or other sound) in which a smart audio device is configured to wake up in response to the detection ("hearing") of a sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "wake up" means that the device enters a state of waiting (in other words, listening) for a sound command. In some instances, what may be referred to as a "wake word" in this document may include more than one word, such as a phrase.

[0017] In this document, the term "wake word detector" refers to a device (or software, including instructions for configuring the device to continuously search for alignments between real-time sound (e.g., speech) features and a training model. Typically, a wake word event is triggered whenever the wake word detector determines that the probability of detecting a wake word exceeds a predefined threshold. For example, this threshold might be a predetermined threshold adjusted to provide a reasonable trade-off between a false acceptance rate and a false rejection rate. Following a wake word event, the device may enter a state (which may be referred to as a "wake-up" state or an "attention" state) in which the device listens for commands and passes the received commands to a larger, more computationally intensive recognizer. Summary of the Invention

[0018] Some embodiments relate to a method for managing the playback of multiple audio streams by at least one (e.g., all or some) of a set of smart audio devices and / or by at least one (e.g., all or some) of another set of speakers.

[0019] One type of embodiment relates to a method for managing playback by at least one (e.g., all or some) of a plurality of coordinated (arranged) smart audio devices. For example, a set of smart audio devices present in a user's home (system) can be orchestrated to handle various simultaneous use cases, including flexibly rendering audio for playback by all or some of the smart audio devices (i.e., by all or some of the speakers of the smart audio devices).

[0020] Orchestrending smart audio devices (e.g., handling various simultaneous use cases at home) can involve simultaneously playing one or more audio program streams on a set of interconnected speakers. For example, a user might be listening to an Atmos movie soundtrack (or other object-based audio program) through a set of speakers (e.g., contained in or controlled by a set of smart audio devices), and then the user could speak a command (e.g., a wake word followed by a command) to an associated smart audio device (e.g., a smart assistant). In this case, the audio played back by the system can be corrected (according to some embodiments) to spatially distort the program (e.g., an Atmos mix) away from the speaker's (the user speaking) location and direct the corresponding response from the smart audio device (e.g., a voice assistant) to a speaker near the speaker. This can provide significant benefits compared to simply reducing the playback volume of the audio program content in response to the detection of a command (or the corresponding wake word). Similarly, a user might want to use speakers to get cooking tips in the kitchen while playing the same program (e.g., an Atmos soundtrack) in an adjacent open-plan living space. In this scenario, according to some embodiments, the playback of a program (e.g., an Atmos audio track) can be distorted and moved away from the kitchen, and cooking prompts can be played through speakers located near or within the kitchen. Additionally, cooking prompts played in the kitchen can be dynamically adjusted (according to some embodiments) to be heard by someone in the kitchen, and their sound can be louder than any program (e.g., an Atmos audio track) that might seep into the living space.

[0021] Some embodiments are multi-stream rendering systems configured to implement the example use cases described above, as well as many other contemplated example use cases. In one type of embodiment, the audio rendering system may be configured to render multiple audio program streams for simultaneous playback (and / or simultaneous streaming) on ​​multiple arbitrarily placed loudspeakers, wherein at least one of the program streams is a spatial mix, and the rendering (or rendering and playback) of the spatial mix is ​​dynamically adjusted in response to (or in combination with) the simultaneous playback (or rendering and playback) of one or more additional program streams.

[0022] Aspects of some implementations include a system configured (e.g., programmed) to perform any embodiment of the disclosed methods or steps thereof, and a tangible, non-transitory computer-readable medium (e.g., a disk or other tangible storage medium) that implements non-transitory storage of data, the tangible, non-transitory computer-readable medium storing code for performing any embodiment of the disclosed methods or steps thereof (e.g., code executable to perform any embodiment of the disclosed methods or steps thereof). For example, some embodiments may be or include a programmable general-purpose processor, digital signal processor, or microprocessor that is programmed with software or firmware to and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system including input devices, memory, and a processing subsystem programmed (and / or otherwise configured) to perform embodiments of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0023] At least some aspects of this disclosure can be implemented via apparatus. For example, one or more apparatuses may be able to perform at least partially the methods disclosed herein. In some embodiments, the apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof.

[0024] In some implementations, the control system includes or implements at least two rendering modules. According to some examples, the control system may include or implement N rendering modules, where N is an integer greater than 2.

[0025] In some examples, the first rendering module is configured to receive a first audio program stream via an interface system. In some instances, the first audio program stream includes a first audio signal arranged to be reproduced by at least some speakers in the environment. In some examples, the first audio program stream includes first spatial data, which includes channel data and / or spatial metadata. According to some examples, the first rendering module is configured to render the first audio signal for reproduction via speakers in the environment, thereby producing a first rendered audio signal.

[0026] In some implementations, the second rendering module is configured to receive a second audio program stream via an interface system. In some instances, the second audio program stream includes a second audio signal arranged to be reproduced by at least some speakers in the environment. In some examples, the second audio program stream includes second spatial data, which includes channel data and / or spatial metadata. According to some examples, the second rendering module is configured to render the second audio signal for reproduction via speakers in the environment, thereby producing a second rendered audio signal.

[0027] According to some examples, the first rendering module is configured to modify the rendering process for the first audio signal at least in part based on at least one of the second audio signal, the second rendered audio signal, or a characteristic thereof, to generate a modified first rendered audio signal. In some embodiments, the second rendering module is further configured to modify the rendering process for the second audio signal at least in part based on at least one of the first audio signal, the first rendered audio signal, or a characteristic thereof, to generate a modified second rendered audio signal.

[0028] In some embodiments, the audio processing system includes a mixing module configured to mix the modified first rendered audio signal and the modified second rendered audio signal to produce a mixed audio signal. In some examples, the control system is further configured to provide the mixed audio signal to at least some speakers in the environment.

[0029] According to some examples, the audio processing system may include one or more additional rendering modules. In some instances, each of the one or more additional rendering modules may be configured to receive an additional audio program stream via an interface system. The additional audio program stream may include an additional audio signal arranged to be reproduced by at least one speaker in the environment. In some instances, each of the one or more additional rendering modules may be configured to render the additional audio signal for reproduction via at least one speaker in the environment, thereby producing an additional rendered audio signal. In some instances, each of the one or more additional rendering modules may be configured to modify the rendering process for the additional audio signal at least in part based on at least one of the first audio signal, the first rendered audio signal, the second audio signal, the second rendered audio signal, or characteristics thereof, to produce a modified additional rendered audio signal. In some such examples, the mixing module may be further configured to mix the modified additional rendered audio signal with at least the modified first rendered audio signal and the modified second rendered audio signal to produce the mixed audio signal.

[0030] In some implementations, modifying the rendering process for the first audio signal may involve distorting the rendering of the first audio signal away from the rendering position of the second rendered audio signal. Alternatively or additionally, modifying the rendering process for the first audio signal may involve modifying the loudness of one or more of the first rendered audio signals in response to the loudness of the second audio signal or one or more of the second rendered audio signals.

[0031] According to some examples, correcting the rendering process for the second audio signal may involve distorting the rendering of the second audio signal away from the rendering position of the first rendered audio signal. Alternatively or additionally, correcting the rendering process for the second audio signal may involve correcting the loudness of one or more of the second rendered audio signals in response to the loudness of the first audio signal or one or more of the first rendered audio signals. According to some implementations, correcting the rendering process for the first audio signal and / or the second audio signal may involve performing spectral correction, audibility-based correction, and / or dynamic range correction.

[0032] In some examples, the audio processing system may include a microphone system comprising one or more microphones. In some such examples, a first rendering module may be configured to modify the rendering process for a first audio signal based at least in part on a first microphone signal from the microphone system. In some such examples, a second rendering module may be configured to modify the rendering process for a second audio signal based at least in part on the first microphone signal.

[0033] According to some examples, the control system may be further configured to estimate the location of a first sound source based on the first microphone signal and to correct the rendering process for at least one of the first audio signal or the second audio signal based at least in part on the location of the first sound source. In some examples, the control system may be further configured to determine whether the first microphone signal corresponds to ambient noise and to correct the rendering process for at least one of the first audio signal or the second audio signal based at least in part on whether the first microphone signal corresponds to ambient noise.

[0034] In some examples, the control system may be configured to determine whether a first microphone signal corresponds to human speech and to modify the rendering process for at least one of a first audio signal or a second audio signal, at least in part, based on whether the first microphone signal corresponds to human speech. According to some such examples, modifying the rendering process for the first audio signal may involve reducing the loudness of the first rendered audio signal reproduced by a speaker near the first sound source location, compared to the loudness of the first rendered audio signal reproduced by a speaker remote from the first sound source location.

[0035] According to some examples, the control system can be configured to determine that the first microphone signal corresponds to a wake-up word, to determine a response to the wake-up word, and to control at least one speaker near the first sound source location to reproduce the response. In some examples, the control system can be configured to determine that the first microphone signal corresponds to a command, to determine a response to the command, to control at least one speaker near the first sound source location to reproduce the response, and to execute the command. According to some examples, the control system can be further configured to, after controlling at least one speaker near the first sound source location to reproduce the response, revert to an uncorrected rendering process for the first audio signal.

[0036] In some implementations, the control system may be configured to obtain loudness estimates of the reproduced first audio program stream and / or the reproduced second audio program stream, at least in part, based on the first microphone signal. According to some examples, the control system may be further configured to modify the rendering process for at least one of the first audio signal or the second audio signal, at least in part, based on the loudness estimate. In some instances, the loudness estimate may be a perceived loudness estimate. According to some such examples, modifying the rendering process may involve altering at least one of the first audio signal or the second audio signal to maintain the perceived loudness of the first audio signal and / or the second audio signal in the presence of interference signals.

[0037] In some examples, the control system can be configured to determine that the first microphone signal corresponds to human speech and to reproduce the first microphone signal in one or more speakers near a location in the environment different from the location of the first sound source. According to some such examples, the control system can be further configured to determine whether the first microphone signal corresponds to a child's cry. In some such examples, the location of the environment may correspond to an estimated location of a caregiver.

[0038] According to some examples, the control system can be configured to obtain loudness estimates of a reproduced first audio program stream and / or a reproduced second audio program stream. In some such examples, the control system can be further configured to modify the rendering process for the first audio signal and / or the second audio signal, at least in part, based on the loudness estimate. According to some examples, the loudness estimate may be a perceived loudness estimate. Modifying the rendering process may involve altering at least one of the first audio signal or the second audio signal to maintain its perceived loudness in the presence of interfering signals.

[0039] In some implementations, rendering the first audio signal and / or rendering the second audio signal may involve flexible rendering to a speaker at an arbitrary location. In some such examples, the flexible rendering may involve centroid amplitude translation or flexible virtualization.

[0040] At least some aspects of this disclosure can be implemented via one or more audio processing methods. In some instances, the methods(s) can be implemented at least in part by control systems such as those disclosed herein. Some such methods involve receiving a first audio program stream by a first rendering module, the first audio program stream comprising a first audio signal arranged to be reproduced by at least some speakers in the environment. In some examples, the first audio program stream includes first spatial data, the first spatial data comprising channel data and / or spatial metadata. Some such methods involve rendering the first audio signal by the first rendering module for reproduction via speakers in the environment, thereby producing a first rendered audio signal.

[0041] Some such methods involve receiving a second audio program stream by a second rendering module. In some examples, the second audio program stream includes a second audio signal arranged to be reproduced by at least one speaker in the environment. Some such methods involve rendering the second audio signal by the second rendering module for reproduction via at least one speaker in the environment, thereby producing a second rendered audio signal.

[0042] Some such methods involve modifying the rendering process for the first audio signal by the first rendering module at least in part based on at least one of the second audio signal, the second rendered audio signal, or characteristics thereof, to generate a modified first rendered audio signal. Some such methods involve modifying the rendering process for the second audio signal by the second rendering module at least in part based on at least one of the first audio signal, the first rendered audio signal, or characteristics thereof, to generate a modified second rendered audio signal. Some such methods involve mixing the modified first rendered audio signal and the modified second rendered audio signal to generate a mixed audio signal, and providing the mixed audio signal to at least some speakers in the environment.

[0043] According to some examples, modifying the rendering process for a first audio signal may involve distorting the rendering of the first audio signal away from the rendering position of the second rendered audio signal and / or modifying the loudness of one or more of the first rendered audio signals in response to the loudness of the second audio signal or one or more of the second rendered audio signals.

[0044] In some examples, modifying the rendering process for the second audio signal may involve distorting the rendering of the second audio signal away from the rendering position of the first rendered audio signal and / or modifying the loudness of one or more of the second rendered audio signals in response to the loudness of the first audio signal or one or more of the first rendered audio signals.

[0045] According to some examples, correcting the rendering process for the first audio signal may involve performing spectral correction, audibility-based correction, and / or dynamic range correction.

[0046] Some methods may involve the first rendering module modifying the rendering process for the first audio signal based at least in part on a first microphone signal from a microphone system. Some methods may involve the second rendering module modifying the rendering process for the second audio signal based at least in part on the first microphone signal.

[0047] Some methods may involve estimating the location of a first sound source based on the first microphone signal and modifying the rendering process for at least one of a first audio signal or a second audio signal based at least in part on the location of the first sound source.

[0048] Some or all of the operations, functions, and / or methods described herein can be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Therefore, some innovative aspects of the subject matter described in this disclosure can be implemented in non-transitory media on which software is stored.

[0049] For example, the software may include instructions for controlling one or more devices to perform a method involving receiving a first audio program stream by a first rendering module, the first audio program stream comprising a first audio signal arranged to be reproduced by at least some speakers in the environment. In some examples, the first audio program stream includes first spatial data, the first spatial data including channel data and / or spatial metadata. Some such methods involve rendering the first audio signal by the first rendering module for reproduction via speakers in the environment, thereby producing a first rendered audio signal.

[0050] Some such methods involve receiving a second audio program stream by a second rendering module. In some examples, the second audio program stream includes a second audio signal arranged to be reproduced by at least one speaker in the environment. Some such methods involve rendering the second audio signal by the second rendering module for reproduction via at least one speaker in the environment, thereby producing a second rendered audio signal.

[0051] Some such methods involve modifying the rendering process for the first audio signal by the first rendering module at least in part based on at least one of the second audio signal, the second rendered audio signal, or characteristics thereof, to generate a modified first rendered audio signal. Some such methods involve modifying the rendering process for the second audio signal by the second rendering module at least in part based on at least one of the first audio signal, the first rendered audio signal, or characteristics thereof, to generate a modified second rendered audio signal. Some such methods involve mixing the modified first rendered audio signal and the modified second rendered audio signal to generate a mixed audio signal, and providing the mixed audio signal to at least some speakers in the environment.

[0052] According to some examples, modifying the rendering process for a first audio signal may involve distorting the rendering of the first audio signal away from the rendering position of the second rendered audio signal and / or modifying the loudness of one or more of the first rendered audio signals in response to the loudness of the second audio signal or one or more of the second rendered audio signals.

[0053] In some examples, modifying the rendering process for the second audio signal may involve distorting the rendering of the second audio signal away from the rendering position of the first rendered audio signal and / or modifying the loudness of one or more of the second rendered audio signals in response to the loudness of the first audio signal or one or more of the first rendered audio signals.

[0054] According to some examples, correcting the rendering process for the first audio signal may involve performing spectral correction, audibility-based correction, and / or dynamic range correction.

[0055] Some methods may involve the first rendering module modifying the rendering process for the first audio signal based at least in part on a first microphone signal from a microphone system. Some methods may involve the second rendering module modifying the rendering process for the second audio signal based at least in part on the first microphone signal.

[0056] Some methods may involve estimating the location of a first sound source based on the first microphone signal and modifying the rendering process for at least one of a first audio signal or a second audio signal based at least in part on the location of the first sound source.

[0057] Details of one or more embodiments of the subject matter described herein are set forth in the following figures and description. Other features, aspects, and advantages will become apparent from the description, figures, and claims. Note that the relative dimensions in the following figures may not be drawn to scale. Attached Figure Description

[0058] Figure 1A This is a block diagram illustrating examples of components of an apparatus capable of implementing various aspects of the present disclosure.

[0059] Figure 1B This is a block diagram of the minimum version of the embodiment.

[0060] Figure 2A Another (more capable) embodiment with additional features is described.

[0061] Figure 2B It outlines what can be achieved by, for example Figure 1A , Figure 1B or Figure 2A The flowchart shows an example of a method performed by the apparatus or system.

[0062] Figure 2C and Figure 2D This is a diagram illustrating a set of speaker activation and object rendering positions.

[0063] Figure 2E It outlines what can be achieved by, for example Figure 1A The flowchart shows an example of a method performed by the apparatus or system.

[0064] Figure 2F This is a diagram showing speaker activation in an example embodiment.

[0065] Figure 2G This is a diagram showing the object rendering locations in the example embodiment.

[0066] Figure 2H This is a diagram showing speaker activation in an example embodiment.

[0067] Figure 2I This is a diagram showing the object rendering locations in the example embodiment.

[0068] Figure 2J This is a diagram showing speaker activation in an example embodiment.

[0069] Figure 2H This is a diagram showing speaker activation in an example embodiment.

[0070] Figure 2I This is a diagram showing the object rendering locations in the example embodiment.

[0071] Figure 2J This is a diagram showing speaker activation in an example embodiment.

[0072] Figure 2K This is a diagram showing the object rendering locations in the example embodiment.

[0073] Figure 3A and Figure 3B An example floor plan of the connected living spaces is shown.

[0074] Figure 4A and Figure 4B An example of a multi-stream renderer that provides simultaneous playback of spatial music mixing and voice assistant responses is shown.

[0075] Figure 5A , Figure 5B and Figure 5C The illustration shows a third example use case of the disclosed multi-stream renderer.

[0076] Figure 6 It shows Figure 1B The example shown is a frequency domain / transform domain renderer in the image.

[0077] Figure 7 It shows Figure 2A The example shown is a frequency domain / transform domain renderer in the image.

[0078] Figure 8 An implementation of a multi-stream rendering system with an audio stream loudness estimator is shown.

[0079] Figure 9A An example of a multi-stream rendering system configured for cross-gradient rendering of multiple rendered streams is shown.

[0080] Figure 9B This is a diagram indicating the points where the speaker is activated in an example embodiment.

[0081] Figure 10 It is a graph based on trilinear interpolation between points indicating speaker activation, as shown in the example.

[0082] Figure 11 A floor plan of the listening environment is depicted, which in this example is a living space.

[0083] Figure 12A , Figure 12B , Figure 12C and Figure 12D Showing the target Figure 11 The examples shown are of multiple different listening positions and orientations in a living space, which are used to flexibly render spatial audio with reference to spatial patterns.

[0084] Figure 12E An example of reference space pattern rendering is shown when two listeners are in different locations within the listening environment.

[0085] Figure 13A An example of a graphical user interface (GUI) for receiving user input related to the listener's location and orientation is shown.

[0086] Figure 13B A distributed spatial rendering mode according to an example embodiment is described.

[0087] Figure 14A A partially distributed spatial rendering mode is described based on an example.

[0088] Figure 14B A fully distributed spatial rendering mode is described based on an example.

[0089] Figure 15 Example rendering locations on a 2D plane for a centroid amplitude translation (CMAP) and flexible virtualization (FV) rendering system are depicted.

[0090] Figure 16A , Figure 16B and Figure 16C It shows Figure 15 Distributed spatial patterns represented in the middle and Figure 16D Various examples of intermediate distributed space patterns between distributed space patterns are represented in the figure.

[0091] Figure 16D An example of distortion is depicted, which is applied to Figure 15 All rendering points are used to achieve a fully distributed rendering mode.

[0092] Figure 17 An example of a GUI that users can use to select the rendering mode is shown.

[0093] Figure 18 It is a flowchart outlining an example of a method that can be performed by those devices or systems disclosed herein.

[0094] Figure 19 An example of the geometric relationship between three audio devices in an environment is shown.

[0095] Figure 20 It shows Figure 19 Another example of the geometric relationship between three audio devices in an environment is shown in the image.

[0096] Figure 21A It shows Figure 19 and Figure 20 The two triangles depicted in the image do not correspond to any other audio devices or environmental features.

[0097] Figure 21B An example of estimating the interior angles of a triangle formed by three audio devices is shown.

[0098] Figure 22 It outlines what can be achieved by, for example Figure 1A The flowchart shows an example of a method performed by the apparatus.

[0099] Figure 23 An example is shown in which each audio device in the environment is the vertex of multiple triangles.

[0100] Figure 24 An example of a part of the forward alignment process is provided.

[0101] Figure 25 An example of multiple audio device position estimations that have already occurred during the forward alignment process is shown.

[0102] Figure 26 An example of a part of the reverse alignment process is provided.

[0103] Figure 27 An example of multiple audio device position estimations that have already occurred during the reverse alignment process is shown.

[0104] Figure 28 A comparison between the estimated audio device location and the actual audio device location is shown.

[0105] Figure 29 It outlines what can be achieved by, for example Figure 1A The flowchart shows an example of a method performed by the apparatus.

[0106] Figure 30A It shows Figure 29 Examples of some boxes.

[0107] Figure 30B Additional examples of determining listener angle orientation data are shown.

[0108] Figure 30C Additional examples of determining listener angle orientation data are shown.

[0109] Figure 30D It is shown according to the reference Figure 30C The described method is an example of determining the appropriate rotation of the audio device coordinates.

[0110] Figure 31 This is a block diagram illustrating examples of components of a system capable of implementing various aspects of this disclosure.

[0111] Figure 32A , Figure 32B and Figure 32C An example of playback limit thresholds and corresponding frequencies is shown.

[0112] Figure 33A and Figure 33B This is a diagram illustrating an example of dynamically range compressed data.

[0113] Figure 34 An example of a spatial area for listening to the environment is shown.

[0114] Figure 35 It shows Figure 34 An example of a loudspeaker within a spatial area.

[0115] Figure 36 Showing the coverage Figure 35 Examples of spatial zones and nominal spatial locations on speakers.

[0116] Figure 37 It is a flowchart outlining an example of a method that can be performed by those devices or systems disclosed herein.

[0117] Figure 38A , Figure 38B and Figure 38C It shows the relationship with Figure 2C and Figure 2D Examples of corresponding loudspeaker participation values.

[0118] Figure 39A , Figure 39B and Figure 39C It shows the relationship with Figure 2F and Figure 2G Examples of corresponding loudspeaker participation values.

[0119] Figure 40A , Figure 40B and Figure 40C It shows the relationship with Figure 2H and Figure 2I Examples of corresponding loudspeaker participation values.

[0120] Figure 41A , Figure 41B and Figure 41C It shows the relationship with Figure 2J and Figure 2K Examples of corresponding loudspeaker participation values.

[0121] Figure 42 This is a diagram of the environment, which in this example is a living space.

[0122] In the various figures, the same reference numerals and names indicate similar elements. Detailed Implementation

[0123] Flexible rendering is a technique for rendering spatial audio on any number of speakers placed arbitrarily. With the widespread deployment of smart audio devices (e.g., smart speakers) in homes, there is a need for flexible rendering techniques that allow consumers to use smart audio devices to perform flexible rendering of audio and to play back such rendered audio.

[0124] Several techniques have been developed to implement flexible rendering, including Centroid Amplitude Translation (CMAP) and Flexible Virtualization (FV). Both techniques treat the rendering problem as one of the minimization of a cost function, which consists of two terms: a first term simulating the desired spatial impression the renderer attempts to achieve, and a second term allocating the cost for activating the speakers. To date, this second term has focused on creating sparse solutions, where only speakers very close to the desired spatial location of the audio being rendered are activated.

[0125] Some embodiments of this disclosure are methods for managing the playback of multiple audio streams by at least one (e.g., all or some) of a group of smart audio devices (or by at least one (e.g., all or some) of another group of speakers).

[0126] One type of embodiment relates to a method for managing playback by at least one (e.g., all or some) of a plurality of coordinated (arranged) smart audio devices. For example, a set of smart audio devices present in a user's home (system) can be orchestrated to handle various simultaneous use cases, including flexibly rendering audio for playback by all or some of the smart audio devices (i.e., by all or some of the speakers of the smart audio devices).

[0127] Orchestrending smart audio devices (e.g., handling various simultaneous use cases in the home) can involve simultaneously playing one or more audio program streams on a set of interconnected speakers. For example, a user might be listening to a movie Atmos soundtrack (or other object-based audio program) on a set of speakers, but then the user might give a command to an associated smart assistant (or other smart audio device). In this case, the audio played back by the system can be corrected (according to some embodiments) to distort the spatial presentation of the Atmos mix away from the speaker's (the user speaking) location and away from the nearest smart audio device, while simultaneously distorting the corresponding response of the smart audio device (the voice assistant) towards the speaker's location. This can provide significant benefits compared to simply reducing the playback volume of the audio program content in response to the detection of a command (or the corresponding wake word). Similarly, a user might want to use speakers in the kitchen to get cooking prompts while playing the same Atmos soundtrack in adjacent open-plan living spaces. In this context, according to some examples, the Atmos track can be distorted away from the kitchen and / or the loudness of one or more rendered signals of the Atmos track can be adjusted in response to the loudness of one or more rendered signals of the cooking cue track. Additionally, in some implementations, cooking cue played in the kitchen can be dynamically adjusted to be heard by people in the kitchen, and its sound can be louder than any Atmos track that might seep into the living space.

[0128] Some embodiments relate to multi-stream rendering systems configured to implement the example use cases described above, as well as many other contemplated example use cases. In one type of embodiment, the audio rendering system can be configured to simultaneously play multiple audio program streams on multiple arbitrarily placed loudspeakers, wherein at least one of the program streams is a spatial mix, and the rendering of the spatial mix is ​​dynamically adjusted in response to (or in combination with) the simultaneous playback of one or more additional program streams.

[0129] In some embodiments, a multi-stream renderer can be configured to implement the scenarios described above, as well as many other situations where simultaneous playback of multiple audio program streams must be managed. Some implementations of a multi-stream rendering system can be configured to perform the following operations:

[0130] ● Simultaneously render and play back multiple audio program streams on multiple arbitrarily placed loudspeakers, wherein at least one of the program streams is a spatial mix.

[0131] The term "program stream" refers to a collection of one or more audio signals intended to be listened to as a whole. Examples include music, movie soundtracks, podcasts, live voice calls, and synthesized voice responses from intelligent assistants.

[0132] ○ "Spatial mixing" is a program stream (not just mono) designed to deliver different signals to the listener's left and right ears. Examples of audio formats used for spatial mixing include stereo, 5.1 and 7.1 surround sound, and object audio formats such as Dolby Atmos and Ambisonics.

[0133] ○ "Rendering" a program stream refers to the process of actively distributing one or more associated audio signals across multiple loudspeakers to achieve a specific perceptual impression.

[0134] ● Dynamically modify the rendering of at least one spatial mix based on the rendering of one or more of the additional program streams. Examples of such modification to the rendering of spatial mixes include, but are not limited to, those mentioned above.

[0135] ○ Correct the relative activation of multiple loudspeakers based on the relative activation of the loudspeakers associated with the rendering of at least one of one or more additional program streams.

[0136] ○ Distort the intended spatial balance of the spatial mix based on the spatial properties of the rendering of at least one of one or more additional program streams.

[0137] ○ The loudness or audibility of the spatial mix is ​​adjusted based on the loudness or audibility of at least one of the one or more additional program streams.

[0138] Figure 1AThis is a block diagram illustrating examples of components of an apparatus capable of implementing various aspects of the present disclosure. According to some examples, apparatus 100 may be or may include a smart audio device configured to perform at least some of the methods disclosed herein. In other embodiments, apparatus 100 may be or may include another device configured to perform at least some of the methods disclosed herein, such as a laptop computer, cellular phone, tablet device, smart home hub, etc. In some such embodiments, apparatus 100 may be or may include a server. In some embodiments, apparatus 100 may be configured to implement a device that may be referred to herein as an "audio session manager".

[0139] In this example, device 100 includes an interface system 105 and a control system 110. In some embodiments, interface system 105 may be configured to communicate with one or more devices executing or configured to execute a software application. Such a software application may be referred to herein as an "application" or simply an "app". In some embodiments, interface system 105 may be configured to exchange control information and associated data related to the application. In some embodiments, interface system 105 may be configured to communicate with one or more other devices in an audio environment. In some examples, the audio environment may be a home audio environment. In some embodiments, interface system 105 may be configured to exchange control information and data associated with audio devices in the audio environment. In some examples, the control information and associated data may relate to one or more applications with which device 100 is configured to communicate.

[0140] In some embodiments, the interface system 105 may be configured to receive an audio program stream. The audio program stream may include audio signals arranged to be reproduced by at least some speakers in the environment. The audio program stream may include spatial data such as channel data and / or spatial metadata. In some embodiments, the interface system 105 may be configured to receive input from one or more microphones in the environment.

[0141] Interface system 105 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some embodiments, interface system 105 may include one or more wireless interfaces. Interface system 105 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, interface system 105 may include a control system 110 and a memory system (such as...). Figure 1AOne or more interfaces between the optional memory systems 115 shown. However, in some instances, the control system 110 may include a memory system.

[0142] The control system 110 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0143] In some implementations, the control system 110 may reside in more than one device. For example, a portion of the control system 110 may reside in a device within one of the environments depicted herein, and another portion of the control system 110 may reside in a device outside the environment, such as a server, mobile device (e.g., a smartphone or tablet), etc. In other examples, a portion of the control system 110 may reside in a device within one of the environments depicted herein, and another portion of the control system 110 may reside in one or more other devices in the environment. For example, control system functionality may be distributed across multiple smart audio devices in the environment, or may be shared by an orchestration device (such as a device that may be referred to herein as a smart home hub) and one or more other devices in the environment. In some such examples, the interface system 105 may also reside in more than one device.

[0144] In some implementations, the control system 110 may be configured to perform at least partially the methods disclosed herein. According to some examples, the control system 110 may be configured to implement a method for managing the playback of multiple audio streams on multiple speakers.

[0145] Some or all of the methods described herein can be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media can include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. One or more non-transitory media can, for example, be located on... Figure 1A In the optional memory system 115 and / or control system 110 shown. Therefore, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media on which software is stored. For example, the software may include instructions for controlling at least one device to process audio data. For example, the software may be provided by, for example, Figure 1A The control system 110 and other control system components perform the operation.

[0146] In some examples, device 100 may include Figure 1AThe optional microphone system 120 is shown. The optional microphone system 120 may include one or more microphones. In some embodiments, the one or more microphones may be part of or associated with another device, such as a speaker in a speaker system, a smart audio device, etc. In some examples, the device 100 may not include the microphone system 120. However, in some such embodiments, the device 100 may still be configured to receive microphone data from one or more microphones in an audio environment via the interface system 110.

[0147] According to some embodiments, device 100 may include Figure 1A The optional loudspeaker system 125 is shown in the diagram. The optional loudspeaker system 125 may include one or more loudspeakers, which may also be referred to herein as “loudspeakers.” In some examples, at least some of the loudspeakers of the optional loudspeaker system 125 may be arbitrarily positioned. For example, at least some of the loudspeakers of the optional loudspeaker system 125 may be placed in locations that do not correspond to any standard-defined loudspeaker layout (such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, Hamasaki 22.2, etc.). In some such examples, at least some of the loudspeakers of the optional loudspeaker system 125 may be placed in locations that are convenient for space (e.g., where there is space to accommodate the loudspeakers), but not in any standard-defined loudspeaker layout. In some examples, the device 100 may not include the loudspeaker system 125.

[0148] In some embodiments, device 100 may include Figure 1A The optional sensor system 129 is shown. The optional sensor system 129 may include one or more cameras, touch sensors, gesture sensors, motion detectors, etc. According to some embodiments, the optional sensor system 129 may include one or more cameras. In some embodiments, the camera may be a stand-alone camera. In some examples, one or more cameras of the optional sensor system 129 may reside in a smart audio device, which may be a single-purpose audio device or a virtual assistant. In some such examples, one or more cameras of the optional sensor system 129 may reside in a TV, mobile phone, or smart speaker. In some examples, the device 100 may not include the sensor system 129. However, in some such embodiments, the device 100 may still be configured to receive sensor data from one or more sensors in the audio environment via the interface system 110.

[0149] In some embodiments, device 100 may include Figure 1AThe optional display system 135 is shown. The optional display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some instances, the optional display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples where the device 100 includes the display system 135, the sensor system 129 may include a touch sensor system and / or gesture sensor system for proximity to one or more displays of the display system 135. According to some such embodiments, the control system 110 may be configured to control the display system 135 to present one or more graphical user interfaces (GUIs).

[0150] According to some such examples, device 100 may be or may include a smart audio device. In some such embodiments, device 100 may be or may include a wake word detector. For example, device 100 may be or may include a virtual assistant.

[0151] Figure 1B This is a block diagram of a minimal version of the embodiment. It depicts N program streams (N≥1), where the first is explicitly labeled as spatial, and its corresponding audio signal set is fed through a corresponding renderer, each individually configured to play back its corresponding program stream through a common set of M arbitrarily spaced loudspeakers (M≥2). The renderer may also be referred to herein as a "rendering module". The rendering module and mixer 130a can be implemented via software, hardware, firmware, or some combination thereof. In this example, the rendering module and mixer 130a are implemented via a control system 110a, which is referenced above. Figure 1AAn example of the described control system 110 is provided. The output of each of N renderers is fed across a set of M loudspeakers, summed across all N renderers, for simultaneous playback on the M loudspeakers. According to this embodiment, information regarding the layout of the M loudspeakers within the listening environment is provided to all renderers, indicated by dashed lines returned from the loudspeaker frames, so that the renderers can be correctly configured for playback via loudspeakers. This layout information may or may not be sent from one or more loudspeakers themselves, depending on the specific implementation. According to some examples, the layout information may be provided by one or more smart loudspeakers configured to determine the relative position of each of the M loudspeakers in the listening environment. Some such automatic positioning methods may be based on a direction-of-arrival (DOA) method or a time-of-arrival (TOA) method. In other examples, the layout information may be determined by another device and / or input by a user. In some examples, loudspeaker specification information regarding the capabilities of at least some of the M loudspeakers within the listening environment may be provided to all renderers. Such loudspeaker specifications may include impedance, frequency response, sensitivity, rated power, the number and location of individual drivers, etc. According to this example, rendering information from one or more additional program streams is fed into the renderer of the main spatial stream, allowing the rendering to be dynamically adjusted based on this information. This information is represented by a dashed line running from render box 2 to render box N and back up to render box 1.

[0152] Figure 2A Another (more capable) embodiment with additional features is depicted. In this example, the rendering module and mixer 130b are implemented via a control system 110b, which is referenced above. Figure 1AAn example of the described control system 110. In this version, the dashed lines traveling up and down among all N renderers represent the idea that any one of the N renderers can contribute to the dynamic correction of any of the remaining N-1 renderers. In other words, the rendering of any one of the N program streams can be dynamically corrected based on a combination of one or more renderings of any of the remaining N-1 program streams. Additionally, any one or more program streams can be spatial mixes, and the rendering of any program stream, whether spatial or not, can be dynamically corrected based on any of the other program streams. For example, as described above, loudspeaker layout information can be provided to the N renderers. In some examples, loudspeaker specification information can be provided to the N renderers. In some implementations, microphone system 120a can include a set of K microphones (K≥1) within the listening environment. In some examples, the microphones can be attached to or associated with one or more loudspeakers. These microphones can feed both their captured audio signals (represented by solid lines) and additional configuration information (e.g., their location) (represented by dashed lines) back to the set of N renderers. Any of the N renderers can then be dynamically adjusted based on the additional microphone input. Various examples are provided in this article.

[0153] Examples of information obtained from microphone input and subsequently used to dynamically correct any of the N renderers include, but are not limited to:

[0154] ● Detection of specific words or phrases spoken by users of the system.

[0155] ●Estimation of the location of one or more users in the system.

[0156] ●Estimation of the loudness of any combination of N program streams at a specific location in the listening space.

[0157] ●Estimation of the loudness of other ambient sounds (such as background noise) in the listening environment.

[0158] Figure 2B It outlines what can be achieved by, for example Figure 1A , Figure 1B or Figure 2A The flowchart illustrates an example of a method performed by an apparatus or system. As with other methods described herein, the blocks of method 200 need not be performed in the indicated order. Furthermore, this method may include more or fewer blocks than those shown and / or described. The blocks of method 200 may be performed by one or more devices, which may be (or may include) a control system, such as... Figure 1A , Figure 1B and Figure 2AThe control system 110, control system 110a or control system 110b shown and described above, or one of other disclosed control system examples.

[0159] In this embodiment, block 205 relates to receiving a first audio program stream via an interface system. In this example, the first audio program stream includes a first audio signal arranged to be reproduced by at least some speakers in the environment. Here, the first audio program stream includes first spatial data. According to this example, the first spatial data includes channel data and / or spatial metadata. In some examples, block 205 relates to controlling a first rendering module of the system that receives the first audio program stream via the interface system.

[0160] According to this example, block 210 relates to rendering a first audio signal for reproduction via ambient speakers, thereby generating a first rendered audio signal. For example, as described above, some examples of method 200 involve receiving loudspeaker layout information. For example, as described above, some examples of method 200 involve receiving loudspeaker specification information. In some examples, the first rendering module may generate the first rendered audio signal at least in part based on the loudspeaker layout information and / or the loudspeaker specification information.

[0161] In this example, block 215 relates to receiving a second audio program stream via an interface system. In this embodiment, the second audio program stream includes a second audio signal arranged to be reproduced by at least some speakers in the environment. According to this example, the second audio program stream includes second spatial data. The second spatial data includes channel data and / or spatial metadata. In some examples, block 215 relates to controlling a second rendering module of the system that receives the second audio program stream via the interface system.

[0162] According to this embodiment, block 220 relates to rendering a second audio signal for reproduction via ambient speakers, thereby generating a second rendered audio signal. In some examples, the second rendering module may generate the second rendered audio signal at least in part based on received loudspeaker layout information and / or received loudspeaker specification information.

[0163] In some instances, some or all of the speakers in an environment can be positioned arbitrarily. For example, at least some of the speakers in an environment may be placed in locations that do not correspond to any standard-defined speaker layout (such as Dolby 5.1, Dolby 7.1, Hamasaki 22.2, etc.). In some such examples, at least some of the speakers in an environment may be placed in locations convenient for the environment's furniture, walls, etc. (e.g., where there is space to accommodate the speakers), without adhering to any standard-defined speaker layout.

[0164] Therefore, some implementation blocks 210 or 220 may involve flexible rendering to speakers at arbitrary locations. Some such implementations may involve centroid amplitude translation (CMAP), flexible virtualization (FV), or a combination of both. At a high level, both techniques render a set of one or more audio signals (each audio signal having an associated desired perceived spatial location) for playback on two or more speakers in a set, wherein the relative activation of the set of speakers is a model of the perceived spatial location of the audio signals being played back by the speakers and a function of the proximity of the desired perceived spatial location of the audio signals to the speaker locations. The model ensures that the listener hears the audio signals near their desired spatial location, and the proximity term controls which speakers are used to achieve that spatial impression. Specifically, the proximity term favors activating speakers that are close to the desired perceived spatial location of the audio signals. For both CMAP and FV, this functional relationship can be conveniently derived from a cost function, written as the sum of two terms, one for the spatial aspect and one for the proximity:

[0165]

[0166] Here, collection This indicates the positions of a group of M loudspeakers. Let g represent the desired perceived spatial location of the audio signal, and g represent the M-dimensional vector of speaker activations. For CMAP, each activation in the vector represents the gain of each speaker, while for FV, each activation represents a filter (in the second case, g can be equivalently viewed as a vector of complex values ​​at a specific frequency, and different g values ​​are computed across multiple frequencies to form filters). The optimal vector of activations is found by minimizing the cost function across the activations:

[0167]

[0168] Under certain definitions of the cost function, it is difficult to control the absolute level of the optimal activation generated by the above minimization, although g opt The relative levels between the components are appropriate. To resolve this issue, g can be executed. opt Subsequent normalization is performed to control the absolute level of activation. For example, it might be desirable to normalize the vector to have a unit length, which conforms to the commonly used constant power shift rule:

[0169]

[0170] The exact behavior of the flexible rendering algorithm depends on the C of the cost function. spatial and C proximity The specific construction of these two items. For CMAP, C spatialIt is derived from a model that places the perceptual spatial location of an audio signal played from a set of loudspeakers into the associated activation gain g of those loudspeakers. i The centroids of the positions of these loudspeakers, weighted by the elements of vector g:

[0171]

[0172] Equation 3 is then manipulated to represent the space cost of the squared error between the desired audio position and the audio position produced by the activated loudspeaker:

[0173]

[0174] For FV, the spatial term of the cost function is defined differently. The goal is to generate values ​​at the listener's left and right ears that correspond to the location of the audio object. The corresponding binaural response b. Conceptually, b is a 2×1 vector of filters (one filter for each ear), but it is more convenient to consider it as a complex 2×1 vector at a specific frequency. Continuing this representation at specific frequencies, the desired binaural response can be obtained from a set of HRTF indices based on the object's location:

[0175]

[0176] Meanwhile, the 2×1 binaural response e generated by the loudspeaker at the listener's ear is modeled as an M×1 vector g multiplied by a 2×M acoustic transmission matrix H and the complex loudspeaker activation values:

[0177] e = Hg (6)

[0178] The acoustic transmission matrix H is a set based on the loudspeaker positions. Modeled relative to the listener's position. Finally, the spatial component of the cost function is defined as the squared error between the desired binaural response (Equation 5) and the binaural response produced by the amplifier (Equation 6):

[0179]

[0180] Conveniently, the space terms of the cost functions for CMAP and FV defined in both Equations 4 and 7 can be rearranged into matrix quadratic functions as functions of the loudspeaker activation g:

[0181]

[0182] Here, A is an M×M square matrix, B is a 1×M vector, and C is a scalar. The rank of matrix A is 2, and therefore, when M>2, there exist infinitely many loudspeaker activations g with spatial error terms equal to zero. A second term C is introduced into the cost function. proximityThis uncertainty was eliminated, and a specific solution with perceptually beneficial properties compared to other possible solutions was generated. For both CMAP and FV, C... proximity Constructed to make position Far from the desired audio signal location The activation of speakers located close to the desired position is penalized more than that of speakers located close to the desired position. This configuration produces an optimal set of sparse speaker activations, where only speakers located close to the desired audio signal are significantly activated, effectively resulting in spatial reproduction of the audio signal, which is perceptibly more robust to listener movement around the set of speakers.

[0183] Therefore, the second term C of the cost function proximity It can be defined as a distance-weighted sum of the squares of the absolute values ​​of speaker activation. This can be concisely represented in matrix form as follows:

[0184]

[0185] Where D is a diagonal matrix representing the desired audio location and the distance penalty between each speaker:

[0186]

[0187] The distance penalty function can take many forms, but the following are useful parameterizations:

[0188]

[0189] in, d0 is the Euclidean distance between the desired audio location and the speaker location, and α and β are adjustable parameters. Parameter α indicates the global intensity of the penalty; d0 corresponds to the spatial range of the distance penalty (speakers at or beyond a distance of approximately d0 will be penalized), and β explains the abruptness of the penalty initiation at a distance of d0.

[0190] Combining the two terms of the cost function defined in equations 8 and 9a, we obtain the overall cost function:

[0191] C(g)=g * AB+Bg+C+g * Dg = g * (A+D)g+Bg+C (10)

[0192] Setting the derivative of the cost function with respect to g to zero and solving for g yields the optimal loudspeaker activation solution:

[0193]

[0194] Typically, the optimal solution in Equation 11 can produce speaker activations with negative values. For CMAP constructions of flexible renderers, such negative activations may be undesirable, and therefore Equation (11) can be minimized while keeping all activations positive.

[0195] Figure 2C and Figure 2D This is a diagram illustrating a set of speaker activation and object rendering positions as examples. In these examples, the speaker activation and object rendering positions correspond to speaker positions of 4, 64, 165, -87, and -4 degrees. Figure 2C Speaker activations 245a, 250a, 255a, 260a, and 265a are shown, including the optimal solution of Equation 11 for these specific speaker locations. Figure 2D Individual speaker positions are drawn as squares 267, 270, 272, 274, and 275, which correspond to speaker activations 245a, 250a, 255a, 260a, and 265a, respectively. Figure 2D The ideal object position (in other words, the position where the audio object is to be rendered) for a large number of possible object angles is also shown as point 276a, and the corresponding actual rendering position for these objects is shown as point 278a, connected to the ideal object position by a dashed line 279a.

[0196] One type of embodiment relates to a method for rendering audio for playback by at least one (e.g., all or some) of a plurality of coordinated (arranged) smart audio devices. For example, a set of smart audio devices present in a user's home (system) can be orchestrated to handle various simultaneous use cases, including flexibly rendering (according to embodiments) audio for playback by all or some of the smart audio devices (i.e., by all or some of the speakers of the smart audio devices). Numerous interactions with the system are considered, which require dynamic adjustments to the rendering. Such adjustments may, but do not necessarily, focus on spatial fidelity.

[0197] Some embodiments are methods for rendering audio for playback by at least one (e.g., all or some) of a set of smart audio devices (or by at least one (e.g., all or some) of another set of speakers). Rendering may include minimizing a cost function, wherein the cost function includes at least one dynamic speaker activation term. Examples of such dynamic speaker activation terms include (but are not limited to):

[0198] ● The proximity of the speaker to one or more listeners;

[0199] ● The proximity of the speaker to an attractive or repulsive force;

[0200] ● The audibility of the speaker with respect to certain locations (e.g., the listener's location or a nursery);

[0201] ● The speaker's capabilities (e.g., frequency response and distortion);

[0202] ● The speaker is synchronized with other speakers;

[0203] ●Wake word performance; and

[0204] ●Echo canceller performance.

[0205] Multiple dynamic speaker activations can enable at least one of a variety of behaviors, including spatially distorting the audio presentation away from a particular smart audio device, so that the microphone of the particular smart audio device can hear the speaker better, or so that secondary audio streams can be heard better from the multiple speakers of the smart audio device.

[0206] Some embodiments implement rendering for playback using the speakers of multiple smart audio devices in a coordinated (arranged) manner. Other embodiments implement rendering for playback using the speakers of another set of speakers.

[0207] Pairing a flexible rendering method (implemented according to some embodiments) with a set of wireless smart speakers (or other smart audio devices) can produce a highly capable and easy-to-use spatial audio rendering system. When considering interaction with such a system, it is clearly desirable to dynamically adjust the spatial rendering to optimize for additional objectives that may arise during system use. To achieve this, one set of embodiments enhances existing flexible rendering algorithms (where speaker activation is a function of previously disclosed spatial and proximity terms) with one or more additional dynamically configurable features that depend on one or more properties of the audio signal being rendered, one or more properties of the speaker group, and / or other external inputs. According to some embodiments, the cost function of existing flexible rendering given in Equation 1 is augmented with these one or more additional dependencies according to the following equation:

[0208]

[0209] In equation 12, the term This indicates additional cost items, and One or more properties representing a set of audio signals being rendered (e.g., object-based audio programs). This represents one or more properties of a set of speakers that are rendering audio, and This represents one or more additional external inputs. Each item... The return cost, as a function of the activation g, is generally derived from a set of parameters related to one or more properties of the audio signal, one or more properties of the speaker, and / or external inputs. This indicates that a set... It should be understood that... At least includes from or One of the elements of any one of them.

[0210] Examples include, but are not limited to:

[0211] ●The expected perceived spatial location of the audio signal;

[0212] ● The level of the audio signal (which may vary over time); and / or

[0213] ● The spectrum of the audio signal (which may vary over time).

[0214] Examples include, but are not limited to:

[0215] ● The location of the loudspeaker in the listening space;

[0216] ● The frequency response of the loudspeaker;

[0217] ● Reproduction level limitations of the loudspeaker;

[0218] ● Parameters of the speaker's internal dynamic processing algorithm, such as limiter gain;

[0219] ● Measurement or estimation of acoustic transmission from each loudspeaker to other loudspeakers;

[0220] ●Measurements of the performance of the echo canceller on the loudspeaker; and / or

[0221] ● The speakers are in relative synchronization with each other.

[0222] Examples include, but are not limited to:

[0223] ●Replay the location of one or more listeners or speakers in the playback space;

[0224] ● Measurement or estimation of acoustic transmission from each loudspeaker to the listening position;

[0225] ● Measurement or estimation of acoustic transmission from the speaker to the loudspeaker assembly;

[0226] ● Replay the locations of other landmarks in the space; and / or

[0227] ● Measurements or estimates of acoustic transmission from each loudspeaker to some other landmark in the playback space;

[0228] Using the new cost function defined in Equation 12, the optimal activation set can be found by minimizing and possibly post-normalizing with respect to g as previously specified in Equations 2a and 2b.

[0229] Figure 2E It outlines what can be achieved by, for example Figure 1A The flowchart illustrates an example of a method performed by the apparatus or system shown herein. As with other methods described herein, the blocks of method 280 need not be performed in the indicated order. Furthermore, this method may include more or fewer blocks than those shown and / or described. The blocks of method 280 may be performed by one or more devices, which may be (or may include) a control system, such as... Figure 1A The control system 110 shown in the figure.

[0230] In this embodiment, block 285 relates to receiving audio data by a control system and via an interface system. In this example, the audio data includes one or more audio signals and associated spatial data. According to this embodiment, the spatial data indicates a desired perceived spatial location corresponding to the audio signals. In some instances, the desired perceived spatial location may be explicit, for example, indicated by location metadata such as Dolby Atmos location metadata. In other instances, the desired perceived spatial location may be implicit, for example, a hypothetical location associated with a channel according to Dolby 5.1, Dolby 7.1, or other channel-based audio formats. In some examples, block 285 relates to a rendering module of the control system that receives the audio data via an interface system.

[0231] According to this example, block 290 relates to a control system rendering audio data for reproduction via a set of loudspeakers in the environment, thereby generating a rendered audio signal. In this example, rendering each of one or more audio signals included in the audio data involves determining the relative activation of the set of loudspeakers in the environment by optimizing a cost function. According to this example, when played back on said set of loudspeakers in the environment, the cost is a function of a model of the perceived spatial location of the audio signal. In this example, the cost is also a function of a measure of the proximity of the expected perceived spatial location of the audio signal to the location of each loudspeaker in the set. In this implementation, the cost is also a function of one or more additional dynamically configurable functions. In this example, the dynamically configurable functionality is based on one or more of the following: the proximity of the loudspeaker to one or more listeners; the proximity of the loudspeaker to an attraction location, where attraction is a factor that favors the activation of a relatively higher loudspeaker closer to the attraction location; the proximity of the loudspeaker to a repulsion location, where repulsion is a factor that favors the activation of a relatively lower loudspeaker closer to the repulsion location; the capability of each loudspeaker relative to other loudspeakers in the environment; the synchronization of the loudspeaker with respect to other loudspeakers; wake word performance; or echo canceller performance.

[0232] In this example, box 295 relates to providing a rendered audio signal to at least some of the loudspeakers in the set of loudspeakers in the environment via an interface system.

[0233] According to some examples, a model of perceived spatial location can produce a binaural response corresponding to the location of an audio object at the listener's left and right ears. Alternatively or additionally, a model of perceived spatial location can place the perceived spatial location of an audio signal played from a set of loudspeakers at the centroid of the location of the set of loudspeakers, weighted by the associated activation gain of the loudspeakers.

[0234] In some examples, one or more additional dynamically configurable features may be based at least in part on the levels of one or more audio signals. In some instances, one or more additional dynamically configurable features may be based at least in part on the spectrum of one or more audio signals.

[0235] Some examples of method 280 involve receiving loudspeaker layout information. In some examples, one or more additional dynamically configurable features may be based at least in part on the location of each loudspeaker in the environment.

[0236] Some examples of method 280 involve receiving loudspeaker specification information. In some examples, one or more additional dynamically configurable features may be based at least in part on the capabilities of each loudspeaker, which may include one or more of the following: frequency response, playback level limit, or parameters of one or more loudspeaker dynamic processing algorithms.

[0237] According to some examples, one or more additional dynamic configurable functions may be based at least in part on measurements or estimates of acoustic transmission from each loudspeaker to the other loudspeakers. Alternatively or additionally, one or more additional dynamic configurable functions may be based at least in part on the listener or speaker locations of one or more people in the environment. Alternatively or additionally, one or more additional dynamic configurable functions may be based at least in part on measurements or estimates of acoustic transmission from each loudspeaker to the listener or speaker location. The estimation of acoustic transmission may be based, for example, at least in part on walls, furniture, or other objects that may reside between each loudspeaker and the listener or speaker location.

[0238] Alternatively or additionally, one or more additional dynamic configurable functions may be based at least in part on the object location of one or more non-loudspeaker objects or landmarks in the environment. In some such implementations, one or more additional dynamic configurable functions may be based at least in part on measurements or estimates of acoustic transmission from each loudspeaker to the object location or landmark location.

[0239] Flexible rendering can be implemented to achieve many new and useful behaviors by employing one or more appropriately defined additional cost items. All example behaviors listed below pertain to penalizing certain loudspeakers under conditions deemed undesirable. The end result is that these loudspeakers are less activated in the spatial rendering of a set of audio signals. In many of these cases, one might consider simply turning down the undesirable loudspeakers, regardless of any modifications to the spatial rendering, but this strategy can significantly degrade the overall balance of the audio content. For example, certain components of the mix might become completely inaudible. On the other hand, for the disclosed embodiments, integrating these penalties into the core optimizations of the rendering allows the rendering to adapt and perform the best possible spatial rendering using the remaining, less penalized loudspeakers. This is a more elegant, adaptable, and efficient solution.

[0240] Example use cases include, but are not limited to:

[0241] ● Provides a more balanced spatial presentation around the listening area

[0242] It has been found that spatial audio is best presented across loudspeakers that are approximately equidistant from the intended listening area. Costs can be configured such that loudspeakers that are significantly closer or farther from the listening area than the average distance to the listening area are penalized, thus reducing the activation of those loudspeakers.

[0243] ● Move the audio away from or toward the listener or speaker.

[0244] If a user of the system is trying to speak to the system or an intelligent voice assistant associated with the system, the cost of creating a loudspeaker that is closer to the speaker may be worthwhile. This way, these loudspeakers are activated less, allowing their associated microphones to hear the speaker better.

[0245] In order to provide a more intimate experience for individual listeners, i.e., to minimize the playback level for others in the listening space, speakers located far from the listener may be severely penalized so that only the speakers closest to the listener are most significantly activated;

[0246] ●Move the audio source away from or toward a landmark, district, or area.

[0247] Certain locations near the listening space may be considered sensitive, such as a nursery, crib, office, reading area, or study area. In such cases, the cost of using speakers close to that location, area, or zone can be penalized.

[0248] Alternatively, for the same (or similar) situation described above, the loudspeaker system may have already generated measurements of the acoustic transmission from each loudspeaker to the nursery, particularly when one of the loudspeakers (with an attached or associated microphone) resides within the nursery. In this case, instead of using the physical proximity of the loudspeakers to the nursery, it is possible to construct a system that penalizes the cost of the loudspeakers by using the measured acoustic transmission to the room; and / or

[0249] ●Optimal use of speaker capabilities

[0250] The capabilities of different loudspeakers can vary significantly. For example, a popular smart speaker may contain only a single 1.6” full-range driver with limited low-frequency capability. On the other hand, another smart speaker may contain a more capable 3” woofer. These capabilities are typically reflected in the speaker’s frequency response, and thus, a set of responses associated with the speaker can be utilized in the cost item. At a given frequency, a speaker with weaker capabilities relative to other speakers (as measured by its frequency response) is penalized and therefore activated to a lesser extent. In some implementations, this frequency response value can be stored by the smart loudspeaker and then reported to a computing unit responsible for optimizing flexible rendering;

[0251] Many loudspeakers contain more than one driver, each responsible for playing a different frequency range. For example, a popular smart loudspeaker is a two-way design, containing a woofer for lower frequencies and a tweeter for higher frequencies. Typically, such loudspeakers include crossover circuitry for dividing the full-range playback audio signal into appropriate frequency ranges and sending them to the corresponding drivers. Alternatively, such loudspeakers can provide flexible renderer playback access for each individual driver, along with information about the capabilities of each individual driver, such as frequency response. By applying the cost terms described above, in some examples, the flexible renderer can automatically establish a crossover between the two drivers based on their relative capabilities at different frequencies.

[0252] The frequency response examples described above focus on the inherent capabilities of a loudspeaker but may not accurately reflect the capabilities of a loudspeaker placed in a listening environment. In some cases, such as when measuring a loudspeaker's frequency response at the intended listening location, this can be obtained through calibration procedures. Such measurements can be used instead of pre-calculated responses to better optimize loudspeaker usage. For example, a loudspeaker may be inherently very capable at a particular frequency, but its placement (e.g., behind a wall or piece of furniture) might result in a very limited response at the intended listening location. Capturing this response and feeding it into a measurement with an appropriate cost term can prevent significant activation of such a loudspeaker;

[0253] Frequency response is only one aspect of a loudspeaker's playback capability. Many smaller loudspeakers begin to distort and then reach their offset limits as the playback level increases, especially at lower frequencies. To reduce this distortion, many loudspeakers implement dynamic processing that limits the playback level below certain thresholds that can vary with frequency. It makes sense to reduce the signal level in the limited loudspeaker and transfer that energy to the other, less burdened loudspeakers when the loudspeaker is close to or at these thresholds while other loudspeakers involved in flexible rendering are not. According to some embodiments, this behavior can be automated by appropriately configuring the associated cost item. This cost item may involve one or more of the following:

[0254] ■ Monitor global playback volume in relation to the loudspeaker's limit threshold. For example, loudspeakers with volume levels close to their limit threshold may be penalized more.

[0255] ■ Monitor dynamic signal levels (which may also vary with frequency) related to the loudspeaker's limiting threshold. For example, loudspeakers with monitored signal levels close to their limiting threshold may be subject to more penalties;

[0256] ■ Parameters that directly monitor the dynamic processing of the loudspeaker, such as gain limiting. In some such examples, parameters indicating more limiting may result in loudspeakers being penalized more; and / or

[0257] ■ Monitor the actual instantaneous voltage, current, and power delivered by the amplifier to the loudspeaker to determine if the loudspeaker is operating within its linear range. For example, a loudspeaker operating less linearly may be penalized more severely;

[0258] Smart speakers with integrated microphones and interactive voice assistants typically employ some type of echo cancellation to reduce the level of the audio signal played by the speaker, picked up by the recording microphone. The greater the reduction, the greater the speaker's chance of hearing and understanding speakers in the space. If the residual of the echo canceller is consistently high, this can indicate that the speaker is being driven into a non-linear region where predicting the echo path becomes challenging. In such cases, it can be meaningful to transfer signal energy away from the speaker, and thus, it can be beneficial to consider the cost of the echo canceller's performance. Such a cost could allocate high costs to speakers with poorly performing echo cancellers.

[0259] To achieve predictable imaging when rendering spatial audio across multiple speakers, playback across a set of loudspeakers typically needs to be reasonably synchronized over time. This is given for wired loudspeakers, but synchronization can be challenging and the end result variable for a large number of wireless loudspeakers. In such cases, it may be possible for each loudspeaker to report its relative degree of synchronization with the target, and this degree can then be fed into a synchronization cost term. In some such examples, loudspeakers with lower levels of synchronization may be penalized more and thus excluded from rendering. Additionally, certain types of audio signals may not require tight synchronization, such as components of audio mixes designed for diffusion or non-directional transmission. In some implementations, components can be tagged with metadata, and the synchronization cost term can be adjusted to reduce the penalty.

[0260] The following describes an example of an embodiment.

[0261] Similar to the proximity costs defined in equations 9a and 9b, each new cost function term... It is also convenient to express this as a weighted sum of the squares of the absolute values ​​of speaker activation:

[0262]

[0263] Among them, W j Weight The diagonal matrix describes the cost associated with speaker i for activation term j:

[0264]

[0265] Combining Equations 13a and 13b with the matrix quadratic versions of the CMAP and FV cost functions given in Equation 10 yields a potentially beneficial implementation of the general extended cost function (in some embodiments) given in Equation 12:

[0266] C(g)=g * Ag+Bg+C+g * Dg+∑ j g * W j g = g * (A+D+∑ j W j )g+Bg+C (14)

[0267] With this definition of the new cost function term, the overall cost function remains a matrix quadratic, and the activation g of the optimal group can be found through the differentiation of Equation 14. opt To produce:

[0268]

[0269] Weight term W ij Each of them is considered as a given consecutive penalty value for each of the loudspeakers. The function is useful. In one example embodiment, the penalty value is the distance from the (to be rendered) object to the considered loudspeaker. In another example embodiment, the penalty value represents the frequency that a given loudspeaker cannot reproduce. Based on this penalty value, the weight term W... ij It can be parameterized as:

[0270]

[0271] Where, α j Let represent the pre-factor (which takes into account the global strength of the weighting terms), where τ j This represents the penalty threshold (approximately to or exceeding the penalty threshold, at which point the weight term becomes significant), and where f j (x) represents a monotonically increasing function. For example, with... The weight term has the following form:

[0272]

[0273] Where, α j β j τ j These are adjustable parameters that indicate the global strength of the penalty, the abruptness of the penalty's initiation, and the severity of the penalty, respectively. Care should be taken when setting these adjustable values ​​to ensure that the cost item C... j Compared to any other additional cost item and C spatial and C proximity The relative influence applies to achieving the desired outcome. For example, empirically, if one wants a particular punishment to clearly dominate other punishments, its intensity α is... j Setting it to approximately ten times the next maximum penalty intensity might be appropriate.

[0274] If all loudspeakers are penalized, it is usually convenient to subtract the minimum penalty from all weights in post-processing so that at least one loudspeaker is not penalized:

[0275] w ij →w′ ij =w ij -min i (w ij (18)

[0276] As described above, many possible use cases can be realized using the new cost function term described herein (and similar new cost function terms employed in other embodiments). More specific details are then described using the following three examples: moving audio toward the listener or speaker, moving audio away from the listener or speaker, and moving audio away from a landmark.

[0277] In the first example, what will be referred to herein as "attraction" is used to pull audio toward a location, which in some examples could be the location of a listener or speaker, a landmark, furniture, etc. This location may be referred to herein as an "attraction location" or "attractor location." As used herein, "attraction" is a factor that favors relatively higher amplifier activation closer to the attraction location. According to this example, the weight w ij Using the form of Equation 17, the continuous penalty value p ij The distance from the i-th speaker to the fixed attractor position The distance is given, and the threshold τ j The maximum value among these distances across all speakers is given:

[0278]

[0279]

[0280] To illustrate the use case of "pulling" audio toward the listener or speaker, specifically set α j =20, β j =3, and will Set as a vector corresponding to the listener / speaker's position at 180 degrees. α j β j and These values ​​are merely examples. In some implementations, α j It can be in the range of 1 to 100 and β j It can be in the range of 1 to 25.

[0281] Figure 2F This is a diagram showing speaker activation in an example embodiment. In this example, Figure 2F The diagram shows the optimal solution for speaker activations 245b, 250b, 255b, 260b, and 265b, which includes the cost function for the same speaker locations in Figures 1 and 2, plus the cost function calculated by w. ij The appeal of the expression. Figure 2G This is a graph showing the object rendering locations in the example embodiment. In this example, Figure 2GThe diagram shows the corresponding ideal object position 276b for a large number of possible object angles and the corresponding actual rendering position 278b for those objects, connected to the ideal object position 276b by a dashed line 279b. The actual rendering position 278b faces a fixed position. The tilt orientation illustrates the influence of attractor weights on the optimal solution of the cost function.

[0282] In the second and third examples, "repulsion" is used to "push" audio away from a location, which can be the listener's location, the speaker's location, or other locations such as landmarks, furniture, etc. In some examples, repulsion can be used to push audio away from an area or zone of the listening environment, such as an office area, a reading area, a bed or bedroom area (e.g., a crib or bedroom), etc. According to some such examples, a specific location can be used as a representation of a zone or zone. For example, the location representing a crib could be an estimated position of the baby's head, an estimated sound source location corresponding to the baby, etc. This location may be referred to herein as a "repulsion location" or "repulsion position." As used herein, "repulsion" is a factor that favors a relatively lower amplifier activation closer to the repulsion location. According to this example, relative to a fixed repulsion location... Define p ij and τ j Similar to the attraction in Equation 19:

[0283]

[0284]

[0285] To illustrate the use case of pushing audio away from the listener or speaker, specifically set α j =5,β j =2, and will Set as a vector corresponding to the listener / speaker's position at 180 degrees. α j β j and These values ​​are merely examples. As mentioned above, in some examples, α j It can be in the range of 1 to 100 and β j It can be in the range of 1 to 25. Figure 2H This is a diagram showing speaker activation in an example embodiment. According to this example, Figure 2H The diagram shows speaker activations 245c, 250c, 255c, 260c, and 265c, which include the optimal solution of the cost function for the same speaker location as in the previous figure, plus the cost function of w. i j represents the repulsive force. Figure 2I This is a graph showing the object rendering locations in the example embodiment. In this example, Figure 2IThe diagram illustrates the ideal object position 276c for a large number of possible object angles and the corresponding actual rendering position 278c for those objects, which is connected to the ideal object position 276c by a dashed line 279c. The actual rendering position 278c is located away from the fixed position. The tilt orientation illustrates the influence of the repulsion weights on the optimal solution of the cost function.

[0286] The third example use case is to "push" audio away from acoustically sensitive landmarks, such as the door to a room leading to a sleeping baby. Similar to the last example, [the following text is incomplete and requires further context: "to push" audio away from acoustically sensitive landmarks, such as the door to a room leading to a sleeping baby."] Set to the vector corresponding to the 180-degree door position (bottom center of the plot). To achieve a stronger repulsive force and tilt the sound field completely towards the front of the main listening space, set α. j =20, β j =5. Figure 2J This is a diagram showing speaker activation in an example embodiment. Again, in this example, Figure 2J The speaker activations 245d, 250d, 255d, 260d, and 265d are shown, which include the optimal solution for the same set of speaker positions, plus a stronger repulsive force. Figure 2K This is a diagram showing the object rendering locations in the example embodiment. And again, in this example, Figure 2K The diagram illustrates the ideal object position 276d for a large number of possible object angles and the corresponding actual rendering position 278d for those objects, which is connected to the ideal object position 276d by a dashed line 279d. The tilt orientation of the actual rendering position 278d illustrates the effect of stronger repulsion weights on the optimal solution of the cost function.

[0287] Return now Figure 2B In this example, block 225 relates to modifying the rendering process for the first audio signal, at least in part, based on at least one of the second audio signal, the second rendered audio signal, or characteristics thereof, to produce a modified first rendered audio signal. Various examples of modifying the rendering process are disclosed herein. For example, the “characteristics” of the rendered signal may include loudness or audibility estimated or measured at the intended listening location, whether in silence or in the presence of one or more additional rendered signals. Other examples of characteristics include parameters associated with the rendering of the signal, such as the intended spatial location of the constituent signals of the associated program stream, the position of the loudspeaker on which the signal is rendered, the relative activation of the loudspeaker according to the intended spatial location of the constituent signals, and any other parameters or states associated with the rendering algorithm used to generate the rendered signal. In some examples, block 225 may be performed by the first rendering module.

[0288] According to this example, box 230 relates to modifying the rendering process for a second audio signal based at least in part on a first audio signal, a first rendered audio signal, or a characteristic thereof, to produce a modified second rendered audio signal. In some examples, box 230 may be performed by a second rendering module.

[0289] In some implementations, modifying the rendering process for the first audio signal may involve distorting the rendering of the first audio signal away from the rendering position of the second rendered audio signal, and / or modifying the loudness of one or more of the first rendered audio signals in response to the loudness of the second audio signal or one or more of the second rendered audio signals. Alternatively or additionally, modifying the rendering process for the second audio signal may involve distorting the rendering of the second audio signal away from the rendering position of the first rendered audio signal, and / or modifying the loudness of one or more of the second rendered audio signals in response to the loudness of the first audio signal or one or more of the first rendered audio signals. Several examples are provided below with reference to Figure 3 and the following description.

[0290] However, other types of rendering process corrections are within the scope of this disclosure. For example, in some instances, correcting the rendering process for a first or second audio signal may involve performing spectral correction, audibility-based correction, or dynamic range correction. These corrections may or may not be related to loudness-based rendering corrections, depending on the specific example. For instance, in the case described above where the primary spatial stream is rendered in an open-plan living area and a secondary stream including a cooking cues is rendered in an adjacent kitchen, it might be desirable to ensure that the cooking cues remain audible in the kitchen. This can be accomplished by estimating the loudness of the rendered cooking cues stream in the kitchen without interfering with the first signal, then estimating the loudness of the first signal present in the kitchen, and finally dynamically correcting the loudness and dynamic range of both streams across multiple frequencies, thereby ensuring the audibility of the second signal in the kitchen.

[0291] exist Figure 2B In the example shown, box 235 involves at least mixing a modified first rendered audio signal and a modified second rendered audio signal to produce a mixed audio signal. For example, box 235 can be... Figure 2A The mixer 130b shown in the figure is performing.

[0292] According to this example, box 240 involves providing a mixed audio signal to at least some speakers in the environment. Some examples of method 200 involve playing back the mixed audio signal by the speakers.

[0293] like Figure 2BAs shown, some implementations may provide more than two rendering modules. Some such implementations may provide N rendering modules, where N is an integer greater than 2. Therefore, some such implementations may include one or more additional rendering modules. In some such examples, each of the one or more additional rendering modules may be configured to receive an additional audio program stream via an interface system. The additional audio program stream may include an additional audio signal arranged to be reproduced by at least one speaker in the environment. Some such implementations may involve rendering the additional audio signal for reproduction via at least one speaker in the environment to produce an additional rendered audio signal, and modifying the rendering process for the additional audio signal at least in part based on at least one of a first audio signal, a first rendered audio signal, a second audio signal, a second rendered audio signal, or a characteristic thereof to produce a modified additional rendered audio signal. According to some such examples, a mixing module may be configured to mix the modified additional rendered audio signal with at least the modified first rendered audio signal and the modified second rendered audio signal to produce a mixed audio signal.

[0294] As referenced above Figure 1A and Figure 2A As described, some implementations may include a microphone system comprising one or more microphones in a listening environment. In some such examples, a first rendering module may be configured to modify the rendering process for a first audio signal based at least in part on a first microphone signal from the microphone system. The “first microphone signal” may be received from a single microphone or from two or more microphones, depending on the specific implementation. In some such implementations, a second rendering module may be configured to modify the rendering process for a second audio signal based at least in part on the first microphone signal.

[0295] As referenced above Figure 2AIn some instances, the locations of one or more microphones may be known and provided to the control system. According to some such implementations, the control system can be configured to estimate the location of a first sound source based on a first microphone signal and to refine a rendering process for at least one of a first audio signal or a second audio signal, at least in part based on the first sound source location. For example, the first sound source location may be estimated based on DOA data from each of three or more microphones or a group of microphones with known locations, according to a triangulation process. Alternatively or additionally, the first sound source location may be estimated based on the amplitude of signals received from two or more microphones. It may be assumed that the microphone producing the highest amplitude signal is closest to the first sound source location. In some such examples, the first sound source location may be set to the location of the nearest microphone. In some such examples, the first sound source location may be associated with the location of a region, wherein the region is selected by processing signals from two or more microphones through a pre-trained classifier (such as a Gaussian mixer model).

[0296] In some such implementations, the control system may be configured to determine whether a first microphone signal corresponds to ambient noise. Some such implementations may involve modifying the rendering process for at least one of a first audio signal or a second audio signal, at least in part, based on whether the first microphone signal corresponds to ambient noise. For example, if the control system determines that the first microphone signal corresponds to ambient noise, modifying the rendering process for the first audio signal or the second audio signal may involve increasing the level of the rendered audio signal such that the perceived loudness of the signal in the presence of noise at the intended listening location is substantially equal to the perceived loudness of the signal in the absence of noise.

[0297] In some examples, the control system may be configured to determine whether a first microphone signal corresponds to human speech. Some such implementations may involve modifying the rendering process for at least one of a first audio signal or a second audio signal, at least in part, based on whether the first microphone signal corresponds to human speech. For example, if the control system determines that the first microphone signal corresponds to human speech (such as a wake word), modifying the rendering process for the first audio signal or the second audio signal may involve reducing the loudness of the rendered audio signal reproduced by a speaker near the first sound source location compared to the loudness of the rendered audio signal reproduced by a speaker remote from the first sound source location. Modifying the rendering process for the first audio signal or the second audio signal may alternatively or additionally involve modifying the rendering process to distort the intended location of the component signals of the associated program stream away from the first sound source location, and / or penalizing the use of speakers near the first sound source location compared to speakers remote from the first sound source location.

[0298] In some implementations, if the control system determines that the first microphone signal corresponds to human speech, the control system can be configured to reproduce the first microphone signal in one or more speakers near a location in the environment different from the location of the first sound source. In some such examples, the control system can be configured to determine whether the first microphone signal corresponds to a child's cry. According to some such implementations, the control system can be configured to reproduce the first microphone signal in one or more speakers near a location in the environment corresponding to the estimated location of a caregiver (such as a parent, relative, guardian, childcare provider, teacher, nurse, etc.). In some examples, the process of estimating the estimated location of the caregiver can be triggered by a voice command such as "<wake word>, do not wake the baby." The control system will be able to estimate the location of the speaker (caregiver) based on triangulation using DOA information provided by three or more local microphones, etc., according to the location of the nearest smart audio device implementing the virtual assistant. According to some implementations, the control system will have prior knowledge of the location of the baby's room (and / or the listening device therein) and will then be able to perform appropriate processing.

[0299] According to some such examples, the control system can be configured to determine whether a first microphone signal corresponds to a command. In some instances, if the control system determines that the first microphone signal corresponds to a command, the control system can be configured to determine a response to the command and control at least one speaker near the first sound source location to reproduce the response. In some such examples, the control system can be configured to, after controlling at least one speaker near the first sound source location to reproduce the response, revert to an uncorrected rendering process for either the first or second audio signal.

[0300] In some implementations, the control system can be configured to execute commands. For example, the control system may be or may include a virtual assistant configured to control audio devices, televisions, home appliances, etc., according to commands.

[0301] pass Figure 1A , Figure 1B and Figure 2A The definition of this minimal yet more capable multi-stream rendering system, shown in the diagram, enables dynamic management of simultaneous playback of multiple program streams for many useful scenarios. Reference will now be made to... Figure 3A and Figure 3B Describe several examples.

[0302] First, examine the previously discussed example involving simultaneously playing a spatial cinematic soundtrack in the living room and cooking tips in the adjoining kitchen. The spatial cinematic soundtrack is an example of the "first audio stream" mentioned above, and the cooking tip audio is an example of the "second audio stream" mentioned above. Figure 3Aand Figure 3B An example floor plan of connected living spaces is shown. In this example, living space 300 includes a living room in the upper left, a kitchen in the lower center, and a bedroom in the lower right. The boxes and circles 305a to 305h distributed across the living space represent a group of eight loudspeakers placed in locations convenient for the space, but not adhering to any standard layout (arbitrary placement). Figure 3A In this design, only the spatial movie soundtrack is played back, and taking into account the capabilities and layout of the amplifiers, all the amplifiers in the living room 310 and kitchen 315 are used to create an optimized spatial reproduction around a listener 320a sitting on sofa 325 facing television 330. This optimal reproduction of the movie soundtrack is visually represented by cloud-shaped lines 335a located within the boundaries of the activated amplifiers.

[0303] exist Figure 3B In the kitchen 315, cooking cues are simultaneously rendered and played back on a single loudspeaker 305g for a second listener 320b. The visual representation of this second program stream is represented by cloud-like lines 340 emanating from the loudspeaker 305g. If these cooking cues are played back simultaneously without modifying the rendering of the movie soundtrack, such as... Figure 3A As shown, audio from a movie soundtrack emanating from a speaker in or near kitchen 315 would interfere with a second listener's ability to understand the cooking cues. Instead, in this example, the rendering of the spatial movie soundtrack is dynamically adjusted based on the rendering of the cooking cues. Specifically, the rendering of the movie soundtrack is moved away from the speaker near the rendering location of the cooking cues (kitchen 315), where... Figure 3B The smaller, cloud-like line 335b, pushed away from the speaker near the kitchen, visually represents this displacement. In some implementations, if the playback of cooking prompts stops while a movie soundtrack is playing, the rendering of the movie soundtrack can dynamically shift back to where it was. Figure 3A The original optimal configuration is shown in the image. This dynamic shifting in the rendering of spatial cinematic soundtracks can be achieved through many publicly available methods.

[0304] Many spatial audio mixes consist of multiple component audio signals designed to be played at specific locations within a listening space. For example, Dolby 5.1 and 7.1 surround sound mixes consist of six and eight signals, respectively, intended for playback on speakers at specified canonical locations around the listener. Object-based audio formats (e.g., Dolby Atmos) consist of component audio signals with associated metadata describing possible time-varying 3D locations within the listening space where the audio will be rendered. Assuming the renderer of a spatial cinematic soundtrack can render individual audio signals at any location with respect to any group of speakers, this can be achieved by distorting the intended locations of the audio signals within the spatial mix. Figure 3A and Figure 3BThe rendering depicts dynamic displacement. For example, the 2D or 3D coordinates associated with the audio signal can be pushed away from the location of the speaker in the kitchen, or alternatively pulled towards the upper left corner of the living room. The result of this distortion is that speakers near the kitchen are used less, because the distorted location of the spatially mixed audio signal is now farther from that position. While this method does achieve the goal of making the second audio stream more intelligible to a second listener, the cost is a significant alteration to the intended spatial balance of the film soundtrack for the first listener.

[0305] A second method for achieving dynamic shifting in spatial rendering can be implemented using a flexible rendering system. In some such implementations, the flexible rendering system can be CMAP, FV, or a hybrid of both, as described above. Some such flexible rendering systems attempt to reproduce a spatial mix in which all component signals are considered to originate from their intended locations. In some examples, while doing so for each signal of the mix, priority is given to activating speakers near the desired location of that signal. In some implementations, additional factors can be dynamically added to the rendering optimization, which penalizes the use of certain speakers based on other criteria. For example, something that could be called a “push-repulsion” could be dynamically positioned in the kitchen to highly penalize the use of speakers near that location and effectively push away the rendering of spatial movie soundtracks. As used herein, the term “push-repulsion” can refer to a factor corresponding to relatively low speaker activation in a particular location or area of ​​the listening environment. In other words, the phrase “push-repulsion” can refer to a factor that favors the activation of speakers relatively far from the particular location or area corresponding to the “push-repulsion”. However, according to some such implementations, the renderer can still attempt to reproduce the intended spatial balance of the mix using the remaining, less penalized speakers. Thus, this technique can be considered a superior method for achieving dynamic shifting of the render compared to methods that simply distort the intended positions of the component signals of the mix.

[0306] It can be used Figure 1B The minimal version of the multi-stream renderer described in the image is used to implement the scene described, which involves removing the rendering of the spatial cinematic soundtrack from the cooking cues in the kitchen. However, this can be achieved by employing... Figure 2AA more capable system is described in the text to achieve improvements to the scene. While shifting the rendering of the spatial cinematic soundtrack does improve the intelligibility of cooking cues in the kitchen, the cinematic soundtrack can still be clearly audible in the kitchen. Depending on the instantaneous conditions of the two streams, the cooking cues may be masked by the cinematic soundtrack; for example, a loud moment in the cinematic soundtrack masks a gentle moment in the cooking cues. To address this issue, dynamic corrections can be added to the rendering of the cooking cues based on the rendering of the spatial cinematic soundtrack. For example, a method for dynamically altering the audio signal across frequency and time to maintain its perceived loudness in the presence of interfering signals can be implemented. In this scene, an estimate of the perceived loudness of the shifted cinematic soundtrack at the kitchen location can be generated, and said estimate can be fed as an interfering signal into such a process. The temporal and frequency variation levels of the cooking cues can then be dynamically corrected to maintain their perceived loudness above the interfering signal, thereby better maintaining intelligibility for a second listener. The desired estimate of the loudness of the cinematic soundtrack in the kitchen can come from the speaker feed of the rendering of said soundtrack, signals from microphones in or near the kitchen, or a combination thereof. Maintaining the perceived loudness of cooking cues often increases the level of the cues, and in some cases, the overall loudness can become unpleasantly high. To address this, another rendering correction can be employed. The interfering spatial cinematic soundtrack can be dynamically downset if the loudness-corrected cooking cues in the kitchen become too loud. Finally, some external noise sources can simultaneously interfere with the audibility of both program streams; for example, a blender might be used in the kitchen during cooking. Loudness estimates of these ambient noise sources in both the living room and kitchen can be generated by microphones connected to the rendering system. For example, this estimate can be added to the loudness estimate of the soundtrack in the kitchen to influence the loudness correction of the cooking cues. Simultaneously, the rendering of the soundtrack in the living room can be further corrected based on the ambient noise estimate to maintain the perceived loudness of the soundtrack in the living room in the presence of that ambient noise, thereby better maintaining audibility for listeners in the living room.

[0307] As can be seen, this example use case of the disclosed multi-stream renderer employs numerous interrelated modifications to the two program streams to optimize their simultaneous playback. In summary, these modifications to the streams can be listed as follows:

[0308] ●Space Cinema Soundtrack

[0309] o Move the space rendering away from the kitchen based on the cooking tips rendered in the kitchen.

[0310] o Dynamically reduce loudness based on the loudness of cooking cues rendered in the kitchen.

[0311] o Dynamically increase loudness based on an estimate of the loudness of the disruptive mixer noise from the kitchen in the living room.

[0312] ●Cooking Tips

[0313] The loudness is dynamically increased based on a combined loudness estimate of both the movie soundtrack and the noise from the kitchen mixer.

[0314] The second example use case for the disclosed multi-stream renderer involves the playback of simultaneous spatial program streams (such as music) and the response of a smart voice assistant to some user queries. For existing smart speakers where playback is typically constrained to mono or stereo playback on a single device, interaction with a voice assistant generally consists of the following stages:

[0315] 1) Play music

[0316] 2) The user says the voice assistant wake-up word.

[0317] 3) The smart speaker recognizes the wake word and significantly lowers (avoids) the music.

[0318] 4) The user gives a command to the smart assistant (i.e., "play the next song").

[0319] 5) The smart speaker recognizes the command, confirms it by playing a voice response (i.e., "Okay, play the next song") mixed with the escaping music through the speaker, and then executes the command.

[0320] 6) The smart speaker will turn the music back up to its original volume.

[0321] Figure 4A and Figure 4B An example of a multi-stream renderer that provides simultaneous playback of spatial music mixing and voice assistant responses is shown. Some embodiments offer improvements to the event chain described above when playing spatial audio across multiple orchestrated smart speakers. Specifically, the spatial mix can be removed from one or more speakers selected to relay responses from the voice assistant. Creating this spatial mix for the voice assistant response means that the spatial mix can be reduced less, or not at all, compared to the existing configurations listed above. Figure 4A and Figure 4B The scenario is depicted. In this example, the modified event chain can occur as follows:

[0322] 1) Currently streaming spatial music programs to users across a variety of orchestrated smart speakers. Figure 4A (Cloud-like line 335c in the middle).

[0323] 2) User 320c says the wake-up word for the voice assistant.

[0324] 3) One or more smart speakers (e.g., speaker 305d and / or speaker 305f) recognize the wake word and use associated records from microphones associated with one or more smart speakers to determine the location of user 320c or which speaker(s) user 320c is closest to.

[0325] 4) When the expected voice assistant response stream is rendered near that location, move the rendering of the spatial music mix away from the location determined in the previous step. Figure 4B (Cloud-like line 335d in the middle).

[0326] 5) The user speaks commands to the smart assistant (e.g., to a smart speaker running smart assistant / virtual assistant software).

[0327] 6) The smart speaker recognizes the command, synthesizes the corresponding response program stream, and renders the response near the user's location. Figure 4B (Cloud-like line 440 in the middle).

[0328] 7) When the voice assistant completes its response, the rendering of the spatial music program stream will revert to its original state. Figure 4A (Cloud-like line 335c in the middle).

[0329] In addition to optimizing the simultaneous playback of spatial music mixing and voice assistant responses, shifting the spatial music mix can also improve the ability of a set of speakers in step 5 to understand the listener. This is because the music has been shifted away from the speakers near the listener, thereby boosting the speech to other proportions of the relevant microphones.

[0330] Similar to the scenario described for the previous scenario with spatial film mix and cooking prompts, the current scenario can be further optimized beyond the optimization provided by shifting the rendering of the spatial mix based on the voice assistant response. Shifting the spatial mix on its own may not be sufficient to make the voice assistant response fully understandable to the user. A simple solution is to also reduce the spatial mix by a fixed amount, albeit less than what is currently required. Alternatively, the loudness of the voice assistant response stream can be dynamically increased based on the loudness of the spatial music mix stream to maintain the audibility of the response. As an extension, the loudness of the spatial music mix can also be dynamically reduced if this increase on the response stream becomes too large.

[0331] Figure 5A , Figure 5B and Figure 5C The illustration shows a third example use case for the disclosed multi-stream renderer. This example involves managing the simultaneous playback of a spatial music mix program stream and a comfort noise program stream, while attempting to ensure that a baby remains asleep in an adjacent room, but is still able to hear if the baby is crying. Figure 5AThe starting point is depicted, where the spatial music mix (represented by cloud-shaped line 335e) runs across all the speakers in the living room 310 and kitchen 315, playing optimally for many people at a party. Figure 5B In the scene, baby 510 is now trying to sleep in the adjacent bedroom 505, depicted in the lower right corner. To help ensure this, the spatial music mix is ​​dynamically moved away from the bedroom to minimize leakage within it, as depicted by cloud-like line 335f, while still maintaining a reasonable experience for the people at the party. Simultaneously, a second program stream containing soothing white noise (represented by cloud-like line 540) is played from the speaker 305h in the baby's room to mask any remaining leakage of music from the adjacent room. In some examples, to ensure complete masking, the loudness of the white noise stream can be dynamically adjusted based on an estimate of the loudness of the spatial music leaking into the baby's room. This estimate can be generated from the speaker feed of the spatial music rendering, the signal from the microphone in the baby's room, or a combination thereof. Similarly, if the loudness of the spatial music mix becomes too large, its loudness can be dynamically attenuated based on a loudness-modified noise. This is analogous to the loudness handling between the spatial film mix and the cooking cues in the first scene. Finally, the microphone in the baby's room (e.g., a microphone associated with speaker 305h, which in some embodiments may be a smart speaker) can be configured to record audio from the baby (eliminating sounds that may be picked up from spatial music and white noise), and if crying is detected (by machine learning, via pattern matching algorithms, etc.), the combination of these processed microphone signals can then serve as a third program stream that can be simultaneously played back in the vicinity of a listener 320d in the living room 310, who may be a parent or other caregiver. Figure 5C The reproduction of this additional stream is depicted using cloud-like lines 550. In this case, the spatial music mix can be further removed from the speakers playing the baby's cries near the parents, as relative to... Figure 5B The shape of cloud-like line 335f and the modified shape of cloud-like line 335g are shown, and the loudness of the baby's cry program stream can be corrected according to the spatial music stream so that the baby's cry remains audible to the listener 320d. The interconnection corrections for optimizing the simultaneous playback of three program streams considered in this example can be summarized as follows:

[0332] ● Spatial music mixing in the living room

[0333] o Move the spatial rendering away from the baby's room to reduce propagation into said room.

[0334] o Dynamically reduce loudness based on the loudness of the white noise rendered in the baby's room.

[0335] o Renders on speakers near the parents based on the baby's cries, moving spatial rendering away from the parents.

[0336] ●White noise

[0337] o Dynamically increase loudness based on an estimate of the loudness of the music stream permeating into the baby's room.

[0338] ● Record the baby's cries

[0339] o Dynamically increase loudness based on an estimate of the loudness of the music mix at the location of the parent or other caregiver.

[0340] The following describes examples of how some of the mentioned embodiments can be implemented.

[0341] exist Figure 1B In this context, each of the rendering blocks 1...N can be implemented as the same instance of any single-stream renderer, such as the previously mentioned CMAP, FV, or hybrid renderer. Constructing a multi-stream renderer in this way has several convenient and useful properties.

[0342] First, if rendering is performed within this hierarchical layout, and each single-stream renderer instance is configured to operate in the frequency / transform domain (e.g., QMF), then stream mixing can also occur in the frequency / transform domain, and the inverse transform only needs to be run once for M channels. This represents a significant efficiency improvement compared to running N×M inverse transforms and mixing in the time domain.

[0343] Figure 6 It shows Figure 1B The following is an example of a frequency domain / transform domain multistream renderer. In this example, a quadrature mirror analysis filter bank (QMF) is applied to each of the program streams 1 to N before each program stream is received by its corresponding rendering module among rendering modules 1 to N. According to this example, rendering modules 1 to N operate in the frequency domain. After mixer 630a mixes the outputs of rendering modules 1 to N, inverse synthesis filter bank 635a converts the mix to the time domain and provides the mixed speaker feed signal in the time domain to amplifiers 1 to M. In this example, the quadrature mirror filter bank, rendering modules 1 to N, mixer 630a, and inverse filter bank 635a are components of the control system 110c.

[0344] Figure 7 It shows Figure 2A The example shown is a frequency domain / transform domain rendering of a multi-stream renderer. (See also...) Figure 6In this example, before each program stream is received by the corresponding rendering module among rendering modules 1 to N, a quadrature mirror filter bank (QMF) is applied to each of the program streams 1 to N. According to this example, rendering modules 1 to N operate in the frequency domain. In this embodiment, the time-domain microphone signal from microphone system 120b is also provided to the quadrature mirror filter bank, so that rendering modules 1 to N receive the microphone signal in the frequency domain. After mixer 630b mixes the outputs of rendering modules 1 to N, inverse filter bank 635b transforms the mix to the time domain and provides the mixed speaker feed signal in the time domain to amplifiers 1 to M. In this example, the quadrature mirror filter bank, rendering modules 1 to N, mixer 630b, and inverse filter bank 635b are components of control system 110d.

[0345] Another benefit of the hierarchical approach in the frequency domain is that it allows for the calculation of the perceived loudness of each audio stream and the use of this information to dynamically adjust one or more audio streams in other audio streams. To illustrate this embodiment, consider the references above. Figure 3A and Figure 3B The previously mentioned example is described. In this case, there are two audio streams (N=2), a spatial movie soundtrack and cooking cues. There may also be ambient noise from a blender in the kitchen, picked up by one or more of the K microphones.

[0346] After each audio stream s is rendered individually and each microphone i is captured and transformed to the frequency domain, the source excitation signal E can be calculated. s or E i This serves as a time-varying estimate of the perceived loudness of each audio stream s or microphone signal i. In this example, these source excitation signals are for c speakers, for b frequency bands across time t, from the rendered stream or captured microphone via transform coefficients X for the audio stream. s Or the transformation coefficient X for the microphone signal i It is calculated using a frequency-dependent time constant λ. b Smooth:

[0347] E s (b, t, c) = λ b E s (b, t-1, c) + (1-λ) b )|X s (b, t, c)| 2

[0348] (20a)

[0349] E i (b, t, c) = λ b (b)E i (b, t-1, c) + (1-λ)b )|X i (b, t, c)| 2

[0350] (20b)

[0351] The original source excitation is an estimate of the perceived loudness of each flow at a specific location. For spatial flows, this location is... Figure 3B The location is in the middle of cloud-shaped line 335b, while for cooking prompts, the location is in the middle of cloud-shaped line 340. The location of the blender noise picked up by the microphone can be based, for example, on the specific location of the microphone(s) closest to the source of the blender noise.

[0352] The original source excitation must be transformed to the listening position of the (multiple) audio streams to be modified in order to estimate the perceptibility of the original source excitation as noise at the listening position of each target audio stream. For example, if audio stream 1 is a movie soundtrack and audio stream 2 is a cooking cue, then This will be the transformed (noise) excitation. For each frequency band b, according to each loudspeaker c, by using the audibility scaling factor A... xs Applying A from the source audio stream s to the target audio stream x or using A xi The transformation is calculated by applying it from microphone i to the target audio stream x. xs and A xi The value can be determined by using a distance ratio or an estimate of actual audibility, and it can vary over time.

[0353]

[0354]

[0355] In equation 13a, This represents the raw noise excitation calculated for the source audio stream, without reference to the microphone input. In Equation 13b, This represents the original noise excitation calculated from the reference microphone input. Following this example, the original noise excitation is then calculated across currents 1 to N, microphones 1 to K, and output channels 1 to M. or Summing is performed to obtain a total noise estimate for the target flow x.

[0356]

[0357] According to some alternative implementations, by omitting terms in Equation 14 Total noise estimation can be obtained without referencing microphone input.

[0358] In this example, the total raw noise estimate is smoothed to avoid perceptible artifacts caused by excessively rapid correction of the target stream. According to this implementation, the smoothing is based on the concept of using fast start-up and slow release, similar to an audio compressor. The smoothed noise estimate for the target stream x is shown below. In this example, it is calculated as follows:

[0359]

[0360]

[0361] Once a complete noise estimate for stream x is available The previously calculated source excitation signal E can then be reused. x (b, t, c) are used to determine the time-varying gain set G. x (b, t, c) are applied to the target audio stream x to ensure that the target audio stream remains audible in the face of noise. These gains can be calculated using any of a variety of techniques.

[0362] In one embodiment, the loudness function L{·,·} can be applied to the excitation to model various nonlinearities in human loudness perception and compute a specific loudness signal describing the time-varying distribution of perceived loudness over frequency. Applying L{·,·} to the excitation for noise estimation and the rendered audio stream x yields an estimate of the specific loudness for each signal:

[0363]

[0364] L x (b, t, c) = L{E x (b, t, c)} (25b)

[0365] In equation 17a, L xn This represents an estimate of the specific loudness of the noise, and in Equation 17b, L x This represents an estimate of the specific loudness for a rendered audio stream x. These specific loudness signals represent the perceived loudness when the signal is heard alone. However, masking can occur if the two signals are mixed. For example, if the noise signal is much louder than the stream x signal, the noise signal will mask the stream x signal, thus reducing the perceived loudness of the signal relative to its perceived loudness when heard alone. This phenomenon can be modeled using a partial loudness function PL{·, ·} with two inputs. The first input is the excitation of the signal of interest, and the second input is the excitation of the competing (noise) signal. The function returns a partial loudness signal PL, which represents the perceived loudness of the signal of interest in the presence of the competing signal. In the presence of the noise signal, the partial loudness of the stream x signal can then be calculated directly from the excitation signal across frequency band b, time t, and loudspeaker c:

[0366]

[0367] To maintain the audibility of the audio stream signal x in the presence of noise, the gain G can be calculated. x (b, t, c) are applied to audio stream x to increase loudness until the signal is audible over noise, as shown in equations 8a and 8b. Alternatively, if the noise comes from another audio stream s, two gain sets can be calculated. In such an example, the first gain set G... x (b, t, c) will be applied to the audio stream x to increase its loudness, and the second gain set G s (b, t) will be applied to the competing audio stream s to reduce its loudness, such that the combination of gains ensures the audibility of the audio stream x, as shown in equations 9a and 9b. In the two sets of equations, This represents a specific loudness of the source signal in the presence of noise after applying compensated gain.

[0368]

[0369] Make

[0370]

[0371]

[0372] Again, making

[0373]

[0374] In practice, the original gain is further smoothed across frequencies using a smoothing function S{·} before being applied to the audio stream to avoid audible artifacts again. and This represents the final compensation gain for the target audio stream x and the competing audio stream s:

[0375]

[0376]

[0377] In one embodiment, these gains can be applied directly to all rendered output channels of the audio stream. In another embodiment, these gains can alternatively be applied to objects of the audio stream before rendering, for example, using the method described in U.S. Patent Application Publication No. 2019 / 0037333 A1, which is incorporated herein by reference. These methods involve calculating translation coefficients for each audio object associated with each of a plurality of predefined channel coverages based on the spatial metadata of the audio objects. The audio signal can be converted into submixes associated with the predefined channel coverages based on the calculated translation coefficients and the audio objects. Each submix can indicate the sum of the components of a plurality of audio objects associated with one of the predefined channel coverages. Submix gains can be generated by applying audio processing to each submix, and the object gain applied to each audio object can be controlled. The object gain can be a function of the translation coefficients for each audio object and the submix gain associated with each predefined channel coverage. Applying gains to objects has several advantages, particularly when combined with other processing of the stream.

[0378] Figure 8 An implementation of a multi-stream rendering system with an audio stream loudness estimator is shown. Based on this example, Figure 8 The multi-stream rendering system is also configured to implement loudness processing as described in Equations 12a to 21b, and to apply compensated gain within each single-stream renderer. In this example, a quadrature mirror filter bank (QMF) is applied to each of program streams 1 and 2 before each program stream is received by the corresponding rendering module in rendering modules 1 to N. In an alternative example, a quadrature mirror filter bank (QMF) may be applied to each of program streams 1 to N before each program stream is received by the corresponding rendering module in rendering modules 1 to N. According to this example, rendering modules 1 and 2 operate in the frequency domain. In this embodiment, loudness estimation module 805a calculates a loudness estimate for program stream 1, for example, as described above with reference to Equations 12a to 17b. Similarly, in this example, loudness estimation module 805b calculates a loudness estimate for program stream 2.

[0379] In this embodiment, the time-domain microphone signal from microphone system 120c is also provided to the quadrature mirror filter bank, such that loudness estimation module 805c receives the microphone signal in the frequency domain. In this embodiment, loudness estimation module 805c calculates a loudness estimate for the microphone signal, for example, as described above with reference to equations 12b to 17a. In this example, loudness processing module 810 is configured to implement loudness processing as described in equations 18 to 21b, and to apply compensation gain for each single-stream rendering module. In this embodiment, loudness processing module 810 is configured to modify the audio signals of program stream 1 and program stream 2 to maintain the perceived loudness of the audio signals in the presence of one or more interfering signals. In some instances, the control system may determine that the microphone signal corresponds to ambient noise, and the program stream should be boosted above that ambient noise. However, in some examples, the control system may determine that the microphone signal corresponds to a wake word, a command, a child's cry, or other such audio that might be heard by a smart audio device and / or one or more listeners. In some such embodiments, the loudness processing module 810 may be configured to modify the microphone signal in order to maintain the perceived loudness of the microphone signal in the presence of interfering audio signals from program stream 1 and / or program stream 2. Here, the loudness processing module 810 is configured to provide appropriate gain to rendering modules 1 and 2.

[0380] After the mixer 630c mixes the outputs of rendering modules 1 to N, the inverse filter bank 635c converts the mix to the time domain and provides the mixed speaker feed signals in the time domain to amplifiers 1 to M. In this example, the quadrature mirror filter bank, rendering modules 1 to N, mixer 630c, and inverse filter bank 635c are components of the control system 110e.

[0381] Figure 9A An example of a multi-stream rendering system configured for crossfading of multiple rendered streams is shown. In some such embodiments, crossfading of multiple rendered streams is used to provide a smooth experience when rendering configurations are dynamically changed. One example is the aforementioned use case of simultaneously playing spatial program streams (such as music), where a smart voice assistant responds to several queries from the listener, as referenced above. Figure 4A and Figure 4B As described. In this case, it is useful to instantiate additional single-stream renderers using alternating spatial rendering configurations and simultaneously cross-gradient between them, as... Figure 9A As shown in the figure.

[0382] In this example, QMF is applied to program stream 1 before it is received by rendering modules 1a and 1b. Similarly, QMF is applied to program stream 2 before it is received by rendering modules 2a and 2b. In some instances, the output of rendering module 1a may correspond to the desired reproduction of program stream 1 before wake-word detection, while the output of rendering module 1b may correspond to the desired reproduction of program stream 1 after wake-word detection. Similarly, the output of rendering module 2a may correspond to the desired reproduction of program stream 2 before wake-word detection, while the output of rendering module 2b may correspond to the desired reproduction of program stream 2 after wake-word detection. In this embodiment, the outputs of rendering modules 1a and 1b are provided to crossfade module 910a, and the outputs of rendering modules 2a and 2b are provided to crossfade module 910b. For example, the crossfade time may range from hundreds of milliseconds to several seconds.

[0383] After mixer 630d mixes the outputs of crossfade modules 910a and 910b, inverse filter bank 635d converts the mix to the time domain and provides the mixed speaker feed signals in the time domain to amplifiers 1 to M. In this example, the quadrature mirror filter bank, rendering module, crossfade module, mixer 630d, and inverse filter bank 635d are components of control system 110f.

[0384] In some embodiments, the rendering configuration used in each of the single-stream renderers 1a, 1b, 2a, and 2b can be pre-computed. This is particularly convenient and efficient for use cases such as intelligent voice assistants, as the spatial configuration is typically known a priori and does not depend on other dynamic aspects of the system. In other embodiments, pre-compiling the rendering configuration may be impossible or undesirable, in which case the complete configuration of each single-stream renderer must be dynamically computed at runtime.

[0385] Some aspects of the embodiments include the following:

[0386] 1. An audio rendering system that simultaneously plays multiple audio program streams on a plurality of arbitrarily placed loudspeakers, wherein at least one of the program streams is a spatial mix, and the rendering of the spatial mix is ​​dynamically modified in response to the simultaneous playback of one or more additional program streams.

[0387] 2. The system of claim 1, wherein the rendering of any one of the plurality of audio program streams can be dynamically modified based on any combination of one or more of the remaining plurality of audio program streams.

[0388] 3. The system of claim 1 or 2, wherein the modification includes one or more of the following:

[0389] ● Correct the relative activation of multiple loudspeakers based on the relative activation of the loudspeakers associated with the rendering of at least one of the one or more additional program streams;

[0390] ● To distort the intended spatial balance of the spatial mix based on the spatial properties of the rendering of at least one of the one or more additional program streams; or

[0391] ● The loudness or audibility of the spatial mix is ​​adjusted based on the loudness or audibility of at least one of the one or more additional program streams.

[0392] 4. The system of claim 1 or 2, further comprising dynamically adjusting the rendering based on one or more microphone inputs.

[0393] 5. The system of claim 4, wherein the information obtained from the microphone input for correcting rendering includes one or more of the following:

[0394] ● Detection of specific phrases in the user's speech within the system;

[0395] ●Estimation of the location of one or more users in the system;

[0396] ● An estimate of the loudness of any combination of N program streams at a specific location in the listening space; or

[0397] ●Estimation of the loudness of other ambient sounds (e.g., background noise) in the listening environment.

[0398] Other examples of embodiments of the present invention’s system and method for managing the playback of multiple audio streams on multiple speakers (e.g., speakers of a set of orchestrated smart audio devices) include the following:

[0399] 1. An audio system (e.g., an audio rendering system) that simultaneously plays multiple audio program streams on a plurality of arbitrarily placed loudspeakers (e.g., speakers of a group of choreographed smart audio devices), wherein at least one of the program streams is a spatial mix, and the rendering of the spatial mix is ​​dynamically modified in response to (or in combination with) the simultaneous playback of one or more additional program streams.

[0400] 2. The system of claim 1, wherein the modification of the spatial mixing includes one or more of the following:

[0401] ● The rendering of the spatial mix is ​​distorted away from the rendering location of the one or more additional streams, or

[0402] ● The loudness of the spatial mix is ​​modified in response to the loudness of the one or more additional streams.

[0403] 3. The system of claim 1 further relates to dynamically modifying the rendering of the spatial mix based on one or more microphone inputs (i.e., signals captured by one or more microphones of one or more intelligent audio devices, for example, a set of choreographed intelligent audio devices).

[0404] 4. The system of claim 3, wherein at least one of the one or more microphone inputs contains (indicating) human speech. Optionally, the rendering is dynamically adjusted in response to a determined location of the source (human) of the speech.

[0405] 5. The system of claim 3, wherein at least one of the one or more microphone inputs includes ambient noise.

[0406] 6. The system of claim 3, wherein the loudness estimate of the spatial stream or one or more additional streams is obtained from at least one of the one or more microphone inputs.

[0407] One practical consideration in implementing dynamic cost-flexible rendering (according to some embodiments) is complexity. In some cases, given that object locations (the locations of each audio object to be rendered, which can be indicated by metadata) may change multiple times per second, it may be impractical to solve for a unique cost function for each frequency band of each audio object in real time. An alternative approach to reduce complexity at the cost of memory is to use a lookup table that samples the three-dimensional space of all possible object locations. The sampling does not need to be identical across all dimensions. Figure 9B This is a graph indicating the points of speaker activation in an example embodiment. In this example, 15 points are sampled along the x and y dimensions, and 5 points are sampled along the z dimension. Other implementations may include more or fewer samples. According to this example, each point represents M speaker activations for a CMAP or FV solution.

[0408] During runtime, in order to determine the actual activation of each speaker, in some examples, trilinear interpolation between the activations of the 8 most recent speaker points can be used. Figure 10This is a diagram of trilinear interpolation between points indicating speaker activation, based on an example. In this example, the process of continuous linear interpolation includes interpolating each pair of points in the top plane to determine a first interpolation point 1005a and a second interpolation point 1005b; interpolating each pair of points in the bottom plane to determine a third interpolation point 1010a and a fourth interpolation point 1010b; interpolating the first interpolation point 1005a and the second interpolation point 1005b to determine a fifth interpolation point 1015 in the top plane; interpolating the third interpolation point 1010a and the fourth interpolation point 1010b to determine a sixth interpolation point 1020 in the bottom plane; and interpolating the fifth interpolation point 1015 and the sixth interpolation point 1020 to determine a seventh interpolation point 1025 between the top and bottom planes. Although trilinear interpolation is an effective interpolation method, those skilled in the art will understand that trilinear interpolation is only one possible interpolation method that can be used to implement aspects of this disclosure, and other examples may include other interpolation methods.

[0409] In the first example above, where repulsion is used to create an acoustic space for a voice assistant, another important concept is the transition from a rendered scene without repulsion to one with repulsion. To create a smooth transition and give the impression of a dynamically distorted sound field, both a previous set of speaker activations without repulsion and a new set of speaker activations with repulsion are calculated over a period of time, and interpolation is performed between the two.

[0410] An example of audio rendering implemented according to the embodiments is: an audio rendering method, including:

[0411] Render a set of one or more audio signals, each having an associated desired perceived spatial location, on two or more loudspeakers in a group, wherein the relative activation of the group of loudspeakers is a function of: a model of the perceived spatial location of the audio signals played back on the loudspeakers, the proximity of the desired perceived spatial location of the audio object to the location of the loudspeakers, and one or more additional dynamically configurable functions that depend on at least one or more properties of the set of audio signals, one or more properties of the group of loudspeakers, or one or more external inputs.

[0412] refer to Figure 11 An example embodiment is described. As with the other figures provided herein, Figure 11 The types and quantities of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types and quantities of elements. Figure 11A floor plan of a listening environment is depicted, which in this example is a living space. According to this example, environment 1100 includes a living room 1110 in the upper left, a kitchen 1115 in the lower center, and a bedroom 1122 in the lower right. Boxes and circles distributed across the living space represent a set of loudspeakers 1105a to 1105h, at least some of which in some embodiments may be smart speakers, placed in locations convenient to the space, but not adhering to any standard layout (arbitrary placement). In some examples, loudspeakers 1105a to 1105h may be coordinated to implement one or more of the disclosed embodiments. In this example, environment 1100 includes cameras 1111a to 1111e distributed throughout the environment. In some embodiments, one or more smart audio devices in environment 1100 may also include one or more cameras. The one or more smart audio devices may be single-purpose audio devices or virtual assistants. In some such examples, one or more cameras of the optional sensor system 130 may reside in or on a television 1130, in a mobile phone, or in one or more smart speakers (such as loudspeakers 1105b, 1105d, 1105e, or 1105h). Although cameras 1111a to 1111e are not shown in every depiction of environment 1100 presented in this disclosure, in some embodiments, each environment 1100 may still include one or more cameras.

[0413] Figure 12A , Figure 12B , Figure 12C and Figure 12D Showing the target Figure 11 The examples shown are of various listening positions and orientations in a living space, which are used as reference spatial patterns to flexibly render spatial audio. Figures 12A to 12D The feature is depicted at four example listening positions. In each example, arrow 1205 pointing to person 1220a indicates the position of the front sound field (the position facing person 1220a). In each example, arrow 1210a indicates the left surround field and arrow 1210b indicates the right surround field.

[0414] exist Figure 12A In this context, for person 1220a sitting on living room sofa 1225, a reference spatial pattern has been determined and spatial audio has been flexibly rendered. According to some implementation methods, the control system (such as...) Figure 1A The control system 110 can be configured to operate according to the interface system (such as...) Figure 1AThe interface system 105 receives reference spatial pattern data to determine the assumed listening position and / or assumed orientation of the reference spatial pattern. Some examples are described below. In some such examples, the reference spatial pattern data may include data from a microphone system (such as...). Figure 1A Microphone data of the microphone system 120.

[0415] In some such examples, the reference spatial pattern data may include microphone data corresponding to wake words and voice commands (such as "[wake word], make the TV the front sound field"). Alternatively or additionally, the microphone data may be used to triangulate the user's position based on the sound of the user's speech, for example, via direction of arrival (DOA) data. For example, three or more loudspeakers 1105a to 1105e may use microphone data to triangulate the position of person 1220a sitting on living room sofa 1225 based on the sound of person 1220a's speech, via DOA data. The orientation of person 1220a can be assumed based on person 1220a's position: if person 1220a is in… Figure 12A At the position shown, it can be assumed that person 1220a is facing television 1130.

[0416] Alternatively or otherwise, the position and orientation of person 1220a can be determined from the camera system (e.g., Figure 1A The image data from the sensor system 130 is used to determine the location.

[0417] In some examples, the position and orientation of person 1220a can be determined based on user input obtained via a graphical user interface (GUI). According to some such examples, the control system can be configured to control a display device (e.g., the display device of a cellular phone) to present a GUI that allows person 1220a to input their position and orientation.

[0418] Figure 13A An example of a GUI for receiving user input related to the listener's location and orientation is shown. According to this example, the user has previously identified several possible listening locations and corresponding orientations. During the setup process, the speaker position corresponding to each location and orientation has been entered and stored. Some examples are described below. For example, a listening environment layout GUI may have been provided, and the user may have been prompted to touch the position corresponding to the possible listening location and speaker position, and to name the possible listening location. In this example, in... Figure 13A At the time depicted, the user has provided user input regarding their position to the GUI 1300 by touching the virtual "living room sofa" button. Since there are two possible forward-facing positions, considering the L-shaped sofa 1225, the user is prompted to indicate which direction they are facing.

[0419] exist Figure 12B In the design, for person 1220a sitting in living room reading chair 1215, a reference spatial pattern has been determined and spatial audio has been flexibly rendered. Figure 12C In the example, for person 1220a standing next to kitchen counter 1230, a reference spatial pattern has been determined and spatial audio has been flexibly rendered. Figure 12D In this study, a reference spatial pattern has been established for a person 1220a seated at breakfast table 1240, and spatial audio has been flexibly rendered. It can be observed that the front soundstage orientation, as indicated by arrow 1205, does not necessarily correspond to any particular loudspeaker within environment 1100. As the listener's position and orientation change, the responsibility of the loudspeakers for rendering the various components of the spatial mix also changes.

[0420] for Figures 12A to 12D In any of the diagrams, person 1220a hears the spatial mix as expected for each position and orientation shown. However, the experience may be suboptimal for additional listeners in the space. Figure 12E This example illustrates the rendering of the reference space pattern when two listeners are in different locations within the listening environment. Figure 12E A reference space mode rendering is depicted for person 1220a on the sofa and person 1220b standing in the kitchen. In this example, the rendering may be optimal for person 1220a, but person 1220b, given his / her position, will hear signals primarily from the surround field and a small amount from the front field.

[0421] In this context, and in other situations where multiple people may move around in unpredictable ways (e.g., parties), a rendering mode more suited to this distributed audience is needed. Figure 13BA distributed spatial rendering pattern according to an example embodiment is depicted. In this example of the distributed spatial pattern, the front soundstage is now rendered uniformly across the entire listening space, rather than just from the position in front of the listener on the sofa. This distribution of the front soundstage is represented by multiple arrows 1305d surrounding a cloud-like line 1335, all arrows 1305d having the same or approximately the same length. The intended meaning of arrows 1305d is that all the depicted listeners (persons 1220a to 1220f) can hear that portion of the mix equally well, regardless of the position of the multiple listeners. However, if this uniform distribution is applied to all components of the mix, all spatial aspects of the mix will be lost; people 1220a to 1220f will essentially hear mono audio. To maintain some sense of space, the left and right surround components of the mix, represented by arrows 1210a and 1210b respectively, are still rendered spatially. (In many instances, there can be left and right surround, left and right rear surround, overhead surround, and dynamic audio objects with spatial positions within that space. Arrows 1210a and 1210b are intended to represent the left and right portions of all these possibilities.) And to maximize the perceived sense of space, the areas on which these components are spatialized are expanded to more completely cover the entire listening space, including the space previously occupied only by the front soundstage. By... Figure 13B The relatively slender arrows 1210a and 1210b shown in the image are... Figure 12A Comparing the relatively short arrows 1210a and 1210b shown, it can be understood that this extended area, on which the surrounding component is rendered, is located. Furthermore, Figure 12A Arrows 1210a and 1210b, which represent the surround component in the reference spatial mode, shown in the diagram, extend approximately from the side of person 1220a to the rear of the listening environment and do not extend to the front sound field area of ​​the listening environment.

[0422] In this example, care must be taken when implementing a uniform distribution of the front soundstage and an expanded spatialization of the surround components so that the perceived loudness of these components is largely maintained compared to a rendering against a reference spatial pattern. The goal is to alter the spatial impression of these components to optimize for multiple people while still maintaining the relative level of each component in the mix. For example, it would be undesirable if the front soundstage became twice as loud as the surround components due to its uniform distribution.

[0423] To switch between various reference rendering modes and distributed rendering modes in the example embodiments, in some examples, the user can interact with a voice assistant associated with a system of orchestrated speakers. For example, to play audio in reference spatial mode, the user can say a wake word (e.g., “Listen, Dolby”) to the voice assistant, and then say the command “Play [insert content name]” or “Play [insert content name] in personal mode.” Then, based on recordings from various microphones associated with the system, the system can automatically determine the user’s location and orientation, or the predefined area closest to the user among several predefined areas, and begin playing audio in the reference mode corresponding to that determined location. To play audio in distributed spatial mode, the user can say different commands, such as “Play [insert content name] in distributed mode.”

[0424] Alternatively or additionally, the system can be configured to automatically switch between reference mode and distributed mode based on other inputs. For example, the system may have means for automatically determining how many listeners are in the space and their locations. This can be achieved, for example, by monitoring speech activity in the space from associated microphones and / or by using other associated sensors (such as one or more cameras). In this case, the system may also be configured to switch between reference spatial mode (e.g., ...) Figure 12E The depiction in the text) and the fully distributed spatial model (such as Figure 13B The mechanism for continuously changing the rendering between points (as depicted in the diagram) can be calculated as, for example, a function of the number of people reported in the space.

[0425] Figure 12A , Figure 14A and Figure 14B The behavior is illustrated. In Figure 12A In this case, the system only detects a single listener (person 1220a) on the sofa facing the TV, and therefore the rendering mode is set to a reference space mode for that listener's position and orientation. Figure 14A This describes a partially distributed spatial rendering mode based on an example. Figure 14A In the image, two additional people (people 1220e and 1220f) are detected behind person 1220a, and the rendering mode is set at a point between the reference spatial mode and the fully distributed spatial mode. This is depicted as some of the front soundstage (arrows 1305a, 1305b, and 1305c) being pulled back towards the additional listeners (people 1220e and 1220f), but still emphasizing the position of the front soundstage in the reference spatial mode more. Compared to the length of arrows 1305b and 1305c, this emphasis is in Figure 14AThe relative lengths of arrows 1205 and 1305a are indicated by this. Similarly, as indicated by the lengths and positions of arrows 1210a and 1210b, the surround field extends only partially toward the position of the front sound field of the reference spatial pattern.

[0426] Figure 14B A fully distributed spatial rendering mode is described based on an example. In some examples, the system may have detected a large number of listeners (humans 1220a, 1220e, 1220f, 1220g, 1220h, and 1220i) spanning the entire space, and the system may have automatically set the rendering mode to fully distributed spatial mode. In other examples, the rendering mode may have been set based on user input. Fully distributed spatial mode in... Figure 14B The length of arrow 1305d is uniform or substantially uniform, and the length and position of arrows 1210a and 1210b are indicated by the length of arrow 1305d and the length and position of arrows 1210a and 1210b.

[0427] In the previous example, a portion of the spatial mix rendered in a more uniformly distributed manner in distributed rendering mode was designated as the front sound field. This makes sense in many spatial mixing scenarios because traditional mixing practices typically place the most important parts of the mix (such as dialogue in a film and lead vocals, drums, and bass guitar in music) in the front sound field. This is correct for most 5.1 and 7.1 surround sound mixes, as well as stereo content mixed to 5.1 or 7.1 using algorithms such as Dolby Pro-Logic or Dolby Surround, where the front sound field is given by the left, right, and center channels. This is also correct for many object-based audio mixes such as Dolby Atmos, where audio data can be designated as the front sound field based on spatial metadata indicating a (x,y) spatial location with y < 0.5. However, for object-based audio, mixing engineers are free to place audio anywhere in 3D space. Specifically, for object-based music, mixing engineers have begun to break with traditional mixing conventions and place elements considered essential to the mix (such as vocals) in unconventional locations (such as overhead). In this context, it becomes difficult to formulate simple rules to determine which components of the mix are suitable for rendering in a more distributed spatial manner for a distributed rendering mode. Object-based audio already contains metadata associated with each of its constituent audio signals, describing where the signal should be rendered in 3D space. In some implementations, to address the described problem, additional metadata can be added, allowing content creators to flag specific signals as suitable for more distributed spatial rendering in a distributed rendering mode. During rendering, the system can use this metadata to select the components of the mix to apply more distributed rendering. This gives content creators control over how they emit a distributed rendering mode for a particular content segment.

[0428] In some alternative implementations, the control system may be configured to implement a content type classifier to identify one or more elements in the audio data that should be rendered in a more spatially distributed manner. In some examples, the content type classifier may refer to content type metadata (e.g., metadata indicating that the audio data is dialogue, vocals, percussion, bass guitar, etc.) to determine whether the audio data should be rendered in a more spatially distributed manner. According to some such implementations, the content type metadata to be rendered in a more spatially distributed manner may be selectable by a user, for example, via a GUI displayed on a display device based on user input.

[0429] The exact mechanism for rendering one or more elements of a spatial audio mix in a more spatially distributed manner than in the reference spatial mode can vary between different embodiments, and this disclosure is intended to cover all such mechanisms. One example mechanism involves creating multiple copies of each such element, wherein multiple associated rendering locations are more evenly distributed across the listening space. In some embodiments, the number and / or the number of rendering locations for the distributed spatial mode can be user-selectable, while in other embodiments, the number and / or the number of rendering locations for the distributed spatial mode can be preset. In some such embodiments, the user can select multiple rendering locations for the distributed spatial mode, and the rendering locations can be preset, for example, evenly spaced throughout the listening environment. The system then renders all these copies at a set of distributed locations, as opposed to rendering the original single element at its original intended location. According to some embodiments, the copies can be level-corrected such that the perceived level associated with the combined rendering of all copies is the same as or substantially the same as the level of the original single element in the reference rendering mode (e.g., within threshold decibels such as 2dB, 3dB, 4dB, 5dB, 6dB, etc.).

[0430] More sophisticated mechanisms can be implemented in the context of flexible rendering systems like CMAP or FV, or a hybrid of the two. In these systems, each element of the spatial mix is ​​rendered at a specific location in space; associated with each element can be a hypothetical fixed location, such as the canonical location of a channel in a 5.1 or 7.1 surround sound mix, or a time-varying location, such as in the case of object-based audio (like Dolby Atmos).

[0431] Figure 15 Example rendering positions for CMAP and FV rendering systems are depicted on a 2D plane. Each numbered small circle represents an example rendering position, and the rendering system is capable of rendering elements of the spatial mix at any position on or within circle 1500. In this example, the positions marked L, R, C, Lss, Rss, Lrs, and Rrs on circle 1500 represent the fixed specification rendering positions for the seven full-range channels of a 7.1 surround mix: left (L), right (R), center (C), left surround (Lss), right surround (Rss), left rear surround (Lrs), and right rear surround (Rrs). In this context, the rendering positions near L, R, and C are considered the front soundstage. For the reference rendering mode (also referred to herein as the "reference spatial mode"), it is assumed that the listener is located at the center of the large circle and facing the C rendering position. Reference renderings for various listening positions and orientations are depicted below. Figures 12A to 12D Any one of them can Figure 15 The central concept is a superposition of ideas on top of the listener, and Figure 15Additionally, rotation and scaling were performed to align position C with the position of the front sound field (arrow 1205), and Figure 15 The circle 1500 surrounds the cloud-like line 1235. Then, the resulting alignment describes the... Figures 12A to 12D Any of the speakers in Figure 15 The relative proximity of any of the rendering positions in the system. In some implementations, this proximity largely governs the relative activation of the speakers when rendering elements of spatial mixing at specific locations for both CMAP and FV rendering systems.

[0432] When mixing spatial audio in a studio, speakers are typically placed at a uniform distance around the listening position. In most cases, no speakers are located within the boundaries of the resulting circle or hemisphere. This differs when the audio is placed "in the room" (e.g., in...). Figure 15 In rendering, the center of the sound (of the speaker) tends to trigger all surrounding speakers to achieve a "sound from nowhere." In CMAP and FV rendering systems, a similar effect can be achieved by altering the proximity penalty term in the cost function that governs speaker activation. Specifically, for Figure 15 The rendering position around the perimeter of the circle 1500 is subject to a proximity penalty that completely penalizes the use of speakers far from the desired rendering position. Thus, speakers only near the desired rendering position are activated in a large quantity. As the desired rendering position moves towards the center of the circle (radius zero), the proximity penalty decreases to zero, resulting in no speaker priority at the center. The corresponding result for a rendering position with a radius of zero is a completely uniform perceived distribution of audio across the listening space, which is precisely the desired result for certain elements of mixing in the widest spatial rendering mode.

[0433] Given this behavior of CMAP and FV systems at a radius of zero, a more spatially distributed rendering of any element in a spatial mix can be achieved by warping the expected spatial location of any element toward the zero-radius point. This warping can be continuous between the original expected location and the zero radius, thereby providing natural, continuous control between the reference spatial pattern and various distributed spatial patterns. Figure 16A , Figure 16B , Figure 16C and Figure 16D It shows the application to Figure 15 All rendering points are used as examples to implement various distributed spatial rendering modes of distortion. Figure 16D Describes the application Figure 15This is an example of a distortion that uses all rendering points to achieve a fully distributed rendering mode. You can see that points L, R, and C (front soundstage) have been folded to zero radius, thus ensuring they are rendered in a completely uniform way. Additionally, the Lss and Rss rendering points have been pulled along the perimeter of the circle towards the original front soundstage, causing a spatialized surround field (Lss, Rss, Lbs, and Rbs) to envelop the entire listening area. This distortion is applied to the entire rendering space, and you can see the distortion from... Figure 15 All rendering points have been distorted to Figure 16D A new position that corresponds to the distortion of the position in the 7.1 specification. Figure 16D The spatial pattern cited in this paper is an example of what may be referred to as the “most widely distributed spatial pattern” or the “fully distributed spatial pattern”.

[0434] Figure 16A , Figure 16B and Figure 16C It shows Figure 15 Distributed spatial patterns represented in the middle and Figure 16D Various examples of intermediate distributed space patterns between distributed space patterns are represented in the figure. Figure 16B express Figure 15 Distributed spatial patterns represented in the middle and Figure 16D The midpoint between distributed spatial patterns is represented in the diagram. Figure 16A express Figure 15 Distributed spatial patterns represented in the middle and Figure 16B The midpoint between distributed spatial patterns is represented in the diagram. Figure 16C express Figure 16B Distributed spatial patterns represented in the middle and Figure 16D The midpoint between distributed spatial patterns is represented in the diagram.

[0435] Figure 17 An example of a GUI that a user can use to select a rendering mode is shown. According to some implementations, the control system can control a display device (e.g., a cellular phone) to display GUI 1700 or a similar GUI on the screen. The display device may include a sensor system (such as a touch sensor system or a gesture sensor system that is close to the display (e.g., covering the display or below the display)). The control system can be configured to receive user input via GUI 1700 in the form of sensor signals from the sensor system. The sensor signals may correspond to user touches or gestures corresponding to elements of GUI 1700.

[0436] According to this example, the GUI includes a virtual slider 1701 that a user can interact with to select a rendering mode. As indicated by arrow 1703, the user can move the slider in either direction along track 1707. In this example, line 1705 indicates the position of the virtual slider 1701 corresponding to a reference space mode (such as one of the reference space modes disclosed herein). Other embodiments may provide additional features that the user can interact with on the GUI, such as virtual knobs or dials. According to some embodiments, after selecting a reference space mode, the control system may present, as shown in some embodiments, a rendering mode. Figure 13A The GUI shown here may be another such GUI that allows users to select the listener position and orientation for a reference space pattern.

[0437] In this example, line 1725 indicates the most widely distributed spatial pattern (such as...). Figure 13B The position of the virtual slider 1701 corresponding to the distributed spatial pattern shown in the figure. According to this embodiment, lines 1710, 1715, and 1720 indicate the position of the virtual slider 1701 corresponding to the intermediate spatial pattern. In this example, the position of line 1710 is as shown in the figure. Figure 16A The intermediate space pattern corresponds to the intermediate space pattern of the middle space pattern, etc. Here, the position of line 1715 is as follows: Figure 16B The intermediate space pattern corresponds to the intermediate space pattern, etc. In this embodiment, the position of line 1720 is as follows: Figure 16C This corresponds to intermediate space modes, such as intermediate space modes. According to this example, a user can interact with the "Apply" button (e.g., touch the "Apply" button) to instruct the control system to implement the selected rendering mode.

[0438] However, other implementations can provide users with alternative ways to select one of the aforementioned distributed space modes. According to some examples, a user can speak a voice command, such as, "Play [insert content name] in semi-distributed mode." "Semi-distributed mode" can be related to... Figure 17 The distributed mode indicated by line 1715 in GUI 1700 corresponds to this mode. According to some examples, a user can speak a voice command, such as, "Play [insert content name] in quarter-distributed mode." "Quarter-distributed mode" can correspond to the distributed mode indicated by line 1710.

[0439] Figure 18This is a flowchart outlining an example of a method that can be performed by those devices or systems disclosed herein. As with other methods described herein, the blocks of method 1800 need not be performed in the indicated order. In some embodiments, one or more blocks of method 1800 may be performed simultaneously. Furthermore, some embodiments of method 1800 may include more or fewer blocks than those shown and / or described. The blocks of method 1800 may be performed by one or more devices, which may be (or may include) a control system, such as… Figure 1A The control system 110 shown and described above, or one of other disclosed control system examples.

[0440] In this embodiment, block 1805 relates to receiving audio data, including one or more audio signals and associated spatial data, by a control system and via an interface system. In this example, the spatial data indicates the expected perceived spatial location corresponding to the audio signals. Here, the spatial data includes channel data and / or spatial metadata.

[0441] In this example, box 1810 relates to the control system determining the rendering mode. In some instances, determining the rendering mode may involve receiving a rendering mode indication via an interface system. Receiving the rendering mode indication may, for example, involve receiving a microphone signal corresponding to a voice command. In some examples, receiving the rendering mode indication may involve receiving sensor signals corresponding to user input via a graphical user interface. The sensor signals may, for example, be touch sensor signals and / or gesture sensor signals.

[0442] In some implementations, receiving a rendering mode indication may involve receiving an indication of the number of people in the listening area. According to some such examples, the control system can be configured to determine the rendering mode based at least in part on the number of people in the listening area. In some such examples, the indication of the number of people in the listening area may be based on microphone data from a microphone system and / or image data from a camera system.

[0443] according to Figure 18The example shown, block 1815, relates to a control system rendering audio data for reproduction via a set of loudspeakers in the environment, according to a rendering mode determined in block 1810, thereby producing a rendered audio signal. In this example, rendering the audio data involves determining the relative activation of a set of loudspeakers in the environment. Here, the rendering mode is variable between a reference spatial mode and one or more distributed spatial modes. In this embodiment, the reference spatial mode has an assumed listening position and orientation. According to this example, in one or more distributed spatial modes, one or more elements of the audio data are each rendered in a more spatially distributed manner than in the reference spatial mode. In this example, in one or more distributed spatial modes, the spatial positions of the remaining elements of the audio data are distorted such that the spatial positions of the remaining elements span the rendered space of the environment more completely than in the reference spatial mode.

[0444] In some implementations, rendering one or more elements of audio data in a more spatially distributed manner than in a reference space mode may involve creating copies of said one or more elements. Some such implementations may involve rendering all copies simultaneously at a distributed set of locations across an environment.

[0445] According to some implementations, rendering can be based on CMAP, FV, or a combination thereof. Rendering one or more elements of audio data in a more spatially distributed manner than in a reference space mode may involve distorting the rendering position of each of the one or more elements toward a zero radius.

[0446] In this example, box 1820 relates to providing a rendered audio signal by a control system and via an interface system to at least some of the loudspeakers in the set of loudspeakers in the environment.

[0447] According to some implementations, the rendering mode can be selected from a continuum of rendering modes ranging from a reference spatial mode to the most widely distributed spatial modes. In some such implementations, the control system may be further configured to determine the assumed listening position and / or orientation of the reference spatial mode based on reference spatial mode data received via an interface system. According to some such implementations, the reference spatial mode data may include microphone data from a microphone system and / or image data from a camera system. In some such examples, the reference spatial mode data may include microphone data corresponding to a voice command. Alternatively or additionally, the reference spatial mode data may include microphone data corresponding to the location of one or more utterances of a person in the listening environment. In some such examples, the reference spatial mode data may include image data indicating the position and / or orientation of a person in the listening environment.

[0448] However, in some instances, the apparatus or system may include a display device and a sensor system proximate to the display device. The control system may be further configured to control the display device to present a graphical user interface. Receiving reference spatial pattern data may involve receiving sensor signals corresponding to user input via the graphical user interface.

[0449] According to some implementations, one or more elements of the audio data rendered in a more spatially distributed manner may correspond to front field data, musical vocals, dialogue, bass guitar, percussion, and / or other solo or lead instruments. In some instances, the front field data may include the left, right, or center signal of audio data received or upmixed to Dolby 5.1, Dolby 7.1, or Dolby 9.1 format. In some examples, the front field data may include audio data received in Dolby Atmos format and having spatial metadata indicating a (x, y) spatial location, where y < 0.5.

[0450] In some instances, the audio data may include spatial distribution metadata that indicates which elements of the audio data should be rendered in a more spatially distributed manner. In some such examples, the control system may be configured to identify one or more elements of the audio data to be rendered in a more spatially distributed manner based on the spatial distribution metadata.

[0451] Alternatively or additionally, the control system may be configured to implement a content type classifier to identify one or more elements of the audio data to be rendered in a more spatially distributed manner. In some examples, the content type classifier may refer to content type metadata (e.g., metadata indicating that the audio data is dialogue, vocals, percussion, bass guitar, etc.) to determine whether the audio data should be rendered in a more spatially distributed manner. According to some such implementations, the content type metadata to be rendered in a more spatially distributed manner may be selectable by the user, for example, via a GUI displayed on a display device based on user input.

[0452] Alternatively or additionally, the content type classifier can be combined with the rendering system to operate directly on the audio signal. For example, a classifier can be implemented using neural networks trained on various content types to analyze the audio signal and determine whether the audio signal belongs to any content type (vocals, lead guitar, drums, etc.) that might be considered suitable for rendering in a more spatially distributed manner. This classification can be performed continuously and dynamically, and the resulting classification can be continuously and dynamically adjusted to the signal set rendered in a more spatially distributed manner. Some such implementations may involve using techniques such as neural networks to implement such a dynamic classification system according to methods known in the art.

[0453] In some examples, at least one of one or more distributed spatial patterns may involve applying time-varying corrections to the spatial location of at least one element. According to some such examples, the time-varying correction may be a periodic correction. For example, a periodic correction may involve rotating one or more rendering positions around the periphery of the listening environment. According to some such implementations, the periodic correction may involve the rhythm of music reproduced in the environment, the beat of music reproduced in the environment, or one or more other features of audio data reproduced in the environment. For example, some such periodic corrections may involve alternating between two, three, four, or more rendering positions. The alternation may correspond to the beat of music reproduced in the environment. In some implementations, the periodic correction may be selectable based on user input, for example, based on one or more voice commands, based on user input received via a GUI, etc.

[0454] Figure 19 An example of the geometric relationship between three audio devices in an environment is illustrated. In this example, environment 1900 is a room including a television 1901, a sofa 1903, and five audio devices 1905. According to this example, the audio devices 1905 are located at positions 1 to 5 in environment 1900. In this embodiment, each audio device 1905 includes a microphone system 1920 having at least three microphones and a speaker system 1925 including at least one speaker. In some embodiments, each microphone system 1920 includes a microphone array. According to some embodiments, each audio device 1905 may include an antenna system containing at least three antennas.

[0455] As with other examples disclosed in this article, Figure 19 The types, quantities, and arrangements of the components shown are merely examples. Other embodiments may have components of different types, quantities, and arrangements, such as more or fewer audio devices 1905, audio devices 1905 in different locations, etc.

[0456] In this example, the vertices of triangle 1910a are at positions 1, 2, and 3. Here, triangle 1910a has sides 12, 23a, and 13a. According to this example, the angle between sides 12 and 23 is θ2, the angle between sides 12 and 13a is θ1, and the angle between sides 23a and 13a is θ3. These angles can be determined from DOA data, as described in more detail below.

[0457] In some implementations, only the relative lengths of the triangle sides can be determined. In alternative implementations, the actual lengths of the triangle sides can be determined. According to some such implementations, the actual lengths of the triangle sides can be estimated based on TOA data, for example, based on the arrival times of sounds generated by an audio device located at one vertex of the triangle and detected by an audio device located at another vertex of the triangle. Alternatively or additionally, the length of the triangle sides can be estimated based on electromagnetic waves generated by an audio device located at one vertex of the triangle and detected by an audio device located at another vertex of the triangle. For example, the length of the triangle sides can be estimated based on the signal strength of electromagnetic waves generated by an audio device located at one vertex of the triangle and detected by an audio device located at another vertex of the triangle. In some implementations, the length of the triangle sides can be estimated based on the phase shift of the detected electromagnetic waves.

[0458] Figure 20 It shows Figure 19 Another example of the geometric relationship between three audio devices in an environment is shown. In this example, the vertices of triangle 1910b are at positions 1, 3, and 4. Here, triangle 1910b has sides 13b, 14, and 34a. According to this example, the angle between sides 13b and 14 is θ4, the angle between sides 13b and 34a is θ5, and the angle between sides 34a and 14 is θ6.

[0459] By comparison Figure 11 As shown in Figure 12, it can be observed that the length of side 13a of triangle 1910a should be equal to the length of side 13b of triangle 1910b. In some embodiments, the side length of a triangle (e.g., triangle 1910a) can be assumed to be correct, and the length of the side shared by adjacent triangles will be constrained to that length.

[0460] Figure 21A It shows Figure 19 and Figure 20 The two triangles depicted in the image do not correspond to any other audio devices or environmental features. Figure 21A Estimates of the side lengths and angular orientations of triangles 1910a and 1910b are shown. Figure 21AIn the example shown, the length of side 13b of triangle 191Ob is constrained to be the same as the length of side 13a of triangle 1910a. The lengths of the other sides of triangle 1910b are scaled proportionally to the changes in the length of side 13b. The resulting triangle 1910b' is... Figure 21A It is shown as adjacent to triangle 1910a.

[0461] According to some implementations, the side lengths of other triangles adjacent to triangles 1910a and 1910b can be determined in a similar manner until the locations of all audio devices in environment 1900 have been determined.

[0462] Some examples of audio device location can be presented as follows. Each audio device can report the DOA of each other audio device in the environment (e.g., a room) based on the sound produced by each other audio device in the environment. The Cartesian coordinates of the i-th audio device can be represented as x. i =[x i y i ] T The superscript T indicates the vector transpose. Given M audio devices in an environment, i = {1...M}.

[0463] Figure 21B An example of estimating the interior angles of a triangle formed by three audio devices is shown. In this example, the audio devices are i, j, and k. The DOA of the sound source emanating from device j, as observed from device i, can be represented as θ. ji The DOA of the sound source emitted from device k, as observed from device i, can be represented as θ. ki .exist Figure 21B In the example shown, θ ji and θ ki It is measured from axis 2105a, the orientation of which is arbitrary, and the axis may, for example, correspond to the orientation of audio device i. The interior angle α of triangle 2110 can be expressed as α = θ. ki -θ ji It can be observed that the calculation of interior angle α does not depend on the orientation of axis 2105α.

[0464] exist Figure 21B In the example shown, θ ij and θ kj It is measured from axis 2105b, the orientation of which is arbitrary and can correspond to the orientation of audio device j. The interior angle b of triangle 2110 can be expressed as b = θ. ij -θ kj Similarly, in this example, θ jk and θ ikIt is measured from axis 2105c. The interior angle c of triangle 2110 can be expressed as c = θ. jk -θ ik .

[0465] In the presence of measurement error, a + b + c ≠ 180°. Robustness can be improved by predicting each angle from the other two angles and averaging the results, for example, as shown below:

[0466]

[0467] In some implementations, edge lengths (A, B, C) can be calculated (up to scaling errors) by applying a sine rule. In some examples, an arbitrary value, such as 1, can be assigned to an edge length. For example, by setting A = 1 and setting the vertex... If placed at the origin, the positions of the remaining two vertices can be calculated as follows:

[0468]

[0469] However, arbitrary rotations are acceptable.

[0470] According to some implementations, the process of triangle parameterization can be repeated for all possible subsets of three audio devices in the environment, within a size of... Enumerate in the superset ζ. In some examples, T l This can represent the first triangle. Triangles may not be enumerated in any particular order, depending on the implementation. Due to possible errors in DOA and / or side length estimation, triangles may overlap and may not be perfectly aligned.

[0471] Figure 22 It outlines what can be achieved by, for example Figure 1A The flowchart illustrates an example of a method performed by the apparatus shown. As with other methods described herein, the blocks of method 2200 need not be performed in the indicated order. Furthermore, this method may include more or fewer blocks than those shown and / or described. In this embodiment, method 2200 involves estimating the position of a speaker in the environment. The blocks of method 2200 may be performed by one or more devices, which may be (or may include) Figure 1A The device 100 shown in the figure.

[0472] In this example, box 2205 involves obtaining the direction of arrival (DOA) data for each of a plurality of audio devices. In some examples, the plurality of audio devices may include all audio devices in the environment, such as Figure 19 All audio devices shown in the figure are 1905.

[0473] However, in some instances, multiple audio devices may only include a subset of all audio devices in the environment. For example, multiple audio devices may include all smart speakers in the environment, but exclude one or more other audio devices in the environment.

[0474] DOA data can be obtained in various ways, depending on the specific implementation. In some instances, determining DOA data may involve determining the DOA data of at least one of a plurality of audio devices. For example, determining DOA data may involve receiving microphone data from each of a plurality of audio device microphones corresponding to a single audio device among the plurality of audio devices, and determining the DOA data of the single audio device at least in part based on the microphone data. Alternatively or additionally, determining DOA data may involve receiving antenna data from one or more antennas corresponding to a single audio device among the plurality of audio devices, and determining the DOA data of the single audio device at least in part based on the antenna data.

[0475] In some such examples, a single audio device can determine its own DOA data. According to some implementations, each of a plurality of audio devices can determine its own DOA data. However, in other implementations, another device, which may be local or remote, can determine the DOA data of one or more audio devices in the environment. According to some implementations, a server can determine the DOA data of one or more audio devices in the environment.

[0476] According to this example, box 2210 involves determining the interior angles of each of a plurality of triangles based on DOA data. In this example, each of the plurality of triangles has vertices corresponding to the audio device positions of the three audio devices. Some such examples have been described above.

[0477] Figure 23 An example is shown where each audio device in the environment is a vertex of multiple triangles. The sides of each triangle correspond to the distance between two audio devices 1905.

[0478] In this implementation, box 2215 relates to determining the side length of each side of each triangle. (The sides of a triangle may also be referred to herein as “edges.”) According to this example, the side length is based at least in part on the interior angles. In some instances, the side lengths can be calculated by determining a first length of the first side of the triangle and determining the lengths of the second and third sides of the triangle based on the interior angles of the triangle. Some such examples have been described above.

[0479] According to some such implementations, determining the first length may involve setting the first length to a predetermined value. However, in some examples, determining the first length may be based on time-of-arrival data and / or received signal strength data. In some implementations, the time-of-arrival data and / or received signal strength data may correspond to sound waves detected by a second audio device in the environment from a first audio device in the environment. Alternatively or additionally, the time-of-arrival data and / or received signal strength data may correspond to electromagnetic waves (e.g., radio waves, infrared waves, etc.) detected by a second audio device in the environment from a first audio device in the environment.

[0480] According to this example, box 2220 involves performing a forward alignment process that aligns each of the plurality of triangles in a first order. According to this example, the forward alignment process produces a forward alignment matrix.

[0481] Based on some such examples, it is expected that the triangle will have an edge (x) i x j Alignment is done in a way that matches adjacent edges, for example, as shown below. Figure 21A As shown in the diagram and described above. Let ε be of size... The set of all edges. In some such implementations, box 2220 may involve traversing ε and aligning the common edges of the triangle in a forward order by forcing the edges to align with the edges of previously aligned edges.

[0482] Figure 24 An example of a part of the forward alignment process is provided. Figure 24 The numbers 1 to 5, shown in bold, correspond to the locations of the audio devices shown in Figures 1, 2, and 5. Figure 24 The order of the forward alignment process shown and described herein is merely an example.

[0483] In this example, as in Figure 21A In this case, the length of side 13b of triangle 1910b is forced to be the same as the length of side 13a of triangle 1910a. Figure 24 The resulting triangle 1910b' is shown, where the same interior angles are maintained. According to this example, the length of side 13c of triangle 1910c is also forced to be the same as the length of side 13a of triangle 1910a. Figure 24 The resulting triangle 1910c' is shown in the figure, where the same interior angles are maintained.

[0484] Next, in this example, the length of side 34b of triangle 1910d is forced to be the same as the length of side 34a of triangle 1910b'. Furthermore, in this example, the length of side 23b of triangle 1910d is forced to be the same as the length of side 23a of triangle 1910a. Figure 24The resulting triangle 1910d' is shown in Figure 5, where the same interior angles are maintained. Based on some such examples, the remaining triangles shown in Figure 5 can be processed in the same way as triangles 1910b, 1910c, and 1910d.

[0485] The result of the forward alignment process can be stored in a data structure. According to some examples, the result of the forward alignment process can be stored in a forward alignment matrix. For instance, the result of the forward alignment process can be stored in a matrix. In the diagram, N indicates the total number of triangles.

[0486] Multiple audio device position estimates will occur when the determination of DOA data and / or initial side lengths contains errors. These errors typically increase during the forward alignment process.

[0487] Figure 25 An example of multiple audio device position estimates that have already occurred during the forward alignment process is shown. In this example, the forward alignment process is based on a triangle with the seven audio device positions as vertices. Here, the triangle is not perfectly aligned due to additional errors in the DOA estimation. Figure 25 The positions of numbers 1 through 7 shown correspond to the estimated audio device positions generated by the forward alignment process. In this example, the audio device position estimate marked "1" is consistent, but the audio device position estimates for audio devices 6 and 7 show significant differences, as shown in the relatively large areas where numbers 6 and 7 are located.

[0488] return Figure 22 In this example, box 2225 relates to a reverse alignment process that aligns each of a plurality of triangles in a second order, which is the reverse of the first order. According to some implementations, the reverse alignment process may involve traversing ε in the same but reverse order as before. In alternative examples, the reverse alignment process may not be exactly the reverse of the order of operations of the forward alignment process. According to this example, the reverse alignment process produces a reverse alignment matrix, which can be represented herein as...

[0489] Figure 26 An example of a part of the reverse alignment process is provided. Figure 26 The numbers 1 to 5 shown in bold are... Figure 19 Figure 21 and Figure 23 The locations of the audio devices shown correspond to those in the diagram. Figure 26 The order of the reverse alignment process shown and described herein is merely an example.

[0490] exist Figure 26In the example shown, triangle 1910e is based on audio device positions 3, 4, and 5. In this implementation, it is assumed that the side length (or "edge") of triangle 1910e is correct, and the side lengths of adjacent triangles are forced to be consistent with it. According to this example, the length of side 45b of triangle 1910f is forced to be consistent with the length of side 45a of triangle 1910e. Figure 26 The resulting triangle 1910f' is shown, where the interior angles remain the same. In this example, the length of side 35b of triangle 1910c is forced to be the same as the length of side 35a of triangle 1910e. Figure 26 The resulting triangle 1910c is shown, where the interior angles remain the same. Based on some such examples, Figure 23 The remaining triangles shown can be processed in the same way as triangles 1910c and 1910f until the reverse alignment process has included all the remaining triangles.

[0491] Figure 27 An example of multiple audio device position estimates that have already occurred during the reverse alignment process is shown. In this example, the reverse alignment process is based on [the method described above]. Figure 25 The triangle describes the locations of the seven audio device positions with the same vertices. Figure 27 The positions of numbers 1 through 7 shown correspond to the estimated audio device positions generated by the reverse alignment process. Again, the triangles are not perfectly aligned due to additional errors in the DOA estimation. In this example, the audio device position estimates labeled 6 and 7 are consistent, but the audio device position estimates for audio devices 1 and 2 show greater differences.

[0492] return Figure 22 Box 2230 relates to generating a final estimate of the location of each audio device based at least in part on the values ​​of the forward alignment matrix and the reverse alignment matrix. In some examples, generating the final estimate of the location of each audio device may involve translating and scaling the forward alignment matrix to produce a translated and scaled forward alignment matrix, and translating and scaling the reverse alignment matrix to produce a translated and scaled reverse alignment matrix.

[0493] For example, by moving the centroid to the origin and forcing the unit Frobenius norm (e.g., and Use this to fix translation and scaling.

[0494] According to some such examples, generating a final estimate of the position of each audio device may also involve generating a rotation matrix based on a translated and scaled forward alignment matrix and a translated and scaled reverse alignment matrix. The rotation matrix may include multiple estimated audio device positions for each audio device. For example, the optimal rotation between forward and reverse alignment can be found through singular value decomposition. In some such examples, generating the rotation matrix may involve performing singular value decomposition on the translated and scaled forward alignment matrix and the translated and scaled reverse alignment matrix, for example, as follows:

[0495]

[0496] In the aforementioned equations, U represents the matrix respectively. Let V denote the left singular vector of the matrix and ∑ denote the singular value matrix. The aforementioned equation produces the rotation matrix R = VU. T Matrix product VU T Generate a rotation matrix such that Optimally rotated to be with Alignment.

[0497] Based on some examples, when determining the rotation matrix R = VU T Then, the alignment can be averaged, for example, as follows:

[0498]

[0499] In some implementations, generating a final estimate of the location of each audio device may also involve averaging the estimated audio device locations for each audio device to produce a final estimate of the location of each audio device. Various disclosed implementations have proven robust, even when the DOA data and / or other calculations include significant errors. For example, due to overlapping vertices from multiple triangles, Containing the same nodes A total of several estimates are calculated. The final estimate is obtained by averaging across common nodes.

[0500] Figure 28 A comparison between the estimated audio device location and the actual audio device location is shown. Figure 28 In the example shown, the audio device location is the same as in the reference above. Figure 17 and Figure 19 The audio device positions estimated during the described forward alignment and reverse alignment processes correspond to each other. In these examples, the error in the DOA estimation has a standard deviation of 15 degrees. Nevertheless, the final estimate of each audio device position (each of the final estimates in...) Figure 28 (represented by "x") and the actual audio device location (each of the actual audio device locations is in...) Figure 28The circle in the middle corresponds well.

[0501] Much of the foregoing discussion pertains to the automatic positioning of audio devices. The following discussion expands upon some of the methods briefly described above for determining listener position and listener angular orientation. In the foregoing description, the term "rotation" was used in essentially the same way as the term "orientation" used below. For example, the "rotation" mentioned above could refer to a global rotation of the final speaker geometry, rather than the rotation of a single triangle during the process described above with reference to Figure 14 and below. This global rotation or orientation can be determined by the listener angular orientation, such as the direction the listener is looking, the direction the listener's nose is pointing, etc.

[0502] Various satisfactory methods for estimating listener position are described below. However, estimating listener angular orientation can be challenging. Some relevant methods are described in detail below.

[0503] Determining the listener's position and angular orientation allows for the realization of desired features, such as audio equipment directional positioning relative to the listener. Knowing the listener's position and angular orientation allows for determining, for example, which speakers in the environment are in front of the listener, which are behind, and which are closer to the center (if any).

[0504] After establishing a correlation between the audio device location and the listener's location and orientation, some implementations may involve providing audio device location data, audio device angular orientation data, listener location data, and listener angular orientation data to an audio rendering system. Alternatively or additionally, some implementations may involve an audio data rendering process that is at least partially based on the audio device location data, audio device angular orientation data, listener location data, and listener angular orientation data.

[0505] Figure 29 It outlines what can be achieved by, for example Figure 1A The flowchart illustrates an example of a method performed by the apparatus shown. As with other methods described herein, the blocks of method 2900 need not be performed in the indicated order. Furthermore, this method may include more or fewer blocks than those shown and / or described. In this example, the blocks of method 2900 are performed by a control system, which may be (or may include) Figure 1A The control system 110 is shown in the figure. As described above, in some embodiments, the control system 110 may reside in a single device, while in other embodiments, the control system 110 may reside in two or more devices.

[0506] In this example, box 2905 relates to obtaining the direction of arrival (DOA) data for each of multiple audio devices in the environment. In some examples, the multiple audio devices may include all audio devices in the environment, such as... Figure 27 All audio devices shown in the figure are 1905.

[0507] However, in some instances, multiple audio devices may only include a subset of all audio devices in the environment. For example, multiple audio devices may include all smart speakers in the environment, but exclude one or more other audio devices in the environment.

[0508] DOA data can be obtained in various ways, depending on the specific implementation. In some instances, determining DOA data may involve determining the DOA data of at least one of a plurality of audio devices. In some examples, DOA data can be obtained by controlling each of a plurality of loudspeakers in an environment to reproduce a test signal. For example, determining DOA data may involve receiving microphone data from each of a plurality of audio device microphones corresponding to a single audio device among a plurality of audio devices, and determining the DOA data of the single audio device at least in part based on the microphone data. Alternatively or additionally, determining DOA data may involve receiving antenna data from one or more antennas corresponding to a single audio device among a plurality of audio devices, and determining the DOA data of the single audio device at least in part based on the antenna data.

[0509] In some such examples, a single audio device can determine its own DOA data. According to some implementations, each of a plurality of audio devices can determine its own DOA data. However, in other implementations, another device, which may be local or remote, can determine the DOA data of one or more audio devices in the environment. According to some implementations, a server can determine the DOA data of one or more audio devices in the environment.

[0510] according to Figure 29 The example shown, box 2910, relates to generating audio device location data via a control system, based at least in part on DOA data. In this example, the audio device location data includes an estimate of the audio device location for each audio device referenced in box 2905.

[0511] Audio device location data can be, for example, coordinates in a coordinate system (such as a Cartesian, spherical, or cylindrical coordinate system). This coordinate system may be referred to herein as the audio device coordinate system. In some such examples, the audio device coordinate system may be oriented with reference to one of the audio devices in the environment. In other examples, the audio device coordinate system may be oriented with reference to an axis defined by a line between two audio devices in the environment. However, in still other examples, the audio device coordinate system may be oriented with reference to another part of the environment (such as a television, a wall in a room, etc.).

[0512] In some examples, box 2910 may refer to the above reference. Figure 22 The process described. According to some such examples, box 2910 may involve determining the interior angles of each of a plurality of triangles based on DOA data. In some instances, each of the plurality of triangles may have vertices corresponding to the audio device positions of the three audio devices. Some such methods may involve determining the side length of each side of each triangle based at least in part on the interior angles.

[0513] Some such methods may involve performing a forward alignment process that aligns each of a plurality of triangles in a first order to produce a forward alignment matrix. Some such methods may involve performing a reverse alignment process that aligns each of the plurality of triangles in a second order that is the reverse of the first order to produce a reverse alignment matrix. Some such methods may involve generating a final estimate of the position of each audio device based at least in part on the values ​​of the forward alignment matrix and the reverse alignment matrix. However, in some embodiments of method 2900, block 2910 may involve applying additional methods beyond those referenced above. Figure 22 Methods other than those described.

[0514] In this example, box 2915 relates to listener location data, determined via a control system, indicating the listener's location within the environment. For example, the listener location data may be referenced to an audio device coordinate system. However, in other examples, the coordinate system may be oriented with reference to the listener or a portion of the environment (such as a television, a room wall, etc.).

[0515] In some examples, box 2915 may involve prompting a listener (e.g., via audio cues from one or more loudspeakers in the environment) to speak one or more utterances and estimating the listener's location based on DOA data. The DOA data may correspond to microphone data obtained from multiple microphones in the environment. The microphone data may correspond to the detection of one or more utterances by the microphones. At least some microphones may be co-localized with the loudspeakers. According to some examples, box 2915 may involve a triangulation process. For example, box 2915 may involve triangulating the user's speech by finding the intersections between DOA vectors passing through the audio device, for example, as referenced below. Figure 30A As described. According to some implementations, block 2915 (or another operation of method 2900) may involve locating the origin of the audio device coordinate system and the origin of the listener coordinate system together after determining the listener's location. Locating the origin of the audio device coordinate system and the origin of the listener coordinate system together may involve transforming the audio device location from the audio device coordinate system to the listener coordinate system.

[0516] According to this embodiment, block 2920 relates to determining listener angular orientation data indicating listener angular orientation via a control system. For example, the listener angular orientation data may be obtained with reference to a coordinate system (such as an audio device coordinate system) used to represent listener position data. In some such examples, the listener angular orientation data may be obtained with reference to the origin and / or axes of the audio device coordinate system.

[0517] However, in some implementations, listener angular orientation data can be obtained with reference to an axis defined by the listener's position and another point in the environment (such as a television, audio device, wall, etc.). In some such implementations, the listener's position can be used to define the origin of the listener's coordinate system. In some such examples, the listener angular orientation data can be obtained with reference to the axis of the listener's coordinate system.

[0518] This document discloses various methods for implementing box 2920. According to some examples, the listener's angular orientation can correspond to the listener's viewing direction. In some such examples, the listener's viewing direction can be inferred, for example, by assuming the listener is viewing a specific object (such as a television), referring to listener position data. In some such implementations, the listener's viewing direction can be determined based on the listener's position and the television's position. Alternatively or additionally, the listener's viewing direction can be determined based on the listener's position and the television speaker's position.

[0519] However, in some examples, the listener's viewing direction can be determined based on listener input. According to some such examples, listener input may include inertial sensor data received from a device held by the listener. The listener can use said device to point to a location in the environment, for example, a location corresponding to the direction the listener is facing. For instance, the listener can use said device to point to a loudspeaker (the loudspeaker that is reproducing sound). Therefore, in such examples, the inertial sensor data may include inertial sensor data corresponding to the loudspeaker that is emitting the sound.

[0520] In some such instances, listener input may include an indication of the audio device selected by the listener. In some examples, the audio device indication may include inertial sensor data corresponding to the selected audio device.

[0521] However, in other examples, instructions for audio devices can be made based on one or more of the listener's utterances (e.g., "The TV is in front of me now.", "Speaker 2 is in front of me now.", etc.). Other examples of determining listener orientation data based on one or more of the listener's utterances are described below.

[0522] according to Figure 29 The example shown, block 2925, relates to determining audio device angular orientation data via a control system, the audio device angular orientation data indicating the audio device angular orientation of each audio device relative to the listener position and the listener angular orientation. According to some such examples, block 2925 may relate to rotating the audio device coordinates about a point defined by the listener position. In some embodiments, block 2925 may relate to transforming audio device position data from an audio device coordinate system to a listener coordinate system. Some examples are described below.

[0523] Figure 30A It shows Figure 29 Examples of some boxes. According to some such examples, the audio device location data includes an estimate of the audio device location of each of audio devices 1 to 5 with reference to audio device coordinate system 3007. In this embodiment, audio device coordinate system 3007 is a Cartesian coordinate system with the location of the microphone of audio device 2 as its origin. Here, the x-axis of audio device coordinate system 3007 corresponds to line 3003 between the microphone locations of audio device 2 and audio device 1.

[0524] In this example, the listener's location is determined by prompting a listener 3005 seated on sofa 1903 (e.g., via an audio prompt from one or more loudspeakers in environment 3000a) to speak one or more utterances 3027 and by estimating the listener's location based on time-of-arrival (TOA) data. The TOA data corresponds to microphone data obtained from multiple microphones in the environment. In this example, the microphone data corresponds to the detection of one or more utterances 3027 by microphones from at least some (e.g., three, four, or all five) of the audio devices 1 to 5.

[0525] Alternatively or additionally, the listener's location is determined based on DOA data provided by the microphones of at least some (e.g., two, three, four, or all five) of the audio devices 1 to 5. According to some such examples, the listener's location can be determined based on the intersection of lines 3009a, 3009b, etc., corresponding to the DOA data.

[0526] According to this example, the listener's position corresponds to the origin of the listener coordinate system 3020. In this example, the listener's angular orientation data is indicated by the y' axis of the listener coordinate system 3020, which corresponds to the line 3013a between the listener's head 3010 (and / or the listener's nose 3025) and the speaker 3030 of the television 101. Figure 30A In the example shown, line 3013a is parallel to the y' axis. Therefore, angle This represents the angle between the y-axis and the y'-axis. In this example, Figure 29 The frame 2925 can involve rotating the audio device coordinates by an angle around the origin of the listener coordinate system 3020. Therefore, although the origin of the audio device coordinate system 3007 is shown as... Figure 30A Corresponding to audio device 2 in the listener coordinate system 3020, but some implementations involve rotating the audio device coordinates by an angle around the origin of the listener coordinate system 3020. Previously, the origin of the audio device coordinate system 3007 and the origin of the listener coordinate system 3020 were jointly located. This joint location can be performed through a coordinate transformation from the audio device coordinate system 3007 to the listener coordinate system 3020.

[0527] In some examples, the location of speaker 3030 and / or television 1901 can be determined by emitting sound from the speaker and estimating its location based on DOA and / or TOA data. This can correspond to sound detection by microphones of at least some (e.g., 3, 4, or all 5) of the audio devices 1 to 5. Alternatively or additionally, the location of speaker 3030 and / or television 1901 can be determined by prompting the user to approach the television and locating the user's speech using DOA and / or TOA data. This can correspond to sound detection by microphones of at least some (e.g., 3, 4, or all 5) of the audio devices 1 to 5. This method may involve triangulation. Such examples can be advantageous in cases where speaker 3030 and / or television 1901 do not have associated microphones.

[0528] In some other examples, where speaker 3030 and / or television 1901 do indeed have associated microphones, the positions of speaker 3030 and / or television 1901 can be determined according to TOA or DOA methods (such as the DOA method disclosed herein). According to some of these methods, the microphone can be co-located with speaker 3030.

[0529] According to some embodiments, the speaker 3030 and / or television 1901 may have an associated camera 3011. The control system may be configured to capture an image of the listener's head 3010 (and / or the listener's nose 3025). In some such examples, the control system may be configured to determine a line 3013a between the listener's head 3010 (and / or the listener's nose 3025) and the camera 3011. Listener angular orientation data may correspond to line 3013a. Alternatively or additionally, the control system may be configured to determine the angle between line 3013a and the y-axis of the audio device coordinate system.

[0530] Figure 30B Additional examples are shown for determining listener angular orientation data. According to this example, the listener position is already... Figure 29 Defined in box 2915. Here, the control system controls the loudspeaker of environment 3000b to render audio object 3035 to various locations within environment 3000b. In some such examples, the control system may cause the loudspeaker to render audio object 3035 such that audio object 3035 appears to rotate around listener 3005, for example, by rendering audio object 3035 such that audio object 3035 appears to rotate around the origin of listener coordinate system 3020. In this example, curved arrow 3040 shows a portion of the trajectory of audio object 3035 as it rotates around listener 3005.

[0531] According to some such examples, the listener 3005 may provide user input indicating when the audio object 3035 is in the direction the listener 3005 is facing (e.g., saying "stop"). In some such examples, the control system may be configured to determine a line 3013b between the listener's position and the position of the audio object 3035. In this example, line 3013b corresponds to the y' axis of the listener's coordinate system that indicates the direction the listener 3005 is facing. In alternative embodiments, the listener 3005 may provide user input indicating when the audio object 3035 is in front of the environment, in a TV position in the environment, in a position of the audio device, etc.

[0532] Figure 30C Additional examples are shown for determining listener angular orientation data. According to this example, the listener position is already... Figure 29 The image is defined in box 2915. Here, the listener 3005 is using the handheld device 3045 to provide input about the listener's viewing direction by pointing the handheld device 3045 at the television 1901 or the speaker 3030. In this example, the dashed outlines of the handheld device 3045 and the listener's arm indicate the time before the listener 3005 points the handheld device 3045 at the television 1901 or the speaker 3030, when the listener 3005 points the handheld device 3045 at the audio device 2. In other examples, the listener 3005 may have already pointed the handheld device 3045 at another audio device, such as audio device 1. According to this example, the handheld device 3045 is configured to determine an angle α between the audio device 2 and the television 1901 or the speaker 3030, which approximates the angle between the audio device 2 and the listener 3005's viewing direction.

[0533] In some examples, the handheld device 3045 may be a cellular phone that includes an inertial sensor system and a wireless interface configured to communicate with the control system of the audio equipment in the control environment 3000c. In some examples, the handheld device 3045 may run an application or "app" configured to control the handheld device 3045 to perform necessary functions, for example, by providing user prompts (e.g., via a graphical user interface), by receiving input indicating that the handheld device 3045 is pointing in a desired direction, by saving corresponding inertial sensor data and / or transmitting corresponding inertial sensor data to the control system of the audio equipment in the control environment 3000c, etc.

[0534] According to this example, the control system (which may be the control system of the handheld device 3045 or the control system of the audio device in the control environment 3000c) is configured to determine the orientation of lines 3013c and 3050 based on inertial sensor data (e.g., gyroscope data). In this example, line 3013c is parallel to axis y' and can be used to determine the listener's angular orientation. According to some examples, the control system may determine the appropriate rotation of the audio device coordinates about the origin of the listener coordinate system 3020 based on the angle α between the viewing direction of the audio device 2 and the listener 3005.

[0535] Figure 30D It is shown according to the reference Figure 30C The described method is an example of determining the appropriate rotation of the audio device coordinates. In this example, the origin of the audio device coordinate system 3007 is co-located with the origin of the listener coordinate system 3020. After the process 2915 in which the listener position is determined, it becomes possible to co-locate the origin of the audio device coordinate system 3007 and the origin of the listener coordinate system 3020. Co-locating the origin of the audio device coordinate system 3007 and the origin of the listener coordinate system 3020 may involve transforming the audio device position from the audio device coordinate system 3007 to the listener coordinate system 3020. (As already referenced above...) Figure 30C The angle α is defined as described. Therefore, angle α corresponds to the desired orientation of audio device 2 in the listener coordinate system 3020. In this example, angle β corresponds to the orientation of audio device 2 in the audio device coordinate system 3007. In this example, it is the angle β-α. Indicates the rotation necessary to align the y-axis of the audio device coordinate system 3007 with the y'-axis of the listener coordinate system 3020.

[0536] In some implementations... Figure 29 The method may involve controlling at least one audio device in an environment based at least in part on the corresponding audio device location, the corresponding audio device angular orientation, listener location data, and listener angular orientation data.

[0537] For example, some implementations may involve providing audio device location data, audio device angular orientation data, listener location data, and listener angular orientation data to an audio rendering system. In some examples, the audio rendering system may be controlled by a control system (such as...). Figure 1AThe control system 110) is implemented. Some implementations may involve controlling the audio data rendering process based at least in part on audio device location data, audio device angular orientation data, listener location data, and listener angular orientation data. Some such implementations may involve providing loudspeaker acoustic capability data to the rendering system. The loudspeaker acoustic capability data may correspond to one or more loudspeakers in the environment. The loudspeaker acoustic capability data may indicate the orientation of one or more drivers, the number of drivers, or the driver frequency response of one or more drivers. In some examples, the loudspeaker acoustic capability data may be retrieved from memory and then provided to the rendering system.

[0538] One type of embodiment relates to a method for rendering audio for playback and / or audio replay of at least one (e.g., all or some) of a plurality of coordinated (arranged) smart audio devices. For example, a set of smart audio devices present in a user's home (system) can be orchestrated to handle various simultaneous use cases, including flexibly rendering audio for playback by all or some of the smart audio devices (i.e., by all or some of the smart audio devices' speakers). Numerous interactions with the system are considered, which require dynamic adjustments to the rendering and / or playback. Such adjustments may, but do not necessarily, focus on spatial fidelity.

[0539] In a scenario where spatial audio mixing (e.g., rendering one or more audio streams) is performed for playback by a smart audio device (or another set of speakers) within a set of smart audio devices, the type of speakers (e.g., in or coupled to a smart audio device) may differ, and therefore the corresponding acoustic capabilities of the speakers can vary significantly. Figure 3A In one example of the audio environment shown, amplifiers 305d, 305f, and 305h could be smart speakers with a single 0.6-inch speaker. In this example, amplifiers 305b, 305c, 305e, and 305f could be smart speakers with a 2.5-inch woofer and a 0.8-inch tweeter. According to this example, amplifier 305g could be a smart speaker with a 5.25-inch woofer, three 2-inch midrange speakers, and a 1.0-inch tweeter. Here, amplifier 305a could be a speaker enclosure with sixteen 1.1-inch beam drivers and two 4-inch woofers. Therefore, the low-frequency capabilities of smart speakers 305d and 305f would be significantly less than those of the other amplifiers in environment 200, especially those with 4-inch or 5.25-inch woofers.

[0540] Figure 31This is a block diagram illustrating examples of components of a system capable of implementing various aspects of this disclosure. As with the other figures provided herein, Figure 31 The types and quantities of elements shown are provided as examples only. Other implementations may include more, fewer, and / or different types and quantities of elements.

[0541] According to this example, system 3100 includes a smart home hub 3105 and loudspeakers 3125a to 3125m. In this example, the smart home hub 3105 includes... Figure 1A An example of the control system 110 shown and described above is presented. According to this embodiment, the control system 110 includes a listening environment dynamic processing configuration data module 3110, a listening environment dynamic processing module 3115, and a rendering module 3120. Some examples of the listening environment dynamic processing configuration data module 3110, the listening environment dynamic processing module 3115, and the rendering module 3120 are described below. In some examples, the rendering module 3120' can be configured for both rendering and listening environment dynamic processing.

[0542] As indicated by the arrows between the smart home hub 3105 and the loudspeakers 3125a to 3125m, the smart home hub 3105 also includes Figure 1A An example of the interface system 105 shown and described above is presented. According to some examples, the smart home hub 3105 may be... Figure 3A This is a portion of environment 300 shown. In some instances, the smart home hub 3105 may be implemented by a smart speaker, smart TV, cellular phone, laptop computer, etc. In some implementations, the smart home hub 3105 may be implemented by software (e.g., via a downloadable software application or "app"). In some instances, the smart home hub 3105 may be implemented in each of loudspeakers 3125a to 3125m, all of which operate in parallel to generate the same processed audio signal from module 3120. According to some such examples, in each loudspeaker, rendering module 3120 may then generate one or more speaker feeds associated with each loudspeaker or group of loudspeakers, and these speaker feeds may be provided to each speaker dynamics processing module.

[0543] In some instances, loudspeakers 3125a to 3125m may include Figure 3A The loudspeakers 305a to 305h are used, while in other examples, loudspeakers 3125a to 3125m may be or may include other loudspeakers. Therefore, in this example, system 3100 includes M loudspeakers, where M is an integer greater than 2.

[0544] Smart speakers, and many other electrodynamic speakers, typically employ some form of internal dynamic processing to prevent speaker distortion. This dynamic processing is usually associated with a signal limiting threshold (e.g., a frequency-varying threshold), below which the signal level is dynamically maintained. For example, the Dolby Audio Conditioner, one of several algorithms in the Dolby Audio Processing (DAP) audio post-processing suite, provides this processing. In some instances, but not typically via the smart speaker's dynamic processing module, dynamic processing may also involve applying one or more compressors, gates, expanders, avoiders, etc.

[0545] Therefore, in this example, each of the loudspeakers 3125a to 3125m includes a corresponding loudspeaker dynamics processing (DP) module A to M. The loudspeaker dynamics processing module is configured to apply separate loudspeaker dynamics processing configuration data to each individual loudspeaker in the listening environment. For example, loudspeaker DP module A is configured to apply separate loudspeaker dynamics processing configuration data suitable for loudspeaker 3125a. In some examples, the separate loudspeaker dynamics processing configuration data may correspond to one or more capabilities of a single loudspeaker, such as the ability of the loudspeaker to reproduce audio data within a specific frequency range and at a specific level without perceptible distortion.

[0546] When rendering spatial audio across a group of heterogeneous loudspeakers (e.g., speakers of a smart audio device or speakers coupled to a smart audio device) that may each have different playback limitations, care must be taken when performing dynamic processing on the entire mix. A simple solution is to render the spatial mix as a speaker feed for each participating loudspeaker and then allow the dynamic processing module associated with each loudspeaker to operate independently on its corresponding loudspeaker feed based on that loudspeaker's limitations.

[0547] While this method will prevent distortion from each speaker, it may dynamically shift the spatial balance of the mix in a perceptibly dispersed manner. For example, see reference... Figure 3ASuppose a television program is playing on television 330, and the corresponding audio is being reproduced by a loudspeaker in environment 300. Assume that during the television program, the audio associated with a stationary object (such as a heavy machinery unit in a factory) is intended to be rendered to a specific location within environment 300. Further assume that because loudspeaker 305b has a much stronger ability to reproduce sound in the low-frequency range, the dynamic processing module associated with loudspeaker 305d will reduce the audio level in the low-frequency range by a much greater amount than the dynamic processing module associated with loudspeaker 305b. If the volume of the signal associated with the stationary object fluctuates, then at higher volumes, the dynamic processing module associated with loudspeaker 305d will cause the audio level in the low-frequency range to be reduced by a much greater amount than the same audio level reduced by the dynamic processing module associated with loudspeaker 305b. This difference in level will cause a noticeable change in the position of the stationary object. Therefore, an improved solution is needed.

[0548] Some embodiments of this disclosure are systems and methods for rendering (or rendering and playing back) spatial audio mixes (e.g., rendering one or more audio streams) for playback by at least one (e.g., all or some) of the intelligent audio devices in a set of intelligent audio devices (e.g., a set of coordinated intelligent audio devices) and / or by at least one (e.g., all or some) of the speakers in another set of speakers. Some embodiments are methods (or systems) for such rendering (e.g., including the generation of speaker feeds) and the playback of the rendered audio (e.g., playback of the generated speaker feeds). Examples of such embodiments include the following:

[0549] Systems and methods for audio processing may include rendering audio (e.g., rendering a spatial audio mix, such as by rendering one or more audio streams) for playback by at least two speakers (e.g., all or some of the speakers in a set of speakers), including by:

[0550] (a) Combine individual loudspeaker dynamic processing configuration data (such as limit thresholds (playback limit thresholds)) to determine the listening environment dynamic processing configuration data (such as combined thresholds) of multiple loudspeakers;

[0551] (b) Using listening environment dynamics configuration data (e.g., combined thresholds) from multiple loudspeakers, perform dynamic processing on audio (e.g., multiple audio streams indicating spatial audio mixing) to generate processed audio; and

[0552] (c) Render the processed audio as a speaker feed.

[0553] According to some implementation methods, process (a) can be performed by, for example... Figure 31The listening environment dynamic processing configuration data module 3110 and other modules shown in the diagram are used to perform this operation. The smart home hub 3105 can be configured to obtain individual loudspeaker dynamic processing configuration data for each of the M loudspeakers via an interface system. In this embodiment, the individual loudspeaker dynamic processing configuration data includes a separate loudspeaker dynamic processing configuration dataset for each of the multiple loudspeakers. According to some examples, the individual loudspeaker dynamic processing configuration data for one or more loudspeakers may correspond to one or more capabilities of the one or more loudspeakers. In this example, each individual loudspeaker dynamic processing configuration dataset includes at least one type of dynamic processing configuration data. In some examples, the smart home hub 3105 can be configured to obtain the individual loudspeaker dynamic processing configuration dataset by querying each of the loudspeakers 3125a to 3125m. In other embodiments, the smart home hub 3105 can be configured to obtain the individual loudspeaker dynamic processing configuration dataset by querying a data structure of previously obtained individual loudspeaker dynamic processing configuration datasets stored in memory.

[0554] In some examples, process (b) can be derived from, for example... Figure 31 The listening environment dynamic processing module 3115 and other modules are used to execute this. Some detailed examples of processes (a) and (b) are described below.

[0555] In some examples, the rendering of process (c) can be performed by, for example... Figure 31 The rendering module 3120 or rendering module 3120', etc., are used to perform the processing. In some embodiments, audio processing may involve:

[0556] (d) Perform dynamic processing on the rendered audio signal according to individual loudspeaker dynamic processing configuration data for each loudspeaker (e.g., limit the loudspeaker feed according to a playback limit threshold associated with the corresponding loudspeaker, thereby generating a limited loudspeaker feed). For example, process (d) can be performed by... Figure 31 The dynamic processing modules A through M shown in the figure are executed.

[0557] The loudspeaker may include (or be coupled to) at least one (e.g., all or some) of the loudspeakers in a group of smart audio devices. In some implementations, to generate the restricted loudspeaker feed in step (d), the loudspeaker feed generated in step (c) may be processed by a second stage of dynamic processing (e.g., by the associated dynamic processing system of each loudspeaker), for example, to generate the loudspeaker feed before final playback by the loudspeaker. For example, the loudspeaker feed (or a subset or portion thereof) may be provided to the dynamic processing system of each different loudspeaker (e.g., the dynamic processing subsystem of the smart audio device, wherein the smart audio device includes or is coupled to an associated loudspeaker among the loudspeakers), and the processed audio output from each of the dynamic processing systems may be used to generate a loudspeaker feed for the associated loudspeaker among the loudspeakers. After loudspeaker-specific dynamic processing (in other words, dynamic processing performed independently for each loudspeaker), the processed (e.g., dynamically restricted) loudspeaker feed may be used to drive the loudspeaker to cause playback of sound.

[0558] The first stage of dynamic processing (in step (b)) can be designed to reduce perceptually dispersed shifts in spatial balance, which would result if steps (a) and (b) were omitted, and the dynamically processed (e.g., limited) speaker feed generated by step (d) is generated in response to the original audio (rather than in response to the processed audio generated in step (b)). This prevents unwanted shifts in the spatial balance of the mix. The second stage of dynamic processing, operating on the rendered speaker feed from step (c), can be designed to ensure no speaker distortion, since the dynamic processing in step (b) may not necessarily guarantee that the signal level has been reduced below a threshold for all speakers. In some examples, a combination of individual speaker dynamic processing configuration data (e.g., a combination of thresholds in the first stage (step (a))) can involve (e.g., including) averaging individual speaker dynamic processing configuration data (e.g., limited thresholds) across speakers (e.g., across smart audio devices) or obtaining the minimum individual speaker dynamic processing configuration data (e.g., limited thresholds) across speakers (e.g., across smart audio devices).

[0559] In some implementations, when the first stage of dynamic processing (in step (b)) operates on audio indicating spatial mixing (e.g., audio of an object-based audio program, including at least one object channel and optionally including at least one speaker channel), the first stage can be implemented according to techniques for audio object processing using spatial zones. In this case, individual loudspeaker dynamic processing configuration data (e.g., combined limiting thresholds) associated with each zone can be obtained by (or as) a weighted average of individual loudspeaker dynamic processing configuration data (e.g., individual speaker limiting thresholds), and this weighting can be given or determined at least in part by the spatial proximity of each speaker to the zone and / or its position within the zone.

[0560] In the example embodiment, assume there are multiple M speakers (M≥2), where each speaker is indexed by a variable i. Associated with each speaker i is a set of playback limit thresholds T for frequency variations. i [f], where the variable f represents an index to a finite set of frequencies for a specified threshold. (Note that if the size of the frequency set is one, the corresponding single threshold can be considered wideband, applying across the entire frequency range). These thresholds are used by each loudspeaker in its own independent dynamic processing function to limit the audio signal below a threshold T. i [f], for specific purposes, such as preventing speaker distortion or preventing speakers from playing beyond a certain level that is considered offensive in their vicinity.

[0561] Figure 32A , Figure 32B and Figure 32C Examples of playback limit thresholds and corresponding frequencies are shown. For example, the frequency range shown may span the range of frequencies audible to the average human (e.g., 20 Hz to 20 kHz). In these examples, the playback limit thresholds are indicated by the vertical axis of graphs 3200a, 3200b, and 3200c, which is labeled "horizontal threshold" in these examples. The playback limit threshold / horizontal threshold increases in the direction of the arrow on the vertical axis. For example, the playback limit threshold / horizontal threshold may be expressed in decibels. In these examples, the horizontal axis of graphs 3200a, 3200b, and 3200c indicates frequencies that increase in the direction of the arrow on the horizontal axis. The playback limit thresholds indicated by curves 3200a, 3200b, and 3200c can be implemented, for example, by a separate dynamic processing module of a loudspeaker.

[0562] Figure 32A Graph 3200a shows a first example of the playback limit threshold as a function of frequency. Curve 3205a indicates the playback limit threshold for each corresponding frequency value. In this example, at the bass frequency f... bAt this point, input level T i The received input audio will be processed by the dynamic processing module to output level T. o Output. For example, the low frequency f. b It can be in the range of 60Hz to 250Hz. However, in this example, at higher frequencies f... t At this point, input level T i The received input audio will be processed by the dynamic processing module at the same level, input level T. i Output. For example, high frequency f t This can be achieved in the range above 1280Hz. Therefore, in this example, curve 3205a corresponds to a dynamic processing module that applies a significantly lower threshold to bass frequencies than to treble frequencies. Such a dynamic processing module could be suitable for amplifiers without a woofer (e.g., Figure 3A (305d loudspeaker).

[0563] Figure 32B Chart 3200b shows a second example of the playback limit threshold as a function of frequency. Curve 3205b indicates that... Figure 32A The same low frequency f shown in the figure b At this point, input level T i The received input audio will be processed by the dynamic processing module at a higher output level T. o Output. Therefore, in this example, curve 3205b corresponds to a dynamic processing module that does not apply a threshold lower than that of curve 3205a to the bass frequencies. Such a dynamic processing module can be applied to amplifiers with at least a small woofer (e.g., Figure 3A (Amplifier 305b).

[0564] Figure 32C Chart 3200c shows a second example of the playback limit threshold as a function of frequency. Curve 3205c (a straight line in this example) indicates that... Figure 32A The same low frequency f shown in the figure b At this point, input level T i The received input audio will be output at the same level by the dynamic processing module. Therefore, in this example, curve 3205c corresponds to a dynamic processing module that can be applied to a loudspeaker capable of reproducing a wide range of frequencies, including bass frequencies. It will be observed that, for simplicity, the dynamic processing module can approximate curve 3205c by implementing curve 3205d, which applies the same threshold to all indicated frequencies.

[0565] Known rendering systems such as Centroid Amplitude Translation (CMAP) or Flexible Virtualization (FV) can be used to render spatial audio mixes with multiple speakers. The rendering system generates speaker feeds from the component parts of the spatial audio mix, one for each of the multiple speakers. In some previous examples, the speaker feeds are then defined by a threshold of T for each speaker. i The associated dynamic processing functions of [f] are processed independently. Without the benefit of this disclosure, the rendering scene described may result in a dispersion shift in the perceived spatial balance of the rendered spatial audio mix. For example, one of the M speakers (say, on the right side of the listening area) may be significantly less capable than the others (e.g., less capable of rendering audio in the bass range), and therefore, at least within a certain frequency range, the threshold T of that speaker will be lower. i [f] may be significantly lower than the threshold of other speakers. During playback, the speaker's dynamic processing module will reduce the level of the right-side component of the spatial mix, with this reduction being significantly greater than the reduction of the left-side component. Listeners are very sensitive to this dynamic shift between the left and right balance of the spatial mix and may find the result to be very diffuse.

[0566] To address this issue, in some examples, individual amplifier dynamic processing configuration data (e.g., playback limit thresholds) for the individual speakers of the listening environment are combined to create listening environment dynamic processing configuration data for all amplifiers in the listening environment. This listening environment dynamic processing configuration data can then be used to perform dynamic processing first within the context of the entire spatial audio mix, before rendering it as a speaker feed. Because this first stage of dynamic processing has access to the entire spatial mix and not just a single speaker feed, processing can be performed in a way that does not cause a dispersion shift in the perceived spatial balance of the mix. Individual amplifier dynamic processing configuration data (e.g., playback limit thresholds) can be combined in a way that eliminates or reduces the amount of dynamic processing performed by any independent dynamic processing function of individual speakers.

[0567] In one example of determining the dynamic processing configuration data for the listening environment, the individual loudspeaker dynamic processing configuration data for a single loudspeaker (e.g., playback limit threshold) can be combined to form the listening environment dynamic processing configuration data (e.g., frequency-varying playback limit threshold) applied to all components of the spatial mix in the first stage of dynamic processing. A single set of (i) components. According to some examples, spatial balance of the mix can be maintained because the constraints for all components are the same. One way to combine individual loudspeaker dynamic processing configuration data (e.g., playback constraint thresholds) is to take the minimum across all loudspeakers i:

[0568]

[0569] This combination essentially eliminates the need for individual dynamic processing of each speaker, as the spatial mix is ​​initially limited to below the threshold of the speaker with the lowest capability at each frequency. However, this strategy can be too aggressive. Many speakers may reproduce below their capabilities, and the combined playback level of all speakers can be offensively low. For example, if... Figure 32A The threshold shown in the diagram is applied to the bass range and... Figure 32C If the corresponding loudspeaker has a specific threshold, then the playback level of the subsequent loudspeaker will be unnecessarily low in the bass range. An alternative combination for determining the dynamic processing configuration data of the listening environment is to average the individual loudspeaker dynamic processing configuration data across all loudspeakers in the listening environment. For example, in the case of a playback limit threshold, the average can be determined as follows:

[0570]

[0571] For this combination, the overall playback level may increase compared to taking the minimum value, because the first stage of dynamic processing limits to a higher level, thereby allowing more capable speakers to reproduce louder. For individual speakers whose limit thresholds fall below the average, their independent dynamic processing functions can still limit their associated speaker feeds if needed. However, the first stage of dynamic processing may reduce the need for such limitations, as some initial constraints have already been applied to the spatial mix.

[0572] Based on some examples of determining the dynamic processing configuration data of the listening environment, tunable combinations can be created that interpolate between the minimum and average values ​​of the individual loudspeaker's dynamic processing configuration data using a tuning parameter α. For example, in a scenario with a playback limitation threshold, the interpolation can be determined as follows:

[0573]

[0574] Other combinations of dynamic processing configuration data for individual loudspeakers are possible, and this disclosure is intended to cover all such combinations.

[0575] Figure 33A and Figure 33B These are graphs illustrating examples of dynamic range compressed data. In graphs 3300a and 3300b, the input signal level, in decibels, is shown horizontally on the horizontal axis, and the output signal level, in decibels, is shown horizontally on the vertical axis. As with other disclosed examples, specific thresholds, ratios, and other values ​​are shown by way of example only and are not limiting.

[0576] exist Figure 33AIn the example shown, the output signal level is equal to the input signal level below a threshold, which in this example is -10 dB. Other examples may involve different thresholds, such as -20 dB, -18 dB, -16 dB, -14 dB, -12 dB, -8 dB, -6 dB, -4 dB, -2 dB, 0 dB, 2 dB, 4 dB, 6 dB, etc. Various examples of compression ratios are shown above the threshold. A ratio N:1 means that above the threshold, for every N dB increase in the input signal, the output signal level will increase by 1 dB. For example, a compression ratio of 10:1 (line 3305e) means that above the threshold, for every 10 dB increase in the input signal, the output signal level will increase by only 1 dB. A compression ratio of 1:1 (line 3305a) means that even above the threshold, the output signal level remains equal to the input signal level. Lines 3305b, 3305c, and 3305d correspond to compression ratios of 3:2, 2:1, and 5:1. Other implementations may provide different compression ratios, such as 2.5:1, 3:1, 3.5:1, 4:3, 4:1, etc.

[0577] Figure 33B An example of an "inflection point" is shown, which controls how the compression ratio changes at or near a threshold, which in this example is 0 dB. According to this example, the compression curve with a "hard" inflection point consists of two straight line segments: segment 3310a reaches the threshold, and segment 3310b is above the threshold. Hard inflection points can be implemented more simply, but may result in artifacts.

[0578] exist Figure 33B The diagram also illustrates an example of a "soft" inflection point. In this example, the soft inflection point spans 10 dB. According to this embodiment, the compression ratio of the compression curve with a soft inflection point is the same as that of the compression curve with a hard inflection point above and below the 10 dB span. Other embodiments may provide a variety of other shapes of "soft" inflection points, which may span more or less decibels, indicate different compression ratios above the span, etc.

[0579] Other types of dynamic range compressed data can include "attack" data and "release" data. An attack is the period of time during which the compressor, for example, reduces its gain in response to an increase in the input to achieve a gain determined by the compression ratio. The attack time for a compressor is typically between 25 milliseconds and 500 milliseconds, but other attack times are also possible. A release is the period of time during which the compressor, for example, increases its gain in response to a decrease in the input to achieve an output gain determined by the compression ratio (or to reach the input level, if the input level has fallen below a threshold). For example, the release time can range from 25 milliseconds to 2 seconds.

[0580] Therefore, in some examples, for each of a plurality of loudspeakers, individual loudspeaker dynamic processing configuration data may include a dynamic range compression dataset. The dynamic range compression dataset may include threshold data, input / output ratio data, attack data, release data, and / or inflection point data. One or more of these types of individual loudspeaker dynamic processing configuration data may be combined to determine the listening environment dynamic processing configuration data. As described above, referring to the combined playback limit threshold, in some examples, the dynamic range compression data may be averaged to determine the listening environment dynamic processing configuration data. In some instances, the minimum or maximum value of the dynamic range compression data may be used to determine the listening environment dynamic processing configuration data (e.g., maximum compression ratio). In other implementations, tunable combinations may be created, for example, via tuning parameters described above with reference to equation (32), interpolating between the minimum and average values ​​of the dynamic range compression data for individual loudspeaker dynamic processing.

[0581] In some of the examples described above, in the first stage of dynamic processing, a single set of dynamic processing configuration data for the listening environment (e.g., combined thresholds) will be used. A single set of spatial mixes is applied to all components of the spatial mix. This implementation maintains the spatial balance of the mix, but can introduce other unwanted artifacts. For example, "spatial evasion" can occur when a very loud portion of the spatial mix in an isolated spatial region causes the entire mix to be downplayed. Other softer components of the mix that are spatially distant from that loud component may be perceived as unnaturally soft. For example, soft background music can appear below the combined threshold in the surround field of the spatial mix. The spatial mix is ​​played horizontally, and therefore the first stage of dynamic processing does not impose restrictions on the spatial mix. Then, a loud gunshot might be briefly introduced at the beginning of the spatial mix (e.g., on screen for a movie soundtrack), and the overall level of the mix increases above the combined threshold. At this point, the first stage of dynamic processing reduces the overall level of the mix back to the threshold. The following. Because the music is spatially separate from the gunshots, this can be perceived as an unnatural evasion within a continuous stream of music.

[0582] To address these issues, some implementations allow for independent or partially independent dynamic processing of different “spatial zones” of the spatial mix. A spatial zone can be considered a subset of the spatial area on which the entire spatial mix is ​​rendered. While much of the following discussion provides examples of dynamic processing based on playback limit thresholds, these concepts are equally applicable to other types of individual loudspeaker dynamic processing configuration data and listening environment dynamic processing configuration data.

[0583] Figure 34 An example of a spatial area for listening to the environment is shown. Figure 34 An example depicting the spatial mixing area (represented by a whole square) is subdivided into three spatial zones: front, center, and surround.

[0584] Although Figure 34 Spatial regions in a spatial mix are often depicted as having hard boundaries, but in practice, it is beneficial to treat the transition from one spatial region to another as continuous. For example, a spatial mix component located in the middle of the left edge of a square could allocate half of its level to the front region and half to the surround region. The signal level from each component of the spatial mix can be allocated and accumulated into each spatial region in this continuous manner. Dynamic processing functions can then operate independently for each spatial region on the overall signal level allocated to it from the mix. For each component of the spatial mix, the results of the dynamic processing from each spatial region (e.g., time-varying gain at each frequency) can then be combined and applied to said component. In some examples, this combination of spatial region results is different for each component and is a function of the allocation of that particular component to each region. The end result is that components of a spatial mix with similar spatial region allocation receive similar dynamic processing, but with independence between spatial regions allowed. Spatial regions can be advantageously selected to prevent undesirable spatial shifts, such as left / right imbalance, while still allowing some independent spatial processing (e.g., to reduce other artifacts such as spatial evasion as described).

[0585] In the first stage of the dynamic processing of this disclosure, techniques for processing spatial mixing by spatial zone can be advantageously employed. For example, different combinations of individual loudspeaker dynamic processing configuration data (e.g., playback limit thresholds) across loudspeaker i can be calculated for each spatial zone. The set of combined zone thresholds can be obtained from... This indicates that index j refers to one of multiple spatial regions. The dynamic processing module can handle this with associated thresholds. Each spatial region can be operated independently, and the results can be applied back to the components of the spatial mix, according to the techniques described above.

[0586] Consider that the rendered spatial signal consists of a total of K individual signals x k [t] constitutes a signal, each of which has an associated desired spatial location (which may be time-varying). A particular method for implementing zone processing involves calculating a description of each audio signal x based on the desired spatial location of the audio signal relative to the zone j. k The time-varying translation gain α of [t]'s contribution to region j kj [t]. These translation gains can be advantageously designed to follow the power-preserving translation law, which requires the sum of the squares of the gains to equal one. Based on these translation gains, the signal s... j[t] can be calculated as the sum of the constituent signals weighted by the translation gain of the signal in that region:

[0587]

[0588] Then, each zone signal s j [t] can be determined by the threshold value. The parameterized dynamic processing function DP is processed independently to generate frequency and time-varying zone-corrected gain G. j :

[0589]

[0590] Then, the area correction gain can be compared with the individual component signal x. k The translation gain of [t] for the region is proportionally combined to calculate the frequency-varying and time-varying correction gain for each of the individual component signals:

[0591]

[0592] These signals correct the gain G k Each component signal can then be applied, for example, using a filter bank, to produce dynamically processed component signals that can then be subsequently rendered as speaker signals.

[0593] Individual loudspeaker dynamic processing configuration data (such as speaker playback limit thresholds) for each spatial zone can be combined in various ways. As an example, spatial zone and speaker dependency weights w can be used. ij [f] Limit the playback threshold of the spatial region. Calculated as the speaker playback limit threshold T i The weighted sum of [f]:

[0594]

[0595] Similar weighting functions can be applied to other types of individual loudspeaker dynamic processing configuration data. Advantageously, the individual loudspeaker dynamic processing configuration data (e.g., playback limit threshold) for a combination of spatial zones can be biased towards the individual loudspeaker dynamic processing configuration data (e.g., playback limit threshold) of the loudspeaker primarily responsible for playing back the spatial mix components associated with that spatial zone. This can be achieved by setting the weight w according to each loudspeaker's responsibility for rendering the spatial mix components associated with that zone for frequency f. ij [f] and thus achieved.

[0596] Figure 35 It shows Figure 34 An example of a loudspeaker within a spatial area. Figure 35 Depicting and Figure 34The same area, but with the positions of five example loudspeakers (loudspeakers 1, 2, 3, 4, and 5) responsible for rendering the spatial mix overlaid. In this example, loudspeakers 1, 2, 3, 4, and 5 are represented by diamonds. In this particular example, loudspeaker 1 is primarily responsible for rendering the central area, loudspeakers 2 and 5 for the front area, and loudspeakers 3 and 4 for the surround area. Weights w can be created based on this one-to-one mapping from loudspeakers to spatial areas. ij [f], however, as with spatial zone-based processing of spatial mixing, a more continuous mapping may be preferred. For example, speaker 4 is very close to the front zone, and the components of the audio mix located between speakers 4 and 5 (although conceptually in the front zone) will likely be primarily played back by the combination of speakers 4 and 5. Thus, it makes sense for individual speaker dynamics configuration data (e.g., playback limit thresholds) for speaker 4 to facilitate individual amplifier dynamics configuration data (e.g., playback limit thresholds) for the combination of the front and surround zones.

[0597] One way to achieve this continuous mapping is to use the weights w ij [f] is set to a speaker participation value that describes the relative contribution of each speaker i in rendering the component associated with spatial region j. Such a value can be obtained directly from a rendering system responsible for rendering the speakers (e.g., from step (c) described above) and a set of one or more nominal spatial locations associated with each spatial region. This set of nominal spatial locations may include a set of locations within each spatial region.

[0598] Figure 36 Showing the coverage Figure 35 Examples of the spatial zones and nominal spatial positions on the speakers. The nominal positions are indicated by numbered circles: the front zone is associated with two positions located at the top corners of the square, the central zone is associated with a single position at the top center of the square, and the surround zone is associated with two positions at the bottom corners of the square.

[0599] To calculate the loudspeaker engagement value for a spatial region, a renderer can be used to render each nominal location associated with the region to generate loudspeaker activations associated with that location. For example, in the case of CMAP, these activations could be the gain of each loudspeaker, or in the case of FV, they could be the complex value of each loudspeaker at a given frequency. Next, for each loudspeaker and region, these activations can be accumulated across each nominal location associated with the spatial region to produce a value g. ij [f]. This value represents the total activation of loudspeaker i used to render the entire nominal set of locations associated with spatial region j. Finally, the loudspeaker engagement value in the spatial region can be calculated as the cumulative activation g normalized by the sum of all these accumulated cross-loudspeaker activations. ij [f]. The weight can then be set to the speaker engagement value:

[0600]

[0601] The described normalization ensures w across all speaker i ij The sum of [f] equals one, which is the ideal property of the weights in Equation 36.

[0602] According to some implementations, the process described above for calculating speaker engagement values ​​and combining thresholds based on these values ​​can be performed as a static process, wherein the resulting combined thresholds are calculated once during a setup procedure to determine the speaker layout and capabilities in the environment. In such a system, it can be assumed that once set up, both the dynamic processing configuration data of the individual loudspeakers and the way the rendering algorithm activates the loudspeakers according to the desired audio signal location remain static. However, in some systems, both of these aspects may change over time, for example, in response to changing conditions in the playback environment, and thus, it may be desirable to update the combined thresholds in a continuous or event-triggered manner according to the process described above to account for such changes.

[0603] Both the CMAP and FV rendering algorithms can be enhanced to adapt to one or more dynamically configurable features in response to changes in the listening environment. For example, regarding... Figure 35 A person located near speaker 3 can utter the wake word of the intelligent assistant associated with the speaker, thus putting the system in a state ready to listen for subsequent commands from that person. When the wake word is spoken, the system can use the microphone associated with the loudspeaker to determine the person's location. With this information, the system can then choose to divert the energy of the audio being played back from speaker 3 to other speakers, allowing the microphone on speaker 3 to hear the person better. In such a scenario, Figure 35 Speaker 2 can essentially "take over" the responsibility of speaker 3 for a period of time, and therefore the speaker participation value in the surround zone changes significantly; the participation value of speaker 3 decreases and the participation value of speaker 2 increases. The zone threshold can then be recalculated, as it depends on the changed speaker participation values. Alternatively, or in addition to these changes in the rendering algorithm, the limiting threshold for speaker 3 can be lowered below its nominal value to prevent speaker distortion. This ensures that any residual audio played from speaker 3 does not increase beyond a certain threshold, which is determined to interfere with the listener's microphone. Since the zone threshold is also a function of the individual speaker threshold, it can also be updated in this case.

[0604] Figure 37This is a flowchart outlining an example of a method that can be performed by means or systems such as those disclosed herein. As with other methods described herein, the blocks of method 3700 need not be performed in the indicated order. In some embodiments, one or more blocks of method 3700 may be performed simultaneously. Furthermore, some embodiments of method 3700 may include more or fewer blocks than those shown and / or described. The blocks of method 3700 may be performed by one or more devices, which may be (or may include) a control system, such as… Figure 1A The control system 110 shown and described above, or one of other disclosed control system examples.

[0605] According to this example, block 3705 relates to obtaining individual loudspeaker dynamic processing configuration data for each of a plurality of loudspeakers in a listening environment by a control system and via an interface system. In this embodiment, the individual loudspeaker dynamic processing configuration data includes a separate loudspeaker dynamic processing configuration dataset for each of the plurality of loudspeakers. According to some examples, the individual loudspeaker dynamic processing configuration data for one or more loudspeakers may correspond to one or more capabilities of one or more loudspeakers. In this example, each individual loudspeaker dynamic processing configuration dataset includes at least one type of dynamic processing configuration data.

[0606] In some instances, box 3705 may involve obtaining a separate loudspeaker dynamic processing configuration dataset from each of a plurality of loudspeakers in the listening environment. In other instances, box 3705 may involve obtaining a separate loudspeaker dynamic processing configuration dataset from a data structure stored in memory. For example, the separate loudspeaker dynamic processing configuration dataset may have been previously obtained, for example, as part of a setup procedure for each loudspeaker and stored in the data structure.

[0607] According to some examples, individual loudspeaker dynamic processing configuration datasets can be proprietary. In some such examples, individual loudspeaker dynamic processing configuration datasets may have been previously estimated based on individual loudspeaker dynamic processing configuration data for loudspeakers with similar characteristics. For example, box 3705 may involve a loudspeaker matching process to determine the most similar loudspeakers from a data structure indicating multiple loudspeakers and a corresponding individual loudspeaker dynamic processing configuration dataset for each of the multiple loudspeakers. The loudspeaker matching process may be based on, for example, a comparison of the size of one or more woofers, tweeters, and / or midrange loudspeakers.

[0608] In this example, block 3710 relates to the control system determining listening environment dynamic processing configuration data for a plurality of loudspeakers. According to this embodiment, determining the listening environment dynamic processing configuration data is based on a separate loudspeaker dynamic processing configuration dataset for each of the plurality of loudspeakers. Determining the listening environment dynamic processing configuration data may involve, for example, combining the individual loudspeaker dynamic processing configuration data of the dynamic processing configuration dataset by averaging one or more types of individual loudspeaker dynamic processing configuration data. In some instances, determining the listening environment dynamic processing configuration data may involve determining a minimum or maximum value of one or more types of individual loudspeaker dynamic processing configuration data. According to some such embodiments, determining the listening environment dynamic processing configuration data may involve interpolating between the minimum or maximum value of one or more types of individual loudspeaker dynamic processing configuration data and an average value.

[0609] In this embodiment, block 3715 relates to receiving audio data, including one or more audio signals and associated spatial data, by a control system and via an interface system. For example, the spatial data may indicate a predicted perceived spatial location corresponding to an audio signal. In this example, the spatial data includes channel data and / or spatial metadata.

[0610] In this example, box 3720 relates to the dynamic processing of audio data by a control system based on dynamic processing configuration data of the listening environment to generate processed audio data. The dynamic processing of box 3720 may involve any of the disclosed dynamic processing methods disclosed herein, including but not limited to applying one or more playback limit thresholds, compressing data, etc.

[0611] Here, box 3725 relates to the rendering of processed audio data by a control system for reproduction via a set of loudspeakers comprising at least some of a plurality of loudspeakers, thereby generating a rendered audio signal. In some examples, box 3725 may relate to applying a CMAP rendering process, an FV rendering process, or a combination of both. In this example, box 3720 is executed prior to box 3725. However, as described above, box 3720 and / or box 3710 may be at least partially based on the rendering process of box 3725. Boxes 3720 and 3725 may relate to performing the process referenced above. Figure 31 The process described in the listening environment dynamic processing module and rendering module 3120.

[0612] According to this example, box 3730 relates to providing a rendered audio signal to a set of loudspeakers via an interface system. In one example, box 3730 may relate to a smart home hub 3105 providing the rendered audio signal to loudspeakers 3125a to 3125m via its interface system.

[0613] In some examples, method 3700 may involve performing dynamic processing on the rendered audio signal based on individual loudspeaker dynamic processing configuration data for each of a set of loudspeakers to which the rendered audio signal is provided. For example, refer again Figure 31 The dynamic processing modules A to M can perform dynamic processing on the rendered audio signal based on the individual loudspeaker dynamic processing configuration data of loudspeakers 3125a to 3125m.

[0614] In some implementations, individual loudspeaker dynamic processing configuration data may include a playback limit threshold dataset for each of a plurality of loudspeakers. In some such examples, the playback limit threshold dataset may include playback limit thresholds for each of a plurality of frequencies.

[0615] In some instances, determining the dynamic processing configuration data for the listening environment may involve determining a minimum playback limit threshold across multiple loudspeakers. In some examples, determining the dynamic processing configuration data for the listening environment may involve averaging the playback limit thresholds to obtain an average playback limit threshold across multiple loudspeakers. In some such examples, determining the dynamic processing configuration data for the listening environment may involve determining a minimum playback limit threshold across multiple loudspeakers and interpolating between the minimum playback limit threshold and the average playback limit threshold.

[0616] According to some implementations, averaging the playback limit thresholds may involve determining a weighted average of the playback limit thresholds. In some such examples, the weighted average may be based at least in part on characteristics of the rendering process implemented by the control system, such as the characteristics of the rendering process of box 3725.

[0617] In some implementations, dynamic processing of audio data can be based on spatial regions. Each spatial region may correspond to a subset of the listening environment.

[0618] According to some such implementations, dynamic processing can be performed individually for each spatial zone. For example, dynamic processing configuration data for determining the listening environment can be performed individually for each spatial zone. For example, a dynamic processing configuration dataset for combining multiple loudspeakers can be performed individually for each of one or more spatial zones. In some examples, performing a dynamic processing configuration dataset for combining multiple loudspeakers individually for each of one or more spatial zones can be based at least in part on activating the loudspeakers according to a rendering process based on the desired audio signal position across one or more spatial zones.

[0619] In some examples, the dynamic processing configuration dataset, which is combined across multiple speakers for each of one or more spatial zones, may be based at least in part on the speaker engagement value of each speaker in each of the one or more spatial zones. Each speaker engagement value may be based at least in part on one or more nominal spatial locations within each of the one or more spatial zones. In some examples, the nominal spatial locations may correspond to the canonical locations of channels in a Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, or Dolby 9.1 surround sound mix. In some such implementations, each speaker engagement value is based at least in part on the activation of each speaker corresponding to the rendering of audio data at each of the one or more nominal spatial locations within each of the one or more spatial zones.

[0620] According to some such examples, the weighted average of the playback limit threshold may be based at least in part on the activation of the loudspeaker during the rendering process based on the audio signal near the spatial zone. In some instances, the weighted average may be based at least in part on the loudspeaker participation value of each loudspeaker in each spatial zone. In some such examples, each loudspeaker participation value may be based at least in part on one or more nominal spatial locations within each spatial zone. For example, the nominal spatial locations may correspond to the canonical positions of channels in a Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, or Dolby 9.1 surround sound mix. In some implementations, each loudspeaker participation value may be based at least in part on the activation of each loudspeaker corresponding to the rendered audio data at each of the one or more nominal spatial locations within each spatial zone.

[0621] According to some implementations, rendering processed audio data may involve determining the relative activation of a group of loudspeakers based on one or more dynamically configurable functions. See below for reference. Figure 10 Examples are described below. One or more dynamically configurable functions can be based on one or more properties of the audio signal, one or more properties of a group of loudspeakers, or one or more external inputs. For example, one or more dynamically configurable functions can be based on: the proximity of the loudspeaker to one or more listeners; the proximity of the loudspeaker to an attraction location, where attraction is a factor that favors the activation of a relatively higher loudspeaker closer to the attraction location; the proximity of the loudspeaker to a repulsion location, where repulsion is a factor that favors the activation of a relatively lower loudspeaker closer to the repulsion location; the capability of each loudspeaker relative to other loudspeakers in the environment; the synchronization of the loudspeakers with respect to other loudspeakers; wake word performance; or echo canceller performance.

[0622] In some examples, the relative activation of the speaker can be based on a cost function of the following: a model of the perceived spatial location of the audio signal when played back on the speaker, a measure of the proximity of the expected perceived spatial location of the audio signal to the speaker location, and one or more dynamically configurable features.

[0623] In some examples, minimizing the cost function (including at least one dynamic speaker activation term) can result in the deactivation of at least one speaker (in the sense that each such speaker does not play relevant audio content) and the activation of at least one speaker (in the sense that each such speaker plays at least some rendered audio content). The (multiple) dynamic speaker activation terms can enable at least one of a variety of behaviors, including spatially distorting the audio presentation away from a particular smart audio device, such that the microphone of said particular smart audio device can better hear the speaker or that secondary audio streams can be better heard from the (multiple) speakers of the smart audio device.

[0624] According to some implementations, for each of a plurality of loudspeakers, the individual loudspeaker dynamic processing configuration data may include a dynamic range compressed dataset. In some instances, the dynamic range compressed dataset may include one or more of threshold data, input / output ratio data, attack data, release data, or inflection point data.

[0625] As mentioned above, in some embodiments, the following can be omitted. Figure 37 At least some blocks of method 3700 are shown. For example, in some embodiments, blocks 3705 and 3710 are performed during the setup process. After determining the dynamic processing configuration data of the listening environment, in some embodiments, steps 3705 and 3710 are not performed again during the "runtime" operation unless the type and / or arrangement of the speakers in the listening environment changes. For example, in some embodiments, an initial check may be performed to determine whether any amplifiers have been added or disconnected, whether any amplifier positions have changed, etc. If so, steps 3705 and 3710 can be performed. If not, steps 3705 and 3710 are not performed again before the "runtime" operation, which may involve blocks 3715-3730.

[0626] Figure 38A , Figure 38B and Figure 38C It shows the relationship with Figure 2C and Figure 2D Examples of corresponding loudspeaker participation values. Figure 38A , Figure 38B and Figure 38C In the middle, the angle is -4.1 and Figure 2D The speaker position corresponds to 272, and the angle is 4.1. Figure 2DThe speaker position corresponds to 274, and the angle is -87. Figure 2D The speaker position corresponds to 267, and the angle is 63.6. Figure 2D The speaker position corresponds to 275 degrees, and the angle is 165.4 degrees. Figure 2D The speaker position corresponds to 270. These amplifier input values ​​are relative to the reference. Figures 34 to 37 Examples of weights related to the described spatial region. Based on these examples, Figure 38A , Figure 38B and Figure 38C The loudspeaker participation values ​​shown are... Figure 34 The participation of each loudspeaker in each spatial zone shown corresponds to: Figure 38A The loudspeaker participation values ​​shown correspond to the participation of each loudspeaker in the central area. Figure 38B The loudspeaker participation values ​​shown correspond to the participation of each loudspeaker in the front left and front right zones, and Figure 38C The loudspeaker participation values ​​shown correspond to the participation of each loudspeaker in the back area.

[0627] Figure 39A , Figure 39B and Figure 39C It shows the relationship with Figure 2F and Figure 2G Examples of corresponding loudspeaker participation values. Figure 39A , Figure 39B and Figure 39C In the middle, the angle is -4.1 and Figure 2D The speaker position corresponds to 272, and the angle is 4.1. Figure 2D The speaker position corresponds to 274, and the angle is -87. Figure 2D The speaker position corresponds to 267, and the angle is 63.6. Figure 2D The speaker position corresponds to 275 degrees, and the angle is 165.4 degrees. Figure 2D The speaker position corresponds to 270. Based on these examples, Figure 39A , Figure 39B and Figure 39C The loudspeaker participation values ​​shown are... Figure 34 The participation of each loudspeaker in each spatial zone shown corresponds to: Figure 39A The loudspeaker participation values ​​shown correspond to the participation of each loudspeaker in the central area. Figure 39B The loudspeaker participation values ​​shown correspond to the participation of each loudspeaker in the front left and front right zones, and Figure 39C The loudspeaker participation values ​​shown correspond to the participation of each loudspeaker in the back area.

[0628] Figure 40A , Figure 40B and Figure 40C It shows the relationship with Figure 2H and Figure 2I Examples of corresponding loudspeaker participation values. Based on these examples, Figure 40A , Figure 40B and Figure 40C The loudspeaker participation values ​​shown are... Figure 34 The participation of each loudspeaker in each spatial zone shown corresponds to: Figure 40A The loudspeaker participation values ​​shown correspond to the participation of each loudspeaker in the central area. Figure 40B The loudspeaker participation values ​​shown correspond to the participation of each loudspeaker in the front left and front right zones, and Figure 40C The loudspeaker participation values ​​shown correspond to the participation of each loudspeaker in the back area.

[0629] Figure 41A , Figure 41B and Figure 41C It shows the relationship with Figure 2J and Figure 2K Examples of corresponding loudspeaker participation values. Based on t...

Claims

1. An audio processing method, comprising: The control system receives first audio data corresponding to a first audio program stream. The first audio data includes one or more first audio signals and first spatial data, wherein the first spatial data indicates the associated desired perceived spatial location of each of the one or more first audio signals. as well as The control system renders the first audio data to generate first rendered audio data for at least two speakers in a set of speakers in the environment, wherein: The rendering involves determining the relative activation of speakers in the group of speakers, based at least in part on the perceived spatial location of the first audio signal played back on the speaker, the proximity of the desired perceived spatial location of the first audio signal to the location of the speaker, and one or more additional dynamically configurable functions, wherein the one or more additional dynamically configurable functions depend on at least one or more attributes of the first audio signal, one or more attributes of the group of speakers, or one or more external inputs; and The additional dynamic configurable features include at least the proximity of one or more speakers to an attractive location or the proximity of one or more speakers to a repulsive location, where attractiveness is a factor that favors activation of a relatively higher speaker that is closer to the attractive location, and repulsiveness is a factor that favors activation of a relatively lower speaker that is closer to the repulsive location.

2. The method according to claim 1, wherein, The additional dynamic configurable features include proximity of one or more speakers to one or more listeners.

3. The method according to claim 1, wherein, The additional dynamically configurable features include audibility of one or more speakers at a location in the environment.

4. The method according to claim 1, wherein, The additional dynamic configurable features include at least one of the following: the ability of one or more speakers or the synchronization of the one or more speakers with respect to one or more other speakers in the environment.

5. The method according to claim 1, wherein, The additional dynamically configurable features include at least one of wake word performance or echo canceller performance.

6. The method according to claim 1, wherein, The rendering includes minimizing a cost function, wherein the cost function includes at least one dynamic speaker activation term.

7. The method according to claim 6, wherein, The cost function is based at least in part on the sum of a term corresponding to spatial audio and a term corresponding to the proximity of the loudspeaker to the desired perceived spatial location of the first audio signal in one or more of the first audio signals.

8. The method according to claim 6, wherein, The cost function is based at least in part on centroid amplitude translation (CMAP), flexible virtualization (FV), or a combination of CMAP and FV.

9. The method according to claim 1, further comprising: The control system receives second audio data corresponding to the second audio program stream. The second audio data includes one or more second audio signals and second spatial data, wherein the second spatial data indicates the associated desired perceived spatial location of each of the one or more second audio signals. The control system renders the second audio data to generate second rendered audio data; The first rendered audio data and the second rendered audio data are mixed to generate a mixed audio signal; as well as The mixed audio signal is provided to at least some speakers in the environment.

10. The method of claim 9, further comprising: The rendering process of the first audio data is modified at least in part based on at least one of the second audio signal, the second rendered audio data, the characteristics of the second audio data, or the characteristics of the second rendered audio data to produce modified first rendered audio data.

11. The method according to claim 10, wherein, Modifying the rendering process of the first audio signal involves: modifying the rendering of the first audio signal so that the space of the first audio signal is distorted away from or towards the rendering position of the second rendered audio data.

12. An audio processing system, comprising: Interface system; Control system, including: The first rendering module is configured as follows: Via the interface system, first audio data corresponding to a first audio program stream is received, the first audio data including one or more first audio signals and first spatial data, the first spatial data indicating the associated desired perceived spatial location of each of the one or more first audio signals; and The control system renders the first audio data to generate first rendered audio data for at least two speakers in a set of speakers in the environment, wherein: The rendering involves determining the relative activation of speakers in the group of speakers, based at least in part on the perceived spatial location of the first audio signal played back on the speaker, the proximity of the desired perceived spatial location of the first audio signal to the location of the speaker, and one or more additional dynamically configurable functions, wherein the one or more additional dynamically configurable functions depend on at least one or more attributes of the first audio signal, one or more attributes of the group of speakers, or one or more external inputs; and The additional dynamic configurable features include at least the proximity of one or more speakers to an attractive location or the proximity of one or more speakers to a repulsive location, where attractiveness is a factor that favors activation of a relatively higher speaker that is closer to the attractive location, and repulsiveness is a factor that favors activation of a relatively lower speaker that is closer to the repulsive location.

13. The audio processing system according to claim 12, wherein, The additional dynamic configurable features include proximity of one or more speakers to one or more listeners.

14. The audio processing system according to claim 12, wherein, The additional dynamically configurable features include audibility of one or more speakers at a location in the environment.

15. The audio processing system according to claim 12, wherein, The additional dynamic configurable features include at least one of the following: the ability of one or more speakers or the synchronization of the one or more speakers with respect to one or more other speakers in the environment.

16. The audio processing system according to claim 12, wherein, The additional dynamically configurable features include at least one of wake word performance or echo canceller performance.

17. The audio processing system according to claim 12, wherein, The rendering includes minimizing a cost function, wherein the cost function includes at least one dynamic speaker activation term.

18. The audio processing system according to claim 17, wherein, The cost function is based at least in part on the sum of a term corresponding to spatial audio and a term corresponding to the proximity of the loudspeaker to the desired perceived spatial location of the first audio signal in one or more of the first audio signals.

19. The audio processing system according to claim 17, wherein, The cost function is based at least in part on centroid amplitude translation (CMAP), flexible virtualization (FV), or a combination of CMAP and FV.

20. The audio processing system according to claim 12, further comprising: The second rendering module is configured as follows: The interface system receives second audio data corresponding to the second audio program stream, the second audio data including one or more second audio signals and second spatial data, the second spatial data indicating the associated desired perceived spatial location of each of the one or more second audio signals; as well as Render the second audio data to generate second rendered audio data; as well as The mixing module is configured to mix the first rendered audio data and the second rendered audio data to generate a mixed audio signal, wherein the control system is further configured to provide the mixed audio signal to at least some speakers in the environment.

21. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions that cause one or more processors to perform the method according to any one of claims 1-11.

22. A computer program product comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Processing object-based audio signals

    US20190037333A1

  • Loudness-based compensation for background noise

    US20080267427A1

  • Multi-output antenna

    US20140159971A1