Management of the playback of multiple audio streams through multiple speakers
The method and system for managing multiple audio streams in smart audio devices dynamically adjust playback across speakers using techniques like CMAP and FV, addressing the challenge of diverse usage scenarios and user commands for enhanced audio clarity and interaction.
Patent Information
- Application Number
- JP2024176251
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-21
- Filing Date
- 2024-10-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-07-27
AI Technical Summary
Existing audio systems struggle to manage the flexible and simultaneous playback of multiple audio streams across multiple speakers, particularly in smart audio devices, failing to adapt to diverse usage scenarios and user commands effectively.
A method and system for managing playback of multiple audio streams by coordinating smart audio devices, involving techniques like Center of Mass Amplitude Panning and Flexible Virtualization, which dynamically modify the rendering of audio streams based on microphone inputs and other audio streams to optimize playback across arbitrarily placed speakers.
Enables flexible and adaptive audio playback across multiple speakers, enhancing user experience by adjusting spatial presentation and loudness in response to user commands and environmental conditions, providing improved audio clarity and user interaction.
Smart Images

Figure 0007714764000095 
Figure 0007714764000096 
Figure 0007714764000097
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 992,068, filed on March 19, 2020; U.S. Provisional Patent Application No. 62 / 949,998, filed on December 18, 2019; European Patent Application No. 19217580.0, filed on December 18, 2019; Spanish Patent Application No. P201930702, filed on July 30, 2019; U.S. Provisional Patent Application No. 62 / 971,421, filed on February 7, 2020; U.S. Provisional Patent Application No. 62 / 705,410, filed on June 25, 2020; U.S. Provisional Patent Application No. 62 / 880,111, filed on July 30, 2019; U.S. Provisional Patent Application No. 62 / 704,754, filed on May 27, 2020; U.S. Provisional Patent Application No. 62 / 705,896, filed on July 21, 2020; U.S. Provisional Patent Application No. 62 / 880,114, filed on July 30, 2019; U.S. Provisional Patent Application No. 62 / 705,351, filed on June 23, 2020; U.S. Provisional Patent Application No. 62 / 880,115, filed on July 30, 2019; and U.S. Provisional Patent Application No. 62 / 705,143, filed on June 12, 2020. Each of these applications is hereby incorporated by reference in its entirety.
[0002] Technical Field The present disclosure relates to systems and methods for audio playback and rendering for some or all speakers (e.g., each activated speaker) of a set of speakers.
Background Art
[0003] Audio devices, including but not limited to smart audio devices, are widely deployed and are becoming a common feature in many homes. Existing systems and methods for controlling audio devices provide advantages, but improved systems and methods would be desirable.
[0004] Notation and Nomenclature Throughout this disclosure, including the claims, the terms "speaker" and "loudspeaker" are used synonymously to represent any sound-emitting transducer (or set of transducers) driven by a single speaker feed. A typical set of headphones includes two speakers. A speaker can be implemented to include multiple transducers (e.g., a woofer and a tweeter) that can be driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feeds can undergo different processing in different circuit branches coupled to different transducers.
[0005] Throughout this disclosure, including the claims, the expression "perform an operation on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to a signal or data) is used in a broad sense to indicate performing the operation directly on the signal or data or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing before performing the operation).
[0006] Throughout this disclosure, including the claims, the term "system" is used in a broad sense to denote an apparatus, a system, or a subsystem. For example, a subsystem implementing a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other X - M inputs are received from an external source) can also be referred to as a decoder system.
[0007] Throughout this disclosure, including the claims, the term "processor" is used in a broad sense to refer to a system or device that is programmable or otherwise configurable (e.g., using software or firmware) to perform operations on data (e.g., audio, video, or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chip sets), digital signal processors programmed and / or otherwise configured to perform pipelined processing on audio or other voice data, programmable general-purpose processors or computers, and programmable microprocessor chips or chip sets.
[0008] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean a direct or indirect connection. Thus, when a first device is coupled to a second device, that connection can be through a direct connection or through an indirect connection via other devices and connections.
[0009] As used herein, a "smart device" is an electronic device that is generally configured to communicate with one or more other devices (or networks) via various wireless protocols such as Bluetooth, Zigbee, near field communication, Wi-Fi, light fidelity (Li-Fi), 3G, 4G, 5G, etc., and that can operate in a somewhat interactive and / or autonomous manner. Some notable types of smart devices are smartphones, smart cars, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smartwatches, smart bands, smart keychains, smart audio devices. The term "smart device" may also refer to a device that exhibits certain characteristics of ubiquitous computing such as artificial intelligence.
[0010] As used herein, the expression "smart audio device" refers to a smart device that is either a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of a virtual assistant function). A single-purpose audio device is a device (e.g., a television (TV) or a mobile phone) that includes or is coupled to at least one microphone (and optionally includes or is coupled to at least one speaker and / or at least one camera), and is designed to achieve mostly or primarily a single purpose. For example, a television can typically (and is considered to be able to) play audio from program material, but in most cases, modern televisions run some operating system on which applications, including a TV viewing application, operate locally. Similarly, the audio input / output of a mobile phone may do many things, but these are served by applications that operate on the phone. In this sense, a single-purpose audio device with speakers and microphones is often configured to run local applications and / or services that directly use the speakers and microphones. Some single-purpose audio devices can be configured to be grouped to achieve audio playback across zones or user-configured areas.
[0011] One common type of multi-purpose audio device is an audio device that implements at least some aspects of a virtual assistant function, although other aspects of the virtual assistant function may be implemented by one or more other devices, such as one or more servers to which the multi-purpose audio device is configured to communicate. Such multi-purpose audio devices may be referred to herein as "virtual assistants." A virtual assistant is a device (e.g., a smart speaker or voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally includes or is coupled to at least one speaker and / or at least one camera). In some examples, the virtual assistant can provide the ability to utilize multiple devices (different from the virtual assistant itself) for applications that are enabled in some sense by the cloud or otherwise not fully implemented within or on the virtual assistant itself. In other words, at least some aspects of the virtual assistant function, such as the speech recognition function, may be implemented (at least in part) by one or more servers or other devices to which the virtual assistant can communicate via a network such as the Internet. Virtual assistants may sometimes cooperate, for example, in a discrete, conditionally defined manner. For example, two or more virtual assistants can cooperate in the sense that the one that is most confident in having heard a wake word responds to that word. Connected virtual assistants can, in some implementations, form a kind of constellation, which may be managed by one main application that may be a virtual assistant (or implement it).
[0012] Here, the "wake word" is used in a broad sense to mean any sound (e.g., a word uttered by a human or some other sound), and the smart audio device is configured to wake up in response to the detection of that sound ("listening") using at least one microphone included in or coupled to the smart audio device, or at least one other microphone. In this context, "waking up" means entering a state where the device waits for a voice command (i.e., listens for whether there is a voice command). In some cases, what may be referred to herein as a "wake word" may include multiple words, e.g., a phrase.
[0013] Here, the expression "wake word detector" refers to a device (or software including instructions for configuring a device) configured to continuously search for an alignment between real-time audio (e.g., speech) features and a trained model. Typically, a wake word event is triggered whenever the wake word detector determines that the probability of the wake word being detected exceeds a predetermined threshold. For example, the threshold may be a predetermined threshold adjusted to give a reasonable compromise between the false acceptance rate and the false rejection rate. Following the wake word event, the device may enter a state (which may be referred to as the "awake" state or the "watching" state) where it waits for a command and passes the received command to a larger, more computationally intensive recognizer. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0014] Some embodiments relate to a method for managing the playback of multiple audio streams by at least one (e.g., all or part) of a set of smart audio devices and / or at least one (e.g., all or part) of another set of speakers of a smart audio device.
[0015] One class of embodiments includes a method for managing playback by at least one (e.g., all or part) of a plurality of coordinated (orchestrated) smart audio devices. For example, a set of smart audio devices present (within the system) in a user's home can be orchestrated to handle various simultaneous usage scenarios, including flexible rendering of audio for playback by all or part of the smart audio devices (i.e., by speakers of all or part of the smart audio devices).
[0016] (For example, in a home environment to handle various simultaneous usage scenarios) orchestrating smart audio devices may involve the simultaneous playback of one or more audio program streams through a set of interconnected speakers. For example, a user may be listening to an Atmos sound track of a movie (or other object-based audio program) through a set of speakers (e.g., those included in a set of smart audio devices or controlled by a set of smart audio devices), and in this case, the user may utter commands (e.g., wake word and subsequent commands) to a related smart audio device (e.g., a smart assistant). In this situation, the spatial presentation of the program (e.g., Atmos mix) may be warped away from the position of the speaker (the user speaking), and the playback of the corresponding response of the smart audio device (e.g., of the voice assistant) may be directed to a response at a speaker closer to the speaker, and the audio playback by the system may be modified (according to some embodiments). This can provide significant advantages compared to simply reducing the volume of the playback of the audio program content in response to the detection of a command (or corresponding wake word). Similarly, a user may wish to use the speakers to obtain cooking hints in the kitchen while the same program (e.g., Atmos sound track) is being played in an adjacent open living space. In this case, according to some embodiments, the playback of the program (e.g., Atmos sound track) can be warped away from the kitchen, and the cooking hints can be played from speakers near or within the kitchen. Further, the cooking hints played in the kitchen can be louder than any program (e.g., Atmos sound track) that may leak from the living space and can be dynamically adjusted (according to some embodiments) to be heard by a person in the kitchen.
[0017] Some embodiments are multiple-stream rendering systems configured to implement the above exemplary use cases and many others contemplated. In one class of embodiments, an audio rendering system may be configured to render (and / or play simultaneously) multiple audio program streams for simultaneous playback through a plurality of arbitrarily placed loudspeakers, at least one of the program streams being a spatial mix, and the rendering (or rendering and playback) of the spatial mix being dynamically modified in response to (or in relation to) the simultaneous playback (or rendering and playback) of one or more additional program streams.
[0018] Aspects of some embodiments include a system configured (e.g., programmed) to execute any embodiment of the disclosed methods or steps thereof, and a tangible, non-transitory computer-readable medium (e.g., a disk or other tangible storage medium) implementing non-transitory storage of data storing code (e.g., code executable to execute) for executing any embodiment of the disclosed methods or steps thereof. For example, some embodiments may be or may include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data including embodiments of the disclosed methods or steps thereof, or may include them. Such a general-purpose processor may be part of or included in a computer system including a processing subsystem programmed (and / or otherwise configured) to execute embodiments of the disclosed methods (or steps thereof) in response to an input device, memory, and data presented thereto.
[0019] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more apparatuses may be capable of implementing, at least in part, the methods disclosed herein. In some implementations, the apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.
[0020] In some implementations, the control system includes or implements at least two rendering modules. According to some examples, the control system may include or implement N rendering modules, where N is an integer greater than 2.
[0021] In some examples, the first rendering module is configured to receive a first audio program stream via the interface system. In some cases, the first audio program stream includes a first audio signal scheduled to be reproduced by at least some speakers of the environment. In some examples, the first audio program stream includes first spatial data including channel data and / or spatial metadata. According to some examples, the first rendering module is configured to render the first audio signal for reproduction via the speakers of the environment and generate a first rendered audio signal.
[0022] In some implementations, the second rendering module is configured to receive a second audio program stream via an interface system. In some examples, the second audio program stream includes a second audio signal scheduled to be played by at least some of the speakers of the environment. In some examples, the second audio program stream includes second spatial data including channel data and / or spatial metadata. According to some examples, the second rendering module is configured to render the second audio signal for playback via the speakers of the environment to generate a second rendered audio signal.
[0023] According to some examples, the first rendering module is configured to modify a rendering process for the first audio signal and generate a modified first rendered audio signal based at least in part on at least one of the second audio signal, the second rendered audio signal, or their characteristics. According to some examples, the second rendering module is further configured to modify a rendering process for the second audio signal and generate a modified second rendered audio signal based at least in part on at least one of the first audio signal, the first rendered audio signal, or their characteristics.
[0024] In some implementations, the audio processing system includes a mixing module configured to mix the modified first rendered audio signal and the modified second rendered audio signal to generate a mixed audio signal. In some examples, the control system is further configured to provide the mixed audio signal to at least some of the speakers of the environment.
[0025] According to some examples, the audio processing system may include one or more additional rendering modules. In some examples, each of the one or more additional rendering modules may be configured to receive an additional audio program stream via an interface system. The additional audio program stream may include additional audio signals scheduled to be played by at least one speaker of the environment. In some examples, each of the one or more additional rendering modules may be configured to render the additional audio signals for playback via at least one speaker of the environment and generate an additional rendered audio signal. In some examples, each of the one or more additional rendering modules may be configured to modify the rendering process for the additional audio signals, at least in part, based on at least one of a first audio signal, a first rendered audio signal, a second audio signal, a second rendered audio signal, or their characteristics, to generate a modified additional rendered audio signal. In some such examples, a mixing module may be further configured to mix the modified additional rendered audio signal with at least the modified first rendered audio signal and the modified second rendered audio signal to generate the mixed audio signal.
[0026] In some implementations, modifying the rendering process for the first audio signal may include warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal. Alternatively or additionally, modifying the rendering process for the first audio signal may include modifying the loudness of one or more of the first rendered audio signals in response to the loudness of one or more of the second audio signal or the second rendered audio signal.
[0027] According to some examples, modifying the rendering process for the second audio signal may include warping the rendering of the second audio signal away from the rendering position of the first rendered audio signal. Alternatively or additionally, modifying the rendering process for the second audio signal may include modifying the loudness of one or more of the second rendered audio signals in response to the loudness of one or more of the first audio signal or the first rendered audio signal. According to some implementations, modifying the rendering process for the first audio signal and / or the second audio signal may include performing spectral modification, audibility-based modification, and / or dynamic range modification.
[0028] In some examples, the audio processing system may include a microphone system that includes one or more microphones. In some such examples, the first rendering module may be configured to modify a rendering process for a first audio signal based at least in part on a first microphone signal from the microphone system. In some such examples, the second rendering module may be configured to modify a rendering process for a second audio signal based at least in part on the first microphone signal.
[0029] According to some examples, the control system may be further configured to estimate a first sound source position based on the first microphone signal and modify a rendering process for at least one of the first audio signal or the second audio signal based at least in part on the first sound source position. In some examples, the control system may determine whether the first microphone signal corresponds to ambient noise and modify a rendering process for at least one of the first audio signal or the second audio signal based at least in part on whether the first microphone signal corresponds to ambient noise.
[0030] In some examples, the control system may determine whether the first microphone signal corresponds to a human voice and be configured to modify a rendering process for at least one of the first audio signal or the second audio signal based at least in part on whether the first microphone signal corresponds to a human voice. According to some such examples, modifying the rendering process for the first audio signal may include reducing the loudness of a first rendered audio signal played by a speaker closer to the first sound source compared to the loudness of a first rendered audio signal played by a speaker farther from the first sound source.
[0031] According to some examples, the control system may be configured to determine that the first microphone signal corresponds to a wake word, determine a response to the wake word, and control at least one speaker near the first sound source location to play the response. In some examples, the control system may be configured to determine that the first microphone signal corresponds to a command, determine a response to the command, control at least one speaker near the first sound source location to play the response, and execute the command. According to some examples, the control system may be further configured to return to an unmodified rendering process for the first audio signal after controlling at least one speaker near the first sound source location to play the response.
[0032] In some implementations, the control system may be configured to derive a loudness estimate for at least the first audio program stream being played and / or the second audio program stream being played, based at least in part on the first microphone signal. According to some examples, the control system may be further configured to modify a rendering process for at least one of the first audio signal or the second audio signal, based at least in part on the loudness estimate. In some cases, the loudness estimate may be a perceived loudness estimate. According to some such examples, modifying the rendering process may include changing at least one of the first audio signal or the second audio signal to maintain the perceived loudness of the first audio signal and / or the second audio signal in the presence of interfering signals.
[0033] In some examples, the control system may be configured to determine that the first microphone signal corresponds to a human voice and reproduce the first microphone signal at one or more speakers close to a location in the environment that is different from the first sound source position. According to some such examples, the control system may be further configured to determine whether the first microphone signal corresponds to a child's cry. In some such examples, the location in the environment may correspond to the estimated position of the caregiver.
[0034] According to some examples, the control system may be configured to derive a loudness estimate for the first audio program stream being reproduced and / or the second audio program stream being reproduced. In some such examples, the control system may be further configured to modify the rendering process for the first audio signal and / or the second audio signal, at least in part, based on the loudness estimate. According to some examples, the loudness estimate may be a perceived loudness estimate. Modifying the rendering process may include changing at least one of the first audio signal or the second audio signal to maintain the perceived loudness in the presence of interfering signals.
[0035] In some implementations, rendering the first audio signal and / or rendering the second audio signal can include flexible rendering to arbitrarily placed speakers. In some such examples, the flexible rendering may include Center of Mass Amplitude Panning or Flexible Virtualization.
[0036] At least some aspects of the present disclosure may be implemented by one or more audio processing methods. In some examples, the method may be implemented, at least in part, by a control system such as those disclosed herein. Some such methods include receiving, by a first rendering module, a first audio program stream, the first audio program stream including a first audio signal scheduled to be reproduced by at least some speakers of an environment. In some examples, the first audio program stream includes first spatial data including channel data and / or spatial metadata. Some such methods include rendering, by the first rendering module, the first audio signal for reproduction via the speakers of the environment to generate a first rendered audio signal.
[0037] Some such methods include receiving, by a second rendering module, a second audio program stream. In some examples, the second audio program stream includes a second audio signal scheduled to be reproduced by at least one speaker of the environment. Some such methods include rendering, by the second rendering module, the second audio signal for reproduction via at least one speaker of the environment to generate a second rendered audio signal.
[0038] Some such methods include modifying a rendering process for a first audio signal, at least in part, based on at least one of a second audio signal, a second rendered audio signal, or their characteristics, by a first rendering module, and generating a modified first rendered audio signal. Some such methods include modifying a rendering process for a second audio signal, at least in part, based on at least one of the first audio signal, the first rendered audio signal, or their characteristics, by a second rendering module, and generating a modified second rendered audio signal. Some such methods include mixing the modified first rendered audio signal and the modified second rendered audio signal to generate a mixed audio signal, and providing the mixed audio signal to at least some speakers of an environment.
[0039] According to some examples, modifying the rendering process for the first audio signal may include warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal, and / or modifying one or more loudnesses of the first rendered audio signal in response to one or more loudnesses of the second audio signal or the second rendered audio signal.
[0040] In some examples, modifying the rendering process for the second audio signal may include warping the rendering of the second audio signal away from the rendering position of the first rendered audio signal, and / or modifying one or more loudnesses of the second rendered audio signal in response to one or more loudnesses of the first audio signal or the first rendered audio signal.
[0041] According to some examples, modifying the rendering process for the first audio signal may include performing spectral modification, audibility-based modification, and / or dynamic range modification.
[0042] Some methods may include modifying the rendering process for the first audio signal by a first rendering module based at least in part on a first microphone signal from a microphone system. Some methods may include modifying the rendering process for the second audio signal by a second rendering module based at least in part on the first microphone signal.
[0043] Some methods may include estimating a first sound source position based on the first microphone signal and modifying the rendering process for at least one of the first audio signal or the second audio signal based at least in part on the first sound source position.
[0044] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include, but are not limited to, memory devices such as random access memory (RAM) devices, read-only memory (ROM) devices, etc., like those described herein. Thus, some innovative aspects of the subject matter described in this disclosure can be implemented in non-transitory media having software stored thereon. Instructions for controlling one or more devices to perform a method may include receiving, by a first rendering module, a first audio program stream, the first audio program stream including a first audio signal scheduled to be played by at least one speaker of an environment. In some examples, the first audio program stream includes first spatial data including channel data and / or spatial metadata. Some such methods include rendering, by the first rendering module, the first audio signal for playback via speakers of the environment to generate a first rendered audio signal.
[0045] Some such methods include receiving, by a second rendering module, a second audio program stream. In some examples, the second audio program stream includes a second audio signal scheduled to be played by at least one speaker of the environment. Some such methods include rendering, by the second rendering module, the second audio signal for playback via at least one speaker of the environment to generate a second rendered audio signal.
[0046] Some such methods include modifying, by a first rendering module, a rendering process for a first audio signal based at least in part on at least one of a second audio signal, a second rendered audio signal, or their characteristics, and generating a modified first rendered audio signal. Some such methods include modifying, by a second rendering module, a rendering process for a second audio signal based at least in part on at least one of the first audio signal, the first rendered audio signal, or their characteristics, and generating a modified second rendered audio signal. Some such methods include mixing the modified first rendered audio signal and the modified second rendered audio signal to generate a mixed audio signal, and providing the mixed audio signal to at least some speakers of the environment.
[0047] According to some examples, modifying the rendering process for the first audio signal may include warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal and / or modifying the loudness of one or more of the first rendered audio signals in response to the loudness of one or more of the second audio signal or the second rendered audio signal.
[0048] In some examples, modifying the rendering process for the second audio signal may include warping the rendering of the second audio signal away from the rendering position of the first rendered audio signal and / or modifying the loudness of one or more of the second rendered audio signals in response to the loudness of one or more of the first audio signal or the first rendered audio signal.
[0049] According to some examples, modifying a rendering process for a first audio signal may include performing spectral modification, audibility-based modification, and / or dynamic range modification.
[0050] Some methods may include modifying, by a first rendering module, a rendering process for a first audio signal based at least in part on a first microphone signal from a microphone system. Some methods may include modifying, by a second rendering module, a rendering process for a second audio signal based at least in part on the first microphone signal.
[0051] Some methods may include estimating a first sound source position based on the first microphone signal and modifying a rendering process for at least one of the first audio signal or the second audio signal based at least in part on the first sound source position.
[0052] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions in the following figures may not be drawn to scale.
Brief Description of the Drawings
[0053]
Figure 1A
Figure 1B
Figure 2A
Figure 2B
Figure 2C
Figure 2D
Figure 2E
Figure 2F
Figure 2G
Figure 2H
Figure 2I
Figure 2J
Figure 2K
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 5C
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12A
Figure 12B
Figure 12C
Figure 12D
Figure 12E
Figure 13A
Figure 13B
Figure 14A
Figure 14B
Figure 15
Figure 16A
Figure 16B
Figure 16C
Figure 16D
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21A
Figure 21B
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30A
Figure 30B
Figure 30C
Figure 30D
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38A
Figure 38B
Figure 38C
Figure 39A
Figure 39B
Figure 39C
Figure 40A
Figure 40B
Figure 40C
Figure 41A
Figure 41B
Figure 41C
Figure 42
[0054] Similar reference numbers and designations in the various drawings indicate similar elements. **DETAILED DESCRIPTION OF THE INVENTION**
[0055] Flexible rendering is a technique for rendering spatial audio through any number of arbitrarily arranged speakers. With the spread of smart audio devices (e.g., smart speakers) in homes, there is a need to implement flexible rendering technology that allows consumers to perform flexible rendering of audio and playback of the rendered audio using smart audio devices.
[0056] To achieve flexible rendering, several techniques have been developed, including Center of Mass Amplitude Panning (CEAP, CMAP) and Flexible Virtualization (FV). Both of these techniques cast the rendering problem as a cost function minimization problem. The cost function consists of two terms: a first term that models the desired spatial impression that the renderer is trying to achieve, and a second term that assigns a cost to activating speakers. To date, this second term has focused on creating a sparse solution where only speakers close to the desired spatial position of the rendered audio are activated.
[0057] Some embodiments relate to a method for managing playback of multiple streams of audio by at least one (e.g., all or part) of a set of smart audio devices (or at least one (e.g., all or part) of another set of speakers).
[0058] One class of embodiments includes a method of managing playback by at least one (e.g., all or part) of a plurality of coordinated (orchestrated) smart audio devices. For example, a set of smart audio devices present (within the system) in a user's home can be orchestrated to handle various simultaneous usage scenarios, including flexible rendering of audio for playback by all or part of the smart audio devices (i.e., by speakers of all or part of the smart audio devices).
[0059] Orchestrating smart audio devices in a home (e.g., to handle diverse simultaneous usage scenarios) may involve the simultaneous playback of one or more audio program streams through an interconnected set of speakers. For example, a user may be listening to an Atmos sound track of a movie (or other object-based audio program) through a set of speakers, and in this case, the user may issue a command to a related smart assistant (or other smart audio device). In this situation, the system may (in some embodiments) modify the audio playback such that the spatial presentation of the Atmos mix is warped away from the position of the speaker (the user speaking) and away from the nearest smart audio device, while simultaneously warping the playback of the corresponding response of the smart audio device (of the voice assistant) towards the position of the speaker. This can provide significant advantages compared to simply reducing the volume of the playback of the audio program content in response to the detection of a command (or corresponding wake word). Similarly, a user may desire to use the speakers to obtain cooking hints in the kitchen while the same Atmos sound track is being played in an adjacent open living space. In this case, according to some examples, the Atmos sound track can be warped away from the kitchen and / or the loudness of one or more rendered signals of the Atmos sound track can be modified in response to the loudness of one or more rendered signals of the cooking hint sound track. Further, in some implementations, the cooking hint played in the kitchen can be made louder than any of the Atmos sound track that may be bleeding into the living space and dynamically adjusted to be heard by the person in the kitchen.
[0060] Some embodiments relate to a multi - stream rendering system configured to implement the above - described exemplary use cases and numerous others that may be contemplated. In one class of embodiments, an audio rendering system may be configured to simultaneously play a plurality of audio program streams through a plurality of arbitrarily - placed loudspeakers, at least one of the program streams being a spatial mix, and the rendering of the spatial mix being dynamically modified in response to (or in relation to) the simultaneous playback of one or more additional program streams.
[0061] In some embodiments, a multi - stream renderer may be configured to implement the scenarios described above and many other cases where the simultaneous playback of a plurality of audio program streams must be managed. Some implementations of a multi - stream rendering system may be configured to perform the following operations: ● Simultaneously render and play a plurality of audio program streams through a plurality of arbitrarily - placed loudspeakers. At least one of the program streams is a spatial mix. ○ The term "program stream" refers to a collection of one or more audio signals intended to be listened to together as a whole. Examples include a music selection, a movie soundtrack, a podcast, a live voice call, a synthetic voice response from a smart assistant, etc. ○ A spatial mix is a program stream intended to deliver different (more than mono) signals to the listener's left and right ears. Examples of audio formats for spatial mixes include stereo, 5.1 and 7.1 surround sound, object - based audio formats such as Dolby Atmos, and ambisonics. ○ The rendering of a program stream refers to the process of actively distributing the relevant one or more audio signals across a plurality of loudspeakers to achieve a particular perceptual impression. ● Dynamically modify the rendering of the at least one spatial mix as a function of the rendering of one or more of the additional program streams. Examples of such modifications to the rendering of the spatial mix include, but are not limited to, the following: ○ Modify the relative activation of a plurality of loudspeakers as a function of the relative activation of the loudspeakers associated with the rendering of at least one of the one or more additional program streams. ○ Warp the intended spatial balance of the spatial mix as a function of the spatial characteristics of the rendering of at least one of the one or more additional program streams. ○ Modify the loudness or audibility of the spatial mix as a function of the loudness or audibility of at least one of the one or more additional program streams.
[0062] FIG. 1A is a block diagram showing an example of the components of an apparatus that can implement various aspects of the present disclosure. According to some examples, apparatus 100 may be a smart audio device configured to execute at least a portion of the methods disclosed herein, or may include a smart audio device. In other implementations, apparatus 100 may be or may include another device configured to execute at least some of the methods disclosed herein, such as a laptop computer, a cellular phone, a tablet device, a smart home hub, etc. In some such implementations, apparatus 100 may be or may include a server. In some implementations, apparatus 100 may be configured to implement what may be referred to herein as an "audio session manager".
[0063] In this example, the apparatus 100 includes an interface system 105 and a control system 110. The interface system 105 may be configured to communicate with one or more devices that, in some implementations, are running or are configured to run a software application. Such a software application may be referred to herein as an "application" or simply an "app". The interface system 105 may be configured to exchange control information and related data regarding the application in some implementations. The interface system 105 may be configured to communicate with one or more other devices in an audio environment in some implementations. The audio environment may be a home audio environment in some examples. The interface system 105 may be configured to exchange control information and related data with an audio device in the audio environment in some implementations. The control information and related data may be related to one or more applications that the apparatus 100 is configured to communicate using in some examples.
[0064] The interface system 105 may be configured to receive an audio program stream in some implementations. The audio program stream may include audio signals scheduled to be reproduced by at least some speakers in the environment. The audio program stream may include spatial data such as channel data and / or spatial metadata. The interface system 105 may be configured to receive input from one or more microphones within the environment in some implementations.
[0065] The interface system 105 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus interfaces). According to some implementations, the interface system 105 may include one or more wireless interfaces. The interface system 105 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 105 may include one or more interfaces between the control system 110 and a memory system such as the optional memory system 115 shown in FIG. 1A, but in some cases, the control system 110 may include a memory system.
[0066] The control system 110 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gates or transistor logic, and / or discrete hardware components.
[0067] In some implementations, control system 110 may be present in more than one device. For example, a portion of control system 110 may be present within a device in one of the environments shown herein, and another portion of control system 110 may be present within a device outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer). In other examples, a portion of control system 110 may be present within a device in one of the environments shown herein, and another portion of control system 110 may be present within one or more other devices of the environment. For example, the functions of the control system may be distributed across multiple smart audio devices of the environment, or may be shared by an orchestration device (e.g., something that may be referred to herein as a smart home hub) and one or more other devices of the environment. Interface system 105 may also be present in more than one device in some such examples.
[0068] In some implementations, control system 110 may be configured, at least in part, to execute the methods disclosed herein. According to some examples, control system 110 may be configured to implement a method for managing the playback of multiple audio streams through multiple speakers.
[0069] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non - transitory media. Such non - transitory media may include memory devices such as, but not limited to, random access memory (RAM) devices, read - only memory (ROM) devices, and the like as described herein. The one or more non - transitory media may be present, for example, in the optional memory system 115 and / or control system 110 shown in FIG. 1A. Thus, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non - transitory media storing software. The software may include, for example, instructions for controlling at least one device to process audio data. The software may be executable by one or more components of a control system such as, for example, the control system 110 of FIG. 1A.
[0070] In some examples, device 100 may include the optional microphone system 120 shown in FIG. 1A. The optional microphone system 120 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device such as a speaker of a speaker system, a smart audio device, etc. In some examples, device 100 may not include the microphone system 120, but in some such implementations, device 100 may still be configured to receive microphone data for one or more microphones in an audio environment via the interface system 110.
[0071] According to some implementations, device 100 may include an optional loudspeaker system 125 as shown in FIG. 1A. The optional speaker system 125 may include one or more speakers. A loudspeaker is also referred to as a "speaker" in this document. In some examples, at least some of the speakers of the optional speaker system 125 may be arbitrarily arranged. For example, at least some of the speakers of the optional speaker system 125 may be arranged in positions that do not conform to any loudspeaker layout defined by standards such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, and Hamasaki 22.2. In some such examples, at least some of the loudspeakers of the optional speaker system 125 may be arranged in a space-convenient position (e.g., a position where there is space to accommodate the loudspeaker), but may also be in a position that is not in any loudspeaker layout defined by a standard. In some examples, device 100 may not include a loudspeaker system 125.
[0072] In some implementations, device 100 may include an optional sensor system 129 as shown in FIG. 1A. The optional sensor system 129 may include one or more cameras, touch sensors, gesture sensors, motion detectors, etc. According to some implementations, the optional sensor system 129 may include one or more cameras. In some implementations, the camera may be a stand-alone camera. In some examples, one or more cameras of the optional sensor system 129 may be present within a smart audio device that may be a single-purpose audio device or a virtual assistant. In some such examples, one or more cameras of the optional sensor system 129 may be present in a television, a mobile phone, or a smart speaker. In some examples, device 100 may not include the sensor system 129. However, in some such implementations, device 100 may still be configured to receive sensor data about one or more sensors in the audio environment via the interface system 110.
[0073] In some implementations, device 100 may include an optional display system 135 as shown in FIG. 1A. The optional display system 135 may include one or more displays such as one or more light-emitting diode (LED) displays. In some cases, the optional display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples where device 100 includes the display system 135, the sensor system 129 may include a touch sensor system and / or a gesture sensor system proximate to one or more displays of the display system 135. According to some such implementations, the control system 110 may be configured to control the display system 135 to present one or more graphical user interfaces (GUIs).
[0074] According to some such examples, the device 100 may be a smart audio device or may include a smart audio device. In some such implementations, the device 100 may be a wake word detector or may include a wake word detector. For example, the device 100 may be a virtual assistant or may include a virtual assistant.
[0075] Figure 1B is a block diagram of a minimal version of an embodiment. Depicted are N program streams (N≥2), with the first one explicitly labeled as spatial. The corresponding collection of audio signals is fed through the corresponding renderers, which are each configured for the reproduction of the corresponding program stream through a common set of M arbitrarily spaced loudspeakers (M≥2). The renderers may be referred to as “rendering modules”. The rendering module and mixer 130a may be implemented via software, hardware, firmware, or some combination thereof. In this example, the rendering module and mixer 130a are implemented via a control system 110a, which is an example of the control system 110 described above with reference to FIG. 1A. Each of the N renderers outputs a set of M loudspeaker feeds. The loudspeaker feeds are summed across all N renderers for simultaneous reproduction through the M loudspeakers. According to this implementation, information regarding the layout of the M loudspeakers in the listening environment is provided to all renderers, as indicated by the dashed lines feeding back from the loudspeaker block, such that the renderers can be appropriately configured for reproduction through those speakers. This layout information may or may not be transmitted from one or more of the speakers themselves, depending on the specific implementation. According to some examples, the layout information may be provided by one or more smart speakers configured to determine the relative position of each of the M loudspeakers in the listening environment. Some of such automatic location determination methods can be based on the direction of arrival (DOA) method or the time of arrival (TOA) method. In other examples, this layout information may be determined by another device and / or input by the user.In some examples, loudspeaker specification information regarding at least some of the capabilities of M loudspeakers in an acoustic environment may be provided to all renderers. Such loudspeaker specification information can include impedance, frequency response, sensitivity, power rating, number and location of individual drivers, and the like. According to this example, information from one or more renderings of an additional program stream is fed to a renderer of the primary spatial stream, whereby the rendering can be dynamically modified as a function of the information. This information is represented by a dashed line returning from rendering blocks 2 through N upward to rendering block 1.
[0076] Figure 2A shows another (more capable) embodiment with additional features. In this example, the rendering module and mixer 130b are implemented via a control system 110b, which is an example of the control system 110 described above with reference to FIG. 1A. In this version, the dashed line that moves up and down between all N renderers represents the idea that any one of the N renderers can contribute to the dynamic modification of any one of the remaining N - 1 renderers. In other words, the rendering of any one of the N program streams can be dynamically modified as a function of the combination of one or more renderings of any one of the remaining N - 1 program streams. Further, any one or more of the program streams may be spatial mixes, and the rendering of any program stream may be dynamically modified as a function of any of the other program streams, whether or not it is spatial. Loudspeaker layout information may be provided to the N renderers, for example, as described above. In some examples, loudspeaker specification information may be provided to the N renderers. In some implementations, the microphone system 120a may include a set of K microphones (K≧1) within the listening environment. In some examples, the microphones may be attached to or associated with one or more of the loudspeakers. These microphones can feedback both the captured audio signals represented by solid lines and additional configuration information (e.g., their positions) represented by dashed lines to the set of N renderers. Any one of the N renderers can then be dynamically modified as a function of this additional microphone input. Various examples are provided herein.
[0077] Examples of information derived from the microphone input and then used to dynamically modify any one of the N renderers include, but are not limited to, the following: · Detection of the utterance of specific words or phrases by the user of the system. ·Estimation of the location of one or more users of the system. ·Estimated value of the loudness of any combination of N program streams at a particular location within the listening space. ·Estimated value of the loudness of other ambient sounds such as background noise in the listening environment.
[0078] Figure 2B is a flowchart outlining an example of a method that can be performed by an apparatus or system as shown in Figures 1A, 1B, or 2A. The blocks of method 200 are not necessarily performed in the order shown, similar to other methods described herein. Further, such a method may include more or fewer blocks than those illustrated and / or described. The blocks of method 200 may be performed by one or more devices that are (or may include) one example of a control system such as control system 110, control system 110a, or control system 110b shown in Figures 1A, 1B, and 2A or other disclosed control system examples.
[0079] In this implementation, block 205 includes receiving a first audio program stream via an interface system. In this example, the first audio program stream includes a first audio signal scheduled to be reproduced by at least some of the speakers in the environment. Here, the first audio program stream includes first spatial data. According to this example, the first spatial data includes channel data and / or spatial metadata. In some examples, block 205 includes the first rendering module of the control system receiving the first audio program stream via the interface system.
[0080] According to this example, block 210 includes rendering a first audio signal for playback via the speakers of the environment and generating a first rendered audio signal. Some examples of method 200 include receiving loudspeaker layout information, for example as described above. Some examples of method 200 include receiving loudspeaker specification information, for example as described above. In some examples, the first rendering module may generate the first rendered audio signal based at least in part on the loudspeaker layout information and / or the loudspeaker specification information.
[0081] In this example, block 215 includes receiving a second audio program stream via an interface system. In this implementation, the second audio program stream includes a second audio signal scheduled to be played back by at least some of the speakers of the environment. According to this example, the second audio program stream includes second spatial data. The second spatial data includes channel data and / or spatial metadata. In some examples, block 215 includes the second rendering module of the control system receiving the second audio program stream via the interface system.
[0082] According to this implementation, block 220 includes rendering a second audio signal for playback via the speakers of the environment and generating a second rendered audio signal. In some examples, the second rendering module can generate the second rendered audio signal based at least in part on the received loudspeaker layout information and / or the received speaker specification information.
[0083] In some cases, some or all of the speakers in the environment can be arbitrarily positioned. For example, at least some of the speakers in the environment may be arranged in positions that do not conform to the speaker layouts defined by standards such as Dolby 5.1, Dolby 7.1, Hamasaki 22.2, etc. In some such examples, at least some of the speakers in the environment may be arranged in convenient positions (e.g., positions where there is space to accommodate the speakers) with respect to the furniture, walls, etc. of the environment, but can be in positions that are not in any standard-defined speaker layout.
[0084] Therefore, some implementation blocks 210 or block 220 may be involved in flexible rendering to arbitrarily positioned speakers. Some such implementations may be involved in center of mass amplitude pan (CMAP), flexible virtualization (FV), or a combination of both. At a high level, these techniques render a set of one or more audio signals, each having a desired perceived spatial position associated therewith, for playback through a set of two or more speakers. Here, the relative activation of the set of speakers is a function of a model of the perceived spatial position of the audio signal being played back through those speakers and the proximity of the desired perceived spatial position of the audio signal to the position of the speakers. The model ensures that the audio signal is audible to the listener near its intended spatial position, and the proximity term controls which speakers are used to achieve this spatial impression. In particular, the proximity term favors the activation of speakers that are close to the desired perceived spatial position of the audio signal. For both CMAP and FV, this functional relationship is conveniently derived from a cost function written as the sum of two terms, one for the spatial aspect and one for the proximity:
Number
[0085] Here, the set
Number
number
[0086] In some definitions of the cost function, g opt Although the relative levels between the components of g are appropriate, it is difficult to control the absolute levels of optimal activations that result from the above minimization. To address this issue, g is used to control the absolute levels of activations. opt Subsequent normalization of the vectors may be performed. For example, normalizing vectors to have unit length may be desirable, which aligns with the commonly used constant power pan rule:
number
[0087] The exact behavior of the flexible rendering algorithm is determined by the two terms in the cost function C spatial and C proximity For CMAP, C spatial is the center of mass of the perceived spatial location of an audio signal played from a set of loudspeakers weighted by the locations of those loudspeakers and their associated activation gains (elements of vector g):
Number
[0088] Next, Equation 3 is manipulated to be the spatial cost representing the square of the error between the desired audio position and the audio position generated by the active loudspeaker:
Number
[0089] In FV, the spatial term of the cost function is defined in a different way. Here, the goal is to generate the binaural response b corresponding to the audio object position →o at the listener's left and right ears. Conceptually, b is a 2×1 vector of filters (one filter for each ear), but more conveniently, it is treated as a 2×1 vector of complex-valued numbers at a specific frequency. Continuing with this representation at a specific frequency, the desired binaural response can be obtained from the set of HRTF indices by the object position:
Number
[0090] At the same time, the 2×1 binaural response e generated by the loudspeaker at the listener's ear is modeled as the product of the 2×M acoustic transfer matrix H and the M×1 vector g of complex speaker activation values:
Number
[0091] The acoustic transfer matrix H is modeled based on the set of loudspeaker positions {→s i} with respect to the listener position. Finally, the spatial component of the cost function is defined as the square of the error between the desired binaural response (Equation 5) and the binaural response generated by the loudspeaker (Equation 6): [Number]
[0092] Conveniently, the spatial terms of the cost functions for CMAP and FV defined in Equations 4 and 7 can both be reconfigured as quadratic forms of matrices as a function of the speaker activation g: [Number] where A is an M×M square matrix, B is a 1×M vector, and C is a scalar. The matrix A has rank 2, and thus, for M > 2, there are an infinite number of speaker activations g for which the spatial error term is equal to zero. Introducing the second term C proximity of the cost function removes this indeterminacy and results in a specific solution that has perceptually beneficial properties compared to other possible solutions. For both CMAP and FV, C proximity is constructed such that the activation of speakers with a position →s i away from the desired audio signal position →o receives a greater penalty than the activation of speakers with a position closer to the desired position. This construction gives an optimal set of sparse speaker activations in which only the speakers close to the position of the desired audio signal are significantly activated, and in practice, results in a spatially more robust reproduction of the audio signal with respect to the movement of the listener around the set of speakers.
[0093] For this purpose, the second term C proximity of the cost function can be defined as a distance-weighted sum of the squares of the absolute values of the speaker activations. This can be concisely expressed in matrix form as follows: [Number] where [Number] is the diagonal matrix of the distance penalty between the desired audio position and each speaker:
Number
[0094] The distance penalty function can take many forms, but the following is a useful parameterization.
Number
Number
[0095] Combining the two terms of the cost function defined by Equations 8 and 9a gives the overall cost function.
Number
Number
Number
Number
[0096] In general, the optimal solution of Equation 11 may result in speaker activation with negative values. For the CMAP construction of the flexible renderer, such negative activation may not be desirable, and thus, Equation 11 may be minimized under the condition that all positive activations remain positive.
[0097] Figures 2C and 2D are diagrams showing an exemplary set of exemplary sets of speaker activation and object rendering positions. In these examples, the speaker activation and object rendering positions correspond to speaker positions of 4, 64, 165, -87, and -4 degrees. Figure 2C shows speaker activations 245a, 250a, 255a, 260a, and 265a that constitute the optimal solution for Equation 11 for these particular speaker positions. Figure 2D plots the individual speaker positions as squares 267, 270, 272, 274, and 275 corresponding to speaker activations 245a, 250a, 255a, 260a, and 265a, respectively. Figure 2D also shows the ideal object positions (in other words, the positions where the audio object should be rendered) for a number of possible object angles as dots 276a, and the corresponding actual rendering positions for those objects as dots 278a connected by a dotted line 279a to the ideal object positions.
[0098] One class of embodiments relates to a method of rendering audio for playback by at least one (e.g., all or some) of a plurality of coordinated (orchestrated) smart audio devices. For example, a set of smart audio devices in a user's home (system) may be orchestrated to handle various simultaneous usage scenarios. Such usage scenarios may include rendering of audio (according to an embodiment) for playback by all or some of the smart audio devices (i.e., by all or some of the speakers). Many interactions with the system are contemplated, which require dynamic modification to the rendering. Such modifications may, but need not, focus on spatial fidelity.
[0099] Some embodiments are methods for rendering audio for playback by at least one (e.g., all or some) of a set of smart audio devices (or for playback by at least one (e.g., all or some) of another set of speakers). The rendering may include minimization of a cost function, which includes at least one dynamic speaker activation term. Examples of such dynamic speaker activation terms include, but are not limited to: · Proximity of speakers to one or more listeners; · Proximity of speakers to gravitational or repulsive forces; · Audibility of speakers with respect to some location (e.g., listener location or living room); · Capabilities of the speakers (frequency response, distortion); · Synchronization of speakers with other speakers; · Wake word performance; and · Echo canceller performance.
[0100] The dynamic speaker activation term can enable at least one of various behaviors. Such behaviors include distorting the spatial presentation of audio away from a particular smart audio device to enable its microphone to better hear a speaker, or to enable a secondary audio stream to be better heard from the speaker of the smart audio device.
[0101] Some embodiments implement rendering for playback by speakers of a plurality of coordinated (orchestrated) smart audio devices. Other embodiments implement rendering for playback by another set of speakers (singular or plural).
[0102] Pairing a flexible rendering method (implemented according to some embodiments) with a collection of wireless smart speakers (or other smart audio devices) can provide a highly capable and user-friendly spatial audio rendering system. Considering the interaction with such a system, it becomes apparent that dynamic modifications to spatial rendering may be desirable for optimization for other purposes that may arise during use of the system. To achieve this goal, one class of embodiments augments an existing flexible rendering algorithm with one or more additional dynamically configurable functions that depend on one or more attributes of the audio signal being rendered, a set of speakers, and / or other external inputs. According to some embodiments, the existing flexible rendering cost function given by Equation 1 is augmented using one or more of these additional dependencies as follows.
Number
[0103] In Equation 12, the term
Number
Number
Number
Number
Number
Number
Number
Number
[0104]
Number
[0105]
Number
[0106]
Number
[0107] Using the new cost function defined in Equation 12, as described above in Equations 2a and 2b, the optimal set of activations can be found through minimization with respect to g and possible post-normalization.
[0108] FIG. 2E is a flowchart outlining an example of a method that can be implemented by an apparatus or system as shown in FIG. 1A. The blocks of method 280 are not necessarily implemented in the order shown, similar to other methods described herein. Further, such a method may include more or fewer blocks than those illustrated and / or described. The blocks of method 280 may be executed by one or more devices that may be (or include) a control system such as control system 110 shown in FIG. 1A.
[0109] In this implementation, block 285 involves receiving audio data by a control system via an interface system. In this example, the audio data includes one or more audio signals and associated spatial data. According to this implementation, the spatial data indicates the intended perceived spatial position corresponding to the audio signal. In some cases, the intended perceived spatial position may be explicit, for example, indicated by position metadata such as Dolby Atmos position metadata. In other cases, the intended perceived spatial position may be implicit, for example, the intended perceived spatial position may be an assumed position associated with a channel according to a Dolby 5.1, Dolby 7.1, or other channel-based audio format. In some examples, block 285 involves the rendering module of the control system receiving audio data via the interface system.
[0110] According to this example, block 290 is involved in rendering audio data by a control system for playback via a set of loudspeakers in an environment, generating a rendered audio signal. In this example, rendering each of one or more audio signals included in the audio data involves determining the relative activation of a set of loudspeakers in the environment by optimizing a cost function. According to this example, the cost is a function of a model of the perceived spatial position of the audio signal when played back by a set of loudspeakers in the environment. In this example, the cost is also a function of an indicator of the proximity of the intended perceived spatial position of the audio signal to the position of each loudspeaker of the set of loudspeakers. In this implementation, the cost is also a function of one or more additional dynamically configurable functions. In this example, the dynamically configurable functions are based on one or more of the following: proximity of the loudspeakers to one or more listeners; proximity of the loudspeakers to an attraction position, where the attraction is a factor that favors relatively higher activation of the loudspeakers closer to the attraction position; proximity of the loudspeakers to a repulsion position, where the repulsion is a factor that favors relatively lower activation of the loudspeakers closer to the repulsion position; the ability of each loudspeaker compared to other loudspeakers in the environment; the synchronization of the loudspeakers with respect to other loudspeakers; wake word performance; or echo canceller performance.
[0111] In this example, block 295 is involved in providing the rendered audio signal to at least some of the set of loudspeakers in the environment via an interface system.
[0112] According to some examples, a model of a perceived spatial location can generate a binaural response corresponding to an audio object location at a listener's left and right ears. Alternatively or additionally, a model of a perceived spatial location can place the perceived spatial location of an audio signal reproduced from a set of loudspeakers at the centroid of a mass that weights the position of the set of loudspeakers by associated activation gains of the loudspeakers.
[0113] In some examples, the one or more additional dynamically configurable functions can be based, at least in part, on the level of the one or more audio signals. In some cases, the one or more additional dynamically configurable functions can be based, at least in part, on the spectrum of the one or more audio signals.
[0114] Some examples of method 280 involve receiving speaker layout information. In some examples, the one or more additional dynamically configurable functions can be based, at least in part, on the position of each loudspeaker in the environment.
[0115] Some examples of method 280 involve receiving loudspeaker specification information. In some examples, the one or more additional dynamically configurable functions can be based, at least in part, on the capabilities of each loudspeaker, which can include one or more of frequency response, reproduction level limits, or parameters of one or more loudspeaker dynamics processing algorithms.
[0116] According to some examples, the one or more additional dynamically configurable functions can be based at least in part on measurements or estimations of acoustic transmission from each loudspeaker to other loudspeakers. Alternatively or additionally, the one or more additional dynamically configurable functions can be based at least in part on the position of one or more human listeners or speakers in the environment. Alternatively or additionally, the one or more additional dynamically configurable functions can be based at least in part on measurements or estimations of acoustic transmission from each loudspeaker to the listener or speaker position. The estimated value of the acoustic transmission can be based, for example, at least in part on walls, furniture, or other objects that may be present between each loudspeaker and the listener or speaker position.
[0117] Alternatively or additionally, the one or more additional dynamically configurable functions can be based at least in part on the position of one or more non-loudspeaker objects or landmark objects in the environment. In some such implementations, the one or more additional dynamically configurable functions can be based at least in part on measurements or estimations of acoustic transmission from each loudspeaker to the object position or landmark position.
[0118] By employing one or more appropriately defined additional cost terms to achieve soft rendering, many new useful behaviors can be achieved. All of the exemplary behaviors listed below are created in the form of penalizing certain loudspeakers under certain conditions considered undesirable. The end result is that these loudspeakers are less activated in the spatial rendering of the set of audio signals. In many of these cases, regardless of the modification of the spatial rendering, one may consider simply reducing the undesirable loudspeakers, but such a strategy may significantly degrade the overall balance of the audio content. Certain components of the mix may, for example, become completely inaudible. On the other hand, in the disclosed embodiments, by integrating these penalty impositions into the core optimization of the rendering, the rendering can adapt and perform the best possible spatial rendering using the remaining speakers with lower penalties. This is a much more elegant, adaptable, and effective solution.
[0119] Exemplary use cases include, but are not limited to, the following.
[0120] ● Provide a more balanced spatial presentation around the listening area ○ It has been found that spatial audio is best presented through loudspeakers that are approximately the same distance from the intended listening area. The cost may be structured such that loudspeakers that are significantly closer to or farther from the average distance of the loudspeakers to the listening area are penalized, thereby reducing their activation.
[0121] ● Move the audio away from or towards the listener or speaker ○ When a user of the system attempts to address the system or an associated smart voice assistant, it is beneficial to impose a cost penalty on the louder speaker closer to the speaker. In this way, these louder speakers are less activated, allowing the associated microphones to better hear the speaker. ○ To provide a more intimate experience for a single listener while minimizing the playback level for other listeners in the listening space, speakers far from the listener's position may receive a large penalty. As a result, only the speaker closest to the listener is most prominently activated.
[0122] ● Move audio away from or closer to landmarks, zones, or areas ○ Certain locations in the vicinity of the listening space, such as a baby room, baby bed, office, reading area, study area, etc., may be considered sensitive. In such cases, a cost penalty may be constructed for the use of speakers close to this location, zone, or area. ○ Alternatively, for the same (or similar) cases as above, the speaker system may generate measurements of acoustic transmission from each speaker to the baby room, especially when one of the speakers (equipped with an attached or associated microphone) is present within the baby room itself. In this case, instead of using the physical proximity of the speaker to the baby room, a cost penalty may be constructed for using speakers with a high measured acoustic transmission to the baby room. And / or
[0123] ● Optimal use of speaker capabilities ○ The capabilities of different loudspeakers can vary significantly. For example, a popular smart speaker includes only a single 1.6-inch full-range driver with limited low-frequency capabilities. On the other hand, another smart speaker includes a much more capable 3-inch woofer. These capabilities are generally reflected in the frequency response of the speaker, and thus the set of responses associated with the speaker can be utilized in the cost term. At a particular frequency, a speaker that is less capable compared to other speakers as measured by the frequency response incurs a penalty and is thus activated to a lesser extent. In some implementations, such frequency response values may be stored in the smart loudspeaker and then reported to a computing unit responsible for optimizing the flexible rendering.
[0124] ○ Many speakers include multiple drivers, each responsible for reproducing a different frequency range. For example, a popular smart speaker is a two-way design that includes a woofer for low frequencies and a tweeter for high frequencies. Typically, such a speaker includes a crossover circuit to split the full-range playback audio signal into appropriate frequency ranges and send them to each driver. Alternatively, such a speaker can provide flexible renderer playback access to each individual driver and also provide information about the capabilities of each individual driver, such as the frequency response. By applying a cost term as described above, in some examples, the flexible renderer can automatically construct a crossover between the two drivers based on their relative capabilities at different frequencies.
[0125] ○The above usage example of frequency response focuses on the inherent capabilities of the speaker, but may not accurately reflect the capabilities of the speaker in the listening environment. In certain cases, the frequency response of the speaker measured at the intended listening position may be available through some calibration procedure. Such measurements may be used in place of the pre-calculated response to better optimize the use of the speaker. For example, certain speakers may be inherently very capable at specific frequencies, but due to their placement (e.g., behind a wall or furniture), may produce a very limited response at the intended listening position. Capturing this response and using the measurement as an input to an appropriate cost term can prevent significant activation of such speakers.
[0126] ○Frequency response is only one aspect of the playback capabilities of a loudspeaker. Many small loudspeakers begin to distort as the playback level increases and then reach their excursion limit, especially at low frequencies. To reduce such distortion, many loudspeakers implement dynamics processing that constrains the playback level to be below some variable limit thresholds across frequencies. It makes sense to reduce the signal level of the limiting speaker and direct this energy to other, less burdened speakers if one speaker is approaching or at these thresholds and other speakers participating in the flexible rendering are not. Such behavior can be automatically achieved according to some embodiments by appropriately configuring the relevant cost terms. Such cost terms may relate to one or more of the following: · Monitoring of the global playback volume related to the limiting threshold of the loudspeaker. For example, a loudspeaker whose volume level is closer to its limiting threshold may be subject to a greater penalty; ·Monitoring, possibly of a dynamic signal level that varies over frequency, and also possibly in relation to a loudspeaker's limit threshold that varies over frequency. For example, a loudspeaker with a monitored signal level closer to its limit threshold may be subject to a greater penalty; ·Direct monitoring of parameters of loudspeaker dynamics processing, such as limiting gain. In some such examples, a loudspeaker with a parameter indicating stronger limiting may be subject to a greater penalty; and / or, ·Monitoring of the actual instantaneous voltage, current, and power delivered to a loudspeaker by an amplifier to determine whether the loudspeaker is operating within its linear range. For example, a loudspeaker operating with lower linearity may be subject to a greater penalty.
[0127] ○Smart speakers having an integrated microphone and an interactive voice assistant typically use some type of echo cancellation to reduce the level of the audio signal reproduced from the speaker and picked up by the recording microphone. The greater this reduction, the more likely the speaker is to hear and understand the speaker within the space. If the echo canceller's residual is consistently high, this may be an indication that the speaker is being driven into a non-linear region where prediction of the echo path becomes difficult. In such cases, it would be reasonable to divert signal energy away from that speaker, and thus a cost term taking into account echo canceller performance may be beneficial. Such a cost term may assign a high cost to a speaker whose accompanying echo canceller is performing poorly.
[0128] ○ When rendering spatial audio with multiple loudspeakers, in order to achieve predictable imaging, generally, the playback by a set of loudspeakers needs to be reasonably synchronized over time. In the case of wired loudspeakers, this is natural. However, when there are many wireless loudspeakers, synchronization is difficult and the final result may be variable. In such cases, each loudspeaker may be able to report the relative degree of synchronization with the target, and this degree may be input into the synchronization cost term. In some such examples, loudspeakers with a lower degree of synchronization may be subject to a greater penalty and thus may be excluded from the rendering. Further, for certain types of audio signals, such as components of an audio mix that are intended to be diffuse or non - directional, strict synchronization may not be required. In some implementations, components may be tagged as such using metadata, and the synchronization cost term may be modified so that the penalty is reduced.
[0129] Next, examples of embodiments will be described.
[0130] Similar to the proximity cost defined by Equations 9a and 9b, the terms of the new cost function
Number
Number
Number
Number
[0131] Combining equations 13a and 13b with the matrix quadratic form versions of the CMAP and FV cost functions given in equation 10 results in a potentially useful implementation of the generally extended cost function (of some embodiments) given in equation 12:
Number
[0132] In this definition of the new cost function term, the overall cost function remains in matrix quadratic form, and the optimal set of activations g opt can be found through the differentiation of equation 14 and is as follows.
Number
[0133] Each of the weight terms w ij is usefully considered as a function of a given continuous penalty value
Number
Number
Number
[0134] When all loudspeakers are penalized, it is often convenient in post-processing to subtract the minimum penalty from all weight terms so that at least one of the speakers is not penalized: [Number]
[0135] As described above, there are many possible use cases that can be realized using the new cost function terms described herein (and similar new cost function terms used according to other embodiments). Next, three examples will be used to explain more specific details. That is, moving the audio towards the listener or speaker, moving the audio away from the listener or speaker, and moving the audio away from the landmark.
[0136] In the first example, what is here called "gravity" is used to pull the audio towards a location. That location may, in some examples, be the location of the listener or speaker, a landmark location, a furniture location, etc. In this specification, this location may be referred to as the "gravity location" or the "attractor location". As used in this specification, "gravity" is a factor that favors relatively higher loudspeaker activation in the vicinity closer to the gravity location. According to this example, the weight w ij takes the form of Equation 17, and the continuous penalty value p ij is the distance from the fixed attractor location of the i-th speaker
Number
Number
[0137] To illustrate the use case of "pulling" the audio towards the listener or speaker, specifically, α j = 20, β j = 3 are set,
Number
Number
[0138] In the second and third examples, "repulsive force" is used to "push" the audio away from a position that may be the listener's position, the speaker's position, or other positions such as the position of a landmark, the position of furniture, etc. In some examples, the repulsive force may be used to push the audio away from an area or zone of the auditory environment such as an office area, a reading area, a bed or bedroom area (e.g., a baby bed or a bedroom). According to some such examples, a specific position may be used to represent the zone or area. For example, the position representing the baby's bed may be the estimated position of the baby's head, the estimated sound source position corresponding to the baby, etc. This position may be referred to herein as the "repulsive force position" or "repulsive position". As used herein, "repulsive force" is a factor that promotes relatively lower speaker activation the closer it is to the repulsive force position. According to this example, for a fixed repulsive position
Number
Number
[0139] To illustrate a use case of moving audio away from the listener or speaker, specifically set α j = 5, β j = 2, and
Number
Number
[0140] A third exemplary use case is to “push” audio away from an acoustically sensitive landmark, such as a door to a baby's room while sleeping. Similar to the previous example, →l j is set to a vector corresponding to the 180-degree door position (lower center of the plot). To achieve a stronger repulsive force and completely bias the sound field to the front part of the main listening space, we set α j = 20, β j = 5. FIG. 2J is a graph of speaker activation in an exemplary embodiment. Again, in this example, FIG. 2J shows speaker activations 245d, 250d, 255d, 260d, and 265d that constitute the optimal solution to the same set of speaker positions, with a stronger repulsive force applied. FIG. 2K is a graph of object rendering positions in an exemplary embodiment. Again, in this example, FIG. 2K shows the ideal object position 276d for a number of possible object angles and the corresponding actual rendering positions 278d for those objects connected to the ideal object position 276d by the dotted line 279d. The skewed orientation of the actual rendering positions 278d shows the influence of the stronger repulsive weighting on the optimal solution to the cost function.
[0141] Returning now to FIG. 2B, in this example, block 225 is involved in modifying the rendering process for the first audio signal, at least in part, based on at least one of the second audio signal, the second rendered audio signal, or its characteristics, to generate a modified first rendered audio signal. Various examples of modifying the rendering process are disclosed herein. The "characteristics" of the rendered signal can include, for example, the estimated or measured loudness or audibility at the intended listening position in a silent environment or in the presence of one or more additional rendered signals. Other examples of characteristics include parameters related to the rendering of the signal, such as the intended spatial position of the constituent signals of the associated program stream, the position of the loudspeaker at which the signal is rendered, the relative activation of the loudspeaker as a function of the intended spatial position of the constituent signals, and any other parameters or states related to the rendering algorithm utilized to generate the rendered signal. In some examples, block 225 may be executed by the first rendering module.
[0142] According to this example, block 230 is involved in modifying the rendering process for the second audio signal, at least in part, based on at least one of the first audio signal, the first rendered audio signal, or its characteristics, to generate a modified second rendered audio signal. In some examples, block 230 may be executed by the second rendering module.
[0143] In some implementations, modifying the rendering process for the first audio signal may include warping the rendering of the first audio signal away from the rendering position of the second rendered audio signal and / or modifying one or more loudnesses of the first rendered audio signal in response to one or more loudnesses of the second audio signal or the second rendered audio signal. Alternatively or additionally, modifying the rendering process for the second audio signal may include warping the rendering of the second audio signal away from the rendering position of the first rendered audio signal and / or modifying one or more loudnesses of the second rendered audio signal in response to one or more loudnesses of the first audio signal or the first rendered audio signal. Some examples are provided below with reference to Figure 3 and below.
[0144] However, other types of rendering process modifications are within the scope of the present disclosure. For example, in some instances, modifying the rendering process for the first audio signal or the second audio signal may involve performing spectral modifications, audibility-based modifications, or dynamic range modifications. These modifications may or may not be related to loudness-based rendering modifications, depending on the specific example. For example, in the above-described case where the primary spatial stream is rendered in a living area of an open-plan and the secondary stream consisting of cooking hints is rendered in an adjacent kitchen, it may be desirable to ensure that the cooking hints remain audible in the kitchen. This can be achieved by estimating what the loudness of the stream of cooking hints rendered in the kitchen would be in the absence of an interfering first signal, then estimating the loudness in the presence of the first signal in the kitchen, and finally dynamically modifying the loudness and dynamic range of both streams over a plurality of frequencies to ensure the audibility of the second signal in the kitchen.
[0145] In the example shown in FIG. 2B, block 235 is involved at least in mixing the modified first rendered audio signal and the modified second rendered audio signal to produce a mixed audio signal. Block 235 may be performed, for example, by mixer 130b shown in FIG. 2A.
[0146] According to this example, block 240 is involved in providing the mixed audio signal to at least some of the speakers of the environment. Some examples of method 200 are involved in the reproduction of the mixed audio signal by the speakers.
[0147] As shown in FIG. 2B, some implementations may provide more than two rendering modules. Some such implementations may provide N rendering modules, where N is an integer greater than 2. Thus, some such implementations may include one or more additional rendering modules. In some such examples, each of the one or more additional rendering modules may be configured to receive an additional audio program stream via an interface system. The additional audio program stream may include an additional audio signal scheduled to be played by at least one speaker of the environment. Some such implementations may involve rendering an additional audio signal for playback via at least one speaker of the environment to generate an additional rendered audio signal and modifying the rendering process for the additional audio signal, at least in part, based on at least one of the first audio signal, the first rendered audio signal, the second audio signal, the second rendered audio signal, or their characteristics, to generate a modified additional rendered audio signal. According to some such examples, a mixing module may be further configured to mix the modified additional rendered audio signal with at least the modified first rendered audio signal and the modified second rendered audio signal to generate the mixed audio signal.
[0148] As described above with reference to FIGS. 1A and 2A, some implementations may include a microphone system that includes one or more microphones in an eavesdropping environment. In some such examples, the first rendering module may be configured to modify a rendering process for a first audio signal based at least in part on a first microphone signal from the microphone system. The "first microphone signal" may be received from a single microphone or from two or more microphones, depending on the specific implementation. In some such implementations, the second rendering module may be configured to modify a rendering process for a second audio signal based at least in part on the first microphone signal.
[0149] As described above with reference to FIG. 2A, in some examples, the position of one or more microphones may be known and provided to a control system. According to some such implementations, the control system may be further configured to estimate a first sound source position based on the first microphone signal and modify a rendering process for at least one of the first audio signal or the second audio signal based at least in part on the first sound source position. The first sound source position may be estimated according to a triangulation process, for example, based on DOA data from each of three or more microphones or groups of microphones having known positions. Alternatively or additionally, the first sound source position may be estimated according to the amplitude of signals received from two or more microphones. The microphone that generates the signal with the highest amplitude may be assumed to be closest to the first sound source position. In some such examples, the first sound source position may be set to the position of the closest microphone. In some such examples, the first sound source position may be associated with the position of a zone, where the zone is selected by processing signals from two or more microphones through a pre-trained classifier such as a Gaussian mixture model.
[0150] In some such implementations, the control system may be configured to determine whether the first microphone signal corresponds to ambient noise. Some such implementations may involve modifying the rendering process for at least one of the first audio signal or the second audio signal, at least in part, based on whether the first microphone signal corresponds to ambient noise. For example, if the control system determines that the first microphone signal corresponds to ambient noise, modifying the rendering process for the first audio signal or the second audio signal may involve increasing the level of the rendered audio signal so that the perceived loudness of the signal in the presence of noise at the intended listening position is substantially equal to the perceived loudness of the signal in the absence of noise.
[0151] In some examples, the control system may be configured to determine whether the first microphone signal corresponds to a human voice. Some such implementations may involve modifying the rendering process for at least one of the first audio signal or the second audio signal, at least partially based on whether the first microphone signal corresponds to a human voice. For example, if the control system determines that the first microphone signal corresponds to a human sound such as a wake word, modifying the rendering process for the first audio signal or the second audio signal may involve reducing the loudness of the rendered audio signal played by a speaker closer to the first sound source, as compared to the loudness of the rendered audio signal played by a speaker farther from the first sound source. Modifying the rendering process for the first audio signal or the second audio signal may alternatively involve warping the intended position of the constituent signals of the associated program stream away from the first sound source and / or modifying the rendering process to penalize the use of speakers closer to the first sound source as compared to more speakers away from the first sound source.
[0152] In some implementations, when the control system determines that the first microphone signal corresponds to a human voice, the control system may be configured to reproduce the first microphone signal at one or more speakers close to a location in the environment that is different from the first sound source position. In some such examples, the control system may be configured to determine whether the first microphone signal corresponds to a baby's cry. According to some such implementations, the control system may be configured to reproduce the first microphone signal at one or more speakers close to a location in the environment corresponding to the estimated position of a caregiver such as a parent, relative, guardian, childcare service provider, teacher, nurse, etc. In some examples, the process of estimating the estimated position of the caregiver may be triggered by a voice command such as "<wake word>, don't wake the baby". The control system can estimate the position of the speaker (caregiver) according to the position of the nearest smart audio device implementing the virtual assistant by triangulation based on DOA information provided by three or more local microphones, etc. According to some implementations, the control system has prior knowledge about the position of the baby room (and / or the listening device therein), and thus can perform appropriate processing.
[0153] In some such examples, the control system may be configured to determine whether the first microphone signal corresponds to a command. When the control system determines that the first microphone signal corresponds to a command, in some cases, the control system may be configured to determine a response to the command and control at least one speaker close to the first sound source position to reproduce the response. In some such examples, the control system may be configured to return to the unmodified rendering process for the first audio signal or the second audio signal after controlling at least one speaker close to the first sound source position to reproduce the response.
[0154] In some implementations, the control system may be configured to execute commands. For example, the control system may be, or may include, a virtual assistant configured to control an audio device, a television, home appliances, etc. according to commands.
[0155] This definition of the minimal and more capable multi-stream rendering systems shown in FIGS. 1A, 1B, and 2A enables dynamic management of the simultaneous playback of multiple program streams for a number of useful scenarios. Here, with reference to FIGS. 3A and 3B, some examples are described.
[0156] First, consider the previously discussed example related to the simultaneous playback of a spatial movie soundtrack in the living room and cooking tips in the connected kitchen. The spatial movie soundtrack is an example of the "first audio program stream" described above, and the audio of the cooking tips is an example of the "second audio program stream" described above. FIGS. 3A and 3B show an example of a floor plan of a connected living space. In this example, the living space 300 includes a living room in the upper left, a kitchen in the lower center, and a bedroom in the lower right. The squares and circles 305a - 305h distributed across the living space represent a set of eight loudspeakers arranged in convenient positions in the space but not conforming to a standard-defined layout (arbitrarily arranged). In FIG. 3A, only the spatial movie soundtrack is being played, and all the loudspeakers in the living room 310 and the kitchen 315 are used to generate an optimized spatial playback around the listener 320a sitting on the couch 325 facing the television 330, taking into account the capabilities and layout of the loudspeakers. The optimal playback of this movie soundtrack is visually represented by a cloud 335a within the range of the active loudspeakers.
[0157] In FIG. 3B, cooking hints are simultaneously rendered and played back for a second listener 320b through a single loudspeaker 305g within the kitchen 315. The playback of this second program stream is visually represented by a cloud 340 emerging from the loudspeaker 305g. If these cooking hints were played back simultaneously without modification to the rendering of the movie soundtrack as shown in FIG. 3A, the audio from the movie soundtrack emitted from speakers within or near the kitchen 315 would interfere with the second listener's ability to understand the cooking hints. Instead, in this example, the rendering of the spatial movie soundtrack is dynamically modified as a function of the rendering of the cooking hints. Specifically, the rendering of the movie soundtrack is shifted away from speakers near the rendering location of the cooking hints (the kitchen 315), and this shift is visually represented by a smaller cloud 335b in FIG. 3B that is pushed away from the speakers near the kitchen. If the playback of the cooking hints stops while the movie soundtrack is still playing, in some implementations, the rendering of the movie soundtrack can be shifted back dynamically to its original optimal configuration as seen in FIG. 3A. Such dynamic shifts in the rendering of the spatial movie soundtrack can be achieved through a number of disclosed methods. Many spatial audio mixes include multiple component audio signals designed to be played back at specific locations within the listening space. For example, Dolby 5.1 and 7.1 surround sound mixes consist of 6 signals and 8 signals respectively, which are intended to be played back on speakers at defined canonical positions around the listener. Object-based audio formats such as Dolby Atmos consist of component audio signals and associated metadata that describe the potentially time-varying 3D positions within the listening space where the audio is intended to be rendered.Assume that a spatial movie sound track renderer can render individual audio signals at any position with respect to any set of loudspeakers. Then, the dynamic shifts to the rendering shown in FIGS. 3A and 3B can be achieved by warping the intended positions of the audio signals within the spatial mix. For example, the 2D or 3D coordinates associated with the audio signal are pushed away from the position of the speaker in the kitchen or pulled towards the upper left corner of the living room. The result of such warping is that the speakers near the kitchen are used less. This is because the warped positions of the audio signals in the spatial mix are now further away from this position. This method achieves the goal of making the second audio stream more understandable for the second listener, but it comes at the cost of significantly changing the intended spatial balance of the movie sound track for the first listener.
[0158] A second way to achieve a dynamic shift to spatial rendering can be realized by using a flexible rendering system. In some such implementations, the flexible rendering system may be a CMAP, FV, or a hybrid of both, as described above. Some such flexible rendering systems attempt to reproduce the spatial mix so that all component signals are perceived as coming from their intended locations. While doing so for each signal in the mix, in some examples, activation of loudspeakers proximate to the desired location of that signal is prioritized. In some implementations, additional terms may be dynamically added to the rendering optimization, which penalize the use of certain loudspeakers based on other criteria. In the present example, something called a "repulsive force" is dynamically placed at the location of the kitchen, imposing a high penalty on the use of loudspeakers near this location and effectively deterring the rendering of a spatial movie soundtrack. As used herein, the term "repulsive force" may refer to a factor corresponding to relatively low speaker activation at a particular location or area of the listening environment. In other words, the phrase "repulsive force" may refer to a factor that favors activation of speakers that are relatively farther away from a particular location or area corresponding to the "repulsive force", but according to some such implementations, the renderer may still attempt to reproduce the intended spatial balance of the mix using the remaining, lower penalty speakers. Thus, this technique may be considered a superior way to achieve a dynamic shift in rendering compared to simply warping the intended locations of the component signals of the mix.
[0159] The above-described scenario of shifting the rendering of a spatial movie soundtrack away from cues for cooking in the kitchen can be achieved using the minimal version of the multi-stream renderer shown in Figure 1B. However, by adopting the more capable system shown in Figure 2A, improvements to the scenario can be realized. Shifting the rendering of the spatial movie soundtrack improves the intelligibility of the cues for cooking in the kitchen, but the movie soundtrack may still be prominently audible in the kitchen. Depending on the instantaneous state of both streams, the cooking cues may be masked by the movie soundtrack. For example, a loud instant in the movie soundtrack masks a soft instant in the cooking cues. To address this problem, a dynamic modification to the rendering of the cooking cues may be added as a function of the rendering of the spatial movie soundtrack. For example, a method of dynamically changing the audio signal over frequency and time may be implemented to maintain its perceived loudness in the presence of interfering signals. In this scenario, an estimated value of the perceived loudness of the shifted movie soundtrack at the kitchen location may be generated and fed into such a process as the interfering signal. The time-varying and frequency-varying levels of the cooking cues may then be dynamically modified to maintain their perceived loudness above this interference. Thereby, better intelligibility for a second listener can be maintained. The required estimated loudness value of the movie soundtrack in the kitchen can be generated from the speaker feed of the soundtrack rendering, signals from microphones in or near the kitchen, or a combination thereof. The process of maintaining the perceived loudness of the cooking cues generally increases the level of the cooking cues and, in some cases, the overall loudness may become uncomfortably high. To address this problem, yet another rendering modification may be adopted. The interfering spatial movie soundtrack may be dynamically reduced in response to the volume of the loudness-modified cooking cues in the kitchen becoming too high.Finally, there may be some external noise source that simultaneously interferes with the audibility of both program streams. For example, a blender [food mixer] may be used in the kitchen during cooking. The estimated value of the loudness of this environmental noise source in both the living room and the kitchen may be generated from a microphone connected to the rendering system. This estimated value may be added, for example, to the estimated value of the loudness of the sound track in the kitchen, which affects the loudness correction of cooking hints. At the same time, the rendering of the sound track in the living room may be additionally modified as a function of the environmental noise estimated value in order to maintain the perceived loudness of the sound track in the living room in the presence of this environmental noise, thereby better maintaining the audibility for the listener in the living room.
[0160] As can be seen, this exemplary use case of the disclosed multi-stream renderer employs a number of interconnected modifications to the two program streams in order to optimize the simultaneous playback of the two program streams. In summary, these modifications to the streams can be enumerated as follows: ● Spatial movie sound track ○ The spatial rendering is shifted away from the kitchen as a function of the cooking hint rendered in the kitchen ○ Dynamic reduction of loudness as a function of the loudness of the cooking hint rendered in the kitchen ○ Dynamically increasing the loudness as a function of the estimated value of the loudness of the interfering blender noise in the living room from the kitchen ● Cooking hint ○ Dynamically increasing the loudness as a function of the combined estimated value of the loudness of both the movie sound track and the blender noise in the kitchen.
[0161] A second exemplary use case of the disclosed multi-stream renderer involves the simultaneous playback of a spatial program stream, such as music, along with the response of a smart voice assistant to some query by a user. In existing smart speakers, playback is generally restricted to monaural or stereo playback through a single device, and interaction with the voice assistant typically consists of the following stages: 1) Music playback 2) The user utters the wake word of the voice assistant 3) The smart speaker recognizes the wake word and turns down (ducks) the music by a significant amount 4) The user issues a command to the smart assistant (i.e., "play the next song"). 5) The smart speaker recognizes the command, confirms this by mixing any voice response (i.e., "Understood, playing the next song") into the ducked music and playing it through the speaker, and then executes the command. 6) The smart speaker returns the music to its original loudness.
[0162] Figures 4A and 4B show an example of a multi-stream renderer that provides simultaneous playback of a spatial music mix and a voice assistant response. When playing spatial audio on a number of orchestrated smart speakers, some embodiments provide improvements to the above series of events. Specifically, the spatial mix may be shifted away from one or more of the speakers appropriately selected to relay the response from the voice assistant. Creating this space for the voice assistant response means that the spatial mix may not have to be reduced as much, or not at all, compared to the state of the art described above. Figures 4A and 4B illustrate this scenario. In this example, the modified series of events occurs as follows: 1) A spatial music program stream is played through a number of orchestrated smart speakers for the user cloud 335c of FIG. 4A. 2) User 320c utters the wake word of the voice assistant. 3) One or more smart speakers (e.g., speaker 305d and / or speaker 305f) use the associated recording from the microphone attached to the one or more smart speakers to recognize the wake word and determine the position of user 320c or which speaker(s) user 320c is closest to. 4) The rendering of the spatial music mix is shifted away from the position determined in the previous step, anticipating that the voice assistant response program stream will be rendered near that position (cloud 335d in FIG. 4B). 5) The user utters a command to the smart assistant (e.g., to the smart speaker running the smart assistant / virtual assistant software). 6) The smart speaker recognizes the command, synthesizes the corresponding response program stream, and renders the response near the user's position (cloud 440 in FIG. 4B). 7) The rendering of the spatial music program stream shifts back to its original state when the voice assistant response is complete (cloud 335c in FIG. 4A).
[0163] In addition to optimizing the simultaneous playback of the spatial music mix and the voice assistant response, the shift of the spatial music mix can also improve the ability of the set of speakers to understand the listener in step 5. This is because the music is shifted away from the speakers near the listener, thereby improving the ratio of the voice of the associated microphone to the others.
[0164] [[ID=...]] Similar to what was described for the previous scenario using spatial movie mixing and cooking hints, the current scenario can be further optimized beyond what is provided by shifting the rendering of the spatial mix as a function of the voice assistant response. By itself, shifting the spatial mix may not be sufficient to make the voice assistant response completely understandable to the user. A simple solution, although not as required by current state-of-the-art, is to reduce the spatial mix by a certain amount. Alternatively, the loudness of the voice assistant response program stream can be dynamically increased as a function of the loudness of the spatial music mix program stream to maintain the audibility of the response. As an extension, the loudness of the spatial music mix may also be dynamically cut if this boost process for the response stream becomes too large.
[0165] Figures 5A, 5B, and 5C show a third example of the disclosed multi-stream renderer. This example involves managing the simultaneous playback of a spatial music mix program stream and a comfort noise program stream, while also attempting to ensure that the baby remains asleep in the adjacent room and yet can be heard if the baby cries. Figure 5A shows the starting point where a spatial music mix (represented by cloud 335e) is optimally played across all speakers in the living room 310 and kitchen 315 for a number of people at a party. In Figure 5B, the baby 510 is attempting to sleep in the adjacent bedroom 505 depicted in the lower right. To ensure this, while maintaining a reasonable experience for the party-goers, the spatial music mix is dynamically shifted away from the bedroom to minimize leakage into the bedroom, as shown by cloud 335f. At the same time, a second program stream (represented by cloud 540) containing pleasant white noise is played from speaker 305h in the baby room to mask any remaining leakage from the music in the adjacent room. To ensure complete masking, in some examples, the loudness of this white noise stream can be dynamically adjusted as a function of an estimated value of the loudness of the spatial music leaking into the baby room. This estimated value may be generated from the speaker feeds of the rendering of the spatial music, signals from microphones in the baby room, or a combination thereof. Also, if the loudness of the spatial music mix becomes too high, it may be dynamically attenuated as a function of the noise with adjusted loudness. This is similar to the loudness processing between the spatial movie mix and cooking hints in the first scenario. Finally, a microphone in the baby room (e.g., a microphone associated with speaker 305h, which may be a smart speaker in some implementations) may be configured to record audio from the baby (squelching sounds that can be picked up from the spatial music and white noise), and a combination of these processed microphone signals can then function as a third program stream.If it is detected that the baby is crying (such as through machine learning and via a pattern matching algorithm), the third program stream may be played simultaneously near the listener 320d, who may be a parent or other caregiver, in the living room 310. Figure 5C depicts the playback of this additional stream with the cloud 550. In this case, as shown by the modified shape of the cloud 335g relative to the shape of the cloud 335f in Figure 5B, the spatial music mix may be further shifted away from the speaker near the parent who is playing the baby's crying sound. The program stream of the baby's crying sound may be loudness-corrected as a function of the spatial music stream so that the baby's crying sound remains audible to the listener 320d. The interconnected modifications that optimize the simultaneous playback of the three program streams considered in this example can be summarized as follows:
[0166] ● Spatial music mix in the living room ○ The spatial rendering is shifted away from the baby room to reduce propagation into the baby room. ○ Dynamic reduction of loudness as a function of the loudness of the white noise rendered in the baby room. ○ The spatial rendering is shifted away from the parent in response to the baby's crying sound being rendered on the speaker near the parent. ● White noise ○ Dynamic increase of loudness as a function of the estimated value of the loudness of the music stream leaking into the baby room. ● Recording of the baby's crying sound ○ Dynamic increase of loudness as a function of the estimated value of the loudness of the music mix at the position of the parent or other caregiver.
[0167] Next, examples of how some of the above-described embodiments can be implemented will be described.
[0168] In FIG. 1B, each rendering block 1...N may be implemented as identical instances of any single stream renderer, such as the aforementioned CMAP, FV, or hybrid renderer. Configuring a multi-stream renderer in this way has several convenient and useful properties.
[0169] First, if rendering is performed in this hierarchical arrangement and each single stream renderer instance is configured to operate in the frequency / transform domain (e.g., QMF), stream mixing may also occur in the frequency / transform domain, and the inverse transform only needs to be performed once for M channels. This represents a significant efficiency improvement compared to performing N×M inverse transforms to mix in the time domain.
[0170] FIG. 6 shows an example of the frequency / transform domain of the multi-stream renderer shown in FIG. 1B. In this example, for each of the program streams 1...N, a quadrature mirror analysis filter bank (QMF) is applied before each program stream is received by the corresponding one of the rendering modules 1...N. According to this example, the rendering modules 1...N operate in the frequency domain. After mixer 630a mixes the outputs of the rendering modules 1...N, inverse synthesis filter bank 635a converts the mix to the time domain and provides the time domain mixed speaker feed signal to loudspeakers 1...M. In this example, the quadrature mirror filter bank, the rendering modules 1...N, the mixer 630a, and the inverse filter bank 635a are components of the control system 110c.
[0171] FIG. 7 shows an example of the frequency / transform domain of the multi - stream renderer shown in FIG. 2A. As in FIG. 6, for each of program streams 1 - N, a quadrature mirror analysis filter bank (QMF) is applied before each program stream is received by the corresponding one of rendering modules 1 - N. According to this example, rendering modules 1 - N operate in the frequency domain. In this implementation, the time - domain microphone signal from microphone system 120b is also provided to the quadrature mirror filter bank, and rendering modules 1 - N receive the microphone signal in the frequency domain. After mixer 630b mixes the outputs of rendering modules 1 - N, inverse filter bank 635b converts the mix to the time domain and provides the time - domain mixed speaker feed signal to loudspeakers 1 - M. In this example, the quadrature mirror filter bank, rendering modules 1 - N, mixer 630b, and inverse filter bank 635b are components of control system 110d.
[0172] Another advantage of the hierarchical approach in the frequency domain is the calculation of the perceived loudness of each audio stream and the use of this information when dynamically modifying one or more of the other audio streams. To illustrate this embodiment, consider the above - mentioned example described with reference to FIGS. 3A and 3B. In this case, there are two audio streams (N = 2), a spatial movie sound track, and cooking hints. Also, there may be ambient noise generated by a blender in the kitchen, picked up by one or more of the K microphones.
[0173] After each audio stream s is rendered individually, each microphone i is captured and converted to the frequency domain, the source excitation signal E s or E ican be calculated. This serves as an estimate of the time-varying perceived loudness of each audio stream s or microphone signal i. In this example, these source excitation signals are the conversion coefficients X s from the rendered stream or captured microphone, for the audio stream, and X i for the microphone signal, are calculated over time t for b frequency bands for c loudspeakers and smoothed using the frequency-dependent time constant λ b ;
Number
[0174] The raw source excitation is the estimated perceived loudness of each stream at a particular location. For the spatial stream, that location is at the center of the cloud 335b in Figure 3B, while for the cooking hint stream, it is at the center of the cloud 340. The location for the blender noise picked up by the microphone may be based, for example, on the particular location of the microphone(s) closest to the source of the blender noise.
[0175] The raw source excitation must be converted to the listening position of the audio stream(s) being modified by them in order to estimate how perceptible they are as noise at the listening position of each target audio stream. For example, if audio stream 1 is a movie soundtrack and audio stream 2 is a cooking hint,
Number
[0176] In Equation 13a, ^E xs represents the raw noise excitation calculated for the source audio stream without reference to the microphone input. In Equation 13b, ^E xi represents the raw noise excitation calculated with reference to the microphone input. According to this example, the raw noise excitation ^E xs or ^E xi is then summed over streams 1 - N, microphones 1 - K, and output channels 1 - M to obtain the total noise estimate for target stream x: [Number]
[0177] According to some alternative implementations, the term in Equation 14 [Number] can be omitted to obtain the total noise estimate without reference to the microphone input.
[0178] In this example, the overall raw noise estimate is smoothed to avoid perceptible artifacts that could be caused by modifying the target stream too abruptly. According to this implementation, the smoothing is based on the concept of using a fast attack and a slow release, similar to an audio compressor. The smoothed noise estimate for target stream x [Number] 〔Hereinafter, ~E x may be written before ~ as follows〕 In this example, it is calculated as follows:
Number
[0179] Once the complete noise estimate value ~E x (b, t) for the stream x is obtained, the previously calculated source excitation signal E x (b, t, c) can be reused to determine a set of time-varying gains Gx(b, t, c) for applying to the target audio stream x to ensure that the target audio stream x remains audible above the noise. These gains can be calculated using any of a variety of techniques.
[0180] In one embodiment, a loudness function L{·, ·} can be applied to the excitation to model various non-linearities in human loudness perception and calculate a specific loudness signal that describes the time-varying frequency distribution of the perceived loudness. Applying L{·, ·} to the noise estimate and the excitation for the rendered audio stream x yields an estimate of the specific loudness for each signal:
Number
[0181] In Equation 17a, L xn represents an estimate of the specific loudness of the noise, and in Equation 17b, L xrepresents an estimated value for the specific loudness of the rendered audio stream x. These specific loudness signals represent the perceived loudness when the signal is listened to in isolation. However, when two signals are mixed, masking may occur. For example, if the noise signal is much larger than the stream x signal, the noise signal will mask the stream x signal, thereby reducing the perceived loudness of that signal relative to the perceived loudness of that signal when listened to in isolation. This phenomenon can be modeled by a partial loudness function PL{·,·} that takes two inputs. The first input is the excitation of the signal of interest, and the second input is the excitation of the competing (noise) signal. This function returns a partial specific loudness signal PL that represents the perceived loudness of the signal of interest in the presence of the competing signal. Then, the partial specific loudness of the stream x signal in the presence of the noise signal can be calculated directly from the excitation signals over the frequency band b, time t, and loudspeaker c: [Number]
[0182] To maintain the audibility of the audio stream x signal in the presence of noise, a gain G x (b,t,c) can be calculated as shown in equations 8a and 8b to increase the loudness of the audio stream x until it is audibly louder than the noise. Alternatively, if the noise is from another audio stream s, two sets of gains can be calculated. In such an example, the first one G x (b,t,c) is applied to the audio stream x to increase its loudness, and the second one G s (b,t) is applied to the competing audio stream s to reduce its loudness. Thereby, the combination of those gains guarantees the audibility of the voice stream x. This is shown in equations 9a and 9b. In both sets of equations, ̄PL x(b, t, c) represents the partial - specific loudness of the source signal in the presence of noise after the application of the compensation gain.
Number
[0183] In practice, again to avoid audible artifacts, the raw gain is further smoothed over frequency using a smoothing function S{·} before being applied to the audio stream. ̄G x (b, t, c) and  ̄G s (b, t) represents the final compensation gain for the target audio stream x and the competing audio stream s.
Number
[0184] In one embodiment, these gains can be applied directly to all rendered output channels of the audio stream. In another embodiment, instead, they may be applied to the objects of the audio stream before being rendered. This uses, for example, the method described in U.S. Patent Application Publication No. 2019 / 0037333A1, which is incorporated herein by reference. These methods include calculating a pan coefficient for each audio object related to each of a plurality of predefined channel coverage zones based on the spatial metadata of the audio object. The audio signal may be converted into submixes with respect to the predefined channel coverage zones based on the calculated pan coefficients and the audio object. Each of the submixes may represent the sum of the components of the plurality of audio objects related to one of the predefined channel coverage zones. The submix gain may be generated by applying audio processing to each of the submixes, and the object gain applied to each audio object may be controlled. The object gain may be a function of the pan coefficient for each audio object and the submix gain related to each of the predefined channel coverage zones. Applying the gain to the object has several advantages, especially when combined with other processing of the stream.
[0185] Figure 8 shows an implementation of a multi - stream rendering system having an audio - stream loudness estimator. According to this example, the multi - stream rendering system of Figure 8 is configured to implement loudness processing, such as described by equations 12a - 21b, and compensation gain application, in each single - stream renderer. In this example, a quadrature mirror filter bank (QMF) is applied to each of program streams 1 and 2 before each program stream is received by the corresponding ones of rendering modules 1 and 2. In an alternative example, a quadrature mirror filter bank (QMF) may be applied to each of program streams 1 - N before each program stream is received by the corresponding ones of rendering modules 1 - N. According to this example, rendering modules 1 and 2 operate in the frequency domain. In this implementation, loudness estimation module 805a calculates a loudness estimate for program stream 1, as described above with reference to equations 12a - 17b, for example. Similarly, in this example, loudness estimation module 805b calculates a loudness estimate for program stream 2.
[0186] In this implementation, the time-domain microphone signal from the microphone system 120c is also provided to the quadrature mirror filter bank, such that the loudness estimation module 805c receives the microphone signal in the frequency domain. In this implementation, the loudness estimation module 805c calculates a loudness estimate for the microphone signal as described above with reference to, for example, equations 12b - 17a. In this example, the loudness processing module 810 is configured to implement loudness processing and compensation gain application, such as described in equations 18 - 21b, for each single-stream rendering module. In this implementation, the loudness processing module 810 is configured to modify the audio signals of program stream 1 and program stream 2 to maintain their perceived loudness in the presence of one or more interfering signals. In some cases, the control system can determine that the microphone signal corresponds to ambient noise and that the program stream should be raised above it. However, in some examples, the control system can determine that the microphone signal corresponds to a wake word, command, crying of a child, or other such audio that may need to be heard by the smart audio device and / or one or more listeners. In some such implementations, the loudness processing module 810 may be configured to modify the microphone signal to maintain its perceived loudness in the presence of the interfering audio signal of program stream 1 and / or the audio signal of program stream 2. Here, the loudness processing module 810 is configured to provide appropriate gains to the rendering modules 1 and 2. After the mixer 630c mixes the outputs of the rendering modules 1 - N, the inverse filter bank 635c converts the mix to the time domain and provides the mixed speaker feed signal in the time domain to the loudspeakers 1 - M. In this example, the quadrature mirror filter bank, the rendering modules 1 - N, the mixer 630c, and the inverse filter bank 635c are components of the control system 110e.
[0187] FIG. 9A shows an example of a multi-stream rendering system configured for cross-fading of a plurality of rendered streams. In some such embodiments, cross-fading of a plurality of rendered streams is used to provide a smooth experience when the rendering configuration is dynamically changed. One example is the aforementioned use case of simultaneous playback of a spatial program stream, such as music, with the response of a smart voice assistant to some query by a listener, as described above with reference to FIGS. 4A and 4B. In this case, as shown in FIG. 9A, it is useful to instantiate an extra single-stream renderer with an alternative spatial rendering configuration and cross-fade between them simultaneously.
[0188] In this example, QMF is applied to program stream 1 before it is received by rendering modules 1a and 1b. Similarly, QMF is applied to program stream 2 before it is received by rendering modules 2a and 2b. In some cases, the output of rendering module 1a may correspond to the desired playback of program stream 1 before wake word detection, while the output of rendering module 1b may correspond to the desired playback of program stream 1 after wake word detection. Similarly, the output of rendering module 2a may correspond to the desired playback of program stream 2 before wake word detection, and the output of rendering module 2b may correspond to the desired playback of program stream 2 after wake word detection. In this implementation, the outputs of rendering modules 1a and 1b are provided to cross-fade module 910a, and the outputs of rendering modules 2a and 2b are provided to cross-fade module 910b. The cross-fade time may be in the range of, for example, hundreds of milliseconds to several seconds.
[0189] After mixer 630d mixes the outputs of crossfade modules 910a and 910b, inverse filter bank 635d converts the mix into the time domain and provides the mixed speaker feed signal in the time domain to loudspeakers 1 through M. In this example, the quadrature mirror filter bank, the rendering module, the crossfade modules, mixer 630d, and inverse filter bank 635d are components of control system 110f.
[0190] In some embodiments, it may be possible to precompute the rendering configurations used in each of single stream renderers 1a, 1b, 2a, and 2b. This can be particularly convenient and efficient for use cases such as smart voice assistants because the spatial configuration is often known a priori and does not depend on other dynamic aspects of the system. In other embodiments, it may not be possible or desirable to precompute the rendering configuration, in which case the complete configuration for each single stream renderer must be computed dynamically while the system is operating.
[0191] Some aspects of some embodiments include the following: 1. An audio rendering system for simultaneously playing a plurality of audio program streams through a plurality of arbitrarily arranged loudspeakers, wherein at least one of the program streams is a spatial mix and the rendering of the spatial mix is dynamically modified in response to the simultaneous playback of one or more additional program streams. 2. The system of claim 1, wherein the rendering of any of the plurality of audio program streams can be dynamically modified as a function of any one or more combinations of the remaining plurality of audio program streams. 3. The system of claim 1 or 2, wherein the modification includes one or more of the following Modify the relative activation of the plurality of loudspeakers as a function of the relative activation of a loudspeaker associated with the rendering of at least one of the one or more additional program streams; Warp the intended spatial balance of the spatial mix as a function of the spatial characteristics of the rendering of at least one of the one or more additional program streams; or Modify the loudness or audibility of the spatial mix as a function of the loudness or audibility of at least one of the one or more additional program streams. 4. The system according to claim 1 or 2, further comprising dynamically modifying the rendering as a function of one or more microphone inputs. 5. The system according to claim 4, wherein the information derived from the microphone input used to modify the rendering comprises one or more of the following · Detection of the utterance of a specific phrase by a user of the system; · An estimated value of the position of one or more users of the system; · An estimated value of the loudness of any combination of N program streams at a specific position in the listening space; or · An estimated value of the loudness of other ambient sounds in the listening environment, such as background noise.
[0192] Other examples of embodiments of the system and method of the present invention for managing the playback of multiple audio streams through multiple speakers (e.g., the speakers of a set of orchestrated smart audio devices) include the following: 1. An audio system (e.g., an audio rendering system) for simultaneously playing multiple audio program streams through a plurality of arbitrarily arranged loudspeakers (e.g., the speakers of a set of orchestrated smart audio devices), wherein at least one of the program streams is a spatial mix, and the rendering of the spatial mix is dynamically modified in response to (or in relation to) the simultaneous playback of one or more additional program streams. 2. The system according to claim 1, wherein the modification to the spatial mix includes one or more of the following: · Warping the rendering of the spatial mix away from the rendering positions of the one or more additional streams, or · Modifying the loudness of the spatial mix according to the loudness of the one or more additional streams. 3. The system according to claim 1, further comprising the step of dynamically modifying the rendering of the spatial mix as a function of one or more microphone inputs (i.e., signals captured by one or more microphones of one or more smart audio devices, e.g., a set of orchestrated smart audio devices). 4. The system according to claim 3, wherein at least one of the one or more microphone inputs includes (indicates) human voice. Optionally, the rendering is dynamically modified in response to the determined position of the voice source (human). 5. The system according to claim 3, wherein at least one of the one or more microphone inputs includes environmental noise. 6. The system according to claim 3, wherein an estimated value of the loudness of the spatial stream or the one or more additional streams is derived from at least one of the one or more microphone inputs.
[0193] One practical consideration when implementing dynamic cost flexible rendering (according to some embodiments) is computational complexity. In some cases, considering that the object position (the position for each audio object to be rendered, which may be indicated by metadata) may change many times per second, it may not be feasible to solve the unique cost function for each frequency band for each audio object in real time. An alternative approach that trades off memory for computational complexity is to use a lookup table that samples the three-dimensional space of all possible object positions. The sampling does not have to be the same in all dimensions. FIG. 9B is a graph of points showing speaker activation in an exemplary embodiment. In this example, the x and y dimensions are sampled at 15 points and the z dimension is sampled at 5 points. Other implementations may include more or fewer samples. According to this example, each point represents the M speaker activation for the CMAP or FV solution.
[0194] At runtime, in some examples, tri-linear interpolation between the most recent 8 points of speaker activation may be used to determine the actual activation for each speaker. FIG. 10 is a graph of tri-linear interpolation between points showing speaker activation according to an example. In this example, the process of successive linear interpolation interpolates each pair of points in the upper surface to determine first and second interpolation points 1005a and 1005b, interpolates each pair of points in the lower surface to determine third and fourth interpolation points 1010a and 1010b, interpolates the first and second interpolation points 1005a and 1005b to determine a fifth interpolation point 1015 in the upper surface, interpolates the third and fourth interpolation points 1010a and 1010b to determine a sixth interpolation point 1020 in the lower surface, and interpolates the fifth and sixth interpolation points 1015 and 1020 to determine a seventh interpolation point 1025 between the upper and lower surfaces. Tri-linear interpolation is an effective interpolation method, but one of ordinary skill in the art will understand that tri-linear interpolation is only one possible interpolation method that may be used when implementing aspects of the present disclosure, and other examples may include other interpolation methods.
[0195] For example, in the first example above where repulsive forces are used to create an acoustic space for a voice assistant, another important concept is the transition from a rendering scene without repulsive forces to a scene with repulsive forces. To create a smooth transition and give the impression that the sound field is dynamically distorted, both a previous set of speaker activations without repulsive forces and a new set of speaker activations with repulsive forces are calculated and interpolated over a time period.
[0196] An example of audio rendering implemented according to an embodiment is an audio rendering method comprising: Rendering a set of one or more audio signals, each having a desired perceived spatial position associated therewith, through a set of two or more loudspeakers, wherein relative activation of the set of loudspeakers is a function of a model of the perceived spatial position of the audio signals to be reproduced through those loudspeakers, proximity of the desired perceived spatial position of the audio object to the positions of the loudspeakers, and at least one or more attributes of the set of audio signals, one or more attributes of the set of loudspeakers, or one or more additional dynamically configurable functions that depend on one or more external inputs.
[0197] Referring to FIG. 11, a further example of an embodiment will be described. Similar to the other figures provided herein, the types and numbers of elements shown in FIG. 11 are given merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. FIG. 11 shows a floor plan of a listening environment that is a living space in this example. According to this example, environment 1100 includes a living room 1110 in the upper left, a kitchen 1115 in the lower center, and a bedroom 1122 in the lower right. The squares and circles distributed across the living space represent a set of loudspeakers 1105a - 1105h that are arranged in convenient positions in the space but do not conform to a standard - defined layout (are arbitrarily arranged), some of which may be smart speakers in some implementations. In some examples, loudspeakers 1105a - 1105h may be coordinated to implement one or more of the disclosed embodiments. In this example, environment 1100 includes cameras 1111a - 1111e distributed throughout the environment. In some implementations, one or more smart audio devices within environment 1100 may also include one or more cameras. The one or more smart audio devices may be single - purpose audio devices or virtual assistants. In some such examples, one or more cameras of optional sensor system 130 may be present within a television 1130, within a mobile phone, or within a smart speaker such as one or more of loudspeakers 1105b, 1105d, 1105e, or 1105h. Cameras 1111a - 1111e are not shown in all of the figures of environment 1100 presented in this disclosure, but nevertheless, each of environment 1100 may include one or more cameras in some implementations.
[0198] Figures 12A, 12B, 12C, and 12D show examples of flexibly rendering spatial audio in a reference spatial mode for a plurality of different listening positions and orientations in the living space shown in FIG. 11. FIGS. 12A-12D show this ability at four exemplary listening positions. In each example, arrow 1205 pointing to person 1220a represents the position of the front sound stage (towards which person 1220a is facing). In each example, arrow 1210a represents the left surround field and arrow 1210b represents the right surround field.
[0199] In FIG. 12A, for person 1220a sitting on the couch 1225 in the living room, a reference spatial mode is determined and spatial audio is flexibly rendered. According to some implementations, a control system (such as control system 110 of FIG. 1A) may be configured to determine an assumed listening position and / or an assumed orientation of the reference spatial mode according to reference spatial mode data received via an interface system such as interface system 105 of FIG. 1A. Some examples are described below. In some such examples, the reference spatial mode data may include microphone data from a microphone system (such as microphone system 120 of FIG. 1A).
[0200] In some such examples, the reference spatial mode data may include microphone data corresponding to wake words and voice commands, such as "[wake word], set the TV to the front sound stage". Alternatively or additionally, the microphone data may be used to triangulate the position of the user according to the sound of the user's voice, for example, via direction of arrival (DOA) data. For example, three or more loudspeakers 1105a-1105e may use the microphone data, via the DOA data, to triangulate the position of person 1220a sitting on the couch 1225 in the living room according to the sound of the voice of person 1220a. The orientation of person 1220a may be assumed according to the position of person 1220a: if person 1220a is in the position shown in FIG. 12A, person 1220a may be assumed to be facing the television 1130.
[0201] Alternatively or additionally, the position and orientation of person 1220a may be determined according to image data from a camera system (such as the sensor system 130 of FIG. 1A).
[0202] In some examples, the position and orientation of person 1220a may be determined according to user input obtained via a graphical user interface (GUI). According to some such examples, the control system may be configured to control a display device (such as the display device of a cellular phone) to present a GUI that allows person 1220a to input their position and orientation.
[0203] FIG. 13A shows an example of a GUI for receiving user input regarding the position and orientation of a listener. According to this example, the user has pre-identified several possible listening positions and corresponding orientations. The loudspeaker positions corresponding to each position and the corresponding direction have already been input and stored during the setup process. Some examples are described below. For example, a listening environment layout GUI may be provided, and the user may be prompted to touch the positions corresponding to the possible listening positions and speaker positions and name the possible listening positions. In this example, at the time shown in FIG. 13A, the user has already provided a user input regarding the user's position to the GUI 1300 by touching the virtual button "Living Room Couch". Due to the L-shaped couch 1225, there are two possible forward-facing positions, so the user is prompted to indicate in which direction the user is facing.
[0204] In FIG. 12B, a reference space mode is determined and spatial audio is rendered flexibly for the person 1220a sitting on the reading chair 1215 in the living room. In FIG. 12C, a reference space mode is determined and spatial audio is rendered flexibly for the person 1220a standing next to the kitchen counter 1230. In FIG. 12D, a reference space mode is determined and spatial audio is rendered flexibly for the person 1220a sitting at the breakfast table 1240. As shown by the arrow 1205, it can be observed that the front sound stage orientation does not necessarily correspond to a particular loudspeaker within the environment 1100. As the position and orientation of the listener change, the roles of the various speakers for rendering the different components of the spatial mix also change.
[0205] For any person 1220a in FIGS. 12A - 12D, the person listens to the spatial mix as intended for each of the illustrated positions and orientations. However, this experience may not be optimal for additional listeners within the space. FIG. 12E shows an example of a reference spatial mode rendering when two listeners are in different positions within the listening environment. FIG. 12E shows a reference spatial mode rendering for person 1220a on the couch and person 1220b standing in the kitchen. In this example, the rendering may be optimal for person 1220a, but person 1220b, considering their position, hears mostly signals from the surround field and hardly hears the front sound stage.
[0206] In this case, and in other cases where multiple people are in a space and may move around in an unpredictable manner (e.g., a party), a more appropriate rendering mode is needed for such a dispersed audience. FIG. 13B shows a distributed spatial rendering mode according to an exemplary embodiment. In this example of the distributed space mode, the front sound stage is rendered uniformly across the entire listening space, rather than only from the living room or from positions in front of listeners on the couch. This distribution of the front sound stage is represented by a plurality of arrows 1305d that circle the cloud 1335, and all the arrows 1305d have the same length, or approximately the same length. The intended meaning of the arrows 1305d is that the plurality of listeners shown (persons 1220a - 1220f) can hear this part of the mix equally well regardless of their position. However, if this uniform distribution were applied to all components of the mix, all spatial aspects of the mix would be lost and persons 1220a - 1220f would essentially be listening to monaural audio. To maintain some degree of spatiality, the left and right surround components of the mix, represented by arrows 1210a and 1210b respectively, are still spatially rendered. (Often, there may be left and right surround, left and right rear surround, overhead, and dynamic audio objects with spatial positions within this space. Arrows 1210a and 1210b are intended to represent the left and right portions of all these possibilities.) To maximize the perceived spatiality, the area in which these components are spatialized is expanded to more fully cover the entire listening space, including the space previously occupied only by the front sound stage. This expanded area in which the surround components are rendered can be understood by comparing the relatively long arrows 1210a and 1210b shown in FIG. 13B with the relatively short arrows 1210a and 1210b shown in FIG. 12A. Further, the arrows 1210a and 1210b shown in FIG. 12A represent the surround components in the reference space mode, extending approximately from the side of person 1220a to the back of the listening environment and not extending into the front stage area of the listening environment.
[0207] In this example, when implementing a uniform distribution of the front sound stage and an enhanced spatialization of the surround components, care is taken so that the perceived loudness of these components is mostly maintained as compared to the rendering for the reference spatial mode. The goal is to shift the spatial impression of these components to optimize for multiple people while maintaining the relative levels of each component in the mix. For example, it would not be desirable if the front sound stage became twice as large relative to the surround components as a result of its uniform distribution.
[0208] To switch between the various reference rendering modes and the diffuse rendering modes of the exemplary embodiments, in some examples, the user can interact with a voice assistant associated with the orchestrated speaker system. For example, to play audio in the reference spatial mode, the user may utter a command, such as "Play [insert content name] for me" or "Play [insert content name] in personal mode" following a wake word for the voice assistant (e.g., "Listen Dolby"). Then, based on recordings from various microphones associated with the system, the system may automatically determine the user's position and orientation, or the zone among some predetermined zones that is closest to the user, and may begin playing the audio in the reference mode corresponding to this determined position. To play audio in the diffuse spatial mode, the user may utter a different command, such as "Play [insert content name] in diffuse mode".
[0209] Alternatively or additionally, the system may be configured to automatically switch between a reference mode and a distributed mode based on other inputs. For example, the system may have means for automatically determining the number and location of listeners in the space. This can be achieved, for example, by monitoring the voice activity in the space from the associated microphones and / or through the use of other relevant sensors such as one or more cameras. In this case, the system may be configured to include a mechanism for continuously varying the rendering between a reference spatial mode as shown in FIG. 12E and a fully distributed spatial mode as shown in FIG. 13B. The point on this continuum at which the rendering is set may be calculated, for example, as a function of the number of people reported in the space.
[0210] FIGS. 12A, 14A, and 14B illustrate this behavior. In FIG. 12A, the system detects only a single listener (person 1220a) sitting on the couch facing the TV, and thus the rendering mode is set to the reference spatial mode for this listener position and orientation. FIG. 14A shows an example of a partially distributed spatial rendering mode. In FIG. 14A, two additional people (persons 1220e and 1220f) are detected behind person 1220a, and the rendering mode is set to a point between the reference spatial mode and the fully distributed spatial mode. Here, some of the front sound stage (arrows 1305a, 1305b, and 1305c) is pulled back towards the additional listeners (persons 1220e and 1220f), but still, there is more emphasis on the position of the front sound stage of the reference spatial mode. This emphasis is shown in FIG. 14A by the length of arrow 1305a, which is relatively long compared to the lengths of arrows 1205 and 1305b and 1305c. Also, the surround field is only partially extended towards the position of the front sound stage of the reference spatial mode, as indicated by the lengths and positions of arrows 1210a and 1210b.
[0211] FIG. 14B shows an example of a fully distributed spatial rendering mode. In some examples, the system may detect a number of listeners (persons 1220a, 1220e, 1220f, 1220g, 1220h, and 1220i) spanning the entire space, and the system may automatically set the rendering mode to the fully distributed spatial mode. In other examples, the rendering mode may be set according to user input. The fully distributed spatial mode is shown in FIG. 14B by the uniform or substantially uniform length of arrow 1305d, as well as the length and position of arrows 1210a and 1210b.
[0212] In the foregoing examples, in the distributed rendering mode, the portion of the spatial mix that is rendered using a more uniform distribution is designated as the front sound stage. In many spatial mix contexts, this makes sense because traditional mix implementations typically place the most important parts of the mix, such as dialog for movies, lead vocals for music, drums, bass, etc., in the front sound stage. This holds true for most 5.1 and 7.1 surround sound mixes, as well as for stereo content upmixed to 5.1 or 7.1 using algorithms such as Dolby Pro Logic or Dolby Surround. Here, the front sound stage is provided by the left, right, and center channels. This also applies to many object-based audio mixes, such as Dolby Atmos, where audio data can be designated as the front sound stage according to spatial metadata indicating (x,y) spatial positions where y < 0.5. However, in object-based audio, the mixing engineer has the freedom to place the audio anywhere in 3D space. In particular, in object-based music, the mixing engineer is moving away from traditional mixing norms and is increasingly placing what are considered important parts of the mix, such as lead vocals, in non-traditional positions such as overhead. In such cases, it becomes difficult to construct simple rules for determining which components of the mix are suitable for rendering in a more distributed spatial manner for the distributed rendering mode. Object-based audio already includes metadata associated with each audio signal of its components that describes where in 3D space the signal should be rendered. To address the above problem, in some implementations, additional metadata may be added that allows the content creator to flag certain signals as being suitable for more distributed spatial rendering in the distributed rendering mode. During rendering, the system may use this metadata to select the components of the mix to which the more distributed rendering is applied.This gives the content creator control over how the distributed rendering mode sounds for a particular piece of content.
[0213] In some alternative implementations, the control system may be configured to implement a content type classifier to identify one or more elements of the audio data to be rendered in a more spatially distributed manner. In some examples, the content type classifier may refer to content type metadata (e.g., metadata indicating that the audio data is dialog, vocal, percussion, bass, etc.) to determine whether the audio data should be rendered in a more spatially distributed manner. According to some such implementations, the content type metadata for content to be rendered in a more spatially distributed manner may be selectable by the user, for example, according to user input via a GUI displayed on a display device.
[0214] The exact mechanism used to render one or more elements of a spatial audio mix in a more spatially distributed manner than in the reference spatial mode may vary between different embodiments, and this disclosure is intended to cover all such mechanisms. One exemplary mechanism involves generating multiple copies of each such element with multiple associated rendering positions that are more uniformly distributed throughout the listening space. In some implementations, the rendering position and / or number of rendering positions for the distributed spatial mode may be user-selectable, while in other implementations, the rendering position and / or number of rendering positions for the distributed spatial mode may be preset. In some such implementations, a user may select several rendering positions for the distributed spatial mode, which may be preset, e.g., evenly spaced throughout the listening environment. The system then renders all of these copies at a set of distributed positions rather than the original single element in its original intended location. According to some implementations, the levels of the copies may be modified so that the perceived level associated with the combined rendering of all copies is the same or substantially the same (e.g., within a threshold number of decibels, such as 2 dB, 3 dB, 4 dB, 5 dB, 6 dB, etc.) as the level of the original single element in the reference rendering mode.
[0215] A more elegant mechanism can be implemented in the context of either a CMAP or FV flexible rendering system, or a hybrid of both systems, where each element of a spatial mix is rendered at a specific location in space, and each element may have an associated assumed fixed position, for example the canonical position of a channel in a 5.1 or 7.1 surround sound mix, or a time-varying position as in the case of object-based audio like Dolby Atmos.
[0216] Figure 15 shows exemplary rendering positions for a CMAP and FV rendering system on a 2D plane. Each small, numbered circle represents an exemplary rendering position, and the rendering system can render elements of the spatial mix anywhere above or within circle 1500. The positions on circle 1500 labeled L, R, C, Lss, Rss, Lrs, and Rrs represent the fixed canonical rendering positions of the seven full-range channels of a 7.1 surround mix in this example: left (L), right (R), center (C), left side surround (Lss), right side surround (Rss), left rear surround (Lrs), and right rear surround (Rrs). In this context, the rendering positions near L, R, and C are considered the front sound stage. For the reference rendering mode (also referred to herein as the "reference spatial mode"), the listener is assumed to be located at the center of the large circle facing the C rendering position. For any of FIGS. 12A - 12D showing reference renderings for various listener positions and orientations, it can be conceptualized that the center of FIG. 15 is superimposed over the listener. Here, FIG. 15 is additionally rotated and scaled such that the C position is aligned with the position of the front sound stage (arrow 1205) and circle 1500 of FIG. 15 surrounds cloud 1235. The resulting alignment describes that any of the speakers in FIGS. 12A - 12D is relatively close to any of the rendering positions in FIG. 15. In some implementations, it is this proximity that largely governs the relative activation of the speakers when rendering elements of the spatial mix at specific positions for both the CMAP and FV rendering systems.
[0217] When spatial audio is mixed within a studio, speakers are generally arranged at a uniform distance around the listening position. In most cases, there are no speakers within the resulting circle or hemisphere. When the audio is placed “in the middle of the room” (e.g., the center of FIG. 15), the rendering tends to direct the emission of all the speakers on the circumference in order to achieve “sound nowhere”. In the CMAP and FV rendering systems, a similar effect can be achieved by changing the proximity penalty term of the cost function that governs speaker activation. In particular, for the rendering positions on the circumference of circle 1500 in FIG. 15, the proximity penalty term fully penalizes the use of speakers that are far from the desired rendering position. Thus, only the speakers near the intended rendering position are substantially activated. As the desired rendering position moves towards the center of the circle (radius zero), the proximity penalty term decreases to zero, and as a result, at the center, no speaker is preferred. The corresponding result for a rendering position at radius zero is a perfectly uniform perceived distribution of the audio across the listening space, which is exactly the desired result for certain elements of the mix in the most diffuse spatial rendering mode.
[0218] Given this behavior of the CMAP and FV systems at zero radius, a more spatially distributed rendering of any element of the spatial mix can be achieved by warping its intended spatial location toward the zero-radius point. This warping may be continuous between the original intended location and zero radius, thereby providing natural, continuous control between the reference spatial mode and various distributed spatial modes. Figures 16A, 16B, 16C, and 16D show examples of warping applied to all of the rendering points in Figure 15 to achieve various distributed spatial rendering modes. Figure 16D shows an example of such warping applied to all of the rendering points in Figure 15 to achieve a fully distributed rendering mode. It can be seen that the L, R, and C points (the front soundstage) are collapsed to a zero radius, thereby ensuring rendering in a completely uniform manner. Additionally, the Lss and Rss rendering points are pulled back toward the front soundstage along the periphery of the circle so that the spatialized surround field (Lss, Rss, Lbs, and Rbs) encompasses the entire listening area. This warping is applied to the entire rendering space, and it can be seen that all of the rendering points in Figure 15 have been warped to new positions in Figure 16D that correspond to the warping of the 7.1 canonical positions. The spatial mode referenced in Figure 16D is an example of what is referred to herein as the "most distributed spatial mode" or "fully distributed spatial mode."
[0219] Figures 16A, 16B, and 16C show various examples of an intermediate distributed spatial mode between the distributed spatial mode represented in FIG. 15 and the distributed spatial mode represented in FIG. 16D. FIG. 16B represents an intermediate point between the distributed spatial mode represented in FIG. 15 and the distributed spatial mode represented in FIG. 16D. FIG. 16A represents an intermediate point between the distributed spatial mode represented in FIG. 15 and the distributed spatial mode represented in FIG. 16B. FIG. 16C represents an intermediate point between the distributed spatial mode represented in FIG. 16B and the distributed spatial mode represented in FIG. 16D.
[0220] Figure 17 shows an example of a GUI through which a user can select a rendering mode. According to some implementations, the control system may control a display device (e.g., a cellular phone) to display GUI 1700 or a similar GUI on the display. The display device may include a sensor system (e.g., a touch sensor system, or a gesture sensor system proximate to the display (e.g., above or below the display)). The control system may be configured to receive user input via GUI 1700 in the form of sensor signals from the sensor system. The sensor signals may correspond to user touches or gestures corresponding to elements of GUI 1700.
[0221] According to this example, the GUI includes a virtual slider 1701 with which the user can interact to select a rendering mode. As indicated by arrow 1703, the user can move the slider in either direction along track 1707. In this example, line 1705 indicates the position of virtual slider 1701 that corresponds to a reference spatial mode, such as one of the reference spatial modes disclosed herein. Other implementations may provide other features on the GUI with which the user can interact, such as a virtual knob or dial. According to some implementations, after selecting a reference spatial mode, the control system may present a GUI such as that shown in FIG. 13A or another such GUI that allows the user to select a listener position and orientation for the reference spatial mode.
[0222] In this example, line 1725 indicates the position of virtual slider 1701 corresponding to the most dispersed spatial mode, such as the dispersed spatial mode shown in FIG. 13B. According to this implementation, lines 1710, 1715, and 1720 indicate positions of virtual slider 1701 corresponding to intermediate spatial modes. In this example, the position of line 1710 corresponds to the intermediate spatial mode, such as that shown in FIG. 16A. Here, the position of line 1715 corresponds to the intermediate spatial mode, such as that shown in FIG. 16B. In this implementation, the position of line 1720 corresponds to the intermediate spatial mode, such as that shown in FIG. 16C. According to this example, a user can interact with (e.g., touch) an “apply” button to instruct the control system to implement the selected rendering mode.
[0223] However, other implementations may provide other ways for a user to select one of the aforementioned distributed space modes. According to some examples, the user may utter a voice command, such as "Please play [insert content name] in semi-distributed mode." "Semi-distributed mode" may correspond to the distributed mode indicated by the position of line 1715 in GUI 1700 of FIG. 17. According to some such examples, the user may utter a voice command, such as "Please play [insert content name] in 1 / 4 distributed mode." "1 / 4 distributed mode" may correspond to the distributed mode indicated by the position of line 1710.
[0224] FIG. 18 is a flowchart outlining an example of a method that may be performed by an apparatus or system such as those disclosed herein. The blocks of method 1800 are not necessarily performed in the order shown, similar to other methods described herein. In some implementations, one or more blocks of method 1800 may be performed concurrently. Further, some implementations of method 1800 may include more or fewer blocks than those illustrated and / or described. The blocks of method 1800 may be performed by one or more devices (such as the control system 110 shown in FIG. 1A and described above, or one of the other disclosed control system examples, or may include it).
[0225] In this implementation, block 1805 involves the control system receiving, via an interface system, audio data including one or more audio signals and associated spatial data. In this example, the spatial data indicates the intended perceived spatial position corresponding to the audio signal. Here, the spatial data includes channel data and / or spatial metadata.
[0226] In this example, block 1810 includes determining a rendering mode by a control system. Determining the rendering mode may, in some cases, include receiving a rendering mode instruction via an interface system. Receiving a rendering mode instruction may involve, for example, receiving a microphone signal corresponding to an audio command. In some examples, receiving a rendering mode instruction may include receiving a sensor signal corresponding to a user input via a graphical user interface. The sensor signal may be, for example, a touch sensor signal and / or a gesture sensor signal.
[0227] In some implementations, receiving a rendering mode instruction may include receiving an indication of the number of people in a listening area. According to some such examples, the control system may be configured to determine the rendering mode based at least in part on the number of people in the listening area. In some such examples, the indication of the number of people in the listening area may be based on microphone data from a microphone system and / or image data from a camera system.
[0228] According to the example shown in FIG. 18, block 1815 includes rendering audio data for playback via a set of loudspeakers in the environment according to a rendering mode determined at block 1810 by a control system, and generating a rendered audio signal. In this example, rendering the audio data includes determining the relative activation of a set of loudspeakers within the environment. Here, the rendering mode is variable between a reference space mode and one or more distributed space modes. In this implementation, the reference space mode has an assumed listening position and orientation. According to this example, in the one or more distributed space modes, one or more elements of the audio data are rendered in a more spatially distributed manner than in the reference space mode. In this example, in the one or more distributed space modes, the spatial positions of the remaining elements of the audio data are warped so that they span the rendering space of the environment more completely than in the reference space mode.
[0229] In some implementations, rendering one or more elements of the audio data in a more spatially distributed manner than in the reference space mode may include generating copies of the one or more elements. Some such implementations may include rendering all copies simultaneously at a distributed set of positions across the environment.
[0230] According to some implementations, the rendering may be based on CMAP, FV, or a combination thereof. Rendering one or more elements of the audio data in a more spatially distributed manner than in the reference space mode may include warping the rendering position of each of the one or more elements towards a zero radius.
[0231] In this example, block 1820 includes providing, by a control system, via an interface system, a rendered audio signal to at least some of the set of loudspeakers of the environment.
[0232] According to some implementations, the rendering mode may be selectable from a continuum of rendering modes ranging from a reference spatial mode to the most diffuse spatial mode. In some such implementations, the control system may be further configured to determine an assumed listening position and / or orientation of the reference spatial mode according to reference spatial mode data received via the interface system. According to some such implementations, the reference spatial mode data may include microphone data from a microphone system and / or image data from a camera system. In some such examples, the reference spatial mode data may include microphone data corresponding to a voice command. Alternatively or additionally, the reference spatial mode data may include microphone data corresponding to the position of one or more utterances of a person in the listening environment. In some such examples, the reference spatial mode data may include image data indicating the position and / or orientation of a person in the listening environment.
[0233] However, in some cases, the apparatus or system may include a display device and a sensor system proximate to the display device. The control system may be configured to control the display device to present a graphical user interface. The reception of the reference spatial mode data may include receiving a sensor signal corresponding to a user input via the graphical user interface.
[0234] According to some implementations, the one or more elements of audio data each rendered in a more spatially distributed manner may correspond to front soundstage data, musical vocals, dialogue, bass, percussion, and / or other solo or lead instruments. In some cases, the front soundstage data may include left, right, or center signals of audio data received in or upmixed to a Dolby 5.1, Dolby 7.1, or Dolby 9.1 format. In some examples, the front soundstage data may include audio data received in a Dolby Atmos format and having spatial metadata indicating an (x,y) spatial location, where y<0.5.
[0235] In some cases, the audio data may include spatial distribution metadata that indicates which elements of the audio data should be rendered in a more spatially distributed manner, and in some such examples, the control system may be configured to identify the one or more elements of the audio data that should be rendered in a more spatially distributed manner in accordance with the spatial distribution metadata.
[0236] Alternatively or additionally, the control system may be configured to implement a content type classifier to identify the one or more elements of the audio data to be rendered in a more spatially distributed manner. In some examples, the content type classifier may refer to content type metadata (e.g., metadata indicating that the audio data is dialog, vocal, percussion, bass, etc.) to determine whether the audio data should be rendered in a more spatially distributed manner. According to some such implementations, the content type metadata to be rendered in a more spatially distributed manner may be selectable by the user, for example, according to user input via a GUI displayed on the display device. Alternatively or additionally, the content type classifier may operate directly on the audio signals in combination with the rendering system. For example, the classifier may use a neural network trained on various content types to analyze the audio signals and determine whether they belong to any content type (vocal, lead guitar, drums, etc.) that may be considered suitable for rendering in a more spatially distributed manner. Some such classifications may be performed in a continuous and dynamic manner, and the resulting classification results may adjust a set of signals to be rendered in a more spatially distributed manner in a continuous and dynamic manner. Some such implementations may include the use of techniques such as neural networks to implement such a dynamic classification system according to methods known in the art.
[0237] In some examples, at least one of the one or more distributed spatial modes may include applying a time-varying modification to the spatial position of at least one element. According to some such examples, the time-varying modification may be a periodic modification. For example, the periodic modification may include orbiting one or more rendering positions around the circumference of the listening environment. According to some such implementations, the periodic modification may include the tempo of music played in the environment, the beat of music played in the environment, or one or more other characteristics of the audio data played in the environment. For example, some such periodic modifications may include alternating between two, three, four, or more rendering positions. The alternation may correspond to the beat of music played in the environment. In some implementations, the periodic modification may be selectable according to user input, e.g., according to one or more voice commands, according to user input received via a GUI, etc.
[0238] FIG. 19 illustrates an example of the geometric relationship between three audio devices in an environment. In this example, environment 1900 is a room that includes a television 1901, a sofa 1903, and five audio devices 1905. According to this example, audio devices 1905 are located at positions 1 through 5 in environment 1900. In this implementation, each audio device 1905 includes a microphone system 1920 having at least three microphones and a speaker system 1925 including at least one speaker. In some implementations, each microphone system 1920 includes an array of microphones. According to some implementations, each of audio devices 1905 may include an antenna system including at least three antennas.
[0239] As with other examples disclosed herein, the types, numbers, and arrangements of the elements shown in FIG. 19 are merely examples. Other implementations may have different types, numbers, and arrangements of elements, for example, more or fewer audio devices 1905, audio devices 1905 in different positions, and the like.
[0240] In this example, triangle 1910a has vertices at positions 1, 2, and 3. Here, triangle 1910a has sides 12, 23a, and 13a. According to this example, the angle between sides 12 and 23 is θ2, the angle between sides 12 and 13a is θ1, and the angle between sides 23a and 13a is θ3. These angles may be determined according to DOA data, as described in more detail later.
[0241] In some implementations, only the relative lengths of the sides of the triangle may be determined. In alternative implementations, the actual lengths of the sides of the triangle may be estimated. According to some such implementations, the actual lengths of the sides of the triangle may be estimated according to TOA data, for example, according to the arrival time of a sound generated by an audio device located at one triangle vertex and detected by an audio device located at another triangle vertex. Alternatively or additionally, the lengths of the sides of the triangle may be estimated by electromagnetic waves generated by an audio device located at one triangle vertex and detected by an audio device located at another triangle vertex. For example, the lengths of the sides of the triangle may be estimated according to the signal strength of the electromagnetic waves generated by an audio device located at one triangle vertex and detected by an audio device located at another triangle vertex. In some implementations, the lengths of the sides of the triangle may be estimated according to the detected phase shift of the electromagnetic waves.
[0242] Figure 20 shows another example of the geometric relationship between three audio devices in the environment shown in Figure 19. In this example, triangle 1910b has vertices at positions 1, 3, and 4. Here, triangle 1910b has sides 13b, 14, and 34a. According to this example, the angle between sides 13b and 14 is θ4, the angle between sides 13b and 34a is θ5, and the angle between sides 34a and 14 is θ6.
[0243] By comparing Figures 11 and 12, it can be observed that the length of side 13a of triangle 1910a should be equal to the length of side 13b of triangle 1910b. In some implementations, even if the length of the side of a certain triangle (for example, triangle 1910a) is assumed to be correct, the length of the side shared by the adjacent triangle is restricted by this length.
[0244] Figure 21A shows both of the triangles shown in Figures 19 and 20 without the corresponding audio devices and other features of the environment. Figure 21A shows the estimated values of the side lengths and angle orientations of triangles 1910a and 1910b. In the example shown in Figure 21A, the length of side 13b of triangle 1910b is restricted to the same length as side 13a of triangle 1910a. The lengths of the other sides of triangle 1910b are scaled in proportion to the resulting change in the length of side 13b. The resulting triangle 1910b' is shown adjacent to triangle 1910a in Figure 21A.
[0245] According to some implementations, the lengths of the sides of other triangles adjacent to triangles 1910a and 1910b can all be determined in a similar manner until all the audio device positions in environment 1900 are determined.
[0246] Some examples of audio device positions can be as follows. Each audio device may report the DOA of all other audio devices in the environment (for example, a room) based on the sounds generated by all other audio devices in the environment. The Cartesian coordinates of the i-th audio device are
Number
[0247] Figure 21B shows an example of estimating the interior angles of a triangle formed by three audio devices. In this example, the audio devices are i, j, and k. The DOA of the sound source emitted from device j as observed from device i is θ ji It may be represented as. The DOA of the sound source emitted from device k as observed from device i is θ ki It may be represented as. In the example shown in Figure 21B, θ ji and θ ki are measured from axis 2105a, the orientation of axis 2105a is arbitrary, and axis 2105a may correspond to, for example, the orientation of audio device i. The interior angle a of triangle 2110 is a = θ ki - θ ji It may be represented as. It can be observed that the calculation of the interior angle a does not depend on the orientation of axis 2105a.
[0248] In the example shown in Figure 21B, θ ij and θ kj are measured from axis 2105b, the orientation of axis 2105b is arbitrary, and axis 2105b may correspond to the orientation of audio device j. The interior angle b of triangle 2110 is b = θ ij - θ kj It may be represented as. Similarly, in this example, θ jk and θ ik are measured from axis 2105c. The interior angle c of triangle 2110 is c = θ jk - θ ik It may be represented as.
[0249] In the case of measurement errors, a + b + c ≠ 180°. Robustness can be improved by predicting each angle from the other two angles and averaging them, for example, as follows.
Equation
[0250] In some implementations, the edge lengths (A, B, C) can be calculated (except for scaling errors) by applying the law of sines. In some examples, one of the edge lengths may be assigned an arbitrary value such as 1. For example, let A = 1 and place the vertex
Number
Number
[0251] According to some implementations, the process of triangle parameterization may be repeated for every possible subset of three audio devices in the environment. Those subsets are enumerated in a supersets ζ of size
Number
[0252] FIG. 22 is a flowchart outlining an example of a method that may be performed by an apparatus as shown in FIG. 1A. The blocks of method 2200 are not necessarily performed in the order shown, similar to other methods described herein. Further, such a method may include more or fewer blocks than those illustrated and / or described. In this implementation, method 2200 includes estimating the positions of speakers in the environment. The blocks of method 2200 may be performed by one or more devices that may be (or may include) device 100 shown in FIG. 1A.
[0253] In this example, block 2205 includes obtaining direction of arrival (DOA) data for each audio device of the plurality of audio devices. In some examples, the plurality of audio devices may include all of the audio devices in an environment, such as all of the audio devices 1905 shown in FIG.
[0254] However, in some cases, the plurality of audio devices may include only a subset of all audio devices in the environment, for example, the plurality of audio devices may include all smart speakers in the environment, but not one or more of the other audio devices in the environment.
[0255] The DOA data may be obtained in various ways, depending on the particular implementation. In some cases, determining the DOA data may include determining DOA data for at least one audio device of the plurality of audio devices. For example, determining the DOA data may include receiving microphone data from each microphone of a plurality of audio device microphones corresponding to a single audio device of the plurality of audio devices and determining the DOA data for the single audio device based at least in part on the microphone data. Alternatively or additionally, determining the DOA data may include receiving antenna data from one or more antennas corresponding to a single audio device of the plurality of audio devices and determining the DOA data for the single audio device based at least in part on the antenna data.
[0256] In some such examples, a single audio device itself may determine the DOA data. According to some such implementations, each audio device of the plurality of audio devices may determine its own DOA data. However, in other implementations, another device, which may be a local or remote device, may determine the DOA data for one or more of the audio devices in the environment. According to some implementations, a server may determine the DOA data for one or more of the audio devices in the environment.
[0257] According to this example, block 2210 includes determining the interior angles for each of a plurality of triangles based on the DOA data. In this example, each of the plurality of triangles has vertices corresponding to the audio device positions of three of the audio devices of the audio devices. Some such examples have been described above.
[0258] FIG. 23 shows an example where each audio device in the environment is a vertex of a plurality of triangles. The sides of each triangle correspond to the distance between two audio devices 1905.
[0259] In this implementation, block 2215 includes determining the side lengths of each side of each triangle. (The sides of a triangle may also be referred to herein as "edges".) According to this example, the side lengths are at least partially based on the interior angles. In some cases, the side lengths may be calculated by determining the length of a first side of the triangle and then determining the lengths of the second and third sides of the triangle based on the interior angles of the triangle. Some such examples have been described above.
[0260] According to some such implementations, determining the first length may include setting the first length to a predetermined value. However, in some examples, the determination of the first length may be based on time-of-arrival data and / or received signal strength data. The time-of-arrival data and / or received signal strength data may, in some implementations, correspond to sound waves from a first audio device in the environment detected by a second audio device in the environment. Alternatively or additionally, the time-of-arrival data and / or received signal strength data may correspond to electromagnetic waves (e.g., radio waves, infrared waves, etc.) from a first audio device in the environment detected by a second audio device in the environment.
[0261] According to this example, block 2220 includes performing a forward alignment process that aligns each of the plurality of triangles in the first sequence. According to this example, the forward alignment process generates a forward alignment matrix.
[0262] According to some such examples, the triangles are shown, for example, in FIG. 21A and, as described above, are expected to align in such a way that the edges (x i , x j ) are equal to adjacent edges.
Number
Number
[0263] FIG. 24 gives an example of part of the alignment process. The numbers 1 to 5 shown in bold in FIG. 24 correspond to the audio device positions shown in FIGS. 1, 2, and 5. The sequence of the alignment process shown in FIG. 24 and described herein is merely an example.
[0264] In this example, as shown in FIG. 21A, the length of side 13b of triangle 1910b is forced to match the length of side 13a of triangle 1910a. The resulting triangle 1910b' is shown in FIG. 24, and the same interior angles are maintained. According to this example, the length of side 13c of triangle 1910c is also forced to match the length of side 13a of triangle 1910a. The resulting triangle 1910c' is shown in FIG. 24, and the same interior angles are maintained.
[0265] Next, in this example, the length of side 34b of triangle 1910d is forced to match the length of side 34a of triangle 1910b'. Further, in this example, the length of side 23b of triangle 1910d is forced to match the length of side 23a of triangle 1910a. The resulting triangle 1910d' is shown in FIG. 24, and the same interior angles are maintained. According to some such examples, the remaining triangles shown in FIG. 5 can be processed in the same way as triangles 1910b, 1910c, and 1910d.
[0266] The result of the alignment process may be stored in a data structure. According to some such examples, the result of the alignment process may be stored in an alignment matrix. For example, the result of the alignment process may be stored in a matrix
Number
[0267] If the DOA data and / or the initial side length determination contains errors, multiple estimated values of the audio device position will occur. These errors generally increase during the alignment process.
[0268] FIG. 25 shows an example of multiple estimates of audio device positions generated during the forward alignment process. In this example, the forward alignment process is based on a triangle with seven audio device positions as vertices, where the triangle does not align perfectly due to additive errors in the DOA estimates. The positions numbered 1 through 7 shown in FIG. 25 correspond to estimated audio device positions generated by the forward alignment process. In this example, the audio device position estimate labeled "1" matches, but the audio device position estimates for audio devices 6 and 7 show a larger difference. This is indicated by the relatively large area over which the numbers 6 and 7 are located.
[0269] Returning to FIG. 22, in this example, block 2225 includes an inverse sorting process that sorts each of the plurality of triangles in a second sequence that is the inverse of the first sequence. According to some implementations, the inverse sorting process may include traversing the ε as before, but in reverse order. In alternative examples, the inverse sorting process may not be exactly the reverse of the sequence of operations of the forward sorting process. According to this example, the inverse sorting process generates an inverse sorting matrix, which is herein referred to as
number
[0270] Figure 26 shows an example of a portion of the reverse alignment process. Numbers 1-5 shown in bold in Figure 26 correspond to the audio device positions shown in Figures 19, 21, and 23. The sequence of the reverse alignment process shown in Figure 26 and described herein is merely an example.
[0271] In the example shown in FIG. 26, triangle 1910e is based on audio device positions 3, 4, and 5. In this implementation, the side lengths (or "edges") of triangle 1910e are assumed to be correct, and the side lengths of adjacent triangles are forced to match them. According to this example, the length of side 45b of triangle 1910f is forced to match the length of side 45a of triangle 1910e. The resulting triangle 1910f' has the same interior angles and is shown in FIG. 26. In this example, the length of side 35b of triangle 1910c is forced to match the length of side 35a of triangle 1910e. The resulting triangle 1910c" has the same interior angles and is shown in FIG. 26. According to some such examples, the remaining triangles shown in FIG. 23 can be processed in the same way as triangles 1910c and 1910f until the inverse alignment process encompasses all of the remaining triangles.
[0272] FIG. 27 shows an example of a plurality of estimated values of audio device positions that occurred during the inverse alignment process. In this example, the inverse alignment process is based on the same triangle with the same seven audio device positions as vertices, described above with reference to FIG. 25. The positions numbered 1 to 7 shown in FIG. 27 correspond to the estimated audio device positions generated by the inverse alignment process. Again, these triangles do not align perfectly due to the additive error of the DOA estimates. In this example, the audio device position estimates labeled 6 and 7 match, but the audio device position estimates for audio devices 1 and 2 show a greater difference.
[0273] Returning to FIG. 22, block 2230 includes generating a final estimate of each audio device position, at least in part, based on the values of the forward alignment matrix and the values of the inverse alignment matrix. In some examples, generating a final estimate of the position of each audio device can include translating and scaling the forward alignment matrix to generate a translated and scaled forward alignment matrix, and translating and scaling the inverse alignment matrix to generate a translated and scaled inverse alignment matrix.
[0274] For example, translation and scaling move the center of mass to the origin and enforce a unit Frobenius norm, e.g.
number
[0275] According to some such examples, generating a final estimate of each audio device position may include generating a rotation matrix based on a translated and scaled forward alignment matrix and a translated and scaled reverse alignment matrix. The rotation matrix may include multiple estimated audio device positions for each audio device. The optimal rotation between the forward alignment and the reverse alignment may be found, for example, by singular value decomposition. In some such examples, generating the rotation matrix may include performing singular value decomposition on the translated and scaled forward alignment matrix and the translated and scaled reverse alignment matrix, for example, as follows:
number
[0276] In the above formula, U represents the left singular vector and V is the matrix
number
number
[0277] According to some examples, the rotation matrix R=VU T After determining , the alignments may be averaged, for example, as follows:
number
[0278] In some implementations, generating a final estimate of each audio device location also includes averaging the estimated audio device locations of each audio device to generate a final estimate of each audio device location. Even when the DOA data and / or other calculations include significant errors, the various disclosed implementations have proven to be robust. For example,
Number
Number
[0279] FIG. 28 shows a comparison of the estimated audio device locations with the actual audio device locations. In the example shown in FIG. 28, the audio device locations correspond to the locations estimated during the forward and reverse alignment processes described above with reference to FIGS. 17 and 19. In these examples, the error in the DOA estimation had a standard deviation of 15 degrees. Nevertheless, the final estimates of each audio device location (each represented by an "x" in FIG. 28) closely match the actual audio device locations (each represented by a circle in FIG. 28).
[0280] Much of the foregoing discussion pertains to the auto-location of an audio device. The following discussion elaborates on several methods for determining the listener's position and the listener's angular orientation, as briefly described above. In the foregoing description, the term "rotation" is used essentially in the same way as the term "orientation" is used in the following description. For example, the foregoing "rotation" may refer to a global rotation of the final speaker geometry, rather than the rotation of individual triangles during the process described above with reference to FIG. 14 and below. This global rotation or orientation may be resolved with reference to the listener's angular orientation, for example, by the direction the listener is looking, by the direction the listener's nose is pointing, etc.
[0281] Various satisfactory methods for estimating the listener position are described below. However, estimating the listener angular orientation can sometimes be difficult. Several related methods are detailed below.
[0282] Determining the listener position and the listener angular orientation can enable several desirable features, such as orienting a located audio device with respect to the listener. Knowing the listener's position and angular orientation allows decisions to be made with respect to the listener, such as which speakers in the environment are in front, which are in the back, and which are (if any) near the center.
[0283] After correlating the audio device position with the listener's position and direction, some implementations may involve providing the audio device's position data, the audio device's angular orientation data, the listener's position data, and the listener's angular orientation data to an audio rendering system. Alternatively or additionally, some implementations may involve an audio data rendering process that is at least partially based on the audio device's position data, the audio device's angular orientation data, the listener position data, and the listener angular orientation data.
[0284] FIG. 29 is a flowchart outlining an example of a method that can be performed by an apparatus as shown in FIG. 1A. The blocks of method 2900 are not necessarily performed in the order shown, similar to other methods described herein. Further, such a method may include more or fewer blocks than those shown and / or described. In this example, the blocks of method 2900 may be (or may include) performed by a control system 110 as shown in FIG. 1A. As described above, in some implementations, control system 110 may be present within a single device, and in other implementations, control system 110 may be present within two or more devices.
[0285] In this example, block 2905 involves obtaining direction-of-arrival (DOA) data for each of a plurality of audio devices in the environment. In some examples, the plurality of audio devices may include all of the audio devices in the environment, such as all of audio devices 1905 shown in FIG. 27.
[0286] However, in some cases, the plurality of audio devices may include only a subset of all of the audio devices in the environment. For example, the plurality of audio devices may include all of the smart speakers in the environment, but may not include one or more of the other audio devices in the environment.
[0287] DOA data can be obtained in various ways depending on the specific implementation. In some cases, determining the DOA data may include determining the DOA data for at least one of the plurality of audio devices. In some examples, the DOA data may be obtained by controlling each loudspeaker of a plurality of loudspeakers in the environment to reproduce a test signal. For example, determining the DOA data may involve receiving microphone data from each microphone of a plurality of audio device microphones corresponding to a single audio device of the plurality of audio devices, and at least partially determining the DOA data for the single audio device based on the microphone data. Alternatively or additionally, determining the DOA data may involve receiving antenna data from one or more antennas corresponding to a single audio device of the plurality of audio devices, and at least partially determining the DOA data for the single audio device based on the antenna data.
[0288] In some such examples, a single audio device itself may determine the DOA data. According to some such implementations, each audio device of the plurality of audio devices may determine its own DOA data. However, in other implementations, other devices, which may be local or remote devices, may determine the DOA data for one or more audio devices in the environment. According to some implementations, a server may determine the DOA data for one or more audio devices in the environment.
[0289] According to the example shown in FIG. 29, block 2910 includes generating audio device position data based at least in part on the DOA data via a control system. In this example, the audio device position data includes an estimated value of the audio device position for each audio device referenced at block 2905.
[0290] The audio device position data may be (or may include) coordinates in a coordinate system such as, for example, a Cartesian coordinate system, a spherical coordinate system, or a cylindrical coordinate system. This coordinate system may be referred to herein as the audio device coordinate system. In some such examples, the audio device coordinate system may be oriented with reference to one of the audio devices in the environment. In other examples, the audio device coordinate system may be oriented with reference to an axis defined by a line between two audio devices in the environment. However, in other examples, the audio device coordinate system may be oriented with reference to another part of the environment, such as a television, a wall of a room, etc.
[0291] In some examples, block 2910 may be involved in the process described above with reference to FIG. 22. According to some such examples, block 2910 may be involved in determining the interior angles for each of a plurality of triangles based on the DOA data. In some cases, each triangle of the plurality of triangles may have vertices corresponding to the audio device positions of three audio devices. Some such methods may be involved, at least in part, in determining the side lengths for each side of each triangle based on the interior angles.
[0292] Some such methods may be involved in performing an alignment process to align each of the plurality of triangles in a first sequence to generate an alignment matrix. Some such methods may be involved in performing a reverse alignment process to align each of the plurality of triangles in a second sequence that is the reverse direction of the first sequence to generate a reverse alignment matrix. Some such methods may be involved, at least in part, in generating a final estimated value of the position of each audio device based on the values of the alignment matrix and the reverse alignment matrix. However, in some implementations of method 2900, block 2910 may be involved in applying a method other than that described above with reference to FIG. 22.
[0293] In this example, block 2915 is involved in determining listener position data indicating a listener's position in the environment via a control system. The listener position data may be, for example, position data referenced to an audio device coordinate system. However, in other examples, the coordinate system may be oriented with reference to the listener or with reference to a part of the environment such as a television, a wall of a room, etc.
[0294] In some examples, block 2915 may include prompting the listener to make one or more utterances (e.g., via an audio prompt from one or more loudspeakers in the environment) and estimating the listener's position according to the DOA data. The DOA data may correspond to microphone data acquired by a plurality of microphones in the environment. The microphone data may correspond to the detection of the one or more utterances by the microphones. At least some of the microphones may be co-located with the loudspeakers. According to some examples, block 2915 may include a triangulation process. For example, block 2915 may include triangulating the user's voice by finding the intersection of DOA vectors passing through the audio device, as will be described later with reference to FIG. 30A. According to some implementations, block 2915 (or another operation of method 2900) may include co-locating the origins of the audio device coordinate system and the listener coordinate system after the listener position has been determined. Co-locating the origins of the audio device coordinate system and the listener coordinate system may include transforming the audio device position from the audio device coordinate system to the listener coordinate system.
[0295] According to this implementation, block 2920 includes determining listener angle orientation data indicating listener angle orientation via a control system. The listener angle orientation data may be created with reference to a coordinate system used to represent listener position data, such as, for example, an audio device coordinate system. In some such examples, the listener angle orientation data may be created with reference to the origin and / or axes of the audio device coordinate system.
[0296] However, in some implementations, the listener angle orientation data may be created with reference to an axis defined by the listener position and another point in the environment, such as a television, an audio device, a wall, etc. In some such implementations, the listener position may be used to define the origin of the listener coordinate system. In some such examples, the listener angle orientation data may be created with reference to the axes of the listener coordinate system.
[0297] Various methods for implementing block 2920 are disclosed herein. According to some examples, the listener angle orientation may correspond to the listener's observation direction. In some such examples, for instance, by assuming that the listener is looking at a specific object such as a television, the listener's observation direction may be estimated with reference to the listener position data. In some such implementations, the listener's observation direction may be determined according to the listener position and the television position. Alternatively or additionally, the listener's observation direction may be determined according to the listener position and the position of the sound bar of the television.
[0298] However, in some examples, the listener's viewing direction may be determined according to listener input. According to some such examples, the listener input may include inertial sensor data received from a device held by the listener. The listener may use the device to point to a location in the environment, e.g., a location corresponding to the direction the listener is facing. For example, the listener may use the device to point to a loudspeaker (a speaker that is playing sound). Thus, in such examples, the inertial sensor data may include inertial sensor data corresponding to the loudspeaker.
[0299] In some such instances, the listener input may include an indication of an audio device selected by the listener, which may, in some examples, include inertial sensor data corresponding to the selected audio device.
[0300] However, in other examples, the audio device instructions may be made in accordance with one or more listener utterances (e.g., "The TV is now in front of me," "Speaker 2 is now in front of me," etc.) Other examples of determining listener angular orientation data in response to one or more listener utterances are described below.
[0301] 29, block 2925 includes determining, via the control system, audio device angular orientation data indicating an audio device angular orientation for each audio device relative to the listener position and listener angular orientation. According to some such examples, block 2925 may include rotating audio device coordinates around a point defined by the listener position. In some implementations, block 2925 may include transforming the audio device position data from the audio device coordinate system to the listener coordinate system. Some examples are described below.
[0302] Figure 30A shows examples of some of the blocks of Figure 29. According to some such examples, the audio device position data includes estimated values of the audio device positions for each of audio devices 1 to 5 with reference to the audio device coordinate system 3007. In this implementation, the audio device coordinate system 3007 is a Cartesian coordinate system with the position of the microphone of audio device 2 as the origin. Here, the x-axis of the audio device coordinate system 3007 corresponds to line 3003 between the position of the microphone of audio device 2 and the position of the microphone of audio device 1.
[0303] In this example, the listener position is determined by prompting the listener 3005, shown as sitting on the couch 1903, to make one or more utterances (e.g., via an audio prompt from one or more loudspeakers in the environment 3000a) and estimating the listener position according to the time-of-arrival (TOA) data. The TOA data corresponds to microphone data acquired by a plurality of microphones in the environment. In this example, the microphone data corresponds to the detection of the one or more utterances 3027 by at least some (e.g., 3, 4, or all 5) of the microphones of audio devices 1 to 5.
[0304] Alternatively or additionally, the listener position according to the DOA data provided by at least some (e.g., 2, 3, 4, or all 5) of the microphones of audio devices 1 to 5. According to some such examples, the listener position may be determined according to the intersection of lines 3009a, 3009b, etc. corresponding to the DOA data.
[0305] According to this example, the listener position corresponds to the origin of the listener coordinate system 3020. In this example, the listener angular orientation data is indicated by the y'-axis of the listener coordinate system 3020 corresponding to a straight line 3013a between the listener's head 3010 (and / or the listener's nose 3025) and the sound bar 3030 of the television 101. In the example shown in FIG. 30A, the line 3013a is parallel to the y'-axis. Thus, the angle Θ represents the angle between the y-axis and the y'-axis. In this example, block 2925 of FIG. 29 may include a rotation by the angle Θ of the audio device coordinates about the origin of the listener coordinate system 3020. Thus, the origin of the audio device coordinate system 3007 is shown to correspond to the audio device 2 in FIG. 30A, but some implementations include co-locating the origin of the audio device coordinate system 3007 with the origin of the listener coordinate system 3020 prior to the rotation of the audio device coordinates by the angle Θ about the origin of the listener coordinate system 3020. This co-location may be performed by a coordinate transformation from the audio device coordinate system 3007 to the listener coordinate system 3020.
[0306] In some examples, the position of the sound bar 3030 and / or the television 1901 may be determined by causing the sound bar to emit sound and estimating the position of the sound bar according to DOA and / or TOA data, where the DOA and / or TOA data may correspond to the detection of the sound by at least some (e.g., 3, 4, or all 5) of the microphones of the audio devices 1-5. Alternatively or additionally, the position of the sound bar 3030 and / or the television 1901 may be determined by prompting the user to walk to the location of the television and localizing the user's speech by DOA and / or TOA data, where the DOA and / or TOA data may correspond to the detection of the sound by at least some (e.g., 3, 4, or all 5) of the microphones of the audio devices 1-5. Such methods may involve triangulation. Such examples may be beneficial in situations where the sound bar 3030 and / or the television 1901 do not have associated microphones.
[0307] In some other examples where the soundbar 3030 and / or the television 1901 has an associated microphone, the position of the soundbar 3030 and / or the television 1901 can be determined according to a TOA or DOA method such as the DOA method disclosed herein. According to some such methods, the microphone may be co-located with the soundbar 3030.
[0308] According to some implementations, the soundbar 3030 and / or the television 1901 may have an associated camera 3011. The control system may be configured to capture an image of the listener's head 3010 (and / or the listener's nose 3025). In some such examples, the control system may be configured to determine a straight line 3013a between the listener's head 3010 (and / or the listener's nose 3025) and the camera 3011. The listener angle orientation data may correspond to the straight line 3013a. Alternatively or additionally, the control system may be configured to determine an angle Θ between the straight line 3013a and the y-axis of the audio device coordinate system.
[0309] Figure 30B shows a further example of determining listener angle orientation data. According to this example, the listener position has already been determined at block 2915 of FIG. 29. Here, the control system is controlling the speakers in the environment 3000b to render an audio object 3035 at various positions within the environment 3000b. In some such examples, the control system can cause the speakers to render the audio object 3035 in such a way that the audio object 3035 is felt to rotate around the listener 3005. Such rendering can be, for example, by rendering the audio object 3035 such that the audio object 3035 is felt to rotate around the origin of the listener coordinate system 3020. In this example, the curved arrow 3040 shows a part of the trajectory of the audio object 3035 when the audio object 3035 rotates around the listener 3005.
[0310] According to some such examples, the listener 3005 can provide a user input indicating when the audio object 3035 is in the direction the listener 3005 is facing (e.g., by saying "stop"). In some such examples, the control system may be configured to determine a straight line 3013b between the listener position and the position of the audio object 3035. In this example, the straight line 3013b corresponds to the y'-axis of the listener coordinate system indicating the direction the listener 3005 is facing. In an alternative implementation, the listener 3005 may provide a user input indicating when the audio object 3035 is in the front of the environment, at the TV position of the environment, at the audio device position, etc.
[0311] FIG. 30C shows a further example of determining listener angular orientation data. According to this example, the listener position has already been determined at block 2915 of FIG. 29. Here, the listener 3005 is using the handheld device 3045 to provide an input regarding the viewing direction of the listener 3005 by pointing the handheld device 3045 towards the TV 1901 or the soundbar 3030. The dashed outlines of the handheld device 3045 and the listener's arm indicate that the listener 3005 pointed the handheld device 3045 towards the audio device 2 in this example, at a time point before the listener 3005 pointed the handheld device 3045 towards the TV 1901 or the soundbar 3030. In other examples, the listener 3005 may point the handheld device 3045 towards another audio device such as the audio device 1. According to this example, the handheld device 3045 is configured to determine an angle α between the audio device 2 and the TV 1901 or the soundbar 3030. The angle α approximates the angle between the audio device 2 and the viewing direction of the listener 3005. The handheld device 3045 may be, in some examples, a cellular phone that includes an inertial sensor system and a wireless interface configured to communicate with a control system that controls an audio device of the environment 3000c. In some examples, the handheld device 3045 may execute an application or "app" configured to control the handheld device 3045 to perform the necessary functions. Performing the functions may include, for example, providing a user prompt (e.g., via a graphical user interface) to receive an input indicating that the handheld device 3045 is pointing in a desired direction, saving corresponding inertial sensor data, and / or transmitting the corresponding inertial sensor data to a control system that controls an audio device of the environment 3000c.
[0312] According to this example, a control system (which may be the control system of the handheld device 3045 or the control system that controls the audio device of the environment 3000c) is configured to determine the orientation of the lines 3013c and 3050 according to inertial sensor data, such as gyroscope data. In this example, the line 3013c is parallel to the axis y' and can be used to determine the listener angle orientation. According to some examples, the control system can determine an appropriate rotation about the origin of the listener coordinate system 3020 for the audio device coordinates according to the angle α between the audio device 2 and the observation direction of the listener 3005.
[0313] FIG. 30D shows an example of determining an appropriate rotation for audio device coordinates according to the method described with reference to FIG. 30C. In this example, the origin of the audio device coordinate system 3007 is co-located with the origin of the listener coordinate system 3020. Co-locating the origin of the audio device coordinate system 3007 and the listener coordinate system 3020 becomes possible after the process 2915 in which the listener position is determined. Co-locating the origin of the audio device coordinate system 3007 and the listener coordinate system 3020 may involve converting the audio device position from the audio device coordinate system 3007 to the listener coordinate system 3020. The angle α is determined as described above with reference to FIG. 30C. Thus, the angle α corresponds to the desired orientation of the audio device 2 in the listener coordinate system 3020. In this example, the angle β corresponds to the orientation of the audio device 2 in the audio device coordinate system 3007. In this example, the angle θ, which is β - α, indicates the rotation necessary to align the y-axis of the audio device coordinate system 3007 with the y'-axis of the listener coordinate system 3020.
[0314] In some implementations, the method of FIG. 29 may include controlling at least one of the audio devices in the environment, at least in part, based on corresponding audio device positions, corresponding audio device angular orientations, listener position data, and listener angular orientation data.
[0315] For example, some implementations can include providing audio device location data, audio device angular orientation data, listener location data, and listener angular orientation data to an audio rendering system. In some examples, the audio rendering system may be implemented by a control system such as control system 110 of FIG. 1A. Some implementations may include controlling an audio data rendering process based at least in part on the audio device location data, audio device angular orientation data, listener location data, and listener angular orientation data. Some such implementations may include providing loudspeaker acoustic capability data to the rendering system. The loudspeaker acoustic capability data may correspond to one or more loudspeakers in the environment. The loudspeaker acoustic capability data may indicate the orientation of one or more drivers, the number of drivers, or the driver frequency response of one or more drivers. In some examples, the loudspeaker acoustic capability data may be retrieved from memory and then provided to the rendering system.
[0316] One class of embodiments includes rendering and / or a method of playing audio for playback by at least one (e.g., all or part) of a plurality of coordinated (orchestrated) smart audio devices. For example, a collection of smart audio devices present in a user's home (within the system) can be orchestrated to handle a variety of simultaneous use cases, including flexible rendering of audio for playback by all or part of the smart audio devices (i.e., all or part of the speakers (singular or plural)). Many interactions with the system that require dynamic modification to the rendering and / or playback are envisioned. Such modifications may focus on spatial fidelity, but not necessarily so.
[0317] In the context of performing spatial audio mixing rendering (or rendering and playback) (e.g., rendering of an audio stream or multiple audio streams) for playback by a smart audio device of a collection of smart audio devices (or by another set of speakers), the type of speaker (e.g., within a smart audio device or coupled to a smart audio device) can change, and thus the corresponding acoustic capabilities of the speaker can vary quite significantly. In an example of an audio environment shown in FIG. 3A, the loudspeakers 305d, 305f, and 305h may be smart speakers having a single 0.6-inch speaker. In this example, the loudspeakers 305b, 305c, 305e, and 305f may be smart speakers having a 2.5-inch woofer and a 0.8-inch tweeter. According to this example, the loudspeaker 305g may be a smart speaker equipped with a 5.25-inch woofer, three 2-inch midrange speakers, and a 1.0-inch tweeter. Here, the loudspeaker 305a may be a soundbar having sixteen 1.1-inch beam drivers and two 4-inch woofers. Thus, the low-frequency capabilities of the smart speakers 305d and 305f will likely be quite lower than other loudspeakers within the environment 200, particularly those loudspeakers having 4-inch or 5.25-inch woofers.
[0318] FIG. 31 is a block diagram showing an example of components of a system that can implement various aspects of the present disclosure. Similar to other figures provided herein, the types and numbers of elements shown in FIG. 31 are provided merely as examples. Other implementations may include more elements, fewer elements, and / or different types and numbers of elements.
[0319] According to this example, the system 3100 includes a smart home hub 3105 and loudspeakers 3125a - 3125m. In this example, the smart home hub 3105 is shown in FIG. 1A and includes an instance of the control system 110 described above. According to this implementation, the control system 110 includes an ambient dynamics processing configuration data module 3110, an ambient dynamics processing module 3115, and a rendering module 3120. Some examples of the ambient dynamics processing configuration data module 3110, the ambient dynamics processing module 3115, and the rendering module 3120 will be described below. In some examples, the rendering module 3120' may be configured for both rendering and ambient dynamics processing.
[0320] As suggested by the arrows between the smart home hub 3105 and the loudspeakers 3125a - 3125m, the smart home hub 3105 also includes an instance of the interface system 105 shown in FIG. 1A and described above. According to some examples, the smart home hub 3105 may be part of the environment 300 shown in FIG. 3A. In some cases, the smart home hub 3105 may be implemented by a smart speaker, a smart TV, a cellular phone, a laptop, etc. In some implementations, the smart home hub 3105 may be implemented by software, for example, via a downloadable software application or "app" software. In some cases, the smart home hub 3105 may be implemented in each of the loudspeakers 3125a - m, all operating in parallel to generate the same processed audio signal from the module 3120. According to some such examples, in each loudspeaker, the rendering module 3120 may then generate one or more speaker feeds associated with each loudspeaker or group of loudspeakers, and provide these speaker feeds to each speaker dynamics processing module.
[0321] In some cases, loudspeakers 3125a - 3125m may include loudspeakers 305a - 305h of FIG. 3A. In other examples, loudspeakers 3125a - 3125m may be other loudspeakers or may include other loudspeakers. Thus, in this example, system 3100 includes M loudspeakers, where M is an integer greater than 2.
[0322] Similar to many other powered speakers, smart speakers typically use some type of internal dynamics processing to prevent the speaker from distorting. Such dynamics processing often involves a signal limiting threshold (e.g., a threshold that is variable across frequencies), and the signal level is dynamically held below it. For example, Dolby's audio regulator, which is one of several algorithms in the Dolby Audio Processing (DAP) audio post - processing suite, provides such processing. In some cases, although not typically, dynamics processing may also involve applying one or more compressors, gates, expanders, duckers, etc., via the dynamics processing module of the smart speaker.
[0323] Therefore, in this example, each of the loudspeakers 3125a to 3125m includes a corresponding speaker dynamics processing (DP) module A to M. The speaker dynamics processing module is configured to apply loudspeaker dynamics processing configuration data for each individual loudspeaker in the listening environment. The speaker DP module A, for example, is configured to apply individual loudspeaker dynamics processing configuration data suitable for the loudspeaker 3125a. In some examples, the individual loudspeaker dynamics processing configuration data may correspond to one or more capabilities of the individual loudspeaker. For example, within a specific frequency range, it is the ability of the loudspeaker to reproduce a specific level of audio data without recognizable distortion.
[0324] When spatial audio is rendered across a set of heterogeneous speakers (e.g., speakers of a smart audio device or speakers coupled to a smart audio device) that each potentially have different reproduction limits, care is needed when performing dynamics processing on the overall mix. A simple solution is to render the spatial mix to the speaker feeds of each participating speaker and then allow the dynamics processing module associated with each speaker to act independently on its corresponding speaker feed according to the limits of that speaker.
[0325] This approach keeps each speaker from being distorted, but may dynamically shift the spatial balance of the mix in a perceptually distracting way. For example, referring to FIG. 3A, assume that a television program is shown on television 330 and the corresponding audio is being played by the loudspeakers in environment 300. During the television program, assume that the audio associated with stationary objects (such as a factory heavy machinery unit) is intended to be rendered at a specific location within environment 300. Further, assume that loudspeaker 305b has substantially greater ability to play bass-range sounds, and thus a dynamics processing module associated with loudspeaker 305d substantially reduces the level of bass-range audio below that of a dynamics processing module associated with loudspeaker 305b. When the volume of the signal associated with the stationary object fluctuates, as the volume increases, the dynamics processing module associated with loudspeaker 305d substantially reduces the level of bass-range audio more than the level of the same audio is reduced by the dynamics processing module associated with loudspeaker 305b. This difference in levels changes the apparent position of the stationary object. Thus, an improved solution is needed.
[0326] Some embodiments of the present disclosure are systems and methods for rendering (or rendering and playing) a spatial audio mix (e.g., rendering an audio stream or multiple audio streams) for playback by at least one (e.g., all or part) of the smart audio devices of a set of smart audio devices (e.g., a set of coordinated smart audio devices) and / or at least one (e.g., all or part) of the speakers of another set of speakers. Some embodiments are methods (or systems) for such rendering (e.g., including generation of speaker feeds) and playback of the rendered audio (e.g., playback of the generated speaker feeds). Examples of such embodiments are as follows.
[0327] Systems and methods for audio processing may include rendering audio (e.g., rendering a spatial audio mix by rendering a stream of audio or multiple streams of audio) for playback over at least two speakers (e.g., all or some of the speakers of a collection of speakers), including by: (a) combining individual loudspeaker dynamics processing configuration data (e.g., limiting thresholds for individual loudspeakers) to thereby determine listening environment dynamics processing configuration data (e.g., combined thresholds) for multiple loudspeakers; (b) performing dynamics processing on the audio (e.g., a stream of audio representing a spatial audio mix) using listening environment dynamics processing configuration data (e.g., combined thresholds) for the multiple loudspeakers to generate processed audio; (c) Rendering the processed audio to the speaker feed.
[0328] According to some implementations, process (a) may be executed by a module such as the listening environment dynamics processing configuration data module 3110 shown in FIG. 31. The smart home hub 3105 may be configured to obtain individual loudspeaker dynamics processing configuration data for each of the M loudspeakers via an interface system. In this implementation, the individual loudspeaker dynamics processing configuration data includes an individual loudspeaker dynamics processing configuration data set for each of the plurality of loudspeakers. According to some examples, the individual loudspeaker dynamics processing configuration data for one or more loudspeakers may correspond to one or more capabilities of the one or more loudspeakers. In this example, each of the individual loudspeaker dynamics processing configuration data sets includes at least one type of dynamics processing configuration data. In some examples, the smart home hub 3105 may be configured to obtain an individual loudspeaker dynamics processing configuration data set by querying each of the loudspeakers 3125a-3125m. In other implementations, the smart home hub 3105 may be configured to obtain an individual loudspeaker dynamics processing configuration data set by querying a data structure of the previously obtained individual loudspeaker dynamics processing configuration data set stored in the memory.
[0329] In some examples, process (b) may be executed by a module such as the listening environment dynamics processing module 3115 in FIG. 31. Some detailed examples of processes (a) and (b) are described below.
[0330] In some examples, the rendering of process (c) may be executed by a module such as the rendering module 3120 or the rendering module 3120' in FIG. 31. In some embodiments, the audio processing involves the following: (d) Executing dynamics processing on the rendered audio signal according to the individual loudspeaker-dynamics processing setting data for each loudspeaker (for example, restricting the speaker feed according to a playback limit threshold associated with the corresponding speaker, thereby generating a restricted speaker feed). Process (d) may be executed, for example, by the dynamics processing modules A to M shown in FIG. 31.
[0331] The speaker may be at least one (for example, all or part) of (or coupled to) the smart audio devices of the set of smart audio devices. In some implementations, in order to generate the restricted speaker feed in step (d), the speaker feed generated in step (c) is processed by a second stage of dynamics processing (for example, by the associated dynamics processing system of each speaker) to, for example, generate the speaker feed prior to final playback through the speaker. For example, the speaker feed (or a subset or part thereof) is for each different one of the speakers a dynamics processing system (for example, a dynamics processing subsystem of the smart audio device, where the smart audio device includes or is coupled to the associated ones of those speakers). The processed audio output from each of the dynamics processing systems may be used to generate a speaker feed for the associated ones of the speakers. Following the speaker-specific dynamics processing (i.e., the dynamics processing executed independently for each speaker), the processed (for example, dynamically restricted) speaker feed may be used to drive the speaker to cause playback of the audio.
[0332] The first stage of the dynamics processing (step (b)) can be designed to reduce a perceptually disturbing shift in the spatial balance that would occur if the dynamics-processed (e.g., limited) speaker feeds resulting from step (d) were generated in response to the original audio (rather than in response to the processed audio generated in step (b)). This can prevent an unwanted shift in the spatial balance of the mix. The second stage of the dynamics processing, which acts on the rendered speaker feeds from step (c), may be designed to ensure that no speaker is distorted. This is because the dynamics processing in step (b) may not necessarily guarantee that the signal level has dropped below the threshold for all speakers. Combining individual loudspeaker dynamics processing configuration data (e.g., the combination of thresholds in the first stage (step (a))) involves, in some examples, averaging the individual loudspeaker dynamics processing configuration data (e.g., the limiting threshold) across speakers (e.g., across a smart audio device), or taking the minimum of the individual loudspeaker dynamics processing configuration data (e.g., the limiting threshold) across speakers (e.g., across a smart audio device).
[0333] In some implementations, when the first stage of dynamics processing (step (b)) acts on audio that exhibits spatial mixing (e.g., object-based audio programs that include at least one object channel and optionally at least one speaker channel), this first stage can be implemented according to techniques for audio object processing through the use of spatial zones. In such cases, the combined individual loudspeaker dynamics processing configuration data (e.g., combined limiting thresholds) associated with each zone may be derived by (or as) a weighted average of the individual loudspeaker dynamics processing configuration data (e.g., individual speaker limiting thresholds), where this weighting may be given or determined at least in part by the spatial proximity of each speaker to the zone and / or its position within the zone.
[0334] In one exemplary embodiment, a plurality M of speakers (M≥2) are assumed, where each speaker is indexed by a variable i. Each speaker i is associated with a reproduction limiting threshold T i [f] that varies with frequency. Here, the variable f represents an index to a finite set of frequencies at which the threshold is specified. (Note that if the size of the set of frequencies is 1, the corresponding single threshold is considered broadband and applied across the entire frequency range.) These thresholds are utilized by each speaker in its own independent dynamics processing function to limit the audio signal below the threshold for a particular purpose. The particular purpose may be to prevent the speaker from distorting or to prevent the speaker from reproducing above some level that is considered undesirable in its vicinity.
[0335] Figures 32A, 32B, and 32C show examples of reproduction limit thresholds and corresponding frequencies. The frequency ranges shown can span, for example, the audible frequency range for an average human (e.g., 20 Hz to 20 kHz). In these examples, the reproduction limit threshold is indicated by the vertical axis of graphs 3200a, 3200b, and 3200c, which is labeled "Level Threshold" in these examples. The reproduction limit / level threshold increases in the direction of the arrow on the vertical axis. The reproduction limit / level threshold can be expressed, for example, in decibels. In these examples, the horizontal axes of graphs 3200a, 3200b, and 3200c indicate frequency, which increases in the direction of the arrow on the horizontal axis. The reproduction limit thresholds shown by curves 3200a, 3200b, and 3200c can be implemented, for example, by the dynamics processing module of an individual loudspeaker.
[0336] Graph 3200a of FIG. 32A shows a first example of a reproduction limit threshold as a function of frequency. Curve 3205a indicates the reproduction limit threshold for each corresponding frequency value. In this example, at base frequency f b where the input audio received at input level T i is output by the dynamics processing module at output level T o . The base frequency f b can be, for example, in the range of 60 to 250 Hz. However, in this example, at high frequency f t where the input audio received at input level T i is output by the dynamics processing module at the same level of input level T i . The high frequency f t can be, for example, in the range above 1280 Hz. Thus, in this example, curve 3205a corresponds to a dynamics processing module that applies a significantly lower threshold for the base frequency than for the high frequency. Such a dynamics processing module may be suitable for a loudspeaker without a woofer (e.g., loudspeaker 305d of FIG. 3A).
[0337] Graph 3200b of FIG. 32B shows a second example of a reproduction limit threshold as a function of frequency. Curve 3205b is the same base frequency f shown in FIG. 32A b at which the input audio received at input level T i is output by the dynamics processing module at a higher output level T o . Thus, in this example, curve 3205b corresponds to a dynamics processing module that does not apply a threshold for base frequencies as low as curve 3205a. Such a dynamics processing module is suitable for speakers having at least small woofers (e.g., speaker 305b of FIG. 3A).
[0338] Graph 3200c of FIG. 32C shows a second example of a reproduction limit threshold as a function of frequency. Curve 3205c (which is a straight line in this example) is the same base frequency f shown in FIG. 32A b at which the input audio received at input level T i is output by the dynamics processing module at the same level. Thus, in this example, curve 3205c corresponds to a dynamics processing module that may be appropriate for a loudspeaker capable of reproducing a wide range of frequencies including the base frequency. For simplicity, it can be seen that the dynamics processing module can approximate curve 3205c by implementing curve 3205d that applies the same threshold for all frequencies shown.
[0339] Spatial audio mixing can be rendered for multiple speakers using a known rendering system such as Center of Mass Amplitude Panning (CMAP) or Flexible Virtualization (FV). From the components of the spatial audio mix, the rendering system generates one speaker feed for each of the multiple speakers. In some previous examples, the speaker feeds were then passed through the associated dynamics processing functions of each speaker with a threshold Ti It is processed independently using [f]. Without the benefits of the present disclosure, this described rendering scenario may cause an annoying shift in the perceived spatial balance of the rendered spatial audio mix. For example, one of the M speakers, such as on the right side of the listening area, may be much less capable than the other speakers (e.g., the ability to render bass-range audio), and thus the threshold for that speaker may be significantly lower than the thresholds of the other speakers, at least in certain frequency ranges. During playback, the dynamics processing module for this speaker will significantly reduce the level of the components of the right-side spatial mix compared to the left-side components. Listeners are very sensitive to such dynamic shifts between the left-right balance of the spatial mix and may find the result very annoying.
[0340] To address this problem, in some examples, the individual loudspeaker dynamics processing configuration data (e.g., playback limit thresholds) of the individual speakers in the listening environment are combined to create listening environment dynamics processing configuration data for all the loudspeakers in the listening environment. Then, the listening environment dynamics processing configuration data can be used to perform dynamics processing in the context of the entire spatial audio mix before rendering to the speaker feeds. This first stage of dynamics processing has access to the entire spatial mix rather than just one independent speaker feed, so the processing can be performed in a way that does not impart an annoying shift to the perceived spatial balance of the mix. The individual loudspeaker dynamics processing configuration data (e.g., playback limit thresholds) may be combined in a way that eliminates or reduces the amount of dynamics processing performed by any of the individual speaker's independent dynamics processing functions.
[0341] In one example of determining listening environment dynamics processing configuration data, individual loudspeaker dynamics processing configuration data (e.g., playback limit thresholds) for individual speakers are combined into a single set of listening environment dynamics processing configuration data (e.g., frequency-varying playback limit thresholds) that are applied to all components of the spatial mix in the first stage of dynamics processing.
Number
Number
[0342] Such a combination essentially eliminates the operation of the individual dynamics processing for each speaker. This is because the spatial mix is initially restricted to be below the threshold of the least capable speaker at all frequencies. However, such a strategy can be overly aggressive. Many speakers play at levels lower than they can handle, and the combined playback level of all speakers may be undesirably low. For example, if the threshold in the bass range shown in Figure 32A were applied to the loudspeaker corresponding to the threshold for Figure 32C, the playback level of the latter speaker would be unacceptably low in the bass range. An alternative combination for determining listening environment dynamics processing configuration data is to take the average of the individual loudspeaker dynamics processing configuration data over all speakers in the listening environment. For example, in the context of playback limit thresholds, the average can be determined as follows:
Number
[0343] In this combination, since the first stage of the dynamics processing is limited to a higher level, the overall playback level may increase compared to taking the minimum, thereby enabling a more capable speaker to play at a higher volume. For speakers whose individual limit thresholds are below the average value, their independent dynamics processing function can limit the feed of the associated speaker if necessary. However, since some initial limitations are being applied to the spatial mix in the first stage of the dynamics processing, this may reduce the requirement for this limitation.
[0344] According to some examples of determining the listening environment dynamics processing configuration data, an adjustable combination can be generated that interpolates between the minimum and the average of the individual loudspeaker dynamics processing configuration data through tuning parameters. For example, in the context of the playback limit threshold, the interpolation can be determined as follows:
Equation
[0345] Other combinations of the individual loudspeaker dynamics processing configuration data are possible, and the present disclosure is intended to cover all such combinations.
[0346] Figures 33A and 33B are graphs showing examples of dynamic range compression data. In graphs 3300a and 3300b, the input signal level in decibels is shown on the horizontal axis and the output signal level in decibels is shown on the vertical axis. Similar to the other disclosed examples, specific thresholds, ratios, and other values are shown merely as examples and are not limiting.
[0347] In the example shown in FIG. 33A, the output signal level is equal to the input signal level below the threshold, which is -10 dB in this example. Other examples may relate to different thresholds, such as -20 dB, -18 dB, -16 dB, -14 dB, -12 dB, -8 dB, -6 dB, -4 dB, -2 dB, 0 dB, 2 dB, 4 dB, 6 dB, etc. Above the threshold, various examples of compression ratios are shown. A ratio of N:1 means that above the threshold, the output signal level increases by 1 dB for every N dB increase in the input signal. For example, a compression ratio of 10:1 (line 3305e) means that above the threshold, the output signal level increases by 1 dB for every 10 dB increase in the input signal. A compression ratio of 1:1 (line 3305a) means that even above the threshold, the output signal level remains the same as the input signal level. Lines 3305b, 3305c, and 3305d correspond to compression ratios of 3:2, 2:1, and 5:1. Other implementations can provide different compression ratios, such as 2.5:1, 3:1, 3.5:1, 4:3, 4:1, etc.
[0348] FIG. 33B shows an example of a "knee", which controls how the compression ratio changes at or near the threshold, which is 0 dB in this example. According to this example, a compression curve with a "hard" knee consists of two straight line segments, namely a straight line segment 3310a up to the threshold and a straight line segment 3310b above the threshold. A hard knee is easier to implement but may cause artifacts.
[0349] FIG. 33B also shows an example of a "soft" knee. In this example, the soft knee spans 10 dB. According to this implementation, above and below the 10 dB span, the compression ratio of the compression curve with a soft knee is the same as that of the compression curve with a hard knee. Other implementations can provide various other shapes of "soft" knees, which may span more or fewer decibels and may also show different compression ratios across the span.
[0350] Other types of dynamic range compression data can include "attack" data and "release" data. Attack is the period during which the compressor decreases gain in response to, for example, an increased level in the input in order to reach the gain determined by the compression ratio. The attack time for a compressor generally ranges from 25 milliseconds to 500 milliseconds, although other attack times are also practical. Release is the period during which the compressor increases gain in order to reach the output gain determined by the compression ratio (or the input level if the input level drops below a threshold), for example, in response to a decreased input level. The release time can be, for example, in the range of 25 milliseconds to 2 seconds.
[0351] Thus, in some examples, the individual loudspeaker dynamics processing configuration data can include a dynamic range compression data set for each of a plurality of loudspeakers. The dynamic range compression data set can include threshold data, input / output ratio data, attack data, release data, and / or knee data. One or more of these types of individual loudspeaker dynamics processing configuration data can be combined to determine the listening environment dynamics processing configuration data. As described above with respect to combinations of playback limit thresholds, in some examples, the dynamic range compression data can be averaged to determine the listening environment dynamics processing configuration data. In some cases, the minimum or maximum value of the dynamic range compression data can be used to determine the listening environment dynamics processing configuration data (e.g., the maximum compression ratio). In other implementations, for example, an adjustable combination can be created that interpolates between the minimum and average of the dynamic range compression data for individual loudspeaker dynamics processing via tuning parameters such as those described above with reference to Equation (32).
[0352] In some of the examples described above, a single set of listening environment dynamics processing configuration data (e.g., the combined threshold
Number
Number
Number
[0353] To address such issues, some implementations allow for independent or partially independent dynamic processing for different "spatial zones" of the spatial mix. A spatial zone may be considered a subset of the spatial region where the entire spatial mix is rendered. Much of the following discussion provides examples of dynamic processing based on playback limit thresholds, but these concepts equally apply to other types of individual loudspeaker dynamic processing configuration data and listening environment dynamic processing configuration data.
[0354] Figure 34 shows an example of spatial zones of a listening environment. Figure 34 shows an example of the area of the spatial mix (represented by the entire square), which is subdivided into three spatial zones: front, center, and surround.
[0355] The spatial zones in Figure 34 are drawn with hard boundaries, but in reality, it can be beneficial to treat the transition from one spatial zone to another as continuous. For example, a component of the spatial mix located at the center of the left edge of the square may have half of its level assigned to the front zone and half assigned to the surround zone. The signal levels from each component of the spatial mix can be assigned and accumulated in this continuous manner to each spatial zone. Then, the dynamic processing function can act independently on each spatial zone with respect to the overall signal level assigned to it from the mix. For each component of the spatial mix, the results of the dynamic processing from each spatial zone (e.g., the time-varying gain for each frequency) can then be combined and applied to that component. In some examples, this combination of spatial zone results is different for each component and is a function of the assignment of that particular component to each zone. The end result is that components of the spatial mix with similar spatial zone assignments receive similar dynamic processing, but independence between spatial zones is allowed. Spatial zones can advantageously be selected to prevent unwanted spatial shifts such as left-right imbalance while allowing for spatially independent processing (e.g., to reduce other artifacts such as the spatial ducking described above).
[0356] Techniques for processing spatial mixing for each spatial zone can be advantageously used in the first stage of the dynamics processing of the present disclosure. For example, different combinations of individual loudspeaker dynamics processing configuration data (e.g., playback limit thresholds) across the speakers i may be calculated for each spatial zone. The set of combined zone thresholds is [Number] may be represented by, where the index j refers to one of the plurality of spatial zones. The dynamics processing module may operate independently on each spatial zone with its associated threshold [Number] and the results may be applied back to the components of the spatial mix according to the techniques described above.
[0357] Consider that a spatial signal composed of the sum of K individual component signals x k [t], each having a desired spatial position (possibly time-varying) associated therewith, is rendered. One specific way to implement zone processing involves calculating a time-varying pan gain α k [t] that describes how much each audio signal x kj [t] contributes to zone j as a function of the desired spatial position of the audio signal with respect to the position of the zone. These pan gains can advantageously be designed to follow a power conservation pan rule that requires the sum of the squares of the gains to be equal to 1. From these pan gains, the zone signal s j [t] can be calculated as the sum of the component signals weighted by those pan gains for that zone: [Number] Next, each zone signal is the zone threshold [Number] is independently processed by a dynamics processing function DP parameterized by , and a zone correction gain G that varies with frequency and time j is generated:
Number
Number
Number
[0358] The combination of individual loudspeaker dynamics processing configuration data (e.g., speaker playback limit thresholds) for each spatial zone can be performed in various ways. As an example, the spatial zone playback limit threshold
Number
Number
[0359] The same weighting function can also be applied to other types of individual loudspeaker dynamics processing configuration data. Advantageously, the combined individual loudspeaker dynamics processing configuration data (e.g., playback limit threshold) for a spatial zone may be biased towards the individual loudspeaker dynamics processing configuration data (e.g., playback limit threshold) of the speaker that most contributes to the playback component of the spatial mix associated with that spatial zone. This is the weight w according to the contribution of each speaker to rendering the component of the spatial mix associated with that zone for frequency f ij can be achieved by setting [f].
[0360] FIG. 35 shows an example of loudspeakers within the spatial zone of FIG. 34. FIG. 35 shows the same zone as FIG. 34, but with the positions of five exemplary loudspeakers (Speakers 1, 2, 3, 4, 5) that contribute to rendering the spatial mix overlaid. In this example, Speakers 1, 2, 3, 4, 5 are represented by diamonds. In this particular example, Speaker 1 is primarily responsible for rendering the central zone, Speakers 2 and 5 for the front zone, and Speakers 3 and 4 for the surround zone. Based on this conceptual one-to-one mapping of speakers to spatial zones, the weight w ij [f] can be generated, but as with the spatial zone-based processing of the spatial mix, a more continuous mapping may sometimes be preferred. For example, Speaker 4 is very close to the front zone, and the audio mix components located between Speakers 4 and 5 will likely be reproduced primarily by the combination of Speakers 4 and 5 (although conceptually in the front zone). Thus, it makes sense for the individual loudspeaker dynamics processing configuration data (e.g., playback limit threshold) of Speaker 4 to contribute to the combined individual loudspeaker dynamics processing configuration data (e.g., playback limit threshold) of the front zone as well as the surround zone.
[0361] One way to achieve this continuous mapping is to set the weight w ij [f] equal to the speaker participation value that describes the relative contribution of each speaker i when rendering the components associated with spatial zone j. Such a value may be directly derived from a rendering system (e.g., from step (c) above) responsible for rendering to the speakers and a set of one or more nominal spatial positions associated with each spatial zone. This set of nominal spatial positions may include a set of positions within each spatial zone.
[0362] Figure 36 shows an example of nominal spatial positions superimposed on the spatial zones and speakers of Figure 35. The nominal positions are indicated by numbered circles. That is, two positions located at the upper corners of the square are associated with the front zone, a single position at the center of the upper part of the square is associated with the center zone, and two positions located at the lower corners of the square are associated with the surround zone.
[0363] To calculate the speaker participation value for a spatial zone, each of the nominal positions associated with that zone may be rendered through a renderer to generate the speaker activation associated with that position. These activations may be, for example, gains for each speaker in the case of CMAP or complex number values at a given frequency for each speaker in the case of FV. Next, for each speaker and zone, these activations are accumulated over all the nominal positions associated with the spatial zone to generate a value g ij [f]. This value represents the total activation of speaker i for rendering the entire set of nominal positions associated with spatial zone j. Finally, the speaker participation value in the spatial zone may be calculated as the cumulative activation normalized by the sum of all these accumulated activations over all the speakers. Thereafter, the weight may be set to this speaker participation value:
Equation
[0364] According to some implementations, the above process for calculating the participation values of speakers and combining the thresholds as a function of these values may be performed as a static process. Here, the resulting combined threshold is calculated once during a setup procedure that determines the layout and capabilities of the speakers in the environment. In such a system, once set up, both the dynamic processing configuration data of the individual loudspeakers and the way in which the rendering algorithm activates the loudspeakers as a function of the desired audio signal positions may be assumed to remain static. However, in some systems, both of these aspects may change over time, for example in response to changes in conditions in the playback environment, and thus it may be desirable to update the combined threshold according to the above process either continuously or in an event-triggered manner to account for such variations.
[0365] Both the CMAP and FV rendering algorithms may be extended to adapt to one or more dynamically configurable functions in response to changes in the listening environment. For example, with respect to FIG. 35, a person located near speaker 3 can utter the wake word of the smart assistant associated with the speaker, thereby putting the system in a state ready to listen for subsequent commands from the person. While the wake word is being uttered, the system can use the microphone associated with the loudspeaker to determine the position of the person. Using this information, the system can then choose to divert the energy of the audio reproduced from speaker 3 to other speakers so that the microphone on speaker 3 can better hear the person. In such a scenario, speaker 2 in FIG. 35 may essentially "take over" the role of speaker 3 for a period of time, resulting in a significant change in the speaker participation value for the surround zone, a decrease in the participation value of speaker 3, and an increase in the participation value of speaker 2. Since the zone threshold depends on the changed speaker participation values, it may be recalculated subsequently. Alternatively or additionally to these changes to the rendering algorithm, the limit threshold of speaker 3 may be lowered below the nominal value set to prevent the speaker from distorting. This can prevent the remaining audio reproduced from speaker 3 from increasing beyond some threshold determined to cause interference to the microphone listening to the person. Since the zone threshold is also a function of the individual speaker thresholds, it may also be updated in this case.
[0366] FIG. 37 is a flow diagram outlining an example of a method that may be implemented by an apparatus or system such as those disclosed herein. The blocks of method 3700 are not necessarily executed in the order shown, similar to other methods described herein. In some implementations, one or more blocks of method 3700 may be executed simultaneously. Further, some implementations of method 3700 may include more or fewer blocks than those illustrated and / or described. The blocks of method 3700 may be executed by one or more devices that are (or may include) a control system such as control system 110 shown in and described above with respect to FIG. 1A, or an example of another disclosed control system.
[0367] According to this example, block 3705 involves obtaining, by a control system, individual loudspeaker dynamics processing configuration data for each of a plurality of loudspeakers in an acoustic environment via an interface system. In this implementation, the individual loudspeaker dynamics processing configuration data includes an individual loudspeaker dynamics processing configuration data set for each of the plurality of loudspeakers. According to some examples, the individual loudspeaker dynamics processing configuration data for one or more loudspeakers may correspond to one or more capabilities of the one or more loudspeakers. In this example, each data set of the individual loudspeaker dynamics processing configuration data sets includes at least one type of dynamics processing configuration data.
[0368] In some cases, block 3705 may be involved in obtaining individual loudspeaker dynamics processing configuration data sets from each of a plurality of loudspeakers in the listening environment. In other examples, block 3705 may be involved in obtaining individual loudspeaker dynamics processing configuration data sets from a data structure stored in memory. For example, the individual loudspeaker dynamics processing configuration data sets may have been previously obtained as part of the setup procedure for each loudspeaker and stored in the data structure.
[0369] According to some examples, the individual loudspeaker dynamics processing configuration data sets may be proprietary. In some such examples, the individual loudspeaker dynamics processing configuration data sets may have been previously estimated based on individual loudspeaker dynamics processing configuration data for speakers having similar characteristics. For example, block 3705 may be involved in a speaker matching process that determines the most similar speaker from a data structure indicating a plurality of speakers and corresponding individual loudspeaker dynamics processing configuration data sets for each of the plurality of speakers. The speaker matching process may be based on, for example, a comparison of the size of one or more woofers, tweeters, and / or midrange speakers.
[0370] In this example, block 3710 is involved in determining, by a control system, listening environment dynamics processing configuration data for a plurality of loudspeakers. According to this implementation, the determination of the listening environment dynamics processing configuration data is based on individual loudspeaker dynamics processing configuration data sets for each of the plurality of loudspeakers. Determining the listening environment dynamics processing configuration data may involve combining the individual loudspeaker dynamics processing configuration data of the dynamics processing data sets, for example, by taking an average of one or more types of individual loudspeaker dynamics processing configuration data. In some cases, determining the listening environment dynamics processing configuration data may involve determining a minimum or maximum value of one or more types of individual loudspeaker dynamics processing configuration data. According to some such implementations, determining the listening environment dynamics processing configuration data may involve interpolating between a minimum or maximum value and an average value of one or more types of individual loudspeaker dynamics processing configuration data.
[0371] In this implementation, block 3715 is involved in receiving, by a control system, via an interface system, audio data including one or more audio signals and associated spatial data. For example, the spatial data may indicate an intended perceived spatial position corresponding to the audio signal. In this example, the spatial data includes channel data and / or spatial metadata.
[0372] In this example, block 3720 is involved in performing, by a control system, dynamics processing on the audio data based on the listening environment dynamics processing configuration data to generate processed audio data. The dynamics processing of block 3720 may involve any of the dynamics processing methods of the present disclosure disclosed herein, including but not limited to applying one or more playback limit thresholds, compressed data, and the like.
[0373] Here, block 3725 is involved in rendering processed audio data by a control system via a set of loudspeakers that includes at least some of a plurality of loudspeakers to generate a rendered audio signal. In some examples, block 3725 may be involved in applying a CMAP rendering process, an FV rendering process, or a combination of both. In this example, block 3720 is executed before block 3725. However, as described above, block 3720 and / or block 3710 may be at least partially based on the rendering process of block 3725. Blocks 3720 and 3725 may be involved in performing a process as described above with reference to the listening environment dynamics processing module and the rendering module 3120 of FIG. 31.
[0374] According to this example, block 3730 is involved in providing the rendered audio signal to a set of loudspeakers via an interface system. In one example, block 3730 may be involved in providing the rendered audio signal to loudspeakers 3125a - 3125m by a smart home hub 3105 via its interface system.
[0375] In some examples, method 3700 may be involved in performing dynamics processing on the rendered audio signal according to individual loudspeaker dynamics processing configuration data for each loudspeaker of the set of loudspeakers to which the rendered audio signal is provided. For example, referring again to FIG. 31, dynamics processing modules A - M can perform dynamics processing on the rendered audio signal according to individual loudspeaker dynamics processing configuration data for loudspeakers 3125a - 3125m.
[0376] In some implementations, the individual loudspeaker dynamics processing configuration data may include a playback limit threshold data set for each of a plurality of loudspeakers. In some such examples, the playback limit threshold data set may include playback limit thresholds for each of a plurality of frequencies.
[0377] Determining the listening environment dynamics processing configuration data may, in some cases, involve determining a minimum playback limit threshold across a plurality of loudspeakers. In some examples, determining the listening environment dynamics processing configuration data may involve averaging the playback limit thresholds to obtain an averaged playback limit threshold across a plurality of loudspeakers. In some such examples, determining the listening environment dynamics processing configuration data may involve determining a minimum playback limit threshold across a plurality of loudspeakers and interpolating between the minimum playback limit threshold and the averaged playback limit threshold.
[0378] According to some implementations, averaging the playback limit thresholds may involve determining a weighted average of the playback limit thresholds. In some such examples, the weighted average may be at least partially based on characteristics of the rendering process implemented by the control system, such as the characteristics of the rendering process of block 3725.
[0379] In some implementations, performing dynamics processing on audio data may be based on spatial zones. Each of the spatial zones may correspond to a subset of the listening environment.
[0380] According to some such implementations, the dynamics processing may be performed separately for each of the spatial zones. For example, determining the listening environment dynamics processing configuration data may be performed separately for each of the spatial zones. For example, combining the dynamics processing configuration data sets across multiple loudspeakers may be performed separately for each of one or more spatial zones. In some examples, combining the dynamics processing configuration data sets across multiple loudspeakers separately for each of one or more spatial zones may be at least partially based on the activation of the loudspeakers by a rendering process according to the desired audio signal positions across one or more spatial zones.
[0381] In some examples, combining the dynamics processing configuration data sets across multiple loudspeakers separately for each of one or more spatial zones may be at least partially based on the loudspeaker participation values for each loudspeaker in each of one or more spatial zones. Each loudspeaker participation value may be at least partially based on one or more nominal spatial positions within each of one or more spatial zones. The nominal spatial positions may, in some examples, correspond to the standard positions of the channels within a Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4 or Dolby 9.1 surround sound mix. In some such implementations, each loudspeaker participation value is at least partially based on the activation of each loudspeaker corresponding to the rendering of the audio data at each of one or more nominal spatial positions within each of one or more spatial zones.
[0382] According to some such examples, the weighted average of the playback restriction threshold may be at least partially based on the activation of the loudspeaker by the rendering process as a function of the proximity to the spatial zone of the audio signal. In some cases, the weighted average may be at least partially based on the loudspeaker participation value for each loudspeaker in each spatial zone. In some such examples, each loudspeaker participation value may be at least partially based on one or more nominal spatial positions within each spatial zone. For example, the nominal spatial positions may correspond to the standard positions of the channels in a Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, or Dolby 9.1 surround sound mix. In some implementations, each loudspeaker participation value may be at least partially based on the activation of each loudspeaker corresponding to the rendering of the audio data at each of one or more nominal spatial positions within each spatial zone.
[0383] According to some implementations, rendering the processed audio data may involve determining the relative activation of a set of loudspeakers according to one or more dynamically configurable functions. Some examples are described below with reference to FIG. 10 and below. The one or more dynamically configurable functions may be based on one or more attributes of the audio signal, one or more attributes of the set of loudspeakers, or one or more external inputs. For example, the one or more dynamically configurable functions may be the proximity of the loudspeakers to one or more listeners; the proximity of the loudspeakers to the gravitational position (gravity is a factor that favors relatively higher loudspeaker activation at closer proximity to the gravitational position); the proximity of the loudspeakers to the repulsive force position (repulsive force is a factor that favors relatively lower loudspeaker activation at closer proximity to the repulsive force position); the ability of each loudspeaker relative to other loudspeakers in the environment; the synchronization of the loudspeakers relative to other loudspeakers; the wake word performance; or the echo canceller performance.
[0384] In some examples, the relative activation of the speakers may be based on a cost function of a model of the perceived spatial position of the audio signal when played through the speakers, a measure of the proximity of the intended perceived spatial position of the audio signal to the speaker positions, and one or more dynamically configurable functions.
[0385] In some examples, minimization of the cost function (including at least one dynamic speaker activation term) can lead to deactivation of at least one of the speakers (in the sense that each such speaker does not play the associated audio content) and activation of at least one of the speakers (in the sense that each such speaker plays at least a portion of the rendered audio content). The dynamic speaker activation term(s) can enable at least one of a variety of behaviors, including distorting the spatial presentation of audio away from a particular smart audio device. Thereby, a microphone can better hear the voice of a speaker, or a secondary audio stream can better be heard from the speaker(s) of the smart audio device.
[0386] According to some implementations, the individual loudspeaker dynamics processing configuration data can include a dynamic range compression data set for each of a plurality of loudspeakers. In some cases, the dynamic range compression data set may include one or more of threshold data, input-output ratio data, attack data, release data, or knee data.
[0387] As described above, in some implementations, at least some of the blocks of method 3700 shown in FIG. 37 may be omitted. For example, in some implementations, blocks 3705 and 3710 are executed during the setup process. After the listening environment dynamics processing configuration data is determined, in some implementations, steps 3705 and 3710 are not re-executed during "runtime" operation unless the type and / or placement of the speakers in the listening environment changes. For example, in some implementations, there may be an initial check to determine whether any loudspeakers have been added or removed, or whether the position of any loudspeakers has changed, etc. If so, steps 3705 and 3710 may be performed. If not, steps 3705 and 3710 may not be re-executed before the "runtime" operations that may involve blocks 3715-3730.
[0388] FIGS. 38A, 38B, and 38C show examples of speaker participation values corresponding to the examples of FIGS. 2C and 2D. In FIGS. 38A, 38B, and 38C, the angle -4.1 corresponds to speaker position 272 in FIG. 2D, the angle 4.1 corresponds to speaker position 274 in FIG. 2D, the angle -87 corresponds to speaker position 267 in FIG. 2D, the angle 63.6 corresponds to speaker position 275 in FIG. 2D, and the angle 165.4 corresponds to speaker position 270 in FIG. 2D. These speaker participation values are examples of "weighting" with respect to the spatial zones described with reference to FIGS. 34-37. According to these examples, the loudspeaker participation values shown in FIGS. 38A, 38B, and 38C correspond to the participation of each loudspeaker in each of the spatial zones shown in FIG. 34: the loudspeaker participation values shown in FIG. 38A correspond to the participation of each loudspeaker in the central zone, the loudspeaker participation values shown in FIG. 38B correspond to the participation of each loudspeaker in the front left and right zones, and the loudspeaker participation values shown in FIG. 38C correspond to the participation of each loudspeaker in the rear zone.
[0389] Figures 39A, 39B, and 39C show examples of loudspeaker participation values corresponding to the examples of FIGS. 2F and 2G. In FIGS. 39A, 39B, and 39C, the angle -4.1 corresponds to the speaker position 272 in FIG. 2D, the angle 4.1 corresponds to the speaker position 274 in FIG. 2D, the angle -87 corresponds to the speaker position 267 in FIG. 2D, the angle 63.6 corresponds to the speaker position 275 in FIG. 2D, and the angle 165.4 corresponds to the speaker position 270 in FIG. 2D. According to these examples, the loudspeaker participation values shown in FIGS. 39A, 39B, and 39C correspond to the participation of each loudspeaker in each spatial zone shown in FIG. 34: the loudspeaker participation values shown in FIG. 39A correspond to the participation in the central zone of each loudspeaker, the loudspeaker participation values shown in FIG. 39B correspond to the participation in the front left and right zones of each loudspeaker, and the loudspeaker participation values shown in FIG. 39C correspond to the participation in the rear zone of each loudspeaker.
[0390] Figures 40A, 40B, and 40C show examples of loudspeaker participation values corresponding to the examples of FIGS. 2H and 2I. According to these examples, the loudspeaker participation values shown in FIGS. 40A, 40B, and 40C correspond to the participation of each loudspeaker in each spatial zone shown in FIG. 34. The loudspeaker participation values shown in FIG. 40A correspond to the participation of each loudspeaker in the central zone, the loudspeaker participation values shown in FIG. 40B correspond to the participation of each loudspeaker in the front left and right zones, and the loudspeaker participation values shown in FIG. 40C correspond to the participation of each loudspeaker in the rear zone.
[0391] Figures 41A, 41B, and 41C show examples of loudspeaker participation values corresponding to the examples of FIGS. 2J and 2K. According to these examples, the loudspeaker participation values shown in FIGS. 41A, 41B, and 41C correspond to the participation of each loudspeaker in each spatial zone shown in FIG. 34. The loudspeaker participation values shown in FIG. 41A correspond to the participation of each loudspeaker in the central zone, the loudspeaker participation values shown in FIG. 41B correspond to the participation of each loudspeaker in the front left and right zones, and the lo...
Claims
1. An audio processing method, comprising: receiving, by a control system, first audio data corresponding to a first audio program stream, the first audio data including one or more first audio signals and first spatial data indicating associated desired perceived spatial positions for each of the one or more first audio signals; rendering, by the control system, the first audio data to generate first rendered audio data for at least two loudspeakers of a set of loudspeakers in an environment, wherein the rendering includes determining relative activation of the loudspeakers of the set of loudspeakers based on a proximity of a perceived spatial position of the first audio signal reproduced through the loudspeakers to the desired perceived spatial position of the first audio signal at the positions of the loudspeakers, and at least one or more attributes of the first audio signal, one or more attributes of the set of loudspeakers, or one or more additional dynamically configurable functions that depend on one or more external inputs, wherein the one or more additional dynamically configurable functions include at least a proximity of one or more loudspeakers to an attracting position or a proximity of one or more loudspeakers to a repelling position, the attracting force being a factor that favors relatively higher loudspeaker activation closer to the attracting position, and the repelling force being a factor that favors relatively lower loudspeaker activation closer to the repelling position; A method.
2. The method according to claim 1, wherein the one or more additional dynamically configurable functions include a proximity of one or more loudspeakers to one or more listeners.
3. The method according to claim 1, wherein the one or more additional dynamically configurable functions include an audibility of one or more loudspeakers at a position in the environment.
4. 10. The method of claim 1, wherein the additional dynamically configurable features include at least one of the capabilities of one or more loudspeakers or synchronization of the one or more loudspeakers with one or more other loudspeakers in the environment.
5. The method of claim 1 , wherein the additional dynamically configurable functionality includes at least one of wake word capability or echo canceller capability.
6. The method of claim 1 , wherein the rendering comprises minimizing a cost function, the cost function comprising at least one dynamic speaker activation term.
7. 7. The method of claim 6, wherein the cost function is based at least in part on a sum of a term corresponding to spatial audio and a term corresponding to a proximity of a loudspeaker to a desired perceived spatial location for one of the one or more first audio signals.
8. The method of claim 6 , wherein the cost function is based at least in part on a Center of Mass Amplitude Pan (CMAP), a Flexible Virtualization (FV), or a combination of CMAP and FV.
9. receiving, by a control system, second audio data corresponding to a second audio program stream, the second audio data including one or more second audio signals and second spatial data indicating an associated desired perceived spatial location for each of the one or more second audio signals; rendering, by the control system, the second audio data to generate second rendered audio data; mixing the first rendered audio signal and the second rendered audio signal to generate a mixed audio signal; providing the mixed audio signal to at least some loudspeakers in the environment. The method of claim 1.
10. The method according to claim 9, further comprising modifying a rendering process for the first audio data, based at least in part on at least one of the second audio signal, the second rendered audio data, a characteristic of the second audio data, or a characteristic of the second rendered audio data, to generate modified first rendered audio data.
11. The method according to claim 10, wherein modifying the rendering process for the first audio signal includes modifying the rendering of the first audio signal such that a spatial presentation of the first audio signal is warped away from or towards a rendering position of the second rendered audio data.
12. An audio processing system having: an interface system; and a control system, wherein the control system includes a first rendering module that: receives, via the interface system, first audio data corresponding to a first audio program stream, the first audio data including one or more first audio signals and first spatial data indicating a respective desired perceived spatial position for each of the one or more first audio signals; and is configured to perform, by the control system, rendering the first audio data to generate first rendered audio data for at least two loudspeakers of a set of loudspeakers in an environment. The rendering includes determining the relative activation of the loudspeakers in the set of loudspeakers based on the perceived spatial position of the first audio signal reproduced through the loudspeakers, the proximity of the desired perceived spatial position of the first audio signal to the positions of the loudspeakers, and at least one or more attributes of the first audio signal, one or more attributes of the set of loudspeakers, or one or more additional dynamically configurable functions that depend on one or more external inputs. The additional dynamically configurable functions include at least the proximity of one or more loudspeakers to an attracting position or the proximity of one or more loudspeakers to a repelling position. The attracting force is a factor that favors relatively higher loudspeaker activation closer to the attracting position, and the repelling force is a factor that favors relatively lower loudspeaker activation closer to the repelling position. Audio processing system.
13. The audio processing system according to claim 12, wherein the additional dynamically configurable functions include the proximity of one or more loudspeakers to one or more listeners.
14. The audio processing system according to claim 12, wherein the additional dynamically configurable functions include the audibility of one or more loudspeakers at a position in the environment.
15. The audio processing system according to claim 12, wherein the additional dynamically configurable functions include at least one of the capabilities of one or more loudspeakers or the synchronization of one or more loudspeakers with one or more other loudspeakers in the environment.
16. The audio processing system according to claim 12, wherein the additional dynamically configurable functions include at least one of wake word performance or echo canceller performance.
17. The audio processing system according to claim 12, wherein the rendering includes minimizing a cost function, and the cost function includes at least one dynamic speaker activation term.
18. The cost function is at least partially based on a sum of a term corresponding to spatial audio and a term corresponding to proximity of the loudspeaker to a desired perceived spatial position for one of the one or more first audio signals, the audio processing system according to claim 17.
19. The cost function is at least partially based on center of mass amplitude pan (CMAP), flexible virtualization (FV), or a combination of CMAP and FV, the audio processing system according to claim 17.
20. Receiving second audio data corresponding to a second audio program stream via the interface system, the second audio data including one or more second audio signals and second spatial data indicating associated desired perceived spatial positions for each of the one or more second audio signals; Rendering the second audio data to generate second rendered audio data A second rendering module configured to perform; and A mixing module configured to mix the first rendered audio signal and the second rendered audio signal to generate a mixed audio signal Further comprising The control system is further configured to provide the mixed audio signal to at least some of the loudspeakers of the environment. The audio processing system according to claim 12.
21. A non-transitory computer-readable storage medium storing computer-executable instructions for performing the method according to any one of claims 1 to 11 by one or more processors.
22. A computer program product having instructions for performing the method according to any one of claims 1 to 11 by a processor when executed by the processor.
Citation Information
Patent Citations
System and method for adaptive audio signal generation, coding, and rendering
JP2014522155A
Pan audio objects to any speaker layout
JP2016530792A
Apparatus and method for audio rendering using geometric distance definition
JP2017513387A
Audio signal processing method and apparatus
US20170019746A1