Spatial Audio Rendering Adapting to Signal Level and Speaker Reproduction Threshold
The method addresses spatial audio rendering challenges by dynamically mapping audio signals to speakers based on playback limits, stabilizing spatial imaging and maintaining playback levels across speakers with varying capabilities.
Patent Information
- Application Number
- JP2025504190
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-01
- Filing Date
- 2023-07-21
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2043-07-21
AI Technical Summary
Existing spatial audio rendering systems face challenges in maintaining consistent spatial balance and avoiding perceptual distortions when using speakers with varying playback capabilities, especially when signal levels approach their playback limits.
A method that dynamically maps audio signals to speakers based on their playback limits, redistributing signal energy to maintain intended spatial locations while avoiding distortion by considering time- and frequency-varying speaker capabilities and signal levels.
The method stabilizes spatial imaging and maintains overall playback level by intelligently redistributing audio signals across speakers, minimizing perceptual shifts and distortions, even when speakers have different capabilities.
Smart Images

Figure 2025527172000001_ABST
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims priority to U.S. Provisional Patent Application No. 63 / 392,794, filed July 27, 2022, U.S. Provisional Patent Application No. 63 / 413,923, filed October 6, 2022, and U.S. Provisional Patent Application No. 63 / 505,652, filed June 1, 2023, which are hereby incorporated by reference in their entireties.
[0002] [Technical field] The present disclosure relates to a system and method for rendering audio for playback over a set of speakers. [Background technology]
[0003] Audio devices, including but not limited to smart audio devices, are being widely deployed and are becoming a common feature in many homes. While existing systems and methods for controlling audio devices provide benefits, improved systems and methods may be desirable.
[0004] Notes and Terminology Throughout this disclosure, including the claims, "speaker" and "loudspeaker" are used interchangeably to refer to any sound emitting transducer (or set of transducers) driven by a single speaker supply. A standard headphone set includes two speakers.
[0005] Throughout this disclosure, including the claims, the phrase performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to indicate performing an operation directly on the signal or data, or on a processed version of the signal or data (e.g., a version of the signal that has undergone preliminary filtering or post-processing before the operation is performed).
[0006] Throughout this disclosure, including the claims, the term "system" is used broadly to denote an apparatus, a system, or a subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, of which the subsystem generates M inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.
[0007] Throughout this disclosure, including the claims, the term "processor" is used broadly to denote a system or device that is programmable or otherwise configurable (by software or firmware) to perform operations on data (e.g., audio or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipeline processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.
[0008] Throughout this disclosure, including the claims, the terms "connect" or "connected" are used to mean a direct or indirect connection. Thus, when a first device connects to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.
[0009] As used herein, the expression "smart audio device" refers to a smart device that is a single-purpose audio device or virtual assistant (e.g., a connected virtual assistant). A single-purpose audio device is a device (e.g., a TV or a mobile phone) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker) and is designed primarily or primarily to achieve a single purpose. While TVs can typically play audio from program material (and are considered capable of doing so), in many instances, modern TVs run some kind of operating system on which applications, including television viewing applications, run locally. Similarly, audio input and output on a mobile phone can do many things, but they are serviced by applications running on the phone. In this sense, single-purpose audio devices with speakers and microphones are often configured to run local applications and / or services to use the speakers and microphones directly. Several single-purpose audio devices may be configured to group together to achieve audio playback across a zone or user-configured area.
[0010] A virtual assistant (e.g., a connected virtual assistant) is a device (e.g., a smart speaker or voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker) and may be cloud-enabled in some sense or otherwise provide the ability to utilize multiple devices (apart from the virtual assistant) for applications not implemented in or on the virtual assistant itself. Virtual assistants may sometimes work together, for example, in a discrete and conditionally defined manner. For example, two or more virtual assistants may work together in the sense that one of them, e.g., the one most confident of having heard the activation word, responds to the word. Connected devices may form a kind of constellation, which may be managed by one main application that may be (or implements) a virtual assistant.
[0011] Herein, "wakeword" is used broadly to refer to any sound (e.g., a word spoken by a human being, or any other sound), and the smart audio device is configured to wake up in response to detecting ("hearing") the sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "awake" refers to the device entering a state in which it is waiting (i.e., listening) for a voice command. In some cases, what is referred to herein as a "wakeword" may include multiple words, such as a phrase.
[0012] Here, the expression "hot word detector" refers to a device (or software including instructions for configuring a device) configured to continuously search for mismatches between real-time speech (e.g., speech) features and a trained model. Typically, a wake-up event is triggered whenever the hot word detector determines that the probability of a hot word being detected exceeds a predetermined threshold. For example, the threshold may be a predetermined threshold that is adjusted to provide a good compromise between false accept and false reject rates. Following a hot word event, the device may enter a state (which may be referred to as an "awake" state or an "attentiveness" state) in which it listens for commands and passes received commands to a larger, more computationally intensive recognizer. Summary of the Invention
[0013] At least some aspects of the present disclosure may be implemented via a method, such as an audio processing method. In some examples, the method may be performed, at least in part, by a control system such as those disclosed herein. Some methods may include receiving audio data by the control system and via an interface system. The audio data may include one or more audio signals and associated spatial data. The spatial data may indicate an intended perceived spatial location corresponding to the audio signal. In some examples, the intended perceived spatial location may correspond to a channel of a channel-based audio format, may correspond to metadata, or may correspond to both a channel and metadata. Some methods may include rendering, by the control system, the audio data for playback through a set of two or more speakers of an environment to generate a speaker signal. Some methods may include providing the speaker signal to at least two speakers of the set of speakers of the environment via the interface system.
[0014] According to some examples, rendering each of the one or more audio signals included in the audio data may include mapping each audio signal to a speaker signal. In some examples, the mapping may be a time- and frequency-varying mapping. In some examples, the mapping for each audio signal may be calculated as a function of the intended perceptual spatial location of the audio signal, the physical locations relative to the speakers, and a time- and frequency-varying representation of the speaker signal level relative to the maximum playback limit of each speaker. According to some examples, each mapping may be calculated to approximately achieve the intended perceptual spatial location of the associated audio signal when the speaker signal is played on two or more corresponding speakers located at the associated speaker locations. In some examples, a representation of the speaker signal level relative to the maximum playback limit may be calculated for each audio signal as a function of one or more of the audio signals and their perceptual spatial location. According to some examples, the mapping of audio signals to a particular speaker signal may decrease as the speaker signal level representation relative to the maximum playback limit increases above a threshold, while the mapping may increase to one or more other speakers where the signal level representation relative to the maximum playback limit of the one or more other speakers is below the threshold.
[0015] In some examples, the mapping may be calculated over the entire audible frequency range (e.g., the human audible frequency range), however, in some examples, the mapping may be calculated over a subset of the audible frequency range.
[0016] According to some examples, the mapping may include minimizing a cost function that includes a first term that models how closely the intended perceptual spatial location can be achieved as a function of mapping the audio signal to the speaker signals, and a second term that assigns a cost for activating each speaker. In some such examples, the cost of activating each speaker may be based at least in part on a function of a representation of the speaker signal level relative to a maximum playback limit.
[0017] In some examples, the representation of the speaker signal level relative to the maximum playback limit may correspond to one or more of a digital signal level, a limiter gain, or an acoustic signal level. In some examples, the representation of the speaker signal level relative to the maximum playback limit may be calculated as a difference between a level estimate for each audio signal and a playback limit threshold for each speaker. In some examples, the level estimate for each audio signal may be based at least in part on zone-based rendering of all audio signals. In some examples, the level estimate for each audio signal may be based at least in part on previously calculated speaker signals. In some examples, the level estimate for each audio signal may further depend on the participation of each speaker in multiple spatial zones. Some methods may further include smoothing the level estimate for each audio signal over time, frequency, or both time and frequency.
[0018] According to some examples, the mapping from audio signals to speaker signals may be determined by querying a data structure indexed by the intended perceptual spatial location and level estimate of each audio signal. In some examples, the mapping from audio signals to speaker signals may be determined by interpolating from a set of pre-computed speaker mappings. In some such examples, the set may be indexed by the intended perceptual spatial location and level estimate for each audio signal. In some examples, the set may be indexed by the intended level estimate for each audio signal.
[0019] In some examples, the level estimate for each audio signal may be expressed as a wideband gain multiplied by a spectral shape. According to some examples, the spectral shape may be selected from a plurality of spectral shapes. In some such examples, each spectral shape of the plurality of spectral shapes may correspond to a content type.
[0020] According to some examples, as the representation of the signal level relative to the maximum playback level increases above a threshold, the mapping to one speaker is decreased and the mapping to another speaker is increased.
[0021] In some examples, approximately achieving the intended perceptual spatial location of the associated audio signals can include minimizing a difference between the perceptual spatial location and the intended perceptual spatial location given the available speakers and the associated speaker locations. According to some examples, approximately achieving the intended perceptual spatial location of the associated audio signals includes minimizing a cost function.
[0022] Some examples may further include controlling the degree of reduction in mapping to one speaker and the increase in mapping to another speaker according to one or more of an audio format, a codec, or metadata. Some methods may further include controlling the degree of reduction in mapping to one speaker and the increase in mapping to another speaker according to a knee parameter.
[0023] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include one or more memory devices as described herein, including, but not limited to, one or more random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, various novel aspects of the subject matter described in this disclosure may be implemented via one or more non-transitory media having software stored thereon.
[0024] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of at least partially performing the methods disclosed herein. In some implementations, the device may include an interface system and a control system. The control system may include at least one of a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic element, discrete gate or transistor logic, discrete hardware components, or a combination thereof.
[0025] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. The relative dimensions of the following drawings may not be drawn to scale. [Brief explanation of the drawings]
[0026] [Figure 1] FIG. 1 is a block diagram illustrating example components of a device capable of implementing various aspects of the present disclosure. [Figure 2] 1 shows a plan view of a listening environment, which is a living space in this example. [Figure 3] FIG. 1 is a block diagram illustrating example components of a system in which various aspects of the present disclosure can be implemented. [Figure 4] 1 shows examples of playback limit thresholds and corresponding frequencies. [Figure 5] 10 is a graph showing an example of dynamic range compressed data. [Figure 6] 1 shows examples of spatial zones in a listening environment. [Figure 7] Illustrates examples of speakers within the spatial zones of FIG. [Figure 8] 8 shows an example of nominal spatial locations superimposed on the spatial zones and speakers of FIG. 7. [Figure 9] 10 is a graph of points illustrating object-to-speaker mapping in an exemplary embodiment; [Figure 10] 10 is a graph of trilinear interpolation between points illustrating object-to-speaker mapping according to an example. [Figure 11] Here are some examples of penalties for various knee parameters: [Figure 12] FIG. 1 is a flow diagram outlining an example of a method that may be performed by an apparatus or system as disclosed herein. DETAILED DESCRIPTION OF THE INVENTION
[0027] Spatial audio playback in consumer environments is typically associated with a given number of speakers arranged in given locations, such as those corresponding to Dolby 5.1 or 7.1 surround sound. In such cases, content is created specifically for the associated speakers and encoded as separate channels for each speaker (e.g., Dolby Digital™, Dolby Digital Plus™, etc.). Recently, immersive, object-based spatial audio formats have been introduced (e.g., Dolby Atmos™) that remove this association between content and specific speaker locations. Instead, content can be described as a collection of individual audio objects, each with possibly time-varying metadata that describes the desired perceived location of the audio object in three-dimensional space and, in some instances, other properties of the audio object. During playback, the audio content is converted into speaker feeds by a renderer that adapts to the number and locations of speakers in the playback system. However, many such renderers restrict the position of a set of speakers to one of a set of prescribed layouts (e.g., Dolby 3.1.2, 5.1.2, 7.1.4, 9.1.6 with Dolby Atmos™, etc.).
[0028] Beyond such constrained rendering, methods have been developed that allow object-based audio to be flexibly rendered with a truly arbitrary number of speakers placed in arbitrary positions. These methods generally require the renderer to know the number and physical locations of the speakers in the listening space. For such systems to be practical for the average consumer, an automated method for locating the speakers would be desirable. One such method relies on the use of a number of microphones, possibly co-located with the speakers. By playing an audio signal through the speakers and recording it using a microphone, the distance between each speaker and microphone can be estimated. From these distances, the positions of both the speakers and microphones can later be inferred.
[0029] Concurrent with the introduction of object-based spatial audio into consumer spaces, so-called “smart speakers” such as the Amazon Echo™ product line are experiencing rapid adoption. While the immense popularity of these devices can be attributed to their simplicity and the convenience offered by wireless connectivity and integrated voice interfaces (e.g., Amazon’s Alexa™), the acoustic capabilities of these devices are generally limited, especially with regard to spatial audio. In most cases, these devices are limited to mono or stereo playback. However, combining the aforementioned flexible rendering and automatic localization techniques with multiple orchestrated smart speakers can result in a system with highly sophisticated spatial playback capabilities that is still very easy for consumers to set up. Consumers can place as many or as few speakers as desired in convenient locations without having to wire the speakers for wireless connectivity, and can automatically localize the associated flexible renderer speakers using their built-in microphones.
[0030] One approach to rendering spatial audio through a set of speakers is to map each component signal of a spatial mix across the set of speakers based purely on the assumed or measured locations of the speakers and the intended perceived locations of the component signals. Such an approach is described in U.S. Patent Nos. 9,712,939 and 11,172,318, which are incorporated herein by reference. If there is variation in reproduction capabilities across the set of speakers, using this approach can compromise the perceived quality of the spatial rendering. Many small speakers begin to distort as the reproduction level increases, especially at low frequencies, and then reach their excursion limits.
[0031] To reduce such distortion, each speaker may implement dynamic processing that, in some instances, limits its playback level below these limits in a manner that varies across frequency. When spatial audio rendered using the above methods is played through a set of speakers, each speaker applies its dynamic processing independently, which can result in very different relative modifications to the audio on different speaker feeds. For example, lower-capability speakers generally attenuate audio more than higher-capability speakers at high playback levels. These variations in processing between speakers can dynamically shift the spatial balance of the mix in perceptually confusing ways and can also disrupt the overall relative balance of the mix. For example, if the front soundstage is primarily reproduced by lower-capability speakers, the front soundstage may be attenuated overall relative to the rear soundstage.
[0032] Applicant has developed methods to mitigate some of these problems by intelligently combining playback thresholds between speakers and applying them to spatial zones throughout the spatial audio mixer before rendering the mix to the speaker feeds. Some examples are disclosed herein. Zones can be selected to prevent perceptually confusing left-to-right imaging shifts while maintaining a degree of independence in processing between parts of the audio mix. Some zone-based methods include four zones: front, center, surround, and overhead.
[0033] These zone-based methods help stabilize the spatial imaging of the rendered audio, but in some instances, such zone-based methods can have the undesirable effect of constraining the overall playback level to the least capable device across a set of speakers.
[0034] The present disclosure provides improved rendering methods, including several improved zone-based methods that better utilize more capable speakers within a set of speakers. Improved methods for rendering spatial audio are disclosed in which the dynamic signal level of a spatial audio mix is additionally considered when mapping each component signal of the mix to a speaker feed signal. In some examples, as the level of an audio mix approaches a playback limit threshold for a particular speaker, component mapping to that speaker is reduced in favor of increased mapping to other speakers whose mix levels are further away from the limit thresholds of other speakers. In this way, the overall level of the rendered audio is not constrained by less capable speakers. However, less capable speakers can be used when the audio signal levels are below their limit thresholds.
[0035] Such dynamic rendering systems must be constructed with care to avoid introducing additional perceptual artifacts in the process of maximizing playback levels. For example, consider an individual component of a spatial audio mix whose intended perceived location is "front left." If a speaker is physically located near this intended perceived location, ideally, most of the component signal energy should be mapped to this speaker. However, if the mix's signal level is approaching the playback limits of this speaker, it may be desirable to map most of this component signal to other, more capable speakers to reduce the activation of the first speaker's dynamic processing, thereby better maintaining the overall playback level of the mix. Because signal energy is dynamically diverted to these other speakers, which may not be physically closer to the intended perceived location and therefore less suited to achieving the component signal's intended perceived location, it is necessary to minimize the likelihood of perceiving this diversion as an undesirable spatial shift of the component signal.
[0036] To achieve this minimization, some disclosed methods employ several strategies simultaneously.
[0037] (1) First, in some examples, the mapping of each component signal to a speaker signal takes into account the current signal level of the audio mix and makes a “best effort” to achieve the desired perceptual location of the component audio signal using the speakers that are considered available under the signal level conditions at that particular time. In some examples, making this “best effort” may include what is described herein as approximately achieving the intended perceptual spatial location of the audio signal component. The intended perceptual spatial location may correspond to a channel in a channel-based audio format, positional metadata of an audio object, or both channel and positional metadata. According to some examples, approximately achieving the intended perceptual spatial location of the audio signal component may include, for example, minimizing the difference between the perceptual spatial location and the intended perceptual spatial location given the speakers available in the audio environment, the capabilities of each speaker, and the associated speaker locations. According to some examples, approximately achieving the intended perceptual spatial location of the audio signal component may include minimizing a cost function. In this manner, such a method individually optimizes the spatial mapping of each component signal with respect to signal level conditions. This differs from simpler solutions, for example, which render to speaker signals using the methods described above that optimize spatial imaging but ignore signal level, and then redistribute the energy of the already rendered signal between speakers based on a comparison of the rendered signal level with each speaker's limit threshold.
[0038] (2) Second, in some instances, both the mapping of component signals to loudspeaker feeds and the characterization of signal levels relative to loudspeaker reproduction limits on which this mapping depends can be calculated in a time- and frequency-varying manner. In this way, diversion of any component signal's energy away from its spatially optimal loudspeaker occurs only in frequency regions and at moments when the signal energy is approaching the limit thresholds of these optimal loudspeakers. This approach minimizes the amount of diverged energy, allowing more of any component signal's energy to remain in the loudspeaker that is optimal for its spatial reproduction. Thus, the perceived location of a component signal remains likely to remain at its desired spatial location.
[0039] (3) Finally, a signal level characterization with respect to the speaker reproduction limits on which the component-to-speaker signal mapping depends is calculated for each individual component signal based on one or more component signals in the mix and their intended perceptual location. In this way, signal energy diversion between speakers can be personalized not only to each component signal's desired perceptual location, as outlined in the first strategy above, but also to an estimate of overall signal level that is optimized in some way with respect to that component signal's relationship to other components in the mix. For example, the overall signal level associated with any component signal can be calculated based on the spatial zone of the zone-based method described above. In this way, the perceived left-right balance of the dynamic rendering is stabilized because component signals in similar spatial zones have similar diversions. Furthermore, component signals belonging to substantially different zones may be diversioned differently if the overall levels associated with these zones are significantly different. For example, if the signal level in the surround zone is low, audio signal components that are significantly associated with the surround zone may have little diversion applied to their mapping. At the same time, if the signal level in the front zone is high, components that are significantly associated with the front zone may have more diversion applied to their mapping. This strategy therefore also helps to minimize the amount of energy that is transformed throughout the spatial mix by applying the transformation only to component signals that belong to the required spatial zone.
[0040] FIG. 1 is a block diagram illustrating example components of a device capable of implementing various aspects of the present disclosure. As with other figures provided herein, the number and types of elements shown in FIG. 1 are merely examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 100 may be or include a smart audio device configured to perform at least some of the methods disclosed herein. In other implementations, device 100 may be or include another device, such as a laptop computer, a cellular phone, a tablet device, a smart home hub, etc., configured to perform at least some of the methods disclosed herein. In some such implementations, device 100 may be or include a server.
[0041] In this example, device 100 includes interface system 105 and control system 110. Interface system 105, in some implementations, may be configured to receive audio data. The audio data may include audio signals scheduled to be played by at least some speakers in the environment. The audio data may include one or more audio signals and associated spatial data. The spatial data may include, for example, channel data and / or spatial metadata. Interface system 105 may be configured to provide rendered audio signals to at least some speakers of a set of speakers in the environment. Interface system 105, in some implementations, may be configured to receive input from one or more microphones in the environment.
[0042] Interface system 105 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more universal serial bus (USB) interfaces). According to some implementations, interface system 105 may include one or more wireless interfaces. Interface system 105 may include one or more devices that implement a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, interface system 105 may include one or more interfaces between control system 110 and a memory system, such as optional memory system 115 shown in FIG. 1 . However, in some cases, control system 110 may include a memory system.
[0043] The control system 110 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic elements, discrete gate or transistor logic, and / or discrete hardware components.
[0044] In some implementations, the functionality of control system 110 may reside on more than one device. For example, a portion of control system 110 may reside on a device within one of the environments illustrated herein, while another portion of control system 110 may reside on a device external to the environment, such as a server, a mobile device (e.g., a smartphone, or a tablet computer), etc. In other examples, a portion of control system 110 may reside on a device within one of the environments illustrated herein, while another portion of control system 110 may reside on one or more other devices in the environment. For example, the functionality of the control system may be distributed across multiple smart audio devices in the environment, or may be shared by an orchestration device (such as what is referred to herein as a smart home hub) and one or more other devices in the environment. Interface system 105 may also reside on more than one device in some such examples.
[0045] In some implementations, control system 110 may be configured to perform, at least in part, the methods disclosed herein. According to some examples, control system 110 may be configured to implement a method for managing the playback of multiple audio streams across multiple speakers.
[0046] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. One or more non-transitory media may reside, for example, in optional memory system 115 shown in FIG. 1 and / or in control system 110. Accordingly, various novel aspects of the subject matter described in this disclosure may be embodied in one or more non-transitory media having software stored thereon. The software may include, for example, instructions for controlling at least one device to process audio data. The software may be executable by one or more components of a control system, such as control system 110 of FIG. 1.
[0047] In some examples, device 100 may include optional microphone system 120 shown in FIG. 1. Optional microphone system 120 may include one or more microphones. In some implementations, the one or more microphones may be part of or associated with another device, such as a speaker in a speaker system, a smart audio device, etc.
[0048] According to some implementations, device 100 may include optional speaker system 125 shown in FIG. 1 . Optional speaker system 125 may include one or more speakers. Loudspeakers are sometimes referred to herein as “speakers.” In some embodiments, at least some of the speakers of optional speaker system 125 may be positioned arbitrarily. For example, at least some of the speakers of optional speaker system 125 may be positioned in locations that do not correspond to a speaker layout defined by any standard, such as Dolby 5.1, Dolby 5.1.2, Dolby 7.1, Dolby 7.1.4, Dolby 9.1, or Hamasaki 22.2. In some such examples, at least some of the speakers of optional speaker system 125 may be positioned in locations that are convenient for the space (e.g., where there is space to accommodate the speakers) rather than in a speaker layout defined by any standard.
[0049] In some examples, device 100 may include optional sensor system 130 as shown in FIG. 1 . Optional sensor system 130 may include one or more cameras, touch sensors, gesture sensors, motion detectors, etc. According to some implementations, optional sensor system 130 may include one or more cameras. In some implementations, the camera may be a standalone camera. In some examples, one or more cameras of optional sensor system 130 may reside in a smart audio device, which may be a single-purpose audio device or a virtual assistant. In some examples, one or more cameras of optional sensor system 130 may reside in a TV, a mobile phone, or a smart speaker.
[0050] In some examples, device 100 may include optional display system 135 shown in FIG. Optional display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some cases, optional display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples in which device 100 includes display system 135, sensor system 130 may include a touch sensor system and / or a gesture sensor system proximate one or more displays of display system 135. According to some such implementations, control system 110 may be configured to control display system 135 to present a graphical user interface (GUI), such as one of the GUIs disclosed herein.
[0051] According to some examples, device 100 may be or include a smart audio device. In some such implementations, device 100 may be or include a hot word detector. In some examples, device 100 may be or include a virtual assistant.
[0052] FIG. 2 shows a floor plan of a listening environment, which in this example is a living space. As with other figures provided herein, the number and types of elements shown in FIG. 2 are merely examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to this example, environment 200 includes a living room 210 in the upper left, a kitchen 215 in the lower center, and a bedroom 222 in the lower right. Boxes and circles distributed throughout the living space represent a set of speakers 205a-205h, which in at least some implementations may be smart speakers, positioned in locations convenient for the space but not conforming to a standardized layout (arbitrarily placed). In some examples, speakers 205a-205h may be adapted to implement one or more disclosed embodiments.
[0053] According to some examples, environment 200 may include a smart home hub that implements at least some of the disclosed methods. According to some such embodiments, the smart home hub may include at least a portion of control system 110 described above. In some examples, a smart device (such as a smart speaker, a mobile phone, a smart TV, a device used to implement a virtual assistant, etc.) may implement a smart home hub.
[0054] In this example, environment 200 includes cameras 211a-211e distributed throughout the environment. Depending on the implementation, one or more smart audio devices in environment 200 may also include one or more cameras. One or more smart audio devices may be single-purpose audio devices or virtual assistants. In some such examples, one or more cameras of optional sensor system 130 may be present in television 230, a mobile phone, or a smart speaker, such as one or more of speakers 205b, 205d, 205e, or 205h. While cameras 211a-211e are not shown in all depictions of environment 200 presented in this disclosure, each of environment 200 may nevertheless include one or more cameras in some implementations.
[0055] In flexible rendering, spatial audio may be rendered by any number of arbitrarily arranged speakers. With the proliferation of smart audio devices (such as smart speakers) in the home, there is a need to provide flexible rendering techniques that allow consumers to use their smart audio devices to perform flexible rendering of audio and playback of the audio so rendered.
[0056] Several techniques have been developed to implement flexible rendering, including Center of Mass Amplitude Panning (CMAP) and Flexible Virtualization (FV).
[0057] In the context of rendering (or rendering and playing) a spatial audio mix (e.g., rendering an audio stream or multiple audio streams) for playback by a smart audio device (or by another set of speakers) of a set of smart audio devices, the types of speakers (e.g., within or coupled to a smart audio device) may be different, and therefore their corresponding acoustic capabilities may vary significantly. In the example shown in FIG. 2, speakers 205d, 205f, and 205h are smart speakers with a single 0.6-inch speaker. In this example, speakers 205b, 205c, 205e, and 205f are smart speakers with a 2.5-inch woofer and a 0.8-inch tweeter. According to this example, speaker 205g is a smart speaker with a 5.25-inch woofer, three 2-inch midrange speakers, and a 1.0-inch tweeter. Here, speaker 205a is a soundbar with sixteen 1.1-inch beam drivers and two 4-inch woofers. Therefore, the low-frequency capabilities of smart speakers 205d and 205f are significantly lower than the low-frequency capabilities of other speakers in environment 200, especially those with 4-inch or 5.25-inch woofers.
[0058] <Examples of dynamic processing, including some related zone-based methods> Figure 3 is a block diagram illustrating example components of a system in which various aspects of the present disclosure can be implemented. As with other figures provided herein, the number and types of elements shown in Figure 3 are merely examples. Other implementations may include more, fewer, and / or different types and numbers of elements.
[0059] According to this embodiment, system 300 includes smart home hub 305 and speakers 205a through 205m. In this embodiment, smart home hub 305 includes an instance of control system 110 shown in FIG. 1 and described above. According to this implementation, control system 110 includes listening environment dynamic processing configuration data module 310, listening environment dynamic processing module 315, and rendering module 320. Some examples of listening environment dynamic processing configuration data module 310, listening environment dynamic processing module 315, and rendering module 320 are described below. In some embodiments, rendering module 320′ may be configured for both rendering and listening environment dynamic processing.
[0060] As suggested by the arrows between the smart home hub 305 and the speakers 205a-m, the smart home hub 305 also includes an instance of the interface system 105 shown in FIG. 1 and described above. According to some examples, the smart home hub 305 may be part of the environment 200 shown in FIG. 2. In some examples, the smart home hub 305 may be implemented by a smart speaker, a smart TV, a mobile phone, a laptop, etc. In some implementations, the smart home hub 305 may be implemented by software, such as a downloadable software application or “app” software. In some implementations, the smart home hub 305 may be implemented in each of the speakers 205a-m, all operating in parallel to generate the same processed audio signal from the module 320. According to some such embodiments, in each of the speakers, the rendering module 320 may then generate one or more speaker feeds associated with each speaker or group of speakers and provide these speaker feeds to a respective speaker dynamic processing module.
[0061] In some examples, speakers 205a-205m may include speakers 205a-205h of Figure 2, while in other examples, speakers 205a-205m may be or include other speakers. Thus, in this example, system 300 includes M speakers, where M is an integer greater than 2.
[0062] Smart speakers, and many other powered speakers, typically employ some type of internal dynamic processing to prevent speaker distortion. Such dynamic processing is often associated with a signal limiting threshold (e.g., a limiting threshold that varies between frequencies) at which the signal level is dynamically maintained. For example, Dolby's Audio Regulator, one of several algorithms in the Dolby Audio Processing (DAP) audio post-processing package, provides such processing. In some instances, dynamic processing may also include applying one or more compressors, gates, expanders, duckers, etc., although not typically via the smart speaker's dynamic processing module.
[0063] Thus, in this example, each of the speakers 205a-205m includes a corresponding speaker dynamics processing (DP) module A-M. The speaker dynamics processing modules are configured to apply individual speaker dynamics processing configuration data for each individual speaker in the listening environment. For example, speaker DP module A is configured to apply individual speaker dynamics processing configuration data appropriate for speaker 205a. In some examples, the individual speaker dynamics processing configuration data may correspond to one of more capabilities of an individual speaker, such as the speaker's ability to reproduce audio data within a particular frequency range and at a particular level without noticeable distortion.
[0064] When spatial audio is rendered across a set of heterogeneous speakers (e.g., speakers of a smart audio device or speakers combined with a smart audio device), each with potentially different playback limitations, care must be taken when performing dynamic processing on the overall audio mix. A simple solution is to render the spatial mix to the speaker feeds of each participating speaker, allowing the dynamic processing module associated with each speaker to operate independently on the corresponding speaker feed according to the limitations of that speaker.
[0065] This approach avoids distorting each speaker while potentially dynamically shifting the spatial balance of the mix in a perceptually disruptive way. For example, referring to FIG. 2, assume that a television program is being displayed on television 230 and corresponding audio is being played by speakers in environment 200. Assume that during the television program, audio associated with a fixed object (such as a heavy machinery unit in a factory) is intended to be rendered at location 244. Further, assume that the dynamic processing module associated with speaker 205d reduces the level of the bass audio substantially more than the dynamic processing module associated with speaker 205b. This is because speaker 205b has a substantially higher ability to reproduce bass sounds. If the volume of the signal associated with the fixed object fluctuates, at higher volumes, the dynamic processing module associated with speaker 205d reduces the level of the bass audio substantially more than the dynamic processing module associated with speaker 205b reduces the level of the same audio. This difference in level changes the apparent location of the fixed object. Therefore, an improved solution is needed.
[0066] Some embodiments of the present disclosure are systems and methods for rendering (or rendering and playing) a spatial audio mix (e.g., rendering an audio stream or multiple audio streams) for playback by at least one (e.g., all or some) smart audio device of a set of smart audio devices (e.g., a coordinated set of smart audio devices) and / or by at least one (e.g., all or some) speaker of another set of speakers. Some embodiments are methods (or systems) for such rendering (e.g., including generating speaker feeds), and playback of the rendered audio (e.g., playing the generated speaker feeds). Examples of such embodiments include the following:
[0067] Systems and methods for audio processing may include rendering audio (e.g., rendering an audio stream or multiple audio streams to render a spatial audio mix) for playback over at least two speakers (e.g., all or a portion of a set of speakers), including by:
[0068] (a) Dynamic processing configuration data for each speaker (such as limit thresholds (reproduction limit thresholds) for each speaker) is combined to determine listening environment dynamic processing configuration data for multiple speakers (such as combined thresholds).
[0069] (b) performing dynamic processing on the audio (e.g., one or more audio streams representing a spatial audio mix) using the multiple speaker listening environment dynamic processing configuration data (e.g., combination thresholds) to generate processed audio;
[0070] (c) Rendering the processed audio to the speaker feed.
[0071] According to some implementations, process (a) may be performed by a module such as the listening environment dynamic processing configuration data module 310 shown in FIG. 3 . The smart home hub 305 may be configured to obtain, via the interface system, individual speaker dynamic processing configuration data for each of the M speakers. In this implementation, the individual speaker dynamic processing configuration data includes an individual speaker dynamic processing configuration data set for each speaker of the plurality of speakers. According to some examples, the individual speaker dynamic processing configuration data for one or more speakers may correspond to one or more functions of the one or more speakers. In this example, each of the individual speaker dynamic processing configuration data sets includes at least one type of dynamic processing configuration data. In some examples, the smart home hub 305 may be configured to obtain the individual speaker dynamic processing configuration data set by querying each of the speakers 205 a- 205 m. In other implementations, the smart home hub 305 may be configured to obtain the individual speaker dynamic processing configuration data set by querying a data structure of previously obtained individual speaker dynamic processing configuration data sets stored in memory.
[0072] In some examples, process (b) may be performed by a module such as listening environment dynamic processing module 315 of Figure 3. Some detailed examples of processes (a) and (b) are described below.
[0073] In some examples, the rendering of process (c) may be performed by a module such as rendering module 320 or rendering module 320' of Figure 3. In some embodiments, the audio processing may include:
[0074] (d) performing dynamic processing on the rendered audio signal in accordance with the individual speaker dynamic processing configuration data for each speaker (e.g., limiting the speaker feed in accordance with a playback limit threshold associated with the corresponding speaker, thereby generating a limited speaker feed). Process (d) may be performed, for example, by dynamic processing modules A-M shown in FIG.
[0075] The speakers may include speakers of (or coupled to) at least one (e.g., all or some) smart audio device of the set of smart audio devices. In some implementations, to generate the limited speaker feeds in step (d), the speaker feeds generated in step (c) may be processed by a second stage of dynamic processing (e.g., by an associated dynamic processing system of each speaker) to generate speaker feeds before final playback on the speakers. For example, the speaker feeds (or a subset or portion thereof) may be provided to a dynamic processing system of each different one of the speakers (e.g., a dynamic processing subsystem of a smart audio device that includes or is coupled to an associated one of the speakers), and the processed audio output from each said dynamic processing system may be used to generate a speaker feed for the associated one of the speakers. Following speaker-specific dynamic processing (i.e., dynamic processing performed independently for each speaker), the processed (e.g., dynamically limited) speaker feed may be used to drive the speaker to cause sound reproduction.
[0076] The first stage of dynamic processing (in step (b)) may be designed to reduce perceptually disruptive shifts in spatial balance that could occur if steps (a) and (b) were omitted, and the dynamically processed (e.g., limited) speaker feeds resulting from step (d) may be generated in response to the original audio (rather than in response to the processed audio generated in step (b)). This can prevent undesirable shifts in the spatial balance of the mix. Because the dynamic processing in step (b) may not necessarily ensure that signal levels have been reduced below threshold for all speakers, the second stage of dynamic processing operating on the rendered speaker feeds from step (c) may be designed to ensure that the speakers do not distort. Combining the individual speaker dynamic processing configuration data (e.g., combining thresholds in the first stage (step (a))) may, in some examples, include averaging the individual speaker dynamic processing configuration data (e.g., limit thresholds) across speakers (e.g., across smart audio devices) or taking the minimum of the individual speaker dynamic processing configuration data (e.g., limit thresholds) across speakers (e.g., across smart audio devices).
[0077] In some embodiments, when the first stage dynamic processing (step (b)) operates on audio exhibiting a spatial mix (e.g., audio of an object-based audio program, further comprising at least one object channel and optionally at least one speaker channel), this first stage may be performed according to the technique of audio object processing through the use of spatial zones. In such cases, the combined individual speaker dynamic processing configuration data (e.g., combined limit thresholds) associated with each zone may be derived by (or as) a weighted average of the individual speaker dynamic processing configuration data (e.g., individual speaker limit thresholds), where the weighting may be provided by or determined, at least in part, by each speaker's spatial proximity to the zone and / or position within the zone.
[0078] In an exemplary embodiment, a plurality of M speakers (M≧2) is assumed, where each speaker is indexed by a variable i. Associated with each speaker i is a frequency-variable playback threshold T i [f], where the variable f represents an index into a finite set of frequencies for which thresholds are specified (note that if the size of the frequency set is 1, the corresponding single threshold is considered broadband, applying to the entire frequency range). According to this example, these thresholds are used to limit the audio signal to a threshold T for a specific purpose, such as preventing a speaker from distorting or from playing above a level that is considered objectionable in that vicinity. i To limit the frequency to less than [f], each independent dynamic processing function is utilized by each speaker.
[0079] 4A, 4B, and 4C show examples of playback thresholds and corresponding frequencies. The range of frequencies shown can span the range of frequencies audible to the average human (e.g., 20 Hz to 20 kHz). In these examples, the playback threshold is indicated by the vertical axis of graphs 400a, 400b, and 400c, labeled "Level Threshold" in these examples. The playback threshold / level threshold increases in the direction of the arrows on the vertical axis. The playback threshold / level threshold can be expressed in decibels, for example. In these examples, the horizontal axis of graphs 400a, 400b, and 400c indicates frequency, increasing in the direction of the arrows on the horizontal axis. The playback threshold shown by curves 400a, 400b, and 400c can be implemented, for example, by dynamic processing modules in the individual speakers.
[0080] Graph 400a in FIG. 4A shows a first example of the reproduction threshold as a function of frequency. Curve 405a shows the reproduction threshold for each corresponding frequency value. In this example, the bass frequency f b At input level T i Input audio received at the output level T o The bass frequency f is output by the dynamic processing module. bmay be in the range of 60 to 250 Hz, for example. However, in this example, the high-pitched frequency f t At input level T i The input audio received at level T i The high frequency f t may be in a range higher than 1280 Hz, for example. Thus, in this example, curve 405a corresponds to a dynamic processing module that applies a significantly lower threshold to bass frequencies than to treble frequencies. Such a dynamic processing module may be suitable for speakers that do not have a woofer (e.g., speaker 205d in FIG. 2).
[0081] Curve 400b of Figure 4B shows a second example of the reproduction threshold as a function of frequency. Curve 405b shows the same bass frequency f shown in Figure 1A. b And the input level T i Input audio received at output level T o 2. The curve 405b corresponds to a dynamic processing module that applies a lower threshold to bass frequencies than curve 405a. Such a dynamic processing module may be suitable for loudspeakers with at least small woofers (e.g., loudspeaker 205b in FIG. 2).
[0082] Graph 400c of Figure 4C shows a third example of the reproduction threshold as a function of frequency. Curve 405c (which is a straight line in this example) shows the same bass frequency f shown in Figure 1A. b And the input level T i 405c indicates that input audio received at 405d is output by the dynamic processing module at the same level. Thus, in this example, curve 405c may correspond to a dynamic processing module that is suitable for a speaker that can reproduce a wide range of frequencies, including bass frequencies. For simplicity, it will be appreciated that the dynamic processing module can approximate curve 405c by implementing curve 405d, which applies the same threshold to all frequencies shown.
[0083] A spatial audio mix can be rendered for multiple speakers using known rendering systems such as Center of Mass Amplitude Panning (CMAP) or Flexible Virtualization (FV). From the components of the spatial audio mix, the rendering system generates a speaker feed for each of the multiple speakers. In some previous examples, the speaker feeds are filtered by a threshold T i [f], independently processed by each loudspeaker's associated dynamic processing function. Without the benefit of this disclosure, this described rendering scenario can result in confusing shifts in the perceived spatial balance of the rendered spatial audio mix. For example, one of the M loudspeakers, say the one on the right side of the listening area, may be much less capable (e.g., at rendering audio in the bass range) than the other loudspeakers, and therefore, the threshold T for that loudspeaker may be too low. i [f] may be significantly lower than the threshold of the other loudspeakers, at least in certain frequency ranges. During playback, the dynamic processing module of this loudspeaker will reduce the level of the right-side spatial mix components significantly more than the left-side components. Listeners are very sensitive to such dynamic shifts between the left and right balance of the spatial mix, and may find the result very confusing.
[0084] To address this issue, in some examples, individual speaker dynamic processing configuration data (e.g., playback limit thresholds) for individual speakers in a listening environment are combined to create listening environment dynamic processing configuration data for all speakers in the listening environment. The listening environment dynamic processing configuration data can then be utilized to first perform dynamic processing in the context of the entire spatial audio mix before rendering to the speaker feeds. Because this first-stage dynamic processing has access to the entire spatial mix, as opposed to a single individual speaker feed, it can perform processing in a manner that does not introduce disruptive shifts in the perceived spatial balance of the mix. Individual speaker dynamic processing configuration data (e.g., playback limit thresholds) can be combined in a manner that eliminates or reduces the amount of dynamic processing performed by any of the individual speaker's independent dynamic processing functions.
[0085] In one example of determining listening environment dynamic processing configuration data, individual speaker dynamic processing configuration data (e.g., playback thresholds) for individual speakers can be combined into one set of listening environment dynamic processing configuration data that is applied to all components of the spatial mix in the first stage of dynamic processing. For example, the following frequency-varying playback thresholds:
number
number
[0086] Such a combination essentially eliminates the individual dynamic processing action of each speaker, since the spatial mix is first limited at all frequencies below the threshold of the least capable speaker. However, such a strategy may be overly aggressive. Many speakers may be playing below their capabilities, and the combined playback level of all speakers may be undesirably low. For example, if the bass threshold shown in Figure 4A is applied to the speaker corresponding to the threshold in Figure 4C, the playback level of the latter speaker will be unnecessarily low in the bass range. Another combination for determining dynamic processing configuration data for a listening environment is to take the mean (average) of the individual speaker dynamic processing configuration data across all speakers in the listening environment. For example, in the context of playback limit thresholds, the mean can be determined as follows:
number
[0087] In this combination, the overall playback level may be increased compared to the minimum case, since the first stage dynamic processing is limited to a higher level. This allows more capable speakers to play at a higher volume. For speakers with individual limiting thresholds below the average, their independent dynamic processing functions may still limit their associated speaker feeds as necessary. However, because the first stage dynamic processing has performed some initial limiting on the spatial mix, this limiting requirement may be reduced.
[0088] According to some examples of determining dynamic processing configuration data for a listening environment, adjustable combinations can be created that interpolate between minimum and average values of the dynamic processing configuration data for individual speakers via tuning parameters. For example, in the context of a playback limit threshold, the interpolation can be determined as follows:
number
[0089] Other combinations of individual speaker dynamics processing configuration data are possible, and this disclosure is meant to cover all such combinations.
[0090] 5A and 5B are graphs illustrating example dynamic range compression data. In graphs 500a and 500b, input signal level in decibels is shown on the horizontal axis and output signal level in decibels is shown on the vertical axis. As with other disclosed examples, the particular thresholds, ratios, and other values are shown by way of example only and not by way of limitation.
[0091] In the example shown in FIG. 5A, the output signal level is equal to the input signal level below a threshold, which in this example is −10 dB. Other examples can include different thresholds, such as −20 dB, −18 dB, −16 dB, −14 dB, −12 dB, −8 dB, −6 dB, −4 dB, −2 dB, 0 dB, 2 dB, 4 dB, and 6 dB. Above the threshold, various example compression ratios are shown. An N:1 ratio means that for every N dB increase in the input signal above the threshold, the output signal level increases by 1 dB. For example, a 10:1 compression ratio (line 505e) means that for every 10 dB increase in the input signal above the threshold, the output signal level increases by only 1 dB. A 1:1 compression ratio (line 505a) means that the output signal level remains equal to the input signal level above the threshold. Lines 505b, 505c, and 505d correspond to compression ratios of 3:2, 2:1, and 5:1. Other implementations may provide different compression ratios, such as 2.5:1, 3:1, 3.5:1, 4:3, 4:1, etc.
[0092] Figure 5B shows an example of a "knee" that controls how the compression ratio changes at or near a threshold, which in this example is 0 dB. According to this example, a compression curve with a "hard" knee consists of two straight line segments: a line segment 510a up to the threshold and a line segment 510b above the threshold. A hard knee is simple to implement, but can introduce artifacts.
[0093] FIG. 5B also shows an example of a "soft" knee. In this example, the soft knee spans 10 dB. According to this implementation, above and below the 10 dB range, the compression ratio of the compression curve with the soft knee is the same as the compression ratio of the compression curve with the stiff knee. Other implementations may provide various other shapes of "soft" knees, which may span more or fewer decibels, may exhibit different compression ratios above that range, etc.
[0094] Other types of dynamic range compression data can include "attack" data and "release" data. Attack is the period during which the compressor reduces the gain, e.g., in response to an increased level at the input, to reach a gain determined by the compression ratio. Compressor attack times typically range from 25 milliseconds to 500 milliseconds, although other attack times are possible. Release is the period during which the compressor increases the gain, e.g., in response to a decrease in level at the input, to reach an output gain determined by the compression ratio (or the input level if the input level falls below a threshold). Release times may range, e.g., from 25 milliseconds to 2 seconds.
[0095] Thus, in some examples, the individual speaker dynamic processing configuration data may include a dynamic range compression data set for each speaker of the plurality of speakers. The dynamic range compression data set may include threshold data, input / output ratio data, attack data, release data, and / or knee data. One or more of these types of individual speaker dynamic processing configuration data may be combined to determine the listening environment dynamic processing configuration data. As described above with respect to the combination of playback limit thresholds, in some examples, the dynamic range compression data may be averaged to determine the listening environment dynamic processing configuration data. In some examples, the minimum or maximum value of the dynamic range compression data may be used to determine the listening environment dynamic processing configuration data (e.g., maximum compression ratio). In other implementations, an adjustable combination that interpolates between the minimum and average values of the dynamic range compression data for the individual speaker dynamic processing may be created via an adjustment parameter, for example, as described above with reference to equation (C).
[0096] In some of the examples described above, in the first stage of dynamic processing, a single set of listening environment dynamic processing configuration data (e.g., a single set of the following combined thresholds) is applied to all components of the spatial mix:
number
number
number
[0097] <Example of zone-based method> To address these issues, some implementations allow independent or partially independent dynamic processing for various "spatial zones" of a spatial mix. A spatial zone can be thought of as a subset of the spatial region into which the entire spatial mix is rendered. The following discussion provides an example of dynamic processing based on playback limit thresholds, but the concepts apply equally to other types of individual speaker dynamic processing configuration data and listening environment dynamic processing configuration data.
[0098] Figure 6 shows example spatial zones of a listening environment. Figure 6 shows an example area of a spatial mix (represented by a whole square) subdivided into three spatial zones: front, center, and surround. Other examples may include more spatial zones, fewer spatial zones, different spatial zones, or a combination thereof. For example, some examples may include one or more overhead zones.
[0099] While the spatial zones in Figure 6 are shown with clear boundaries, in practice it is beneficial to treat the transition from one spatial zone to another as continuous. For example, a component of the spatial mix located at the center of the left edge of the square could have half its level assigned to the front zone and half assigned to the surround zone. The signal level from each component of the spatial mix is assigned and accumulated to each spatial zone in this continuous manner. Dynamic processing functions can then operate independently on each spatial zone at the overall signal level assigned from the mix. For each component of the spatial mix, the results of the dynamic processing (e.g., time-varying gain per frequency) from each spatial zone can then be combined and applied to the component. In some instances, this combination of spatial zone results will vary from component to component and be a function of that particular component's assignment to each zone. The net result is that components of the spatial mix with similar spatial zone assignments will receive similar dynamic processing, while allowing for independence between spatial zones. Spatial zones can be advantageously selected to prevent undesirable spatial shifts, such as left-right imbalance, while allowing for some spatially independent processing (e.g., to reduce other artifacts, such as the described spatial ducking).
[0100] Techniques for processing a spatial mix by spatial zones may be advantageously used in the first stage or multiple stages of dynamic processing referenced above (e.g., stage (a), stage (b), or both). For example, a different combination of individual speaker dynamic processing configuration data (e.g., playback limit thresholds) across speakers i may be calculated for each spatial zone. The set of combined zone thresholds may be represented by the following equation:
number
number
[0101] The spatial signal may be expressed as a total of K individual constituent signals x, each with an associated desired spatial location (possibly time varying). k One particular way to implement zonal processing is to compute the time domain of each audio signal x as a function of the desired spatial location of the audio signal relative to the location of the zone. k Time-varying panning gain α that describes how much [t] contributes to zone j kj These panning gains can be advantageously designed to follow the power conservation panning law, which requires that the sum of the squares of the gains equals 1. From these panning gains, the zone signal s j [t] can be calculated as the sum of the constituent signals weighted by the panning gain of that zone:
number
[0102] Each zone signal j [t] is the zone threshold:
number
number
[0103] The frequency and time varying correction gains are then adjusted to each individual constituent signal x by combining zone correction gains that are proportional to the panning gain of the signal for that zone. k For [t] we can calculate:
number
[0104] These signal-modifying gains G k is then dynamically processed, e.g., using a filter bank, to obtain the component signal x̂ k can be applied to each of the constituent signals to generate [t]. k [t] may later be rendered into a speaker signal.
[0105] The combination of the individual speaker dynamic processing configuration data (such as speaker playback thresholds) for each spatial zone can be performed in a variety of ways. For example, spatial zone playback thresholds:
number
number
[0106] Similar weighting functions can be applied to other types of individual speaker dynamic processing configuration data. Advantageously, the combined individual speaker dynamic processing configuration data (e.g., playback thresholds) of a spatial zone can be biased toward the individual speaker dynamic processing configuration data (e.g., playback thresholds) of the speakers primarily responsible for reproducing the component of the spatial mix associated with that spatial zone. This is done by assigning weights w as a function of each speaker responsible for rendering the component of the spatial mix associated with that zone for frequency f. ij This can be achieved by setting [f].
[0107] Figure 7 shows an example of loudspeakers within the spatial zones of Figure 6. Figure 7 shows the same zones as Figure 6, but overlaid with the locations of five exemplary loudspeakers (Speakers 1, 2, 3, 4, and 5) responsible for rendering the spatial mix. In this example, Speakers 1, 2, 3, 4, and 5 are represented by diamonds. In this particular example, Speaker 1 is primarily responsible for rendering the center zone, Speakers 2 and 5 are primarily responsible for rendering the front zone, and Speakers 3 and 4 are primarily responsible for rendering the surround zone. Based on this conceptual one-to-one mapping from speakers to spatial zones, weights w ij While it is possible to create [f], a more continuous mapping may be preferable, similar to spatial zone-based processing of a spatial mix. For example, speaker 4 may be very close to the front zone, and a component of the audio mix located between speakers 4 and 5 (though in the notional front zone) may be reproduced primarily by the combination of speakers 4 and 5. Therefore, it makes sense that speaker 4's individual speaker dynamic processing configuration data (e.g., playback limit threshold) contributes to the combined individual speaker dynamic processing configuration data (e.g., playback limit threshold) for the front and surround zones.
[0108] One way to achieve this continuous mapping is to use the weights w ij The first step is to set [f] equal to a speaker participation value that describes the relative contribution of each speaker i in rendering the component associated with spatial zone j. Such a value may be derived directly from the rendering system responsible for rendering to the speakers (e.g., from step (c) above) and a set of one or more nominal spatial locations associated with each spatial zone. This set of nominal spatial locations may include a set of locations within each spatial zone.
[0109] Figure 8 shows an example of nominal spatial locations overlaid on the spatial zones and speakers of Figure 7. The nominal locations are indicated by numbered circles. The two locations associated with the front zones are in the upper corners of the square, and the two locations associated with the surround zones are in the lower corners of the square.
[0110] To calculate the participation values for the speakers in a spatial zone, each of the nominal positions associated with the zone can be rendered through a renderer to produce activations for the speakers associated with that position. These activations can be, for example, the gain of each speaker in the case of CMAP, or a complex value at a given frequency for each speaker in the case of FV. Then, for each speaker and zone, these activations are accumulated over each of the nominal positions associated with the spatial zone to produce a value g ij This value represents the total activation of speaker i for rendering the entire set of nominal positions associated with spatial zone j. Finally, the speaker participation value in a spatial zone is calculated as the cumulative activation g normalized by the sum of all of these cumulative activations across speakers. ij [f]. A weight may then be set to this speaker participation value.
number
[0111] The normalization described is w across all speakers i. ij It ensures that the sum of [f] equals 1, which is a desirable property for weights in formula (G).
[0112] According to some implementations, the above process for calculating speaker participation values and combining thresholds as a function of these values may be executed as a static process, and the resulting combined thresholds are calculated once during a setup procedure that determines the layout and capabilities of the speakers in the environment. In such a system, after setup, it may be assumed that both the dynamic processing configuration data for individual speakers and the way the rendering algorithm activates the speakers as a function of the desired audio signal positions remain static. However, in certain systems, both of these aspects may change over time, for example, in response to changing conditions in the playback environment, and in order to account for such changes, it may be desirable to update the combined thresholds according to the above process either continuously or in a manner triggered by events.
[0113] <Example of mapping of audio signal components including <effort> signal> The time - and frequency - varying mapping of the component signals of the spatial audio mix, also referred to here as audio objects, can generally be represented by the following equation: [Number]
[0114] Here, the variables t and f represent the time and frequency variations of the audio object signal O i , the speaker signal S j , and the mapping H from object i to speaker signal j ij . The number of audio objects is given by N o , where N o ≥ 2, and the number of speaker signals is given by N s , where N s ≥ 2. The mapping H ijcan be generally thought of as a time-varying filter whose form and application can take many forms, such as a real or complex time-varying gain applied to individual bands of a filter bank such as a quadrature mirror filter (QMF) or a short-time Fourier transform (STFT), or a time-varying finite impulse response (FIR) or infinite impulse response (IIR) filter applied to an audio object in the time domain.
[0115] According to this example, each audio object signal O i Associated with is the following intended perceptual spatial location, which may change over time:
number
number
number
number
[0116] For a given audio object signal i and speaker j, the signal E ij Let (f,t) represent the time- and frequency-varying representation of the speaker signal level associated with object i relative to the maximum playback limit of speaker j. For brevity, this is hereafter referred to as the effort signal. As the level of the effort signal increases, this indicates that rendering object i on speaker j will result in speaker j approaching or exceeding its playback limit threshold. The set {E ik (f,t)} is the sum of all the speakers k=1...N s represents the effort signal associated with object i for
[0117] Mapping H ij The calculation of is then done for a set of object positions:
number
number
number
[0118] In this example, the mapping function M j is the mapping H across all loudspeaker signals ij Calculates the position of all speaker signals:
number
number
number
[0119] The mathematically described operations above encapsulate the high-level idea that, in some instances, as the level of the spatial mix approaches the reproduction threshold of a particular speaker, the mapping of components to that speaker is reduced in order to increase mappings to other speakers whose spatial mix level is further away from their reproduction thresholds.
[0120] Generally, the frequency variation mapping described above is performed across the entire audible frequency range and can employ a frequency resolution consistent with human perception. For example, in one embodiment, the mapping can be calculated for 20 discrete frequency bands with a resolution of approximately 2 Equivalent Rectangular Bandwidths (ERBs). Utilizing such spacing helps maintain the perceptual transparency of the system.
[0121] However, in other embodiments, it may be advantageous to calculate the dynamic mapping only over a subset of the available frequency range of the speakers. For example, the dynamic mapping may be calculated only for frequencies below a threshold frequency, such as 500 Hz, which is the range where there may be differences in speaker capabilities. If all speakers may be equally capable above 500 Hz or another threshold in some instances, the mapping may be independent of signal level in some such instances.
[0122] According to some examples, "approximating" the intended perceptual spatial location of the associated audio signals includes minimizing a cost function. In some such examples, the mapping described by Equation 2a is used to minimize the cost function by minimizing all speaker signals to a given location.
number
number
number
[0123] In Equation 2b, C(g) is the N s We express the cost C as a function of g, which represents a dimensional vector. In Equation 2b, we consider the set:
number
number
number
number
number
number
number
number
number
[0124] Examples of {o^} include, but are not limited to: the desired perceived spatial location of the audio signal, the level of the audio signal (which may vary over time), and / or The spectrum of the audio signal (which may vary over time).
[0125] {s k Examples of ^} include, but are not limited to: The position of the speakers in the listening space, -Speaker frequency response, - Speaker playback level limit, - Parameters of dynamic processing algorithms within the speaker, such as limiter gain, Measurement or estimation of the acoustic transmission from each loudspeaker to the other loudspeakers; Measuring the echo cancellation performance of the loudspeaker, and / or Relative synchronization between speakers.
[0126] Examples of {e^} include, but are not limited to: one or more listener or speaker positions within the playback space, Measurement or estimation of the sound transmission from each loudspeaker to the listening position, Measurement or estimation of acoustic transmission from a talker to a set of loudspeakers, The location of any other landmarks within the playback space, and / or Measurement or estimation of the acoustic transmission from each loudspeaker to any other landmark in the reproduction space.
[0127] effort signal E ij By mapping (f,t) to the speaker-specific activation penalty in Equation 2b, we obtain the mapping function M j The aforementioned goal of s Mapping over H ij More specifically, for any particular frequency f, time t, and object signal i,
number
[0128] effort signal E ij The exact ways in which (f,t) can represent the speaker signal level relative to the maximum playback limit vary, and the inventors consider many possible options. In some embodiments, the effort signal may be or correspond to a digital level that is either the input to or output from a limiter. In other embodiments, the effort signal may be the actual gain applied by the limiter, or a signal indicating that the speaker level has exceeded the playback threshold. In some other implementations, the effort signal may be an acoustic signal (e.g., a measurement of the sound pressure level in decibels at a particular distance (dBSPL)). The acoustic signal may be derived from the digital level using, for example, speaker sensitivity and known analog amplifier gain.
[0129] Whatever the particular form of expression, the effort signal E ijThe construction of (f,t) is a key component to the functionality of many examples in this disclosure because in these examples, the effort signal E ij This is because (f,t) indicates where and when the transfer of signal energy between loudspeakers occurs. As mentioned above, the effort signal may be calculated for each audio object signal as a function of one or more of the audio object signals and their perceptual spatial location, for example as follows:
number
[0130] By considering the entire spatial mix in their construction, signal energy from any part of the spatial mix may be combined into each object's effort signal for various purposes. One of these possible purposes is to combine object signals in a way that represents how likely the audio object signal is to accumulate in the speaker signal. Another possible purpose is to couple the energy redistribution behavior of the present invention between objects with a particular spatial relationship. Some preferred constructions of the effort signal can achieve both of these purposes.
[0131] In considering the preferred configuration of the effort signal, it is useful to revisit their more detailed definition: a time- and frequency-varying representation of the loudspeaker signal level associated with object i relative to the maximum playback limit of loudspeaker j. This notion of a general representation can be literally translated into the following specific configuration of the effort signal:
number
[0132] In Equation 4, the term L i (f,t) represents the time- and frequency-varying signal level associated with object i, and τ j(f) represents the frequency-varying playback limit of speaker j. In some examples, we assume that the playback limit of the speaker, expressed in decibels in this example, as well as the position of the speaker, have already been characterized and provided to the control system. Therefore, the calculation of the effort signal of the audio object across all speakers can be performed using a single level signal L i This simplifies to the calculation of (f,t).
[0133] As mentioned before, the mapping H ij depends on these audio object level signals L i (f,t) should represent how likely it is that the audio object signal will be accumulated in the speaker signal. This is done by the mapping H, as shown in Equation 2a. ij is used to explicitly generate the speaker signals, which is a circular relationship. One solution to this circularity problem is to generate a simple zone-based rendering of the audio object, where no information about the speakers is required. Instead, the signal energy for a finite set of spatial zones is generated by a simple panning rule. In general, a control system can be configured with N z zones can be used. In one useful embodiment, the rendering process can include four zones, which can be front, center, surround, and overhead zones. Other examples can include more or fewer zones. In some such examples, rendering audio objects into zones can be controlled by simple broadband panning rules that are a function of the intended spatial location of the audio objects:
number
[0134] In Equation 5, g il (t) represents the panning gain of audio object i to spatial zone l. The panning gain can be calculated, for example, as follows: zIt is beneficial to be power conserving across a set of zones:
number
[0135] From these zonal panning gains, a time and frequency varying zonal power spectrum Z is calculated, for example, by accumulating the power spectra of the audio objects weighted by the corresponding panning gains, as follows: l (f,t) can be calculated.
number
[0136] These zonal power spectra correspond to the energy distribution of the overall spatial mix within the designed spatial zone. Assuming that loudspeakers are distributed throughout the zones, e.g., each zone preferably contains at least one loudspeaker, these power spectra represent a reasonable approximation of the signal levels that would appear in the loudspeaker signals within the zone.
[0137] Specific to this disclosure, and to the best of the inventor's knowledge, in some instances, the zonal power spectrum can be adjusted to the object level signal L by applying a zonal panning gain for each object. i We can map it back to (f,t), for example:
number
[0138] To the extent that audio object i is panned across multiple zones, L according to Equation 8 i The construction of (f,t) combines the zonal power spectra in a way that is proportional to this panning. This allows L i(f,t) serves as an approximation of the level of the loudspeaker signal at which object i will be rendered. Furthermore, the signal L for objects that mostly belong to the same zone i (f,t) are similar, which means that the dynamic energy redistribution of the present invention behaves similarly for these objects. In some advantageous embodiments, zones can be designed to couple objects across the left / right axis, thereby constraining the energy redistribution behavior to be similar for the left and right of the spatial mix. This zone design helps reduce perceptually confusing left / right image shifts.
[0139] Finally, in some instances, the object level signal L i (f, t) is the effort signal E as shown in Equation 4, for example. ij To calculate (f,t), the speaker threshold τ j Can be combined with (f).
[0140] When a spatial audio mix is composed of channel-based audio (i.e., each audio object signal may correspond to a channel of a multichannel signal such as a Dolby 5.1 or Dolby 5.1.4 signal, and the desired location of each object is a fixed spatial location), the zone-based construction of the level signals described above can be simplified along several dimensions. Because the locations of the object / channel signals are fixed, the mapping of channels to zones may not require a dynamic panning function such as that shown in Equation 5. Furthermore, this mapping can be simplified to a set of one-to-one mappings from channels to zones and vice versa. For example, for a 5.1.4 signal, the front left and right channels may map to the front zone, the center channel may map to the center zone, the left and right surround channels may map to the surround zone, and all four overhead channels may map to the overhead zone. With this mapping, Equation 7 for calculating the zone power spectrum simplifies to summing the power spectra of each channel belonging to that zone, and Equation 8 simplifies to equating each channel's level signal to the zone power spectrum corresponding to the single zone to which that channel is mapped.
[0141] The above methods for calculating object-level signals rely on an approximation of how objects accumulate in the speaker signals. To eliminate this approximation, an alternative construction of the object-level signal can be used that uses direct feedback of the speaker signals from a previous time interval, at the expense of introducing a delay in the activation of the energy redistribution of the present invention. Therefore, this second category of methods can be referred to herein as feedback methods.
[0142] One such alternative implementation is similar to the method described in Equation 8, except that the zone power spectrum is replaced by the power spectrum of the speaker signal from the previous time interval, and the object-to-zone panning gain is replaced by a normalized version of the object-to-speaker mapping from the previous time interval. For example,
number
[0143] The normalization of the object-to-speaker mapping shown in Equation 9b ensures that the weighted combination of speaker signals in Equation 9a is performed in a power-conserving manner. If the object-to-speaker mapping is calculated in an inherently power-conserving manner, as is the case in some advantageous embodiments, this normalization step is not necessary.
[0144] The construction of the object-level signal in Equation 9a combines all speaker signals from the previous time interval in proportion to the mapping of that audio object to each speaker signal. As such, it is a direct representation of the speaker signal levels to which the audio object is mapped. However, one problem with this construction is that it lacks the concept of relating audio objects in similar spatial zones. Without this capability, instability can result in the imaging of the rendered audio.
[0145] As a third example of constructing an audio object level signal, the zone-based processing of the first example can be combined with the more accurate representation of speaker signal levels provided by the feedback method of the second example. In some such alternatives, the audio object level signal can be constructed as a weighted sum of zone signals, as in the first example, but with a zone power spectrum Z l (f,t) is the speaker zone power spectrum V as follows: l It can be replaced with (f,t).
number
[0146] The zone power spectrum is calculated directly from the object signal, whereas the speaker zone power spectrum is calculated as a weighted sum of the speaker signal power spectra from previous time intervals.
number
[0147] Weighting P jl is sometimes referred to herein as the "speaker zone participation value", which means that it is a measure of the portion of the energy of zone signal l that is rendered to speaker j. jl is a set of raw speaker zone participation values P jl It can be derived by normalizing over ^.
number
[0148] Raw speaker zone participation value P jl ^ is the mapping function M j for each zone, representing the spatial extent of that zone, using l Spatial location of the individual:
number
number
[0149] The effort signal E associated with this simulated rendering lj ^(f,t) is, in some instances, the zonal power spectrum Z defined in Equation 7 l It can be calculated from (f,t).
number
[0150] In summary, the object level signal L i Three different methods for constructing (f,t) are described above: 1) a zone-based method, 2) a feedback-based method, and 3) a hybrid method that combines elements of the zone-based and feedback-based methods.
[0151] In all of the methods described above for constructing an object level signal, further processing can be applied to reduce perceptual artifacts. For example, the object level signal L i (f,t) can be smoothed over time, frequency, or both to regularize the resulting energy fluctuations spread across these dimensions.
[0152] As shown in Equation 4, the effort signal E ij By defining (f,t) the intended audio object position:
number
number
number
[0153] 9 is a graph of points showing the mapping of audio objects to speakers as a function of the audio object's x, y, and z coordinates in an exemplary embodiment. In this example, the x and y dimensions are sampled with 15 points, and the z dimension is sampled with 5 points. Other implementations may include more or fewer samples. According to this example, each point is a point j=1...N for audio object i with the (x, y, z) location corresponding to that point. s A mapping H for a set of speakers ij Represents.
[0154] At run time, to determine the actual mapping for each speaker, in some examples, trilinear interpolation between the speaker mappings of the nearest eight points can be used. Figure 10 is a graph of trilinear interpolation between points showing the mapping to speakers, according to an example. In this example, the process of successive linear interpolation includes interpolating each pair of points on the top surface to determine first and second interpolated points 1005a and 1005b, interpolating each pair of points on the bottom surface to determine third and fourth interpolated points 1010a and 1010b, interpolating the first and second interpolated points 1005a and 1005b to determine a fifth interpolated point 1015 on the top surface, interpolating the third and fourth interpolated points 1010a and 1010b to determine a sixth interpolated point 1020 on the bottom surface, and interpolating the fifth and sixth interpolated points 1015 and 1020 to determine a seventh interpolated point 1025 between the top and bottom surfaces. While trilinear interpolation is an effective interpolation method, those skilled in the art will understand that trilinear interpolation is only one possible interpolation method that may be used in implementing aspects of the present disclosure, and that other examples may include other interpolation methods.
[0155] The above method can be further extended to cover varying audio object signal levels by adding a fourth dimension to the lookup table. This lookup along the fourth dimension of signal level is, in some instances, performed independently across a set of frequency bands, L i The exact nature of how (f,t) varies over frequency can be explicitly captured.
[0156] Some alternative channel-based examples include not performing trilinear interpolation of positions (or in some examples, any interpolation), but performing a level of interpolation, such as linear interpolation.
[0157] In another alternative embodiment, the object level signal may be approximated as a wideband gain multiplied by a prototype spectral shape, for example:
number
[0158] In Equation 15, the prototype spectral shape L p (f) is chosen to represent the average spectral shape associated with the content expected to be processed by the system. i (t) is the estimated object level signal L calculated by one of the methods mentioned above. i (f,t) and its approximations:
number
[0159] Here, according to Equation 2b, the effort signal E ij Mapping H as a function of (f,t) ij Returning to the calculation of , during the optimization of the mapping for each object i, we calculate the effort signal by the per-speaker penalty P j(f). Considering the construction of the effort signal in Equation 4, note that the effort signal is less than 0 when the audio object level signal is less than the speaker playback threshold, and greater than 0 when the audio object signal is greater than this threshold. When mapping these effort signals to speaker penalties, it is useful to apply a transformation such that the penalty increases smoothly and monotonically from a value of 0 as the effort signal increases beyond a specified transition point. Specifying this transition point to correspond to an effort value less than 0 means that the penalty begins to take effect before the signal level reaches the speaker's playback threshold. In such instances, it is advantageous to gradually divert energy away from the speaker before the signal level reaches the playback threshold. According to such instances, the knee parameter K can be used to control the transition point and the rate at which the penalty increases as the effort signal increases.
[0160] Figure 11 shows examples of penalty for different knee parameters. Figure 11 shows how the speaker penalty increases for values of the knee parameter ranging from -24 dB to -6 dB.
[0161] It can be seen that below the value of the knee parameter the penalty is 0 (this value does not affect the mapping calculation). According to these examples, the penalty increases monotonically above the knee value. This means that as the effort signal increases, more and more energy is diverted from the associated speaker. Furthermore, it can be seen that the larger the knee value, the more gradually the speaker penalty is activated. In practice, the inventors have determined that knee values in the range of -18 dB to -12 dB work well.
[0162] The knee value K is an example of a parameter that can be used to control the degree of energy diversion between speakers as a function of signal level. In some embodiments, this degree of energy diversion may be fixed for all content processed by the system. In other embodiments, this degree of diversion may be additionally controlled based on the input audio format, codec, or metadata. For example, some codecs may incorporate metadata that indicates that certain components of a mix are suitable for more energy diversion than other components. This metadata, or the basis for the metadata, may be controlled by a content creator, who may specify different knee values for different audio objects in a Dolby Atmos™ mix, for example. Furthermore, in some examples, the degree of energy diversion may depend at least in part on the content format. For example, the degree of energy diversion for the center channel of a Dolby 5.1 mix may be set to be less than the degree of energy diversion for the other channels to ensure that the perceived location of dialogue remains centered.
[0163] 12 is a flow diagram outlining one example of a method that may be performed by an apparatus or system such as those disclosed herein. The blocks of method 1200, as well as other methods described herein, are not necessarily performed in the order presented. In some implementations, one or more blocks of method 1200 may be performed simultaneously. Furthermore, some implementations of method 1200 may include more or fewer blocks than those shown and / or described. The blocks of method 1200 may be performed by one or more devices, which may be (or may include) a control system, such as control system 110 shown in FIG. 1 and described above, or one of the other disclosed example control systems.
[0164] According to this example, block 1205 includes receiving audio data by the control system and via the interface system. In this example, the audio data includes one or more audio signals and associated spatial data, where the spatial data indicates an intended perceived spatial location corresponding to the audio signals. In some examples, the spatial data may be or include spatial metadata of an object-based audio format, such as Dolby Atmos™. In some examples, the intended perceived spatial location, as disclosed herein, may be expressed as follows:
number
[0165] In this example, block 1210 includes rendering, by the control system, the audio data for playback through a set of two or more speakers in the environment to generate speaker signals. According to this example, rendering each of the one or more audio signals included in the audio data includes a time- and frequency-varying mapping of each audio signal to a speaker signal. In this example, the mapping of each audio signal is calculated as a function of the intended perceived spatial location of the audio signal, the physical locations relative to the speakers, and a time- and frequency-varying representation of the speaker signal level relative to the maximum playback limit of each speaker. According to this example, each mapping is calculated to approximately achieve the intended perceived spatial location of the associated audio signal when the speaker signal is played on two or more corresponding speakers located at the associated speaker locations.
[0166] According to some examples, "approximately achieving" the intended perceptual spatial location of the associated audio signal may include minimizing the difference between the perceptual spatial location and the intended perceptual spatial location, given the available speakers and the associated speaker locations. In some examples, approximately achieving the intended perceptual spatial location of the associated audio signal may include minimizing a cost function, such as one of the cost functions disclosed herein. Equation 2b of the present disclosure encompasses many different possibilities depending on the choice of cost term. For example, the additional cost term {s i Depending on the selection of {e^} and {e^}, various implementations can be achieved.
[0167] In this example, block 1210 includes calculating, for each audio signal, a representation of speaker signal levels relative to a maximum playback limit as a function of one or more of the audio signals and their perceptual spatial location. According to this example, block 1210 includes decreasing the mapping of an audio signal to a particular speaker signal as the representation of the speaker signal levels relative to the maximum playback limit increases above a threshold. Further, in this example, block 1210 includes increasing the mapping to one or more other speakers for which the representation of the signal levels relative to the maximum playback limit of the one or more other speakers is below a threshold.
[0168] According to some examples, the mapping may be calculated over the entire frequency range of normal human hearing. However, in some examples, the mapping may be calculated over a subset of the frequency range. According to some examples, the mapping may include minimizing a cost function including a first term that models how closely an intended perceptual spatial location is achieved as a function of mapping the audio signal to the speaker signals, and a second term that assigns a cost for activating each speaker. Equation 2b provides an example of both the first and second terms. In some examples, the cost of activating each speaker may be based at least in part on a function of a representation of the speaker signal level relative to a maximum playback limit.
[0169] In some examples, the representation of the speaker signal level relative to the maximum playback limit may correspond to one or more of a digital signal level, a limiter gain, or an acoustic signal level. In some examples, the representation of the speaker signal level relative to the maximum playback limit may be calculated as a difference between the level estimate for each audio signal and a playback limit threshold for each speaker. In some examples, the level estimate for each audio signal may be based at least in part on zone-based rendering of all audio signals. In some examples, the level estimate for each audio signal may be based at least in part on previously calculated speaker signals. In some such examples, the level estimate for each audio signal may further depend on the participation of each speaker in multiple spatial zones. According to some examples, block 1210 or other blocks of method 1200 may further include smoothing the level estimate for each audio signal over time, frequency, or both time and frequency.
[0170] According to some examples, the mapping from audio signals to speaker signals may be determined by querying a data structure indexed by the intended perceptual spatial location and level estimate of each audio signal. In some examples, the mapping from audio signals to speaker signals may include determining speaker activations. In some such examples, speaker activations may be determined by interpolating from a set of pre-computed speaker activations. According to some such examples, the set may be indexed by the intended perceptual spatial location and level estimate for each audio signal.
[0171] In some examples, the level estimate for each audio signal can be expressed as a wideband gain multiplied by a spectral shape. According to some such examples, the spectral shape can be selected from a plurality of spectral shapes. Each spectral shape of the plurality of spectral shapes can correspond to, for example, a content type. Content types can include, for example, movie content, television program content, podcast content, music performance content, gaming content, etc.
[0172] According to some examples, as the representation of the signal level relative to the maximum playback level increases above a threshold, the mapping to one speaker may be decreased and the mapping to another speaker may be increased. Figure 11 and the corresponding description provide some examples. In some examples, method 1200 may further include controlling the degree of the decrease in the mapping to one speaker and the increase in the mapping to another speaker according to one or more of an audio format, a codec, or metadata. In some examples, method 1200 may further include controlling the degree of the decrease in the mapping to one speaker and the increase in the mapping to another speaker according to a knee parameter.
[0173] In this implementation, block 1215 includes providing speaker signals to at least two speakers of a set of speakers of the environment via an interface system.
[0174] Some implementations of the present disclosure include systems or devices configured (e.g., programmed) to perform any of the disclosed method embodiments, and tangible computer-readable media (e.g., disks) storing code for implementing the disclosed methods or steps thereof. For example, the disclosed systems may be or include programmable general-purpose processors, digital signal processors, or microprocessors programmed with software or firmware and / or otherwise configured to perform any of the various operations on data, including the disclosed method embodiments or steps thereof. Such general-purpose processors may be or include computer systems including input devices, memory, and processing subsystems programmed (and / or otherwise configured) to perform the disclosed method embodiments (or steps thereof) in response to data asserted thereto.
[0175] Some embodiments of the disclosed systems are implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on audio signals, including performing embodiments of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) are implemented as a general-purpose processor (e.g., a personal computer (PC), other computer system, or microprocessor (which may include input devices and memory)) that is programmed with software or firmware and / or otherwise configured to perform any of various operations, including embodiments of the disclosed methods. Alternatively, elements of some embodiments of the disclosed systems are implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform embodiments of the disclosed methods, and the system also includes other elements (e.g., one or more speakers and / or one or more microphones). A general-purpose processor configured to perform embodiments of the disclosed methods may typically be coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device.
[0176] Another aspect of the present disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) that stores code (e.g., executable code) for performing embodiments of the disclosed methods or steps thereof.
[0177] While particular embodiments and applications have been described herein, it will be apparent to those skilled in the art that many variations to the embodiments and applications described herein are possible without departing from the scope of what is described and claimed herein. It should be understood that, although particular forms have been shown and described, the scope of the disclosure is not limited to the specific embodiments described and shown or to the particular methods described.
Claims
1. 1. A method of audio processing, comprising: receiving, by a control system via an interface system, audio data comprising one or more audio signals and associated spatial data, the spatial data indicating intended perceived spatial locations corresponding to the audio signals; rendering, by the control system, the audio data for playback through a set of two or more speakers in an environment to generate speaker signals; rendering each of the one or more audio signals included in the audio data includes mapping each audio signal to the speaker signals, the mapping being time and frequency-varying; a mapping of each audio signal is calculated as a function of the intended perceptual spatial location of the audio signal, the physical locations associated with said speakers, and a time- and frequency-varying representation of the speaker signal level relative to the maximum playback limit of each speaker; each mapping is calculated to approximately achieve the intended perceived spatial location of an associated audio signal when the speaker signal is played on the two or more corresponding speakers located at associated speaker locations; a representation of speaker signal levels relative to a maximum playback limit is calculated for each audio signal as a function of one or more of said audio signals and their perceived spatial locations; a mapping of audio signals to a particular speaker signal is decreased as the representation of a speaker signal level relative to a maximum playback limit increases above a threshold, and the mapping is increased to one or more other speakers whose representation of signal level relative to a maximum playback limit of the one or more other speakers is below a threshold; providing the speaker signals via the interface system to at least two speakers of the set of speakers in the environment; 1. An audio processing method comprising:
2. 2. The audio processing method of claim 1, wherein the mapping is calculated over the entire audible frequency range.
3. 2. The audio processing method of claim 1, wherein the mapping is calculated over a portion of the audible frequency range.
4. 4. The method of claim 1, wherein the mapping comprises minimizing a cost function comprising a first term that models how closely the intended perceptual spatial location is achieved as a function of mapping audio signals to loudspeaker signals, and a second term that assigns a cost for activating each of the loudspeakers.
5. The method of claim 4 , wherein the cost for activating each speaker is based at least in part on a function of the representation of speaker signal level relative to the maximum playback limit.
6. The method of any one of claims 1 to 5, wherein the representation of a speaker signal level relative to the maximum playback limit corresponds to one or more of a digital signal level, a limiter gain, or an acoustic signal level.
7. A method according to any one of claims 1 to 5, wherein the representation of loudspeaker signal levels relative to the maximum playback limit is calculated as the difference between a level estimate for each audio signal and a playback limit threshold for each loudspeaker.
8. The method of claim 7 , wherein the level estimation for each audio signal is based at least in part on a zone-based rendering of all audio signals.
9. 9. The method of claim 7 or 8, wherein the level estimate for each audio signal is based at least in part on previously calculated speaker signals.
10. The method of claim 9 , wherein the level estimate of each audio signal further depends on the participation of each speaker in multiple spatial zones.
11. The method of any one of claims 7 to 10, further comprising smoothing the level estimates of each audio signal over time, frequency, or both time and frequency.
12. 12. The method of any one of claims 7 to 11, wherein the mapping from audio signals to speaker signals is determined by querying a data structure indexed by the intended perceptual spatial position and level estimate of each audio signal.
13. 12. The method of claim 7, wherein the mapping from audio signals to loudspeaker signals is determined by interpolating from a set of pre-calculated loudspeaker mappings, the set being indexed by the intended perceptual spatial position and level estimate of each audio signal.
14. 12. The method of claim 7, wherein the mapping from audio signals to loudspeaker signals is determined by interpolating from a set of pre-calculated loudspeaker mappings, the set being indexed by the intended level estimate for each audio signal.
15. 14. The method of claim 12 or 13, wherein the level estimate for each audio signal is expressed as a wideband gain multiplied by a spectral shape.
16. The method of claim 15 , wherein the spectral shape is selected from a plurality of spectral shapes, each spectral shape of the plurality of spectral shapes corresponding to a content type.
17. A method according to any preceding claim, wherein as the representation of signal level relative to maximum playback level increases above a threshold, the mapping to one loudspeaker is decreased and the mapping to another loudspeaker is increased.
18. 18. The method of any one of claims 1 to 17, further comprising controlling the degree of reduction in mapping to one speaker and the increase in mapping to another speaker according to one or more of the audio format, codec, or metadata.
19. The method of any one of claims 1 to 18, further comprising controlling the degree of decrease in mapping to one loudspeaker and the increase in mapping to another loudspeaker according to a knee parameter.
20. 20. The method of any one of claims 1 to 19, wherein the intended perceptual spatial location corresponds to a channel of a channel-based audio format, corresponds to metadata, or corresponds to both a channel and metadata.
21. 21. The method of any one of claims 1 to 20, wherein approximately achieving the intended perceived spatial position of an associated audio signal comprises minimizing the difference between the perceived spatial position and the intended perceived spatial position given available speakers and associated speaker positions.
22. The method of any one of claims 1 to 21, wherein approximately achieving the intended perceptual spatial location of associated audio signals comprises minimizing a cost function.
23. Apparatus configured to carry out the method of any one of claims 1 to 22.
24. A system configured to carry out the method of any one of claims 1 to 22.
25. One or more non-transitory media storing instructions, the instructions controlling one or more devices to perform the method of any one of claims 1 to 22.
Citation Information
Patent Citations
Adaptable spatial audio playback
WO2021021460A1
Information processing device, information processing method, and program
WO2022064905A1