Rendering of audio objects with apparent size to arbitrary loudspeaker layout
The method addresses the complexity of three-dimensional audio rendering by using metadata to calculate gain values for virtual sources, improving sound localization and immersion in diverse playback environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-03-04
AI Technical Summary
The increasing complexity of authoring and rendering sound in three-dimensional audio environments, particularly in movie theater sound systems, due to the transition from planar to three-dimensional speaker layouts, poses challenges in effectively positioning and rendering audio objects.
A method for rendering audio playback data that generates audio objects without reference to a specific playback environment, using metadata to define position, size, and other attributes, and calculates gain values for each channel based on virtual source positions within a defined volume, allowing for dynamic rendering of audio objects in various playback environments.
Enables efficient and complex audio rendering in multi-dimensional speaker layouts by calculating weighted averages of virtual source gains, enhancing sound localization and immersion in environments like movie theaters.
Smart Images

Figure 2026035629000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to Spanish Patent Application No. P201330461, filed March 28, 2013, and U.S. Provisional Patent Application No. 61 / 833,581, filed June 11, 2013, the contents of each of which are incorporated herein by reference in their entirety.
[0002] Technical Field FIELD OF THE DISCLOSURE This disclosure relates to authoring and rendering audio reproduction data, and in particular to authoring and rendering audio reproduction data for reproduction environments such as movie theater sound reproduction systems. [Background technology]
[0003] Since the introduction of sound to motion pictures in 1927, there has been a steady advancement in the technology used to capture the artistic intent of a film soundtrack and reproduce it in a cinema environment. In the 1930s, synchronized sound on disc gave way to variable-area sound on film, which was further improved in the 1940s by considerations of theatre acoustics and improved speaker design. Along with this came the early introduction of multitrack recording and directional playback (using control tones to move the sound). In the 1950s and 1960s, magnetic stripes on film allowed for multichannel playback in cinemas, introducing surround channels and up to five screen channels in premium theatres.
[0004] In the 1970s, Dolby introduced noise reduction, both in postproduction and on film, along with a cost-effective means of encoding and distributing three screen channels mixed with a mono surround channel. Cinema sound quality was further improved in the 1980s with Dolby Spectral Recording (SR) noise reduction and certification programs such as THX. Dolby brought digital sound to cinemas in the 1990s with the 5.1-channel format, which provided discrete left, center, and right screen channels, left and right surround arrays, and a subwoofer channel for low-frequency effects. Dolby Surround 7.1, introduced in 2010, increased the number of surround channels by dividing the existing left and right surround channels into four "zones."
[0005] As channel counts increase and speaker layouts transition from planar, two-dimensional (2D) arrays to three-dimensional (3D) arrays that include height, the task of authoring and rendering sound becomes increasingly complex. Improved methods and apparatus would be desirable. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] V. Pulkki, Compensating Displacement of Amplitude-Panned Virtual Sources, Audio Engineering Society (AES) International Conference on Virtual, Synthetic and Entertainment Audio [Non-patent document 2] D. de Vries, Wave Field Synthesis, AES Monograph 1999 Summary of the Invention [Means for solving the problem]
[0007] Some aspects of the subject matter described in this disclosure can be implemented in a tool for rendering audio playback data, including audio objects, that are generated without reference to any particular playback environment. As used herein, the term "audio object" may refer to a stream of audio signals and associated metadata. The metadata may indicate at least the position and apparent size of the audio object. However, the metadata may also indicate rendering constraint data, content type data (e.g., dialogue, effects, etc.), gain data, trajectory data, etc. Some audio objects may be static, while other audio objects may have time-varying metadata: such audio objects may move, change size, and / or have other attributes that change over time.
[0008] When an audio object is monitored or played in a playback environment, the audio object may be rendered according to at least the position and size metadata. The rendering process may involve calculating a set of audio object gain values for each channel of a set of output channels. Each output channel may correspond to one or more playback speakers in the playback environment.
[0009] Some implementations described herein involve a "setup" process that may occur prior to rendering any particular audio object. This setup process, sometimes referred to herein as stage 1, may involve defining multiple virtual source positions within a volume within which the audio object can move. As used herein, a "virtual source position" is the location of a static point source. According to such an implementation, the setup process may involve receiving playback speaker position data and pre-calculating virtual source gain values for each of the virtual sources according to the playback speaker position data and the virtual source positions. As used herein, "speaker position data" may include position data indicating the locations of some or all of the speakers in the playback environment. The position data may be provided as absolute coordinates of the playback speaker positions, e.g., Cartesian coordinates, spherical coordinates, etc. Alternatively or additionally, the position data may be provided as coordinates (e.g., Cartesian coordinates or angular coordinates) relative to other playback environment locations, such as the playback environment's acoustic "sweet spot."
[0010] In some implementations, the virtual source gain values may be stored and used during "runtime" when audio playback data is rendered for speakers of the playback environment. During runtime, for each audio object, contributions from virtual source positions within an area or volume defined by the audio object position data and audio object size data may be calculated. The process of calculating the contributions from the virtual source positions may involve calculating a weighted average of multiple pre-calculated virtual source gain values determined during the setup process for virtual source positions within the audio object area or volume defined by the audio object size and position. A set of audio object gain values for each output channel of the playback environment may be calculated based, at least in part, on the calculated virtual source contributions. Each output channel may correspond to at least one playback speaker of the playback environment.
[0011] Thus, some methods described herein involve receiving audio playback data including one or more audio objects. The audio objects may include an audio signal and associated metadata. The metadata may include at least audio object position data and audio object size data. The methods may involve calculating contributions from virtual sources within an audio object region or volume defined by the audio object position data and audio object size data. The methods may involve calculating a set of audio object gain values for each of a plurality of output channels based, at least in part, on the calculated contributions. For example, the playback environment may be a cinema sound system environment.
[0012] The process of calculating the contributions from the virtual sources may involve calculating a weighted average of the virtual source gain values from the virtual sources within the audio object region or volume, where the weights for the weighted average may depend on the position of the audio object, the size of the audio object, and / or the position of each virtual source within the audio object region or volume.
[0013] The methods may also involve receiving playback environment data including playback speaker position data. The methods may also involve defining a plurality of virtual source positions according to the playback environment data and calculating, for each virtual source position, a virtual source gain value for each of the plurality of output channels. In some implementations, each of the virtual source positions may correspond to a position within the playback environment. However, in some implementations, at least some of the virtual source positions may correspond to positions outside the playback environment.
[0014] In some implementations, the virtual source positions may be uniformly spaced along the x-, y-, and z-axes. However, in some implementations, the spacing may not be the same in all directions. For example, the virtual source positions may have a first uniform spacing along the x- and y-axes and a second uniform spacing along the z-axis. The process of calculating a set of audio object gain values for each of the multiple output channels may involve independent calculation of contributions from virtual sources along the x-, y-, and z-axes. In alternative implementations, the virtual source positions may be non-uniformly spaced.
[0015] In some implementations, the process of calculating an audio object gain value for each of the plurality of output channels comprises: o ,y o ,z o The gain value (g) for an audio object of size (s) to be rendered in l (x o ,y o ,z o For example, the audio object gain value (g l (x o ,y o ,z o ;s)) is
number
[0016] According to some such implementations, g l (x vs ,y vs ,z vs )=g l (x vs )g l (y vs )g l (z vs ), where g l (x vs ), g l (y vs ) and g l (z vs ) represents an independent gain function of x, y, and z. In some such implementations, the weighting function may be factorized as follows:
[0017] w(x vs ,y vs ,z vs ;x o ,y o ,z o ;s)=w x (x vs ;x o ;s)w y (y vs ;y o ;s)w z (z vs ;z o ;s) where w x (x vs ;x o ;s), w y (y vs ;y o ;s) and wz (z vs ;z o ;s) is x vs、 y vs and z vs In some such implementations, p may be a function of the audio object size(s).
[0018] Some such methods may involve storing the calculated virtual source gain values in a memory system. The process of calculating contributions from virtual sources within an audio object region or volume may involve retrieving calculated virtual source gain values corresponding to audio object positions and sizes from the memory system and interpolating between the calculated virtual source gain values. The process of interpolating between calculated virtual source gain values may involve: determining a plurality of neighboring virtual source positions near an audio object position; determining a calculated virtual source gain value for each of the neighboring virtual source positions; determining a plurality of distances between the audio object position and each of the neighboring virtual source positions; and interpolating between the calculated virtual source gain values according to the plurality of distances.
[0019] In some implementations, the playback environment data may include playback environment boundary data. The method may involve determining that an audio object region or volume includes an outer region or volume outside a playback environment boundary and applying a fade-out factor based at least in part on the outer region or volume. Some methods may involve determining that an audio object may be within a threshold distance from a playback environment boundary and not providing speaker feed signals to playback speakers on opposite boundaries of the playback environment. In some implementations, the audio object region or volume may be a rectangle, a cuboid, a circle, a sphere, an ellipse, and / or an ellipsoid.
[0020] Some methods may involve decorrelating at least a portion of the audio playback data, for example, decorrelating audio playback data for audio objects having an audio object size above a certain threshold.
[0021] Alternative methods are described herein. Some such methods involve receiving playback environment data including playback speaker position data and playback environment boundary data, and receiving audio playback data including one or more audio objects and associated metadata. The metadata may include audio object position data and audio object size data. These methods may involve determining that an audio object region or volume defined by the audio object position data and audio object size data includes an outer region or volume outside the playback environment boundary, and determining a fade-out factor based, at least in part, on the outer region or volume. These methods may involve calculating a set of gain values for each of a plurality of output channels based, at least in part, on the associated metadata and the fade-out factors. Each output channel may correspond to at least one playback speaker of the playback environment.
[0022] These methods may involve determining that an audio object may be within a threshold distance from one playback environment boundary and not providing a speaker feed signal to playback speakers on the opposite boundary of the playback environment.
[0023] The methods may involve calculating contributions from virtual sources within an audio object region or volume. The methods may involve defining a plurality of virtual source positions according to playback environment data and calculating, for each of the virtual source positions, a virtual source gain for each of a plurality of output channels. The virtual source positions may or may not be uniformly spaced, depending on the specific implementation.
[0024] Some implementations may be embodied in one or more non-transitory media having software stored thereon. The software may include instructions for controlling one or more devices to receive audio playback data including one or more audio objects. The audio objects may include an audio signal and associated metadata. The metadata may include at least audio object position data and audio object size data. The software may include instructions for calculating, for an audio object from the one or more audio objects, contributions from virtual sources within a region or volume defined by the audio object position data and audio object size data, and for calculating a set of audio object gain values for each of a plurality of output channels based at least in part on the calculated contributions. Each output channel may correspond to at least one playback speaker of a playback environment.
[0025] In some implementations, the process of calculating contributions from virtual sources may involve calculating a weighted average of virtual source gain values from virtual sources within an audio object region or volume, where the weights for the weighted average may depend on the position of the audio object, the size of the audio object, and / or the position of each virtual source within the audio object region or volume.
[0026] The software may include instructions for receiving playback environment data including playback speaker position data. The software may include instructions for defining a plurality of virtual source positions according to the playback environment data, and for each virtual source position, calculating a virtual source gain value for each of the plurality of output channels. Each of the virtual source positions may correspond to a position within the playback environment. In some implementations, at least some of the virtual source positions may correspond to positions outside the playback environment.
[0027] According to some implementations, the virtual source positions may be uniformly spaced. In some implementations, the virtual source positions may have a first uniform spacing along the x-axis and the y-axis and a second uniform spacing along the z-axis. The process of calculating a set of audio object gain values for each of the plurality of output channels may involve independent calculation of contributions from virtual sources along the x-axis, the y-axis, and the z-axis.
[0028] Various devices and apparatuses are described herein. Some such devices may include an interface system and a logic system. The interface system may include a network interface. In some implementations, the devices may include a memory device. The interface system may include an interface between the logic system and the memory device.
[0029] The logic system may be adapted to receive audio playback data from the interface system, the audio objects including one or more audio objects. The audio objects may include an audio signal and associated metadata. The metadata may include at least audio object position data and audio object size data. The logic system may be adapted to calculate, for an audio object from the one or more audio objects, a contribution from a virtual source within an audio object region or volume defined by the audio object position data and the audio object size data. The logic system may be adapted to calculate a set of audio object gain values for each of a plurality of output channels based at least in part on the calculated contributions. Each output channel may correspond to at least one playback speaker of a playback environment.
[0030] The process of calculating contributions from virtual sources may involve calculating a weighted average of virtual source gain values from virtual sources within an audio object region or volume. The weights for the weighted average may depend on the position of the audio object, the size of the audio object, and / or each virtual source position within the audio object region or volume. The logic system may be adapted to receive playback environment data from the interface system, including playback speaker position data.
[0031] The logic system may be adapted to define a plurality of virtual source positions according to playback environment data and, for each virtual source position, calculate a virtual source gain value for each of the plurality of output channels. Each of the virtual source positions may correspond to a position within the playback environment. However, in some implementations, at least some of the virtual source positions may correspond to positions outside the playback environment. Depending on the specific implementation, the virtual source positions may or may not be uniformly spaced. In some implementations, the virtual source positions may have a first uniform spacing along the x-axis and the y-axis and a second uniform spacing along the z-axis. The process of calculating the set of audio object gain values for each of the plurality of output channels may involve independent calculation of contributions from virtual sources along the x-axis, y-axis, and z-axis.
[0032] The apparatus may include a user interface. The logic system may be adapted to receive user input, such as audio object size data, via the user interface. In some implementations, the logic system may be adapted to scale the input audio object size data.
[0033] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. It is noted that the relative dimensions of the following drawings may not be drawn to scale. [Brief explanation of the drawings]
[0034] [Figure 1] This figure shows an example of a playback environment with Dolby Surround 5.1 configuration. [Figure 2] This figure shows an example of a playback environment with Dolby Surround 7.1 configuration. [Figure 3] This figure shows an example of a playback environment with a Hamasaki 22.2 surround sound configuration. [Figure 4A] FIG. 1 shows an example of a graphical user interface (GUI) for depicting speaker zones at various heights in a virtual playback environment. [Figure 4B] FIG. 10 is a diagram illustrating another example of a playback environment. [Figure 5A] 1 is a flow chart giving an overview of an audio processing method; [Figure 5B] A flow chart giving an example of the setup process. [Figure 5C] 10 is a flow diagram illustrating an example of a runtime process for calculating gain values for a received audio object according to pre-calculated gain values for virtual source positions. [Figure 6A] FIG. 1 illustrates an example of a virtual source position relative to a playback environment. [Figure 6B] 10A-10C show alternatives for virtual source positions relative to the playback environment. [Figure 6C] FIG. 10 illustrates an example of applying near-field and far-field panning techniques to audio objects at different positions. [Figure 6D] FIG. 10 illustrates an example of applying near-field and far-field panning techniques to audio objects at different positions. [Figure 6E] FIG. 10 illustrates an example of applying near-field and far-field panning techniques to audio objects at different positions. [Figure 6F] FIG. 10 illustrates an example of applying near-field and far-field panning techniques to audio objects at different positions. [Figure 6G] FIG. 1 shows an example of a playback environment with one speaker at each corner of a square with side length equal to 1. [Figure 7] FIG. 10 shows an example of contributions from a virtual source within a region defined by audio object position data and audio object size data. [Figure 8] A and B show an audio object in two positions within the playback environment. [Figure 9]1 is a flow diagram outlining a method for determining a fade-out factor based, at least in part, on how much of an audio object's area or volume extends outside the boundaries of the playback environment. [Figure 10] FIG. 2 is a block diagram providing examples of components of an authoring and / or rendering device. [Figure 11] FIG. 1A is a block diagram illustrating some components that may be used for audio content generation, and FIG. 1B is a block diagram illustrating some components that may be used for audio playback in a playback environment. Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION
[0035] The following description is directed to certain implementations for purposes of describing some novel aspects of the present disclosure and examples of contexts in which these novel aspects may be implemented. However, the teachings herein may be applied in a variety of different ways. For example, while various implementations are described using a specific playback environment, the teachings herein are broadly applicable to other known and future playback environments. Similarly, the described implementations may be implemented in a variety of authoring and / or rendering tools, which may be implemented in a variety of hardware, software, firmware, etc. Thus, the teachings of the present disclosure are not intended to be limited to the implementations shown in the drawings and / or described herein, but rather have broad applicability.
[0036] FIG. 1 shows an example of a playback environment with a Dolby Surround 5.1 configuration. Although Dolby Surround 5.1 was developed in the 1990s, this configuration is still widely deployed in movie theater sound system environments. A projector 105 may be configured to project a video image, such as for a movie, onto a screen 150. Audio playback data may be synchronized with the video image and processed by a sound processor 110. A power amplifier 115 may provide speaker feed signals to speakers in the playback environment 100.
[0037] The Dolby Surround 5.1 configuration includes a left surround array 120 and a right surround array 125, each of which includes a group of speakers collectively driven by a single channel. The Dolby Surround 5.1 configuration also includes separate channels for the left screen channel 130, the center screen channel 135, and the right screen channel 140. A separate channel for the subwoofer 145 is provided for low-frequency effects (LFE).
[0038] In 2010, Dolby provided an improvement to digital cinema sound by introducing Dolby Surround 7.1. FIG. 2 shows an example of a playback environment with a Dolby Surround 7.1 configuration. Digital projector 205 may be configured to receive digital video data and project a video image onto screen 150. Audio playback data may be processed by sound processor 210. Power amplifier 215 may provide speaker feed signals to speakers in playback environment 200.
[0039] The Dolby Surround 7.1 configuration includes a left lateral surround array 220 and a right lateral surround array 225, each of which may be driven by a single channel. Similar to Dolby Surround 5.1, the Dolby Surround 7.1 configuration includes separate channels for the left screen channel 230, the center screen channel 235, the right screen channel 240, and the subwoofer 245. However, Dolby Surround 7.1 increases the number of surround channels by dividing the left and right surround channels of Dolby Surround 5.1 into four zones. That is, in addition to the left lateral surround array 220 and the right lateral surround array 225, separate channels are included for the left rear surround speaker 224 and the right rear surround speaker 226. Increasing the number of surround zones in the playback environment 200 can significantly improve sound localization.
[0040] In an effort to create a more immersive environment, some playback environments may be configured with an increased number of speakers driven by an increased number of channels. Additionally, some playback environments may include speakers that are deployed at various heights, some of which may be above the seating area of the playback environment.
[0041] Figure 3 shows an example of a playback environment with a Hamasaki 22.2 surround sound configuration. Hamasaki 22.2 was developed at the NHK Science and Technology Research Laboratories in Japan as a surround sound component for ultra-high definition television. Hamasaki 22.2 provides 24 speaker channels that can be used to drive speakers arranged in three tiers. In playback environment 300, upper speaker tier 310 can be driven by nine channels. Middle speaker tier 320 can be driven by ten channels. Lower speaker tier 330 can be driven by five channels, two of which are for subwoofers 345a and 345b.
[0042] Thus, the current trend is not only to include more speakers and more channels, but also to include speakers of different heights. As the number of channels increases and speaker layouts transition from 2D to 3D arrays, the task of positioning and rendering sound becomes increasingly difficult. Accordingly, the assignee of the present application has developed various tools and related user interfaces that enhance functionality and / or reduce authoring complexity for 3D audio sound systems. Some of these tools are described in detail with reference to Figures 5A-19D in U.S. Provisional Patent Application No. 61 / 636,102, filed April 20, 2012, and entitled "Systems and Tools for Improved 3D Audio Creation and Presentation" (the "Creation and Presentation" application), the contents of which are incorporated herein by reference.
[0043] 4A shows an example of a graphical user interface (GUI) depicting speaker zones at various heights in a virtual playback environment. GUI 400 may be displayed on a display device, for example, according to instructions from a logic system, signals received from a user input device, etc. Some such devices are described below with reference to FIG. 10.
[0044] As used herein to refer to a virtual playback environment such as virtual playback environment 404, the term "speaker zone" generally refers to a logical construct that may or may not have a one-to-one correspondence with a playback speaker in a real playback environment. For example, a "speaker zone location" may or may not correspond to a specific playback speaker location in a movie theater playback environment. Instead, the term "speaker zone location" may generally refer to a zone in the virtual playback environment. In some implementations, speaker zones in a virtual playback environment may correspond to virtual speakers through the use of virtualization technology such as Dolby Headphone™ (sometimes referred to as Mobile Surround™), which generates a virtual surround sound environment in real time using a set of two-channel stereo headphones. GUI 400 has seven speaker zones 402a at a first elevation and two speaker zones 402b at a second elevation, for a total of nine speaker zones in virtual playback environment 404. In this example, speaker zones 1 through 3 are in the front region 405 of virtual playback environment 404. The front area 405 may correspond, for example, to the area in a cinema playback environment where the screen 150 is located, or in a home where a television screen is located, etc.
[0045] Here, speaker zone 4 generally corresponds to speakers in left region 410, and speaker zone 5 corresponds to speakers in right region 415 of virtual playback environment 404. Speaker zone 6 corresponds to left rear region 412, and speaker zone 7 corresponds to right rear region 414 of virtual playback environment 404. Speaker zone 8 corresponds to speakers in upper region 420a, and speaker zone 9 corresponds to speakers in upper region 420b, which may be a virtual ceiling region such as the region of virtual ceiling 520 shown in Figures 5D and 5E. Thus, as discussed in more detail in the "Making and Representing" application, the locations of speaker zones 1-9 shown in Figure 4A may or may not correspond to the locations of playback speakers in a real playback environment. Furthermore, other implementations may include more or fewer speaker zones and / or heights.
[0046] In various implementations described in the "Creating and Rendering" application, a user interface such as GUI 400 may be used as part of an authoring tool and / or a rendering tool. In some implementations, the authoring tool and / or rendering tool may be implemented via software stored on one or more non-transitory media. The authoring tool and / or rendering tool may be implemented (at least in part) via hardware, firmware, etc., such as the logic systems and other devices described below with reference to FIG. 10. In some authoring implementations, an associated authoring tool may be used to generate metadata about associated audio data. The metadata may include, for example, data indicating the position and / or trajectory of audio objects in three-dimensional space, speaker zone constraint data, etc. The metadata may be generated with respect to speaker zones 402 of a virtual playback environment 404, rather than with respect to a particular speaker layout of a real playback environment. The rendering tool may receive the audio data and associated metadata and calculate audio gain and speaker feed signals for the playback environment. Such audio gain and speaker feed signals may be calculated according to an amplitude panning process. The amplitude panning process can create the perception that a sound is coming from a position P in the playback environment. For example, a speaker feed signal can be expressed as x i (t)=g i x(t) i=1,…,N (Equation 1) may be provided to playback speakers 1 to N of the playback environment according to
[0047] In formula (1), x i (t) represents the speaker feed signal applied to speaker i, and g irepresents the gain factor of the corresponding channel, x(t) represents the audio signal, and t represents time. The gain factor may be determined, for example, according to the amplitude panning methods described in Section 2, pp. 3-4, of "Analog Audio Signal Processing," incorporated herein by reference. In some implementations, the gain may be frequency dependent. In some implementations, a time delay may be introduced by replacing x(t) with x(t-Δt).
[0048] In some rendering implementations, audio playback data generated with reference to speaker zones 402 may be mapped to speaker locations in a wide range of playback environments, which may have a Dolby Surround 5.1 configuration, a Dolby Surround 7.1 configuration, a Hamasaki 22.2 configuration, or other configurations. For example, referring to FIG. 2, a rendering tool may map audio playback data for speaker zones 4 and 5 to left lateral surround array 220 and right lateral surround array 225 in a playback environment with a Dolby Surround 7.1 configuration. Audio playback data for speaker zones 1, 2, and 3 may be mapped to left screen channel 230, right screen channel 240, and center screen channel 235, respectively. Audio playback data for speaker zones 6 and 7 may be mapped to left rear surround speaker 224 and right rear surround speaker 226.
[0049] 4B shows another example playback environment. In some implementations, a rendering tool may map audio playback data for speaker zones 1, 2, and 3 to corresponding screen speakers 455 in playback environment 450. The rendering tool may map audio playback data for speaker zones 4 and 5 to left lateral surround array 460 and right lateral surround array 465, and may map audio playback data for speaker zones 8 and 9 to left overhead speaker 470a and right overhead speaker 470b. Audio playback data for speaker zones 6 and 7 may be mapped to left rear surround speaker 480a and right rear surround speaker 480b.
[0050] In some authoring implementations, the authoring tool may be used to generate metadata about audio objects. As used herein, the term "audio object" may refer to a stream of audio data and associated metadata. The metadata may indicate the audio object's 3D position, the audio object's apparent size, rendering constraints, and content type (e.g., dialogue, effects, etc.). Depending on the implementation, the metadata may include other types of data, such as gain data and trajectory data. Some audio objects may be static, while others may move. Details of an audio object may be authored or rendered according to associated metadata, which may indicate, for example, the audio object's location in three-dimensional space at a given time. When an audio object is monitored or played in a playback environment, the audio object may be rendered according to its position and size metadata in accordance with the playback speaker layout of the playback environment.
[0051] FIG. 5A is a flow diagram providing an overview of audio processing methods. More detailed examples are described below with reference to FIG. 5B et seq. These methods may include more or fewer blocks than shown and described herein, and may not necessarily be performed in the order presented herein. These methods may be performed, at least in part, by devices such as those shown in FIGS. 10-11 and described below. Software may include instructions for controlling one or more devices to perform the methods described herein.
[0052] In the example shown in FIG. 5A, method 500 begins with a setup process that determines virtual source gain values for virtual source positions for a particular playback environment (step 505). FIG. 6A illustrates example virtual source positions for a playback environment. For example, block 505 may involve determining virtual source gain values for virtual source positions 605 relative to playback speaker positions 625 in playback environment 600a. Virtual source positions 605 and playback speaker positions 625 are merely examples. In the example shown in FIG. 6A, virtual source positions 605 are uniformly spaced along the x, y, and z axes. However, in alternative implementations, virtual source positions 605 may be spaced differently. For example, in some implementations, virtual source positions 605 may have a first uniform spacing along the x and y axes and a second uniform spacing along the z axis. In other implementations, virtual source positions 605 may be non-uniformly spaced.
[0053] 6A, the reproduction environment 600a and the virtual source volume 602a are coextensive, so that each of the virtual source positions 605 corresponds to a position within the reproduction environment 600a. However, in alternative implementations, the reproduction environment 600a and the virtual source volume 602a may not be coextensive. For example, at least some of the virtual source positions 605 may correspond to positions outside the reproduction environment 600a.
[0054] 6B shows an alternative example of a virtual source position relative to the reproduction environment, where a virtual source volume 602b extends outside the reproduction environment 600b.
[0055] Returning to FIG. 5A, in this example, the setup process of block 505 occurs prior to rendering any particular audio object. In some implementations, the virtual source gain values determined in block 505 may be stored in a storage system. The stored virtual source gain values may be used during a "runtime" process of calculating audio object gain values for received audio objects according to at least some of the virtual source gain values (block 510). For example, block 510 may involve calculating audio object gain values based, at least in part, on virtual source gain values corresponding to virtual source positions within an audio object region or volume.
[0056] In some implementations, method 500 may include an optional block 515 involving decorrelating the audio data. Block 515 may be part of a runtime process. In some such implementations, block 515 may involve convolution in the frequency domain. For example, block 515 may involve applying a finite impulse response (“FIR”) filter to each speaker feed signal.
[0057] In some implementations, the process of block 515 may or may not be performed depending on the audio object size and / or the artistic intent of the author. According to some such implementations, an authoring tool may link audio object size to decorrelation by indicating (e.g., via a decorrelation flag included in associated metadata) that decorrelation should be turned on when the audio object size is equal to or greater than a certain size threshold and that decorrelation should be turned off when the audio object size is below the size threshold. In some implementations, decorrelation may be controlled (e.g., increased, decreased, or disabled) according to user input regarding the size threshold and / or other input values.
[0058] FIG. 5B is a flow diagram providing an example of a setup process. Thus, all blocks shown in FIG. 5B are examples of processes that may be performed in block 505 of FIG. 5A. Here, the setup process begins with receiving playback environment data (block 520). The playback environment data may include playback speaker position data. The playback environment data may include data representing boundaries of the playback environment, such as walls, ceilings, etc. If the playback environment is a movie theater, the playback environment data may also include an indication of the movie screen position.
[0059] The playback environment data may also include data indicating correlations of output channels with playback speakers of the playback environment. For example, the playback environment may have a Dolby Surround 7.1 configuration as shown in FIG. 2 and described above. Thus, the playback environment data may also include data indicating correlations between the Lss channel and the left side surround speaker 220, between the Lrs channel and the left rear surround speaker 224, etc.
[0060] In this example, block 525 involves defining a virtual source position 605 according to the playback environment data. The virtual source position 605 may be defined within a virtual source volume. In some implementations, the virtual source volume may correspond to a volume within which an audio object can move. As shown in Figures 6A and 6B, in some implementations, the virtual source volume 602 may be coextensive with the volume of the playback environment 600, while in other implementations, at least a portion of the virtual source position 605 may correspond to a position outside of the playback environment 600.
[0061] Furthermore, the virtual source positions 605 may or may not be uniformly spaced within the virtual source volume 602, depending on the particular implementation. In some implementations, the virtual source positions 605 may be uniformly spaced in all directions. For example, the virtual source positions 605 are spaced N x Multiply N y Multiply N zIn some implementations, the value of N may range from 5 to 100. The value of N may depend, at least in part, on the number of playback speakers in the playback environment: it may be desirable to include more than one virtual source position 605 between each playback speaker position.
[0062] In other implementations, the virtual source positions 605 may have a first uniform spacing along the x-axis and y-axis and a second uniform spacing along the z-axis. x Multiply N y Kakeru M z The number of virtual source positions 605 may form a rectangular grid of virtual source positions 605. For example, in some implementations, there may be fewer virtual source positions 605 along the z-axis than along the x-axis or y-axis. In some such implementations, the value of N may range from 10 to 100, while the value of M may range from 5 to 10.
[0063] In this example, block 530 involves calculating a virtual source gain value for each of the virtual source positions 605. In some implementations, block 530 involves calculating a virtual source gain value for each of the multiple output channels of the playback environment for each virtual source position 605. In some implementations, block 530 may involve applying a vector-based amplitude panning (VBAP) algorithm, a pairwise panning algorithm, or similar algorithm to calculate a gain value for a point source located at each virtual source position 605. In other implementations, block 530 may involve applying a separable algorithm to calculate a gain value for a point source located at each virtual source position 605. As used herein, a "separable" algorithm is one in which the gain of a given speaker can be expressed as a product of two or more factors that can be calculated separately for each coordinate of the virtual source position. Examples include algorithms implemented in various existing mixing console panners, including but not limited to the panners implemented in Pro Tools™ software and digital film consoles offered by AMS Neve. Some two-dimensional examples are given below.
[0064] 6C-6F show examples of the application of near-field and far-field panning techniques to audio objects at various locations. Referring first to FIG. 6C, the audio object is substantially outside the virtual playback environment 400a. Therefore, one or more far-field panning methods are applied in this example. In some implementations, the far-field panning method may be based on vector-based amplitude panning (VBAP) formulas known to those skilled in the art. For example, the far-field panning method may be based on the VBAP formulas described in Section 2.3, p. 4, of "VBAP Panning," a non-patent document 1, which is incorporated herein by reference. In alternative implementations, other methods may be used to pan audio objects in the far and near fields, such as methods involving the synthesis of corresponding acoustic plane or spherical waves. Related methods are described in "VBAP Panning," a non-patent document 2, which is incorporated herein by reference.
[0065] 6D, an audio object 610 is inside the virtual playback environment 400a. Therefore, one or more near-field panning methods are applied in this example. Some such near-field panning methods use several speaker zones that surround the audio object 610 within the virtual playback environment 400a.
[0066] FIG. 6G shows an example playback environment with one speaker at each corner of a square with side lengths equal to 1. In this example, the origin (0,0) of the x and y axes coincides with the left (L) screen speaker 130. Thus, the right (R) screen speaker 140 has coordinates (1,0), the left surround (Ls) speaker 120 has coordinates (0,1), and the right surround (Rs) speaker 125 has coordinates (1,1). The audio object position 615 (x,y) is x units to the right of the L speaker and y units from the screen 150. In this example, each of the four speakers receives a cos / sin factor proportional to their distance along the x and y axes. According to some implementations, the gain may be calculated as follows:
[0067] G_l(x)=cos(pi / 2*x) where l=L,Ls G_l(x)=sin(pi / 2*x) where l=R,Rs G_l(x)=cos(pi / 2*y) where l=L,R G_l(x)=sin(pi / 2*y) where l=Ls,Rs .
[0068] The overall gain is the product: G_l(x,y) = G_l(x)G_l(y). In general, these functions depend on all coordinates of all speakers. However, G_l(x) does not depend on the y position of the source, and G_l(y) does not depend on its x position. To illustrate a simple calculation, suppose the audio object position 615 is (0,0), i.e., the position of the L speaker. G_L(x) = cos(0) = 1 and G_L(y) = cos(0) = 1. The overall gain is the product G_L(x,y) = G_L(x)G_L(y) = 1. Similar calculations give us G_Ls = G_Rs = G_R = 0.
[0069] It may be desirable to blend between different panning modes when an audio object enters or exits the virtual playback environment 400a. For example, when an audio object 610 moves from the audio object position 615 shown in FIG. 6C to the audio object position 615 shown in FIG. 6D, or vice versa, a blend of gains calculated according to the near-field panning method and the far-field panning method may be applied. In some implementations, a pair-wise panning law (e.g., an energy-preserving sine or power law) may be used to blend between gains calculated according to the near-field panning method and the far-field panning method. In alternative implementations, the pair-wise panning law may be amplitude-preserving rather than energy-preserving. Thus, rather than the sum of squares equaling one, the sum equals one. It is also possible to process an audio signal using both panning methods independently and then blend the resulting processed signals, for example, to crossfade the two resulting audio signals.
[0070] Returning now to FIG. 5B, regardless of the algorithm used in block 530, the resulting gain values may be stored in a memory system for use during runtime operation (block 535).
[0071] 5C is a flow diagram illustrating an example of a runtime process for calculating gain values for a received audio object according to pre-calculated gain values for virtual source positions. All of the blocks shown in FIG. 5C are examples of processes that may be performed in block 510 of FIG. 5A.
[0072] In this example, the runtime process begins with receipt of audio playback data including one or more audio objects (block 540). An audio object includes an audio signal and associated metadata, which in this example includes at least audio object position data and audio object size data. Referring to FIG. 6A, for example, audio object 610 is defined, at least in part, by audio object position 615 and audio object volume 620a. In this example, the received audio object size data indicates that audio object volume 620a corresponds to the volume of a rectangular prism. However, in the example shown in FIG. 6B, the received audio object size data indicates that audio object volume 620b corresponds to the volume of a sphere. These sizes and shapes are merely examples. In alternative implementations, audio objects may have a variety of other sizes and / or shapes. In some alternative examples, the area or volume of an audio object may be a rectangle, a circle, an ellipse, an ellipsoid, or a spherical sector.
[0073] In this implementation, block 545 involves calculating the contribution from a virtual source within a region or volume defined by the audio object position data and the audio object size data. In the example shown in FIGS. 6A and 6B, block 545 may involve calculating the contribution from a virtual source at a virtual source position 605 that is within audio object volume 620a or audio object volume 620b. If the metadata of an audio object changes over time, block 545 may be executed again according to the new metadata values. For example, if the audio object size and / or audio object position changes, a different virtual source position 605 may fall within audio object volume 620 and / or the virtual object position 605 used in the previous calculation may be a different distance from audio object position 615. In block 545, the corresponding virtual source contribution is calculated according to the new object size and / or position.
[0074] In some examples, block 545 may involve retrieving from a memory system calculated virtual source gain values for virtual source positions corresponding to an audio object position and size, and interpolating between the calculated virtual source gain values. The process of interpolating between the calculated virtual source gain values may involve determining a plurality of neighboring virtual source positions near an audio object position, determining a calculated virtual source gain value for each of the neighboring virtual source positions, determining a plurality of distances between the audio object position and each of the neighboring virtual source positions, and interpolating between the calculated virtual source gain values according to the plurality of distances.
[0075] The process of calculating the contributions from the virtual sources may involve calculating a weighted average of the calculated virtual source gain values for virtual source positions within a region or volume defined by the size of the audio object, where the weights for the weighted average may depend, for example, on the position of the audio object, the size of the audio object and each virtual source position within said region or volume.
[0076] FIG. 7 shows an example of contributions from virtual sources within a region defined by audio object position data and audio object size data. FIG. 7 depicts a cross-section of audio environment 200a taken perpendicular to the z-axis. Thus, FIG. 7 is depicted from the perspective of an observer looking down on audio environment 200a along the z-axis. In this example, audio environment 200a is a cinema sound system environment with the Dolby Surround 7.1 configuration shown in FIG. 2 and described above. Thus, playback environment 200a includes left side surround speaker 220, left rear surround speaker 224, right side surround speaker 225, right rear surround speaker 226, left screen channel 230, center screen channel 235, right screen channel 240, and subwoofer 245.
[0077] Audio object 610 has a size indicated by audio object volume 620b, the rectangular cross-sectional area of which is shown in Figure 7. Given audio object position 615 at the time depicted in Figure 7, the area encompassed by audio object volume 620b in the x-y plane includes 12 virtual source positions 605. Depending on the extent of audio object volume 620b in the z-direction and the spacing of virtual source positions 605 along the z-axis, additional virtual source positions 605 may or may not be encompassed within audio object volume 620b.
[0078] FIG. 7 shows contributions from virtual source positions 605 within a region or volume defined by the size of audio object 610. In this example, the diameter of the circle used to depict each of the virtual source positions 605 corresponds to the contribution from the corresponding virtual source position 605. The virtual source positions 605a closest to audio object position 615 are shown largest, indicating the largest contribution from the corresponding virtual source. The second largest contribution is from the virtual source at virtual source position 605b, which is second-closest to audio object position 615. A smaller contribution is made by virtual source position 605c, which is further away from audio object position 615 but still within audio object volume 620b. Virtual source position 605d, which is outside audio object volume 620b, is shown smallest, indicating that the corresponding virtual source makes no contribution in this example.
[0079] Referring to FIG. 5C, in this example, block 550 involves calculating a set of audio object gain values for each of a plurality of output channels based at least in part on the calculated contributions. Each output channel may correspond to at least one playback speaker of the playback environment. Block 550 may involve normalizing the resulting audio object gain values. For the implementation shown in FIG. 7, for example, each output channel may correspond to a single speaker or a group of speakers.
[0080] The process of calculating an audio object gain value for each of the plurality of output channels comprises: o ,y o ,z o The gain value (g) for an audio object of size (s) to be rendered in l size (x o ,y o ,z o;s)). This audio object gain value is sometimes referred to herein as the "audio object size contribution". In some implementations, the audio object gain value (g l size (x o ,y o ,z o ;s)) may be expressed as follows:
[0081]
number
[0082] In some examples, the exponent p may have a value between 1 and 10. In some implementations, p may be a function of the audio object size s. For example, if s is relatively larger, in some implementations, p may be relatively smaller. According to some such implementations, p may be determined as follows:
[0083] When p=6 s≦0.5 p=6+(-4)(s-0.5) / (s max -0.5) if s>0.5 where s max is the internal scaled up size s internal (see below), an audio object size s=1 may correspond to an audio object with a size (e.g., diameter) equal to one length of the playback environment boundary (e.g., equal to the length of one wall of the playback environment).
[0084] Depending in part on the algorithm(s) used to calculate the virtual source gain values, it may be possible to simplify equation (2) if the virtual source positions are uniformly distributed along an axis and if the weighting and gain functions are separable, e.g., as described above. If these conditions are met, then g l (x vs ,y vs ,z vs ) is g lx (x vs )g ly (y vs )g lz (z vs ) where g lx (x vs ), g ly (y vs ) and g lz (z vs ) represents the gain function independent of the x, y and z coordinates for the position of the virtual source.
[0085] Similarly, w(x vs ,y vs ,z vs;x o ,y o ,z o ;s) is w x (x vs ;x o ;s)w y (y vs ;y o ;s)w z (z vs ;z o ;s), where w x (x vs ;x o ;s), w y (y vs ;y o ;s) and w z (z vs ;z o ;s) is the x for the position of the virtual source 、 represents independent weighting functions for the y and z coordinates. One such example is shown in Figure 7. In this example, w x (x vs ;x o ;s) is a weighting function 710 y (y vs ;y o ;s) may be calculated independently from the weighting function 720. In some implementations, the weighting functions 710 and 720 may be Gaussian functions, while the weighting function w z (z vs ;z o ;s) may be a product of a cosine and a Gaussian function.
[0086] w(x vs ,y vs ,z vs ;x o ,y o ,z o ;s) is w x (x vs ;x o ;s)w y (y vs ;y o ;s)w z (z vs ;z o ;s), equation (2) can be simplified to:
[0087]
number
[0088] In some implementations, the audio object size contribution g l size may be combined with an "audio object near gain" for the audio object position. As used herein, "audio object near gain" is the calculated gain based on the audio object position 615. The gain calculation may be done using the same algorithm used to calculate each of the virtual source gain values. According to some such implementations, a crossfade calculation may be performed between the audio object size contribution and the audio object near gain result, for example, as a function of the audio object size. Such an implementation may provide smooth panning and smooth growth of the audio object, and may allow smooth transitions between minimum and maximum audio object sizes. In one such implementation:
[0089]
number
[0090] According to some implementations, audio object size values may be scaled up in the larger portion of their range of possible values. For example, in some authoring implementations, the user may specify audio object size values s user ∈[0,1], which may be used for the actual size used by the algorithm, or for a larger range, e.g., the range [0,s max ] where s max >1. This mapping can ensure that the gain is truly independent of the object's position when the size is set to maximum by the user. According to some such implementations, such a mapping is user ,s internal ) according to a piecewise linear function connecting s user represents the user-selected audio object size, and s internal represents the corresponding audio object size determined by the algorithm. According to some such implementations, the mapping is performed for pairs of points (0,0), (0.2,0.3), (0.5,0.9), (0.75,1.5) and (1,s max ) according to a piecewise linear function connecting s max =2.8.
[0091] 8A and 8B show an audio object at two positions within the playback environment. In these examples, the audio object volume 620b is a sphere with a radius less than half the length or width of the playback environment 200a. The playback environment 200a is configured according to Dolby 7.1. At the time depicted in FIG. 8A, the audio object position 615 is relatively closer to the center of the playback environment 200a. At the time depicted in FIG. 8B, the audio object position 615 has moved closer to the boundary of the playback environment 200a. In this example, the boundary is the left wall of the theater, which coincides with the position of the left-side surround speaker 220.
[0092] For aesthetic reasons, it may be desirable to modify the audio object gain calculation for audio objects approaching the boundaries of the playback environment. For example, in Figures 8A and 8B, when audio object position 615 is within a certain threshold distance from left boundary 805 of the playback environment, no speaker feed signal is provided to the speaker at the opposite boundary of the playback environment (here, right side surround speaker 225). In the example shown in Figure 8B, when audio object position 615 is within a certain threshold distance (which may be a different threshold distance) from left boundary 805 of the playback environment, no speaker feed signal is provided to left screen channel 230, center screen channel 235, right screen channel 240, or subwoofer 245 if audio object position 615 is further away from the screen than a certain threshold distance.
[0093] In this example, shown in FIGURE 8B, audio object volume 620b includes an area or volume outside left boundary 805. According to some implementations, the fade-out factor for the gain calculation may be based, at least in part, on how much of left boundary 805 is within audio object volume 620b and / or how much of the audio object's area or volume extends outside such boundary.
[0094] 9 is a flow diagram outlining a method for determining a fade-out factor based, at least in part, on how much of an audio object's area or volume extends outside the boundaries of the playback environment. At block 905, playback environment data is received. In this example, the playback environment data includes playback speaker position data and playback environment boundary data. Block 910 involves receiving audio playback data including one or more audio objects and associated metadata. In this example, the metadata includes at least audio object position data and audio object size data.
[0095] In this implementation, block 915 involves determining that the audio object region or volume defined by the audio object position data and the audio object size data includes an outer region or volume outside the playback environment boundary. Block 915 may also involve determining what percentage of the audio object region or volume is outside the playback environment boundary.
[0096] At block 920, a fade-out factor is determined. In this example, the fade-out factor may be based at least in part on the outer region. For example, the fade-out factor may be proportional to the outer region.
[0097] In block 925, a set of audio object gain values may be calculated for each of a plurality of output channels based at least in part on the associated metadata (in this example, audio object position data and audio object size data) and the fade-out factor, where each output channel may correspond to at least one playback speaker of the playback environment.
[0098] In some implementations, the audio object gain calculation may involve calculating contributions from virtual sources within an audio object region or volume. The virtual sources may correspond to multiple virtual source positions, which may be defined with reference to playback environment data. The virtual source positions may or may not be uniformly spaced. For each of the virtual source positions, a virtual source gain value may be calculated for each of the multiple output channels. As noted above, in some implementations, these virtual source gain values may be calculated during a setup process, stored, and retrieved for use during runtime operation.
[0099] In some implementations, a fade-out factor may be applied to all virtual source gain values corresponding to virtual source positions in the playback environment. l size may be modified as follows:
[0100]
number
[0101] In an alternative implementation, g l size may be modified as follows:
[0102]
number
[0103] 10 is a block diagram providing an example of components of an authoring and / or rendering device. In this example, device 1000 includes an interface system 1005. Interface system 1005 may include a network interface, such as a wireless network interface. Alternatively or additionally, interface system 1005 may include a universal serial bus (USB) interface or other such interface.
[0104] The device 1000 includes a logic system 1010. The logic system 1010 may include a processor, such as a general-purpose single-chip or multi-chip processor. The logic system 1010 may include a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or discrete hardware components, or a combination thereof. The logic system 1010 may be configured to control other components of the device 1000. Although interfaces between components of the device 1000 are not shown in FIG. 10, the logic system 1010 may be configured with interfaces for communication with other components. The other components may or may not be configured for communication with each other, as appropriate.
[0105] Logic system 1010 may be configured to perform audio authoring and / or rendering functions, including but not limited to audio authoring and / or rendering functions of the type described herein. In some such implementations, logic system 1010 may be configured to operate (at least in part) according to software stored on one or more non-transitory media. Non-transitory media may include memory associated with logic system 1010, such as random access memory (RAM) and / or read-only memory (ROM). Non-transitory media may include memory of memory system 1015. Memory system 1015 may include one or more suitable types of non-transitory storage media, such as flash memory, hard drives, etc.
[0106] Display system 1030 may include one or more suitable types of displays, depending on the implementation of device 1000. For example, display system 1030 may include a liquid crystal display, a plasma display, a bi-stable display, etc.
[0107] The user input system 1035 may include one or more devices configured to accept input from a user. In some implementations, the user input system 1035 may include a touchscreen overlaying the display of the display system 1030. The user input system 1035 may include a mouse, a trackball, a gesture detection system, a joystick, one or more GUIs and / or menus presented on the display system 1030, buttons, keyboards, switches, etc. In some implementations, the user input system 1035 may include a microphone 1025 via which a user may provide voice commands for the device 1000. A logic system may be configured for voice recognition and for controlling at least some operations of the device 1000 according to such voice commands.
[0108] Power system 1040 may include one or more suitable energy storage devices, such as nickel-cadmium batteries or lithium-ion batteries. Power system 1040 may be configured to receive power from an electrical outlet.
[0109] FIG. 11A is a block diagram illustrating several components that may be used for audio content creation. System 1100 may be used for audio content creation in, for example, a mixing studio and / or a dubbing stage. In this example, system 1100 includes audio and metadata authoring tool 1105 and rendering tool 1110. In this implementation, audio and metadata authoring tool 1105 and rendering tool 1110 include audio connection interfaces 1107 and 1112, respectively, which may be configured for communication via AES / EBU, MADI, analog, etc. Audio and metadata authoring tool 1105 and rendering tool 1110 include network interfaces 1109 and 1117, respectively, which may be configured to send and receive metadata via TCP / IP or any other suitable protocol. Interface 1120 is configured to output audio data to speakers.
[0110] The system 1100 may include, for example, an existing authoring system, such as a Pro Tools™ system, that runs a metadata generation tool (i.e., a panner described herein) as a plug-in. The panner may run on a standalone system (e.g., a PC or mixing console) connected to the rendering tool 1110, or it may run on the same physical device as the rendering tool 1110. In the latter case, the panner and renderer may use a local connection, for example, through shared memory. The panner GUI may also be provided on a tablet device, laptop, etc. The rendering tool 1110 may have a rendering system including a sound processor configured to perform rendering methods such as those described in FIGS. 5A-C and 9. The rendering system may include, for example, a personal computer, laptop, etc., including interfaces for audio input and output and appropriate logic systems.
[0111] 11B is a block diagram illustrating some components that may be used for audio playback in a playback environment (e.g., a movie theater). System 1150, in this example, includes a theater server 1155 and a rendering system 1160. Theater server 1155 and rendering system 1160 include network interfaces 1157 and 1162, respectively, which may be configured to send and receive audio objects via TCP / IP or any other suitable protocol. Interface 1164 is configured to output audio data to speakers.
[0112] Various modifications to the implementations described in this disclosure will be readily apparent to those skilled in the art. The general principles defined herein may be applied to other implementations without departing from the spirit or scope of the disclosure. Thus, the scope of the claims is not intended to be limited to the implementations shown herein, but is to be accorded the widest scope consistent with the disclosure, principles, and novel features disclosed herein.
[0113] Some numbering examples are given below. [Numbered Example 1] receiving audio playback data including one or more audio objects, the audio objects including an audio signal and associated metadata, the metadata including at least audio object position data and audio object size data; calculating, for an audio object from the one or more audio objects, a contribution from a virtual source within an audio object region or volume defined by the audio object position data and the audio object size data; calculating a set of audio object gain values for each of a plurality of output channels based, at least in part, on the calculated contributions, each output channel corresponding to at least one playback speaker of the playback environment; method. [Numbered Example 2] 2. The method of numbered Example 1, wherein calculating contributions from virtual sources comprises calculating a weighted average of virtual source gain values from virtual sources within the audio object region or volume. [Numbered Example 3] 3. The method of numbered Example 2, wherein weights for the weighted average depend on the position of the audio object, the size of the audio object, and each virtual source position within the audio object region or volume. [Numbered Example 4] further comprising receiving playback environment data including playback speaker position data; Method described in numbered Example 1. [Numbered Example 5] defining a plurality of virtual source positions according to the playback environment data; for each virtual source position, calculating a virtual source gain value for each of the plurality of output channels; Method described in numbered Example 4. [Numbered Example 6] 6. The method of Example 5, wherein each virtual source position corresponds to a position within the playback environment. [Numbered Example 7] 6. The method of embodiment 5, wherein at least some of the virtual source positions correspond to positions outside the playback environment. [Numbered Example 8] The method of Example 5, wherein the virtual source positions are uniformly spaced along the x, y, and z axes. Numbered Example 9 The method of Example 5, wherein the virtual source positions have first uniform spacing along the x-axis and y-axis and a second uniform spacing along the z-axis. Numbered Example 10 10. The method of numbered Example 8 or 9, wherein calculating a set of audio object gain values for each of the plurality of output channels includes independent calculation of contributions from virtual sources along x, y, and z axes. Numbered Example 11 The method of Example 5, wherein the virtual source positions are non-uniformly spaced. Numbered Example 12 The step of calculating an audio object gain value for each of the plurality of output channels includes: o ,y o ,z o The gain value (g) for an audio object of size (s) to be rendered in l (x o ,y o ,z o ;s)) and determining a gain value (g l (x o,y o ,z o ;s)) is
number
Claims
1. 1. A method of rendering input audio including audio objects and associated metadata, the metadata including audio object size metadata and audio object position metadata corresponding to the audio objects, the method comprising: determining a plurality of virtual audio objects based on the audio object size metadata and the audio object position metadata corresponding to the audio objects; for each virtual audio object of the plurality of virtual audio objects, determining at least one gain of the corresponding virtual audio object, wherein each gain of the corresponding virtual audio object is based on an object audio metadata gain corresponding to that audio object; and rendering the audio objects to one or more speaker feeds, the audio objects being based on corresponding gains of at least some of the virtual audio objects of the plurality of virtual audio objects. method.
2. A non-transitory medium storing software including instructions for performing the method of claim 1.
3. 1. An apparatus for rendering input audio comprising audio objects and associated metadata, the metadata including audio object size metadata and audio object position metadata corresponding to the audio objects, the apparatus comprising: determining a plurality of virtual audio objects based on the audio object size metadata and the audio object position metadata corresponding to the audio objects; for each virtual audio object of the plurality of virtual audio objects, determining at least one gain of the corresponding virtual audio object, wherein each gain of the corresponding virtual audio object is based on an object audio metadata gain corresponding to that audio object; rendering the audio object to one or more speaker feeds; wherein the processor is configured to render the audio objects based on corresponding gains of at least some of the virtual audio objects of the plurality of virtual audio objects. Device.