Rendering of audio objects with apparent size to arbitrary loudspeaker layouts

By calculating audio object gain values based on virtual source locations and metadata, the method addresses the complexity of rendering audio objects in 3D cinema sound systems, improving sound localization and simplifying the authoring process.

KR102995735B1Active Publication Date: 2026-07-29DOLBY LABORATORIES LICENSING CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2014-03-10
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

The increasing complexity of audio playback environments, particularly in cinema sound systems with 3D speaker layouts, poses challenges in authoring and rendering audio objects effectively.

Method used

A method involving the calculation of audio object gain values based on virtual source locations and metadata, including position, size, and trajectory, to render audio objects in various playback environments, such as cinema sound systems, using weighted averages and interpolation techniques.

Benefits of technology

Enables accurate and efficient rendering of audio objects in complex playback environments, enhancing sound localization and reducing authoring complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 112024103863023-PAT00024_ABST
    Figure 112024103863023-PAT00024_ABST
Patent Text Reader

Abstract

Multiple virtual source locations may be defined for a volume in which audio objects can move. A setup process for rendering audio data may include the steps of receiving playback speaker location data and pre-calculating gain values ​​for each of the virtual sources based on the playback speaker location data and each virtual source location. The gain values ​​may be stored and used during "runtime," while the audio playback data is rendered for the speakers of the playback environment. During runtime, for each audio object, contributions from virtual source locations within a volume or area defined by audio object location data and audio object size data may be calculated. A set of gain values ​​for each output channel of the playback environment may be calculated, at least in part, based on the calculated contributions. Each output channel may correspond to at least one playback speaker of the playback environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Cross-reference regarding related applications

[0002] This application claims priority to Spanish patent application no. P201330461, filed March 28, 2013, and U.S. patent application no. 61 / 833,581, filed June 11, 2013, each incorporated herein by reference in its entirety.

[0003] The present disclosure relates to the authoring and rendering of audio playback data. In particular, the present disclosure relates to authoring and rendering audio playback data for playback environments such as cinema sound playback systems. Background Technology

[0004] Since the introduction of sound in film in 1927, there has been a steady development of technology used to capture the artistic intent of motion picture soundtracks and reproduce them in a cinematic environment. In the 1930s, synchronized sound on discs paved the way for variable-range sound in films, which was further improved in the 1940s with the early introduction of multi-track recording and controllable playback (using control tones to move sounds), leading to the consideration of theater acoustics and improved loudspeaker designs. In the 1950s and 1960s, magnetic striping in films enabled multi-channel playback in theaters, and premium theaters introduced surround channels and up to five screen channels.

[0005] In the 1970s, Dolby introduced noise reduction to both post-production and cinema, along with cost-effective means of encoding and distributing mixes with three screen channels and a mono surround channel. The quality of cinema sound was further improved in the 1980s through Dolby Spectral Recording (SR) noise reduction and certification programs such as THX. During the 1990s, Dolby brought digital sound to cinema with a 5.1 channel format providing separate left, center, and right screen channels, left and right surround arrays, and a subwoofer channel for low-frequency effects. Dolby Surround 7.1, introduced in 2010, increased the number of surround channels by separating the existing left and right surround channels into four "zones."

[0006] As the number of channels increases and loudspeaker layouts transition from flat 2-dimensional (2D) arrays to 3-dimensional (3D) arrays including elevation, the tasks of authoring and rendering sounds are becoming increasingly complex. The problem to be solved

[0007] Improved methods and devices over conventional technologies would be desirable. means of solving the problem

[0008] Some aspects of the subject matter described in this disclosure may be implemented in tools for rendering audio playback data including audio objects created without reference to any particular playback environment. As used herein, the term “audio object” may refer to a stream of audio signals and associated metadata. The metadata may indicate at least the position and apparent size of the audio object. However, the metadata may also indicate rendering constraint data, content type data (e.g., dialogue, effects, etc.), gain data, trajectory data, etc. Some audio objects may be static, while others may have time-varying metadata: these audio objects may be movable, resized, and / or have other attributes that change over time.

[0009] When audio objects are monitored or played in a playback environment, the audio objects may be rendered according to at least the position and size metadata. The rendering process may involve calculating a set of audio object gain values ​​for each channel of a set of output channels. Each output channel may correspond to one or more playback speakers of the playback environment.

[0010] Some implementations described herein involve a “set-up” process that may occur before rendering any specific audio objects. Additionally, the setup process, which may be referred to herein as Stage 1 or Stage 1, may involve defining a number of virtual source locations in a volume in which the audio objects can move. As used herein, “virtual source location” is a location of a static point source. According to these implementations, the setup process may involve receiving playback speaker location data and pre-calculating virtual source gain values ​​for each of the virtual sources based on the playback speaker location data and the virtual source location. As used herein, the term “speaker location data” may include location data indicating the locations of some or all of the speakers in the playback environment. The location data may be provided as absolute coordinates of the playback speaker locations, for example, Cartesian coordinates, spherical coordinates, etc. Alternatively, or additionally, location data may be provided as coordinates for other playback environment locations, such as acoustic "sweet spots" of the playback environment (e.g., Cartesian coordinates or angular coordinates).

[0011] In some implementations, the virtual source gain values ​​may be stored and used during "runtime," while audio playback data is rendered to the speakers of the playback environment. During runtime, for each audio object, contributions from virtual source locations within an area or volume defined by audio object location data and audio object size data may be calculated. The process of calculating contributions from virtual source locations may involve calculating a weighted average of a plurality of pre-calculated virtual source gain values ​​determined during the setup process for virtual source locations within an audio object area or volume defined by the size and location of the audio object. A set of audio object gain values ​​for each output channel of the playback environment may be calculated, at least partially, based on the calculated virtual source contributions. Each output channel may correspond to at least one playback speaker of the playback environment.

[0012] Accordingly, some of the methods described herein involve receiving audio playback data containing one or more audio objects. The audio objects may include audio signals and associated metadata. The metadata may include at least audio object location data and audio object size data. The methods may involve calculating contributions from virtual sources within an audio object region or volume defined by the audio object location data and the audio object size data. The methods may, at least partially, involve calculating a set of audio object gain values ​​for each of a plurality of output channels based on the calculated contributions. Each output channel may correspond to at least one playback speaker in a playback environment. For example, the playback environment may be a cinema sound system environment.

[0013] The process of calculating contributions from virtual sources may involve calculating a weighted average of virtual source gain values ​​from virtual sources within the audio object region or volume. The weights for the weighted average may depend on the position of the audio object within the audio object region or volume, the size of the audio object, and / or the respective virtual source position.

[0014] The above methods may also involve receiving playback environment data including playback speaker location data. The above methods may also involve defining a plurality of virtual source locations according to the playback environment data and, for each of the virtual source locations, calculating a virtual source gain value for each of the plurality of output channels. In some implementations, each of the virtual source locations may correspond to a location within the playback environment. However, in some implementations, at least some of the virtual source locations may correspond to locations outside the playback environment.

[0015] In some implementations, the virtual source locations may be uniformly spaced along the x, y, and z axes. However, in some implementations, the spacing may not be uniform in all directions. For example, the virtual source locations may have a first uniform spacing along the x and y axes and a second uniform spacing along the z axis. The process of calculating the set of audio object gain values ​​for each of the plurality of output channels may involve independent calculations of contributions from the virtual sources along the x, y, and z axes. In alternative implementations, the virtual source locations may be non-uniformly spaced.

[0016] In some implementations, the process of calculating the audio object gain value for each of the plurality of output channels is a position (x o , y o , z o Gain value (g) for an audio object of size(s) to be rendered in ) l (x o , y o , z o ; s)) may involve determining. For example, the audio object gain value (g l (x o , y o , z o ; s)) can be expressed as follows:

[0017] ,

[0018] Here (x vs , y vs , z vs ) indicates the virtual source location, and g ι (x vs , y vs , z vs ) is the virtual source location (x vs , y vs , z vs Represents the gain value for channel (l) for ) and w(x vs , y vs , z vs ; x o , y o , z o ;s) is at least partially, the position of the audio object (x o , y o , z o ), size(s) of the audio object and virtual source position(x vs , y vs , z vs determined based on ), g l (x vs , y vs , z vs Represents one or more weighting functions for ).

[0019] According to some of these implementations, g l (x vs , y vs , z vs ) = g l (x vs )g l (y vs )g l (z vs ) and, here g l (x vs ), g l (y vs ) and g l (z vs ) represents independent gain functions of x, y, and z. In some of these implementations, the weighting functions take the following arguments:

[0020] w(x vs , y vs , z vs ; x o , y o , z o ; s) = w x (x vs ; x o ; s)w y (y vs ; y o ; s)w z (z vs ; z o ; s),

[0021] Here w x (x vs ; x o ; s), w y (y vs ; y o ; s) and w z (z vs ; z o ; s) is x vs , y vs and z vs Represents independent weighting functions. According to some of these implementations, p can be a function of the audio object size(s).

[0022] Some of these methods may involve storing calculated virtual source gain values ​​in a memory system. The process of calculating contributions from virtual sources within an audio object region or volume may involve retrieving calculated virtual source gain values ​​corresponding to the audio object location and size from the memory system, and interpolating among the calculated virtual source gain values. The process of interpolating among the calculated virtual source gain values ​​may involve: determining a plurality of neighboring virtual source locations near the audio object location; determining calculated virtual source gain values ​​for each of the neighboring virtual source locations; determining a plurality of distances between the audio object location and each of the neighboring virtual source locations; and interpolating among the calculated virtual source gain values ​​according to the plurality of distances.

[0023] In some implementations, the playback environment data may include playback environment boundary data. The method may involve determining that an audio object region or volume includes an outer region or volume outside the playback environment boundary and applying a fade-out factor based at least partially on said outer region or volume. Some methods may involve determining that an audio object may be within a threshold distance from the playback environment boundary and not providing any speaker supply signals to playback speakers on the opposing boundary of said playback environment. In some implementations, the audio object region or volume may be a rectangle, a rectangular prism, a circle, a sphere, an ellipse, and / or an ellipsoid.

[0024] Some methods may involve decorrelating at least a portion of the audio playback data. For example, the methods may involve decorrelating audio playback data for audio objects having an audio object size exceeding a threshold.

[0025] Alternative methods are described herein. Some of these methods involve receiving playback environment data, including playback speaker location data and playback environment boundary data, and receiving audio playback data, including one or more audio objects and associated metadata. The metadata may include audio object location data and audio object size data. The methods may involve determining that an audio object region or volume, defined by the audio object location data and the audio object size data, includes an outer region or volume outside the playback environment boundary, and determining a fade-out factor based at least partially on said outer region or volume. The methods may involve calculating a set of gain values ​​for each of a plurality of output channels based at least partially on said associated metadata and said fade-out factor. Each output channel may correspond to at least one playback speaker of said playback environment. The fade-out factor may be proportional to said outer region.

[0026] The above methods may also involve determining that an audio object may be within a critical distance from the boundary of the playback environment and not providing any speaker supply signals to playback speakers on the opposite boundary of the playback environment.

[0027] The above methods may also involve calculating contributions from virtual sources within the audio object region or volume. The methods may involve defining a plurality of virtual source locations according to the playback environment data and, for each of the virtual source locations, calculating a virtual source gain for each of the plurality of output channels. The virtual source locations may be uniformly spaced or not spaced, depending on the particular implementation.

[0028] Some implementations may be represented on one or more non-transient media storing the software. The software may include instructions for controlling one or more devices for receiving audio playback data including one or more audio objects. The audio objects may include audio signals and associated metadata. The metadata may include at least audio object location data and audio object size data. The software may include, for an audio object from the one or more audio objects, calculating contributions from virtual sources within a region or volume defined by the audio object location data and the audio object size data, and at least partially, calculating a set of audio object gain values ​​for each of a plurality of output channels based on the calculated contributions. Each output channel may correspond to at least one playback speaker in the playback environment.

[0029] In some implementations, the process of calculating contributions from virtual sources may involve calculating a weighted average of virtual source gain values ​​from the virtual sources within the audio object region or volume. The weights for the weighted average may depend on the position of the audio object within the audio object region or volume, the size of the audio object, and / or the position of each virtual source.

[0030] The software may include instructions for receiving playback environment data, including playback speaker location data. The software may define a plurality of virtual source locations according to the playback environment data and, for each of the virtual source locations, include instructions for calculating a virtual source gain value for each of the plurality of output channels. Each of the virtual source locations may correspond to a location within the playback environment. In some implementations, at least some of the virtual source locations may correspond to locations outside the playback environment.

[0031] According to some implementations, the virtual source locations may be uniformly spaced. In some implementations, the virtual source locations may have a first uniform interval along the x and y axes and a second uniform interval along the z axis. The process of calculating a set of audio object gain values ​​for each of the plurality of output channels may involve independent calculations of contributions from virtual sources along the x, y, and z axes.

[0032] Various devices and apparatuses are described herein. Some of these devices may include an interface system and a logic system. The interface system may include a network interface. In some implementations, the device may include a memory device. The interface system may include an interface between the logic system and the memory device.

[0033] The logic system may be adapted to receive audio playback data including one or more audio objects from the interface system. The audio objects may include audio signals and associated metadata. The metadata may include at least audio object location data and audio object size data. The logic system may be adapted to calculate contributions from virtual sources within an audio object region or volume defined by the audio object location data and the audio object size data for an audio object from the one or more audio objects. The logic system may be adapted, at least partially, to calculate a set of audio object gain values ​​for each of a plurality of output channels based on the calculated contributions. Each output channel may correspond to at least one playback speaker in the playback environment.

[0034] The process of calculating contributions from virtual sources may involve calculating a weighted average of virtual source gain values ​​from the virtual sources within the audio object region or volume. The weights of the weighted average may depend on the position of the audio object within the audio object region or volume, the size of the audio object, and the position of each virtual source. The logic system may be adapted to receive playback environment data, including playback speaker position data, from the interface system.

[0035] The logic system may define a plurality of virtual source locations according to the playback environment data and, for each of the virtual source locations, be adapted to calculate a virtual source gain value for each of the plurality of output channels. Each of the virtual source locations may correspond to a location within the playback environment. However, in some implementations, at least some of the virtual source locations may correspond to locations outside the playback environment. Depending on the implementation, the virtual source locations may be uniformly spaced or not spaced. In some implementations, the virtual source locations may have a first uniform interval along the x and y axes and a second uniform interval along the z axis. The process of calculating a set of audio object gain values ​​for each of the plurality of output channels may involve independent calculations of contributions from virtual sources along the x, y, and z axes.

[0036] The device may also include a user interface. The logic system may be adapted to receive user input, such as audio object size data, through the user interface. In some implementations, the logic system may be adapted to scale the input audio object size data.

[0037] Details of one or more embodiments of the subject matter described herein are set forth in the following accompanying drawings and description. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Note that the relative dimensions in the following drawings may not be drawn in a constant proportion. Effects of the invention

[0038] According to the present invention, audio playback data for playback environments such as cinema sound playback systems can be authored and rendered. Brief explanation of the drawing

[0039] Figure 1 illustrates an example of a playback environment having a Dolby Surround 5.1 configuration. Figure 2 illustrates an example of a playback environment having a Dolby Surround 7.1 configuration. FIG. 3 illustrates an example of a playback environment having a Hamasaki 22.2 surround sound configuration. FIG. 4A illustrates an example of a graphical user interface (GUI) representing speaker zones at varying elevations in a virtual playback environment. Figure 4B illustrates another example of a regeneration environment. Figure 5A is a flowchart providing an overview of an audio processing method. Figure 5B is a flowchart providing an example of a setup process. FIG. 5C is a flowchart providing an example of a runtime process for calculating gain values ​​for received audio objects based on pre-calculated gain values ​​for virtual source locations. Figure 6A illustrates examples of virtual source locations for a playback environment. Fig. 6B illustrates an alternative example of virtual source locations for a playback environment. FIGS. 6C through 6F illustrate examples of applying near-field and far-field panning techniques to audio objects at different locations. FIG. 6G illustrates an example of a playback environment having one speaker at each corner of a square having an edge length equal to 1. Figure 7 illustrates examples of contributions from virtual sources within an area defined by audio object location data and audio object size data. FIGS. 8A and FIGS. 8B illustrate audio objects at two locations within a playback environment. FIG. 9 is a flowchart outlining a method for determining a fade-out factor based on how much of an area or volume of an audio object extends beyond the boundaries of the playback environment, at least partially. FIG. 10 is a block diagram providing examples of components of a writing and / or rendering device. FIG. 11A is a block diagram showing some components that can be used to generate audio content. FIG. 11B is a block diagram showing some components that can be used for audio playback in a playback environment. In various drawings, similar reference numbers and names indicate similar elements. Specific details for implementing the invention

[0040] The following description relates to specific implementations for the purpose of describing some innovative aspects of the present disclosure, as well as examples of contexts in which these innovative aspects may be implemented. However, the teachings herein may be applied in various different ways. For example, while various implementations have been described for specific playback environments, the teachings herein are broadly applicable to other known playback environments, as well as playback environments that may be introduced in the future. Furthermore, the described implementations may be implemented in various authoring and / or rendering tools, which may be implemented in various hardware, software, firmware, etc. Accordingly, the teachings of the present disclosure are not intended to be limited to the implementations shown in the drawings and / or described herein, but instead have broad applicability.

[0041] FIG. 1 illustrates an example of a playback environment having a Dolby Surround 5.1 configuration. Although Dolby Surround 5.1 was developed in the 1990s, this configuration is still widely deployed in cinema sound system environments. A projector (105) can be configured to project video images onto a screen (150), for example, for movies. Audio playback data is synchronized with the video images and can be processed by a sound processor (110). Power amplifiers (115) can provide speaker supply signals to the speakers of the playback environment (100).

[0042] The Dolby Surround 5.1 configuration includes a left surround array (120) and a right surround array (125), each comprising a group of speakers gang-driven by a single channel. The Dolby Surround 5.1 configuration also includes separate channels for a left screen channel (130), a center screen channel (135), and a right screen channel (140). A separate channel for a subwoofer (145) is provided for low-frequency effects (LFE).

[0043] In 2010, Dolby provided enhancements to digital cinema sound by introducing Dolby Surround 7.1. FIG. 2 illustrates an example of a playback environment having a Dolby Surround 7.1 configuration. A digital projector (205) may be configured to receive digital video data and project video images onto a screen (150). Audio playback data may be processed by a sound processor (210). Power amplifiers (215) may provide speaker supply signals to speakers of the playback environment (200).

[0044] A Dolby Surround 7.1 configuration includes a left side surround array (220) and a right side surround array (225), each of which can be driven by a single channel. Like Dolby Surround 5.1, a Dolby Surround 7.1 configuration includes separate channels for a left screen channel (230), a center screen channel (235), a right screen channel (240), and a subwoofer (245). However, Dolby Surround 7.1 increases the number of surround channels by separating the left and right surround channels of Dolby Surround 5.1 into four zones: in addition to the left side surround array (220) and the right side surround array (225), separate channels are included for left rear surround speakers (224) and right rear surround speakers (226). Increasing the number of surround zones within the playback environment (200) can significantly improve the localization of sound.

[0045] In an effort to create a more immersive environment, some playback environments may be configured with an increased number of speakers driven by an increased number of channels. Furthermore, some playback environments may include speakers placed at various heights, with a portion of which may be located above the seating area of ​​the playback environment.

[0046] FIG. 3 illustrates an example of a playback environment having a Hamasaki 22.2 surround sound configuration. The Hamasaki 22.2 was developed by NHK Science & Technology Research Laboratories in Japan as a surround sound component for ultra-high definition television. The Hamasaki 22.2 provides 24 speaker channels that can be used to drive speakers arranged in three layers. The upper speaker layer (310) of the playback environment (300) can be driven by 9 channels. The middle speaker layer (320) can be driven by 10 channels. The lower speaker layer (330) can be driven by 5 channels, two of which are for subwoofers (345a and 345b).

[0047] Accordingly, current trends involve not only including more speakers and more channels, but also including speakers at different heights. As the number of channels increases and speaker placement transitions from 2D arrays to 3D arrays, the tasks of positioning and rendering sounds are becoming increasingly difficult. Accordingly, the assignee is developing various tools as well as related user interfaces, which increase the capabilities for 3D audio sound systems and / or reduce authoring complexity. Some of these tools are described in detail with reference to FIGS. 5A through 19D of U.S. Provisional Patent Application No. 61 / 636,102, filed April 20, 2012, titled “System and tools for enhanced 3D audio authoring and rendering” (“Authoring and Rendering Application”), which is incorporated herein by reference.

[0048] FIG. 4A illustrates an example of a graphical user interface (GUI) representing speaker zones at variable elevations in a virtual playback environment. The GUI (400) may be displayed on a display device according to instructions from a logic system, for example, according to signals received from user input devices. Some such devices are described below with reference to FIG. 10.

[0049] As used herein with reference to virtual playback environments such as the virtual playback environment (404), the term “speaker zone” generally refers to a logical configuration that may or may not have a one-to-one correspondence with the playback speakers of the actual playback environment. For example, a “speaker zone location” may or may not correspond to a specific playback speaker location in the cinema playback environment. Alternatively, the term “speaker zone location” may generally refer to a zone in the virtual playback environment. In some implementations, the speaker zones of the virtual playback environment may correspond to virtual speakers through the use of virtualization technology, such as Dolby Headphones™ (sometimes called Mobile Surround™), which creates a virtual surround sound environment in real time using a set of 2-channel stereo headphones, for example. In the GUI (400), there are seven speaker zones (402a) at the first elevation and two speaker zones (402b) at the second elevation, creating a total of nine speaker zones in the virtual playback environment (404). In this example, speaker zones (1 to 3) are located in the front area (405) of the virtual playback environment (404). The front area (405) may correspond to the area of ​​the cinema playback environment where the screen (150) is located, for example, in a home area where a television screen is located.

[0050] Here, speaker area (4) generally corresponds to speakers in the left area (410), and speaker area (5) corresponds to speakers in the right area (415) of the virtual playback environment (404). Speaker area (6) corresponds to the left rear area (412), and speaker area (7) corresponds to the right rear area (414) of the virtual playback environment (404). Speaker area (8) corresponds to speakers in the upper area (420a), and speaker area (9) corresponds to speakers in the upper area (420b), which may be the virtual ceiling area. Thus, as described in more detail in the authoring and rendering application, the locations of the speaker areas (1 through 9) shown in FIG. 4A may or may not correspond to the locations of the playback speakers in the actual playback environment. Furthermore, other implementations may include more or fewer speaker areas and / or elevations.

[0051] In the various implementations described in the authoring and rendering application, a user interface such as a GUI (400) may be used as part of the authoring tool and / or rendering tool. In some implementations, the authoring tool and / or rendering tool may be implemented through software stored on one or more non-transient media. The authoring tool and / or rendering tool may be implemented (at least partially) by hardware, firmware, etc., such as the logic system and other devices described below with reference to FIG. 10. In some authoring implementations, the associated authoring tool may be used to generate metadata for the associated audio data. The metadata may include, for example, data indicating the location and / or trajectory of an audio object in three-dimensional space, speaker area constraint data, etc. The metadata may be generated for speaker areas (402) of a virtual playback environment (404) rather than for a specific speaker arrangement in a real playback environment. The rendering tool may receive the audio data and the associated metadata and may calculate audio gains and speaker supply signals for the playback environment. These audio gains and speaker supply signals can be calculated according to an amplitude panning process, which can generate the perception that sound is coming from a location (P) in the playback environment. For example, speaker supply signals can be provided to playback speakers (1 to N) in the playback environment according to the following equation:

[0052] x i (t) = g i x(t), i = 1, ...N (Equation 1)

[0053] In Equation 1, x i (t) represents the speaker supply signal to be applied to speaker (i), and g irepresents the gain factor of the corresponding channel, x(t) represents the audio signal, and t represents time. The gain factors may be determined, for example, according to the amplitude panning methods described in Section 2, pages 3-4 of V. Pulkki’s *Compensating Displacement of Amplitude-Panned Virtual Sources* (AES International Conference on Virtual, Synthetic and Entertainment Audio), which are incorporated herein by reference. In some implementations, the gains may be frequency-dependent. In some implementations, a time delay may be introduced by replacing x(t) with x(t-△t).

[0054] In some rendering implementations, audio playback data generated by referencing speaker zones (402) can be mapped to speaker locations in a wide range of playback environments, which may be in a Dolby Surround 5.1 configuration, a Dolby Surround 7.1 configuration, a Hamasaki 22.2 configuration, or another configuration. For example, referring to FIG. 2, a rendering tool can map audio playback data for speaker zones (4 and 5) to the left side surround array (220) and the right side surround array (225) of a playback environment having a Dolby Surround 7.1 configuration. Audio playback data for speaker zones (1, 2 and 3) can be mapped to the left screen channel (230), the right screen channel (240), and the center screen channel (235), respectively. Audio playback data for speaker zones (6 and 7) can be mapped to the left rear surround speakers (224) and the right rear surround speakers (226).

[0055] FIG. 4B illustrates another example of a playback environment. In some implementations, a rendering tool may map audio playback data for speaker zones (1, 2 and 3) to corresponding screen speakers (455) of a playback environment (450). The rendering tool may map audio playback data for speaker zones (4 and 5) to a left-side surround array (460) and a right-side surround array (465), and audio playback data for speaker zones (8 and 9) to left overhead speakers (470a) and right overhead speakers (470b). Audio playback data for speaker zones (6 and 7) may be mapped to left rear surround speakers (480a) and right rear surround speakers (480b).

[0056] In some authoring implementations, authoring tools may be used to generate metadata for audio objects. As noted above, the term "audio object" may refer to a stream of audio data signals and associated metadata. Metadata may indicate the 3D position of the audio object, the apparent size of the audio object, rendering constraints, as well as content types (e.g., dialogue, effects). Depending on the implementation, metadata may include other types of data, such as gain data, trajectory data, etc. Some audio objects may be static, while others may be movable. Audio object details may be authored or rendered according to associated metadata that indicates, among other things, the position of the audio object in 3-dimensional space at a given time point. When audio objects are monitored or played in a playback environment, the audio objects may be rendered according to their position and size metadata based on the placement of the playback speakers in the playback environment.

[0057] FIG. 5A is a flowchart providing an overview of an audio processing method. More detailed examples are described below with reference to FIG. 5B and subsequent. These methods may include more or fewer blocks than those shown and described herein and are not necessarily performed in the order shown herein. These methods may be performed at least partially by devices such as those shown in FIG. 10 through 11b and described below. In some embodiments, these methods may be implemented at least partially by software stored on one or more non-transient media. The software may include instructions for controlling one or more devices to perform the methods described herein.

[0058] In the example illustrated in FIG. 5A, the method (500) begins with a setup process (block 505) for determining virtual source gain values ​​for virtual source locations for a specific playback environment. FIG. 6A illustrates an example of virtual source locations for a playback environment. For example, block (505) may involve determining virtual source gain values ​​for virtual source locations (605) for playback speaker locations (625) of a playback environment (600a). The virtual source locations (605) and playback speaker locations (625) are merely examples. In the example illustrated in FIG. 6A, the virtual source locations (605) are uniformly spaced along the x, y, and z axes. However, in alternative implementations, the virtual source locations (605) may be spaced differently. For example, in some implementations, the virtual source locations (605) may have a first uniform spacing along the x and y axes and a second uniform spacing along the z axis. In other implementations, virtual source locations (605) may be unevenly spaced.

[0059] In the example illustrated in FIG. 6A, the playback environment (600a) and the virtual source volume (602a) are co-extensive, and thus each of the virtual source locations (605) corresponds to a location within the playback environment (600a). However, in alternative implementations, the playback environment (600) and the virtual source volume (602) may not be co-extensive. For example, at least some of the virtual source locations (605) may correspond to locations outside the playback environment (600).

[0060] FIG. 6B illustrates an alternative example of virtual source locations for a playback environment. In this example, the virtual source volume (602b) extends outside the playback environment (600b).

[0061] Referring again to FIG. 5A, in this example, the setup process of block (505) occurs before any specific audio objects are rendered. In some implementations, the virtual source gain values ​​determined in block (505) may be stored in a storage system. The stored virtual source gain values ​​may be used during a "runtime" process to calculate audio object gain values ​​for received audio objects based on at least some of the virtual source gain values ​​(block 510). For example, block (510) may involve calculating audio object gain values ​​based on virtual source gain values ​​corresponding to virtual source locations within an audio object region or volume, at least partially.

[0062] In some implementations, the method (500) may include an optional block (515) that involves decorrelating audio data. The block (515) may be part of a run-time process. In some of these implementations, the block (515) may involve convolution in the frequency domain. For example, the block (515) may involve applying a finite impulse response ("FIR") filter for each speaker supply signal.

[0063] In some implementations, the processes of block (515) may or may not be performed depending on the audio object size and / or the author's artistic intent. According to some of these implementations, the authoring tool may relate the inverse correlation to the audio object size by indicating that the inverse correlation should be turned on when the audio object size is greater than or equal to a size threshold and that the inverse correlation should be turned off when the audio object size is below the size threshold (e.g., via an inverse correlation flag included in associated metadata). In some implementations, the inverse correlation may be controlled based on user input regarding the size threshold and / or other input values ​​(e.g., increased, decreased, or disabled).

[0064] FIG. 5B is a flowchart providing an example of a setup process. Thus, all of the blocks illustrated in FIG. 5B are examples of processes that can be performed in block (505) of FIG. 5A. Here, the setup process begins with the reception of playback environment data (block 520). The playback environment data may include playback speaker location data. The playback environment data may also include data indicating the boundaries of the playback environment, such as walls, ceilings, etc. If the playback environment is a cinema, the playback environment data may also include an indication of the movie screen location.

[0065] The playback environment data may also include data indicating the correlation between the playback speakers and output channels of the playback environment. For example, the playback environment may have a Dolby Surround 7.1 configuration as illustrated in FIG. 2 and described above. Accordingly, the playback environment data may also include data indicating the correlation between the Lss channel and the left side surround speakers (220), between the Lrs channel and the left rear surround speakers (224), etc.

[0066] In this example, the block (525) involves defining virtual source locations (605) according to playback environment data. The virtual source locations (605) may be defined within a virtual source volume. In some implementations, the virtual source volume may correspond to a volume in which audio objects can move. As illustrated in FIGS. 6A and 6B, in some implementations, the virtual source volume (602) may correspond to a volume of the playback environment (600), whereas in other implementations, at least some of the virtual source locations (605) may correspond to locations outside the playback environment (600).

[0067] Furthermore, the virtual source locations (605) may or may not be uniformly spaced within the virtual source volume (602), depending on the specific implementation. In some implementations, the virtual source locations (605) may be uniformly spaced in all directions. For example, the virtual source locations (605) are N x ×N y ×N z A rectangular grid of virtual source locations (605) can be formed. In some implementations, the value of N may be in the range of 5 to 100. The value of N may depend at least partially on the number of playback speakers in the playback environment: it may be desirable to include two or more virtual source locations (605) between each playback speaker location.

[0068] In other implementations, the virtual source locations (605) may have a first uniform interval along the x and y axes and a second uniform interval along the z axis. The virtual source locations (605) are N x ×N y ×M z A rectangular grid of virtual source locations (605) can be formed. For example, in some implementations, there may be fewer virtual source locations (605) along the z-axis than along the x or y axes. In some of these implementations, the value of N may be in the range of 10 to 100, while the value of M may be in the range of 5 to 10.

[0069] In this example, block (530) involves calculating virtual source gain values ​​for each of the virtual source locations (605). In some implementations, block (530) involves calculating virtual source gain values ​​for each of the multiple output channels of the playback environment for each of the virtual source locations (605). In some implementations, block (530) may involve applying a vector-based amplitude panning ("VBAP") algorithm, a pairwise panning algorithm, or a similar algorithm to calculate gain values ​​for point sources located at each of the virtual source locations (605). In other implementations, block (530) may involve applying a separable algorithm to calculate gain values ​​for point sources located at each of the virtual source locations (605). As used herein, a "separable" algorithm is one in which the gain of a given speaker can be expressed as the product of two or more factors that can be calculated individually for each of the coordinates of the virtual source location. Examples include, but are not limited to, algorithms implemented in various existing mixing console panners, including Pro Tools™ software and panners implemented in digital film consoles provided by AMS Neve. Some two-dimensional examples are provided below.

[0070] FIGS. 6C through 6F illustrate examples of applying near-field and far-field panning techniques to audio objects at different locations. Referring first to FIG. 6C, the audio object is substantially outside the virtual playback environment (400a). Therefore, one or more far-field panning methods will be applied in this instance. In some implementations, far-field panning methods may be based on vector-based amplitude panning (VBAP) equations known to those skilled in the art. For example, far-field panning methods may be based on VBAP equations described in page 4, section 2.3 of V. Pulkki’s Method for Compensating Displacement of Amplitude-Panned Virtual Sources (AES International Conference on Virtual, Synthetic and Entertainment Audio), which is incorporated herein by reference. In alternative implementations, other methods, such as those involving the synthesis of corresponding acoustic planes or spherical waves, may be used to pan far-field and near-field audio objects. D. dE Vreis, Wave Field Synthesis (AES Monograph 1999), incorporated here by reference, describes the relevant methods.

[0071] Now, referring to FIG. 6D, the audio object (610) is inside the virtual playback environment (400a). Therefore, one or more near-field panning methods will be applied in this instance. Some of these near-field panning methods will use multiple speaker zones surrounding the audio object (610) in the virtual playback environment (400a).

[0072] FIG. 6G illustrates an example of a playback environment having one speaker at each corner of a square having an edge length equal to 1. In this example, the origin (0,0) of the xy axis coincides with the left (L) screen speaker (130). Thus, the right (R) screen speaker (140) has coordinates (1,0), the left surround (Ls) speaker (120) has coordinates (0,1), and the right surround (Rs) speaker (125) has coordinates (1,1). The audio object location (615)(x, y) is x units to the right of the L speaker and y units from the screen (150). In this example, each of the four speakers receives a factor cos / sin proportional to their distance along the x and y axes. According to some implementations, the gains can be calculated as follows:

[0073] If 1=L, Ls, then G_1(x) = cos(pi / 2* x)

[0074] If 1=R,Rs, then G_1(x) = sin(pi / 2* x)

[0075] If 1=L,R, then G_1(y) = cos(pi / 2* y)

[0076] If 1=Ls,Rs, then G_1(y) = sin(pi / 2* y).

[0077] The total gain is the product: G_1(x,y) = G_1(x)G_1(y). In general, these functions depend on the coordinates of all speakers. However, G_1(x) does not depend on the y-position of the source, and G_1(y) does not depend on its x-position. To illustrate a simple output, let's assume the audio object position (615) is (0,0), which is the position of speaker L. G_L(x) = cos(0) = 1. G_L(y) = cos(0) = 1. The total gain is the product: G_L(x,y) = G_L(x)G_L(y)=1. Similar outputs lead to G_Ls = G_Rs = G_R = 0.

[0078] It may be desirable to blend between different panning modes when an audio object enters or leaves a virtual playback environment (400a). For example, blending of gains calculated according to near-field panning methods and far-field panning methods may be applied when the audio object (610) moves from the audio object position (615) shown in FIG. 6C to the audio object position (615) shown in FIG. 6D or vice versa. In some implementations, a pair-star panning law (e.g., an energy-conserving sine or power law) may be used to blend between gains calculated according to near-field panning methods and far-field panning methods. In alternative implementations, the pair-star panning law may be amplitude-conserving rather than energy-conserving, and thus the sum is equal to 1 instead of the sum of squares equal to 1. For example, it is also possible to process an audio signal using both panning methods independently and to blend the resulting processed signals to cross-fade the two resulting audio signals.

[0079] Now, moving on to Fig. 5B, regardless of the algorithm used in block (530), the resulting gain values ​​can be stored in a memory system for use during runtime operations (block 535).

[0080] FIG. 5C is a flowchart providing an example of a runtime process for calculating gain values ​​for received audio objects according to pre-calculated gain values ​​for virtual source locations. All of the blocks illustrated in FIG. 5C are examples of processes that can be performed in the block (510) of FIG. 5A.

[0081] In this example, the runtime process begins with the reception of audio playback data containing one or more audio objects (Block 540). The audio objects include audio signals and associated metadata, in this example, including at least audio object location data and audio object size data. Referring to FIG. 6A, for example, an audio object (610) is defined at least partially by an audio object location (615) and an audio object volume (620a). In this example, the received audio object size data indicates that the audio object volume (620a) corresponds to that of a rectangular prism. However, in the example illustrated in FIG. 6B, the received audio object size data indicates that the audio object volume (620b) corresponds to that of a sphere. These sizes and shapes are merely examples; in alternative implementations, audio objects may have various other sizes and / or shapes. In some alternative examples, the area or volume of the audio object may be a rectangular, circular, elliptical, ellipsoidal, or spherical sector.

[0082] In this implementation, block (545) involves calculating contributions from virtual sources within an area or volume defined by audio object location data and audio object size data. In the examples illustrated in FIGS. 6A and 6B, block (545) may involve calculating contributions from virtual sources at virtual source locations (605) within the audio object volume (620a) or audio object volume (620b). If the metadata of the audio object changes over time, block (545) may be re-executed according to the new metadata values. For example, if the audio object size and / or audio object location changes, different virtual source locations (605) may be included within the audio object volume (620) and / or the virtual source locations (605) used in the previous calculation may be at different distances from the audio object location (615). In block (545), the corresponding virtual source contributions will be calculated according to the new audio object size and / or location.

[0083] In some examples, block (545) may involve retrieving virtual source gain values ​​calculated for virtual source locations corresponding to audio object locations and sizes from a memory system, and interpolating between said calculated virtual source gain values. The process of interpolating between the calculated virtual source gain values ​​may involve determining a plurality of neighboring virtual source locations close to the audio object location, determining virtual source gain values ​​calculated for each of said neighboring virtual source locations, determining a plurality of distances between the audio object location and each of said neighboring virtual source locations, and interpolating between said calculated virtual source gain values ​​according to said plurality of distances.

[0084] The process of calculating contributions from virtual sources may involve calculating a weighted average of calculated virtual source gain values ​​for virtual source locations within an area or volume defined by the size of the audio object. The weights for the weighted average may depend, for example, on the location of the audio object within the said area or volume, the size of the audio object, and each virtual source location.

[0085] FIG. 7 illustrates an example of contributions from virtual sources within an area defined by audio object position data and audio object size data. FIG. 7 depicts a cross-section of the audio environment (200a) taken perpendicular to the z-axis. Thus, FIG. 7 is drawn from the perspective of a viewer looking downward into the audio environment (200a) along the z-axis. In this example, the audio environment (200a) is a cinema sound system environment having a Dolby Surround 7.1 configuration as illustrated in FIG. 2 and described above. Thus, the playback environment (200a) includes left side surround speakers (220), left rear surround speakers (224), right side surround speakers (225), right rear surround speakers (226), a left screen channel (230), a center screen channel (235), a right screen channel (240), and a subwoofer (245).

[0086] The audio object (610) has a size indicated by the audio object volume (620b), the rectangular single region of which is illustrated in FIG. 7. When considering the audio object location (615) at the time instance depicted in FIG. 7, 12 virtual source locations (605) are contained within the area contained by the audio object volume (620b) in the xy plane. Depending on the degree of the audio object volume (620b) in the z direction and the spacing of the virtual source locations (605) along the z-axis, additional virtual source locations (605s) may or may not be contained within the audio object volume (620b).

[0087] FIG. 7 illustrates contributions from virtual source locations (605) within an area or volume defined by the size of the audio object (610). In this example, the diameter of the circle used to depict each of the virtual source locations (605) corresponds to the contribution from the corresponding virtual source location (605). The virtual source locations (605a) are closest to the audio object location (615) depicted as the largest and represent the largest contribution from the corresponding virtual sources. The second largest contributions are from virtual sources at the virtual source locations (605b), which are the second closest to the audio object location (615). Smaller contributions are made by the virtual source locations (605c), which are further from the audio object location (615) but are still within the audio object volume (620b). Virtual source locations (605d) outside the audio object volume (620b) are shown as the smallest, which indicates that the corresponding virtual sources in this example do not contribute any.

[0088] Moving to FIG. 5C, in this example, block (550) involves, at least partially, calculating a set of audio object gain values ​​for each of a plurality of output channels based on calculated contributions. Each output channel may correspond to at least one playback speaker in the playback environment. Block (550) may involve normalizing the resulting audio object gain values. For the implementation illustrated in FIG. 7, for example, each output channel may correspond to a single speaker or a group of speakers.

[0089] The process of calculating the audio object gain value for each of multiple output channels is the position (x o , y o , z o Gain value (g) for an audio object of size(s) to be rendered in ) l size (x o , y o , z o ; s)) may involve determining. This audio object gain value may sometimes be referred to here as the "audio object size contribution." According to some implementations, the audio object gain value (g l size (x o , y o , z o ; s)) can be expressed as follows:

[0090] . (Equation 2)

[0091] In Equation 2, (x vs , y vs , z vs ) indicates the virtual source location, and g l (x vs , y vs , z vs ) is the virtual source location (x vs , y vs , z vs Represents the gain value for channel (l) for ) and w(x vs, yv s , z vs ; x o , y o , z o ;s) is at least partially, the position of the audio object (x o , y o , z o ), size(s) of the audio object and virtual source position(x vs , y vs , z vs g determined based on ) l (x vs , y vs , z vs Represents the weighting for ).

[0092] In some examples, the exponent (p) can have a value between 1 and 10. In some implementations, p can be a function of the audio object size (s). For example, if s is relatively large, in some implementations, p can be relatively smaller. According to some of these implementations, p can be determined as follows:

[0093] If s≤0.5, then p = 6

[0094] If s > 0.5, p = 6 + (-4)(s - 0.5) / (s max -0.5)

[0095] Here s max is the internal scale-up size(s internal Corresponding to the maximum value of )(described below), audio object size(s) = 1 may correspond to an audio object having a size (e.g., diameter) equal to the length of one of the boundaries of the playback environment (e.g., the length of one wall of the playback environment).

[0096] It may be possible to simplify Equation 2 by partially depending on the algorithm(s) used to calculate the virtual source gain values, for example, if the virtual source locations are uniformly distributed along the axis as described above and the weighting functions and gain functions are separable. If these conditions are satisfied, g l (x vs , y vs , z vs ) is g lx (x vs )g ly (y vs )g lz (z vs It can be expressed as ), where g lx (x vs ), g lx (y vs ) and g lz (z vs ) represents independent gain functions of x, y, and z coordinates for the location of a virtual source.

[0097] Similarly, w(x vs , y vs , z vs ; x o , y o , z o ;s) is w x (x vs ;x o ;s)w y (y vs ;y o ;s)w z (z vs ;z o It can be considered as ;s), and here w x (x vs ;x o ;s), w y (y vs ;y o ;s) and w z (z vs ;z o ;s) represents independent weighting functions of the x, y, and z coordinates for the location of the virtual source. One such example is illustrated in Fig. 7. In this example, wx (x vs ;x o The weighting function (710), expressed as ;s), is w y (y vs ;x o It can be calculated independently from the weighting function (720) expressed as ;s). In some implementations, the weighting functions (710 and 720) may be Gaussian functions, whereas the weighting function (w z (z vs ;z o ;s)) can be the product of cosine and Gaussian functions.

[0098] w(x vs , y vs , z vs ; x o , y o , z o ;s) is w x (x vs ;x o ;s)w y (y vs ;y o ;s)w z (z vs ;z o If it can be a factor like ;s), then Equation 2 is,

[0099] Simplified as, here

[0100]

[0101] Functions (f) may include all required information regarding virtual sources. If possible object locations are discretized along each axis, each function (f) can be represented as a matrix. Each function (f) may be pre-calculated during the setup process of block (505) (see FIG. 5A) and stored in a memory system, for example, as a matrix or as a lookup table. At runtime (block 510), the lookup tables or matrices may be retrieved from the memory system. The runtime process may involve interpolating between the nearest corresponding values ​​of these matrices, taking into account the audio object locations and sizes. In some implementations, the interpolation may be linear.

[0102] In some implementations, the audio object size contribution (g l size ) can be combined with the "audio object neargain" result for the audio object location. As used herein, the "audio object neargain" is a calculated gain based on the audio object location (615). The gain calculation can be performed using the same algorithm used to calculate each of the virtual source gain values. According to some of these implementations, cross-fading calculation can be performed between the audio object size contribution and the audio object neargain result, for example as a function of the audio object size. These implementations can provide smooth panning and smooth growth of audio objects and allow for smooth transitions between minimum and maximum audio object sizes. In one such implementation,

[0103] , here

[0104]

[0105] Here is previously calculated Represents the normalized version of. In some of these implementations, s xfade = 0.2. However, in alternative implementations, s xfade It can have other values.

[0106] According to some implementations, audio object size values ​​can be scaled up to a larger portion of their range of possible values. In some authoring implementations, for example, the user can scale up to a larger range, e.g., range([0, s max Audio object size values ​​(s) that are mapped to the actual size used by the algorithm up to ]) user It can be exposed to ∈[0.1]), where s max >1. Such mapping can guarantee that when the size is set to the maximum by the user, the gains become truly independent of the object's position. According to some of these implementations, these mappings are pairs of points (s user , s internal It can be performed according to a piece-wise linear function connecting ), where s user represents the size of the user-selected audio object and s internal represents the corresponding audio object size determined by the algorithm. According to some of these implementations, the mapping consists of pairs of points ((0,0), (0.2, 0.3), (0.5, 0.9), (0.75, 1.5) and (1, s max It can be achieved according to an interval linear function connecting )). In one such implementation, s max = 2.8.

[0107] FIGS. 8A and FIGS. 8B illustrate an audio object at two locations within a playback environment. In these examples, the audio object volume (620b) is a sphere having a radius less than half the length or width of the playback environment (200a). The playback environment (200a) is configured according to Dolby 7.1. In the time instance depicted in FIG. 8A, the audio object location (615) is relatively closer to the middle of the playback environment (200a). In the time depicted in FIG. 8B, the audio object location (615) moves closer to the boundary of the playback environment (200a). In this example, the boundary is the left wall of the cinema and coincides with the locations of the left-side surround speakers (220).

[0108] For aesthetic reasons, it may be desirable to change the audio object gain outputs for audio objects reaching the boundaries of the playback environment. In FIGS. 8A and 8B, for example, no speaker supply signals are provided to speakers on the opposite boundary of the playback environment (here, right side surround speakers (225)) when the audio object location (615) is within a threshold distance from the left boundary (805) of the playback environment. In the example illustrated in FIG. 8B, no speaker supply signals are provided to speakers corresponding to the left screen channel (230), center screen channel (235), right screen channel (240), or subwoofer (245) when the audio object location (615) is within a threshold distance (which may be a different threshold distance) from the left boundary (805) of the playback environment, if the audio object location (615) is also greater than a threshold distance from the screen.

[0109] In the example illustrated in FIG. 8B, the audio object volume (620b) includes an area or volume outside the left boundary (805). According to some implementations, the fade-out factor for gain output may be based, at least in part, on how much of the left boundary (805) is within the audio object volume (620b) and / or how much of the area or volume of the audio object extends outside of these boundaries.

[0110] FIG. 9 is a flowchart outlining a method for determining a fade-out factor based, at least partially, on how much of an area or volume of an audio object extends outside the boundaries of the playback environment. In block (905), playback environment data is received. In this example, the playback environment data includes playback speaker position data and playback environment boundary data. Block (910) involves receiving audio playback data including one or more audio objects and associated metadata. In this example, the metadata includes at least audio object position data and audio object size data.

[0111] In this implementation, block (915) involves determining that an audio object region or volume, defined by audio object location data and audio object size data, includes an outer region or volume outside the playback environment boundary. Block (915) may also involve determining which proportion of the audio object region or volume is outside the playback environment boundary.

[0112] In block (920), a fade-out factor is determined. In this example, the fade-out factor may be based at least partially on the outer region. For example, the fade-out factor may be proportional to the outer region.

[0113] In block (925), a set of audio object gain values ​​can be calculated for each of a plurality of output channels at least partially based on the associated metadata (in this example, audio object position data and audio object size data) and a fade-out factor. Each output channel can correspond to at least one playback speaker in the playback environment.

[0114] In some implementations, audio object gain calculations may involve calculating contributions from virtual sources within an audio object region or volume. Virtual sources may correspond to multiple virtual source locations that can be defined by referencing playback environment data. The virtual source locations may be uniformly spaced or not spaced. For each of the virtual source locations, a virtual source gain value may be calculated for each of the multiple output channels. As described above, in some implementations, these virtual source gain values ​​may be calculated and stored during the setup process and then retrieved for use during runtime operations.

[0115] In some implementations, a fade-out factor can be applied to all virtual source gain values ​​corresponding to virtual source locations within the playback environment. In some implementations, can be modified as follows:

[0116] , here

[0117] d bound If ≥ s, fade-out factor = 1.

[0118] d bound If < s, fade-out factor = d bound / s

[0119] d bound represents the minimum distance between the boundary of the playback environment and the audio object location, and represents the contribution of virtual sources along the boundary. For example, referring to Fig. 8B, This may represent the contribution of virtual sources within the audio object volume (620b) and adjacent to the boundary (805). In this example, as in FIG. 6A, there are no virtual sources located outside the playback environment.

[0120] In alternative implementations, can be changed as follows:

[0121] ,

[0122] Here represents audio object gains based on virtual sources located outside the playback environment but within the audio object region or volume. For example, referring to FIG. 8B, It may represent the contribution of virtual sources that are within the audio object volume (620b) and outside the boundary (805). In this example, as in FIG. 6B, there are virtual sources on both the inside and outside of the playback environment.

[0123] FIG. 10 is a block diagram providing examples of components of a writing and / or rendering device. In this example, the device (1000) includes an interface system (1005). The interface system (1005) may include a network interface, such as a wireless network interface. Alternatively, or additionally, the interface system (1005) may include a Universal Serial Bus (USB) interface or another such interface.

[0124] The device (1000) includes a logic system (1010). The logic system (1010) may include a processor, such as a general-purpose single- or multi-chip processor. The logic system (1010) may include a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or discrete hardware components, or a combination thereof. The logic system (1010) may be configured to control other components of the device (1000). Although no interfaces between the components of the device (1000) are shown in FIG. 10, the logic system (1010) may be configured to have interfaces for communication with other components. The other components may or may not be configured to communicate with each other as appropriate.

[0125] The above logic system (1010) may be configured to perform audio authoring and / or rendering functions, including but not limited to the types of audio authoring and / or rendering functions described herein. In some such implementations, the logic system (1010) may be configured to operate (at least partially) according to software stored on one or more non-transient media. The non-transient media may include memory associated with the logic system (1010), such as random access memory (RAM) and / or read-only memory (ROM). The non-transient media may include memory of a memory system (1015). The memory system (1015) may include one or more suitable types of non-transient storage media, such as flash memory, a hard drive, etc.

[0126] The display system (1030) may include one or more suitable types of displays depending on the indication of the device (1000). For example, the display system (1030) may include a liquid crystal display, a plasma display, a bistable display, etc.

[0127] The user input system (1035) may include one or more devices configured to receive input from a user. In some implementations, the user input system (1035) may include a touch screen placed over the display of the display system (1030). The user input system (1035) may include a mouse, a trackball, a gesture detection system, a joystick, one or more GUIs and / or menus provided on the display system (1030), buttons, a keyboard, switches, etc. In some implementations, the user input system (1035) may include a microphone (1025): and the user may provide voice commands to the device (1000) through the microphone (1025). A logic system may be configured to control at least some operations of the device (1000) according to these voice commands and for speech recognition.

[0128] The power system (1040) may include one or more suitable energy storage devices, such as a nickel-cadmium battery or a lithium-ion battery. The power system (1040) may be configured to receive power from an electrical outlet.

[0129] FIG. 11A is a block diagram showing some components that can be used for audio content generation. The system (1100) can be used for audio content generation, for example, in mixing studios and / or dubbing stages. In this example, the system (1100) includes an audio and metadata authoring tool (1105) and a rendering tool (1110). In this implementation, the audio and metadata authoring tool (1105) and the rendering tool (1110) each include audio connection interfaces (1107 and 1112), which can be configured for communication via AES / EBU, MADI, analog, etc. The audio and metadata authoring tool (1105) and the rendering tool (1110) each include network interfaces (1109 and 1117), which can be configured to transmit and receive metadata via TCP / IP or any other suitable protocol. An interface (1120) is configured to output audio data to speakers.

[0130] The system (1100) may include, for example, an existing authoring system such as Pro Tools™ that runs a metadata generation tool (i.e., a panner as described herein) as a plugin. The panner may also run on a standalone system (e.g., a PC or a mixing console) connected to the rendering tool (1110), or on the same physical device as the rendering tool (1110). In the latter case, the panner and renderer may use a local connection, for example, through shared memory. The panner GUI may also be provided on a tablet device, a laptop, etc. The rendering tool (1110) may include a rendering system comprising a sound processor configured to execute rendering methods as described in FIGS. 5A through 5C and FIG. 9. The rendering system may include, for example, a personal computer, a laptop, etc., comprising interfaces for audio input / output and a suitable logic system.

[0131] FIG. 11B is a block diagram showing some components that may be used for audio playback in a playback environment (e.g., a movie theater). The system (1150) includes, in this example, a cinema server (1155) and a rendering system (1160). The cinema server (1155) and the rendering system (1160) each include network interfaces (1157 and 1162), which may be configured to transmit and receive audio objects via TCP / IP or any other suitable protocol. Interface (1164) is configured to output audio data to speakers.

[0132] Various modifications to the embodiments described in this disclosure may be readily apparent to those skilled in the art. The general principles defined herein may be applied to other embodiments without departing from the spirit or scope of this disclosure. Accordingly, the claims are not intended to be limited to the embodiments illustrated herein, but will be consistent with the broadest scope consistent with the disclosure, principles, and novel features disclosed herein. Explanation of the symbols

[0133] 100: Playback Environment 105: Projector 110: Sound processor 115: Power amplifier 120: Left surround array 125: Right surround array 130: Left screen channel 135: Center screen channel 140: Right screen channel 145: Subwoofer 150: Screen 200: Playback Environment 205: Digital Projector 210: Sound Processor 215: Power amplifier 220: Left side surround array 224: Left rear surround speaker 225: Right side surround array 226: Right rear surround speaker 230: Left screen channel 235: Center screen channel 240: Right screen channel 245: Subwoofer 300: Playback Environment 310: Upper speaker layer 320: Middle speaker layer 330: Lower speaker layer 345a, 345b: Subwoofer 400: GUI 402: Speaker Zone 404: Virtual Playback Environment 405: Forward Area 450: Playback Environment 455: Screen Speakers 460: Left side surround array 465: Right side surround array 610: Audio Object 1000: Device 1005: Interface System 1010: Logic System 1015: Memory system 1025: Microphone 1030: Display System 1035: User Input System 1040: Power System 1100: System 1105: Audio and Metadata Authoring Tools 1110: Rendering Tools 1109, 1117: Network interface 1120: Interface 1150: System 1155: Cinema Server 1160: Rendering system 1157, 1162: Network interface 1164: Interface

Claims

Claim 1 A method for rendering input audio comprising an audio object and metadata, wherein the metadata comprises audio object size metadata and audio object location metadata corresponding to the audio object, the method for rendering input audio comprises: receiving the audio object size metadata and the audio object location metadata; receiving zone metadata for zone constraints for one or more speaker supplies; determining at least a virtual audio object based on the input audio, the audio object size metadata, and the audio object location metadata; determining the location of at least the virtual audio object based on at least one of the audio object size metadata and the audio object location metadata; and rendering the audio object to one or more speaker supplies based on the location of at least the virtual audio object, wherein the rendering also comprises rendering the audio object based on the zone metadata. Claim 2 A non-transient medium on which software is recorded, wherein the software comprises instructions for executing the method of claim 1. Claim 3 A device for rendering input audio comprising an audio object and metadata, wherein the metadata comprises audio object size metadata and audio object location metadata corresponding to the audio object, the device for rendering input audio comprising: a receiver receiving the audio object size metadata and the audio object location metadata, wherein the receiver is also configured to receive zone metadata for zone restrictions for one or more speaker supplies; a first processor determining at least a virtual audio object based on the input audio, the audio object size metadata, and the audio object location metadata; a second processor determining the location of at least the virtual audio object based on at least one of the audio object size metadata and the audio object location metadata; and a renderer rendering an audio object to one or more speaker supplies based on the location of at least the virtual audio object, wherein the rendering is also based on the zone metadata. Claim 4 delete Claim 5 delete Claim 6 delete Claim 7 delete Claim 8 delete Claim 9 delete Claim 10 delete Claim 11 delete Claim 12 delete Claim 13 delete