Systems and tools for improved 3D audio creation and presentation

An authoring and rendering tool generates metadata for audio objects in 3D environments, addressing the challenge of complex speaker layouts by using speaker zone constraints, enhancing sound localization and immersion in 3D audio systems.

JP7735499B2Active Publication Date: 2025-09-08DOLBY LABORATORIES LICENSING CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024130681
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-04-20
Filing Date
2024-08-07
Publication Date
2025-09-08
Estimated Expiration
2032-06-27

AI Technical Summary

Technical Problem

As channel counts increase and speaker layouts transition from planar, two-dimensional (2D) arrays to three-dimensional (3D) arrays, the task of positioning and rendering sound becomes increasingly difficult.

Method used

The implementation of an authoring and rendering tool that generates metadata about audio objects, considering speaker zones and playback environments, to render audio playback data effectively in various configurations such as Dolby Surround 5.1, 7.1, and Hamasaki 22.2, using speaker zone constraints and dynamic object mapping.

Benefits of technology

Enhances audio authoring and rendering capabilities, providing improved sound localization and immersion in 3D audio systems by accurately positioning and rendering sound in complex speaker layouts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007735499000001
    Figure 0007735499000001
  • Figure 0007735499000002
    Figure 0007735499000002
  • Figure 0007735499000003
    Figure 0007735499000003
Patent Text Reader

Abstract

To provide an improved tool for authoring and rendering of audio replay data.SOLUTION: Several authoring tools as such permit generalization of audio replay data for a wide variety of replay environments. Audio replay data can be authored by generating meta-data on an audio object. The meta-data may be generated with reference to a speaker zone. During the rendering process, the audio replay data may be replayed according to a replay speaker layout of a specific replay environment.SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 61 / 504,005, filed July 1, 2011, and U.S. Provisional Application No. 61 / 636,076, filed April 20, 2012, both of which are incorporated by reference in their entirety for all purposes.

[0002] technology This disclosure relates to authoring and rendering audio reproduction data, and in particular to authoring and rendering audio reproduction data for reproduction environments such as movie theater sound reproduction systems. [Background technology]

[0003] Since the introduction of sound to motion pictures in 1927, there has been a steady advancement in the technology used to capture the artistic intent of a film soundtrack and reproduce it in a cinema environment. In the 1930s, synchronized sound on disc gave way to variable-area sound on film, which was further improved in the 1940s by considerations of theatre acoustics and improved speaker design. Along with this came the early introduction of multitrack recording and directional playback (using control tones to move the sound). In the 1950s and 1960s, magnetic stripes on film allowed for multichannel playback in cinemas, introducing surround channels and up to five screen channels in premium theatres.

[0004] In the 1970s, Dolby introduced noise reduction, both in postproduction and on film, along with a cost-effective means of encoding and distributing three screen channels mixed with a mono surround channel. Cinema sound quality was further improved in the 1980s with Dolby Spectral Recording (SR) noise reduction and certification programs such as THX. Dolby brought digital sound to cinemas in the 1990s with the 5.1-channel format, which provided discrete left, center, and right screen channels, left and right surround arrays, and a subwoofer channel for low-frequency effects. Dolby Surround 7.1, introduced in 2010, increased the number of surround channels by dividing the existing left and right surround channels into four "zones." [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] V. Pulkki, Compensating Displacement of Amplitude-Panned Virtual Sources, Audio Engineering Society (AES) International Conference on Virtual, Synthetic and Entertainment Audio [Non-patent document 2] D. de Vries, Wave Field Synthesis, AES Monograph 1999 Summary of the Invention [Problem to be solved by the invention]

[0006] As channel counts increase and speaker layouts transition from planar, two-dimensional (2D) arrays to three-dimensional (3D) arrays that include height, the task of positioning and rendering sound becomes increasingly difficult. Improved audio authoring and rendering methods would be desirable. [Means for solving the problem]

[0007] Some aspects of the subject matter described in this disclosure can be implemented in a tool for authoring and rendering audio playback data. Some such authoring tools allow audio playback data to be generalized for a wide variety of playback environments. According to some such implementations, audio playback data is authored by generating metadata about audio objects. The metadata may be generated with reference to speaker zones. During the rendering process, the audio playback data may be played back according to the playback speaker layout of the particular playback environment.

[0008] Some implementations described herein provide an apparatus including an interface system and a logic system. The logic system may be configured to receive, via the interface system, audio playback data including one or more audio objects and associated metadata and playback environment data. The playback environment data may include an indication of a number of playback speakers in the playback environment and an indication of a position of each playback speaker within the playback environment. The logic system may be configured to render the audio objects into one or more speaker feed signals based, at least in part, on the associated metadata and the playback environment data, where each speaker feed signal corresponds to at least one of the playback speakers in the playback environment. The logic system may be configured to calculate speaker gains corresponding to the virtual speaker positions.

[0009] The playback environment may be, for example, a movie theater sound system environment. The playback environment may have a Dolby Surround 5.1 configuration, a Dolby Surround 7.1 configuration, or a Hamasaki 22.2 surround sound configuration. The playback environment data may include playback speaker layout data indicating playback speaker positions. The playback environment data may include playback speaker zone layout data indicating playback speaker areas and playback speaker positions that correspond to the playback speaker areas.

[0010] The metadata may include information for mapping an audio object position to a single playback speaker position. The rendering may involve generating an overall gain based on one or more of a desired audio object position, a distance from the desired audio object position to a reference position, an audio object velocity, or an audio object content type. The metadata may include data for constraining the audio object position to a one-dimensional curve or a two-dimensional surface. The metadata may include trajectory data for the audio object.

[0011] Rendering may involve imposing speaker zone constraints. For example, the device may include a user input system. According to some implementations, rendering may involve applying screen-to-room balance control according to screen-to-room balance control data received from the user input system.

[0012] The apparatus may include a display system, and the logic system may be configured to control the display system to display a dynamic three-dimensional view of the playback environment.

[0013] Rendering may involve controlling the extent of audio objects in one or more of three dimensions. Rendering may involve dynamic object blobbing in response to speaker overload. Rendering may involve mapping audio object positions onto the plane of a speaker array in the playback environment.

[0014] The apparatus may include one or more non-transitory storage media, such as memory devices of a memory system. The memory devices may include, for example, random access memory (RAM), read-only memory (ROM), flash memory, one or more hard drives, etc. The interface system may include an interface between the logic system and one or more such memory devices. The interface system may also include a network interface.

[0015] The metadata may include speaker zone constraint metadata. The logic system may be configured to attenuate selected speaker feed signals by performing the following operations: calculating a first gain that includes a contribution from the selected speaker; calculating a second gain that does not include a contribution from the selected speaker; and blending the first gain with the second gain. The logic system may be configured to determine whether to apply panning rules for the audio object positions or to map the audio object positions to a single speaker position. The logic system may be configured to smooth a transition in speaker gain when transitioning from mapping the audio object positions to a first single speaker position to a second single speaker position. The logic system may be configured to smooth a transition in speaker gain when transitioning between mapping the audio object positions to a single speaker position and applying panning rules for the audio object positions. The logic system may be configured to calculate speaker gains for audio object positions along a one-dimensional curve between the virtual speaker positions.

[0016] Some methods described herein involve receiving audio playback data including one or more audio objects and associated metadata, and receiving playback environment data including an indication of the number of playback speakers in the playback environment. The playback environment data may include an indication of the location of each playback speaker within the playback environment. These methods may involve rendering the audio objects into one or more speaker feed signals based, at least in part, on the associated metadata. Each speaker feed signal may correspond to at least one of the playback speakers in the playback environment. The playback environment may be a cinema sound system environment.

[0017] The rendering may involve generating an overall gain based on one or more of a desired audio object position, a distance from the desired audio object position to a reference position, an audio object velocity, or an audio object content type. The metadata may include data for constraining the audio object position to a one-dimensional curve or a two-dimensional surface. The rendering may involve imposing speaker zone constraints.

[0018] Some implementations may be embodied in one or more non-transitory media having software stored thereon. The software may include instructions for controlling one or more devices to perform the following operations: receive audio playback data including one or more audio objects and associated metadata; receive playback environment data including an indication of the number of playback speakers in the playback environment and an indication of the location of each playback speaker within the playback environment; and render the audio objects into one or more speaker feed signals based at least in part on the associated metadata. Each speaker feed signal may correspond to at least one of the playback speakers in the playback environment. The playback environment may be, for example, a movie theater sound system environment.

[0019] The rendering may involve generating an overall gain based on one or more of a desired audio object position, a distance from the desired audio object position to a reference position, an audio object velocity, or an audio object content type. The metadata may include data for constraining the audio object position to a one-dimensional curve or a two-dimensional surface. The rendering may involve imposing speaker zone constraints. The rendering may involve dynamic object blobbing in response to speaker overload.

[0020] Alternative devices and apparatuses are described herein. Some such apparatuses may include an interface system, a user input system, and a logic system. The logic system may be configured to receive audio data via the interface system, receive a position of the audio object via the user input system or the interface system, and determine a position of the audio object in three-dimensional space. The determination may involve constraining the position to a one-dimensional curve or a two-dimensional surface within three-dimensional space. The logic system may be configured to generate metadata associated with the audio object based, at least in part, on user input received via the user input system. The metadata includes data indicating the position of the audio object in three-dimensional space.

[0021] The metadata may include trajectory data indicating a time-varying position of the audio object in three-dimensional space. The logic system may be configured to calculate the trajectory data according to user input received via the user input system. The trajectory data may include a set of positions in three-dimensional space at multiple points in time. The trajectory data may include an initial position, velocity data, and acceleration data. The trajectory data may include an equation defining the initial position and positions in three-dimensional space and corresponding times.

[0022] The apparatus may include a display system, and the logic system may be configured to control the display system to display the audio object trajectory according to the trajectory data.

[0023] The logic system may be configured to generate speaker zone constraint metadata according to user input received via the user input system. The speaker zone constraint metadata may include data for disabling selected speakers. The logic system may be configured to generate the speaker zone constraint metadata by mapping audio object positions to a single speaker.

[0024] The apparatus may include a sound reproduction system, and the logic system may be configured to control the sound reproduction system at least in part in accordance with said metadata.

[0025] The position of the audio object may be constrained to a one-dimensional curve, and the logic system may be further configured to generate virtual speaker positions along the one-dimensional curve.

[0026] Alternative methods are described herein. Some such methods involve receiving audio data, receiving a position of an audio object, and determining the position of the audio object in three-dimensional space. The determination may involve constraining the position to a one-dimensional curve or a two-dimensional surface within the three-dimensional space. These methods may also involve generating metadata associated with the audio object based, at least in part, on user input.

[0027] The metadata may include data indicating a position of an audio object in three-dimensional space. The metadata may include trajectory data indicating a time-varying position of an audio object in three-dimensional space. Generating the metadata may involve generating speaker zone constraint metadata, for example according to user input. The speaker zone constraint metadata may include data for disabling selected speakers.

[0028] The positions of the audio objects may be constrained to a one-dimensional curve, and the methods may involve generating virtual speaker positions along that one-dimensional curve.

[0029] Other aspects of the present disclosure may be embodied in one or more non-transitory media having software stored thereon. The software may include instructions for controlling one or more devices to perform the following operations: receive audio data; receive a position of an audio object; and determine a position of the audio object in three-dimensional space. The determination may involve constraining the position to a one-dimensional curve or a two-dimensional surface within the three-dimensional space. The software may include instructions for controlling one or more devices to generate metadata associated with the audio object. The metadata may be generated at least in part based on user input.

[0030] The metadata may include data indicating a position of an audio object in three-dimensional space. The metadata may include trajectory data indicating a time-varying position of an audio object in three-dimensional space. Generating the metadata may involve generating speaker zone constraint metadata, for example according to user input. The speaker zone constraint metadata may include data for disabling selected speakers.

[0031] The positions of the audio objects may be constrained to a one-dimensional curve, and the software may include instructions for controlling one or more devices to generate virtual speaker positions along the one-dimensional curve.

[0032] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. It is noted that the relative dimensions of the following drawings may not be drawn to scale. [Brief explanation of the drawings]

[0033] [Figure 1] This figure shows an example of a playback environment with Dolby Surround 5.1 configuration. [Figure 2] This figure shows an example of a playback environment with Dolby Surround 7.1 configuration. [Figure 3] This figure shows an example of a playback environment with a Hamasaki 22.2 surround sound configuration. [Figure 4A] FIG. 1 shows an example of a graphical user interface (GUI) for depicting speaker zones at various heights in a virtual playback environment. [Figure 4B] FIG. 10 is a diagram illustrating another example of a playback environment. [Figure 5A] Figure 1 shows an example of a loudspeaker response corresponding to an audio object whose position is constrained to a two-dimensional plane in three-dimensional space. [Figure 5B] Figure 1 shows an example of a loudspeaker response corresponding to an audio object whose position is constrained to a two-dimensional plane in three-dimensional space. [Figure 5C] Figure 1 shows an example of a speaker response corresponding to an audio object whose position is constrained to a two-dimensional plane in three-dimensional space. [Figure 5D] FIG. 10 illustrates an example of a two-dimensional surface to which an audio object can be constrained. [Figure 5E] FIG. 10 illustrates an example of a two-dimensional surface to which an audio object can be constrained. [Figure 6A] 1 is a flow diagram outlining an example of a process for constraining the position of an audio object to a two-dimensional plane. [Figure 6B] 1 is a flow diagram outlining an example of a process for mapping audio object positions to a single speaker position or a single speaker zone. [Figure 7] 1 is a flow diagram outlining a process for establishing and using a virtual speaker. [Figure 8]1A to 1C are diagrams showing examples of virtual speakers and corresponding speaker responses mapped to line endpoints. [Figure 9] Figures A-C show an example of using a virtual tether to move an audio object. [Figure 10A] Flowchart outlining the process of using a virtual tether to move an audio object. [Figure 10B] 1 is a flow diagram outlining an alternative process for using a virtual tether to move an audio object. [Figure 10C] FIG. 10C illustrates an example of the process outlined in FIG. 10B. [Figure 10D] FIG. 10C illustrates an example of the process outlined in FIG. 10B. [Figure 10E] FIG. 10C illustrates an example of the process outlined in FIG. 10B. [Figure 11] FIG. 10 illustrates an example of applying speaker zone constraints in a virtual playback environment. [Figure 12] FIG. 1 is a flow diagram outlining some examples of applying speaker zone constraints. [Figure 13A] FIG. 10 illustrates an example of a GUI that can switch between two-dimensional and three-dimensional views of a virtual playback environment. [Figure 13B] FIG. 10 illustrates an example of a GUI that can switch between two-dimensional and three-dimensional views of a virtual playback environment. [Figure 13C] FIG. 1 illustrates a combined two-dimensional and three-dimensional rendering of the playback environment. [Figure 13D] FIG. 1 illustrates a combined two-dimensional and three-dimensional rendering of the playback environment. [Figure 13E] FIG. 1 illustrates a combined two-dimensional and three-dimensional rendering of the playback environment. [Figure 14A] 13C-13E is a flow diagram outlining a process for controlling a device to present a GUI such as that shown in FIGS. 13C-13E. [Figure 14B]1 is a flow diagram outlining the process of rendering audio objects for a playback environment. [Figure 15] FIG. 1A shows an example of an audio object and associated audio object width in a virtual playback environment; FIG. 1B shows an example of a spread profile corresponding to the audio object width shown in FIG. [Figure 16] 1 is a flow diagram outlining the process of blobbing an audio object. [Figure 17] A and B show examples of audio objects positioned in a three-dimensional virtual playback environment. [Figure 18] FIG. 10 shows examples of zones corresponding to various pan modes. [Figure 19] 1A-D show examples of applying near-field and far-field panning techniques to audio objects at various positions. [Figure 20] FIG. 10 illustrates speaker zones in a playback environment that can be used in the screen-to-room bias control process. [Figure 21] FIG. 2 is a block diagram providing examples of components of an authoring and / or rendering device. [Figure 22] FIG. 1A is a block diagram illustrating some components that may be used for audio content generation, and FIG. 1B is a block diagram illustrating some components that may be used for audio playback in a playback environment. Reference numbers and characters in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION

[0034] The following description is directed to certain implementations for purposes of describing some novel aspects of the present disclosure and examples of contexts in which these novel aspects may be implemented. However, the teachings herein may be applied in a variety of different ways. For example, while various implementations are described using specific playback environments, the teachings herein are broadly applicable to other known and future playback environments. Similarly, while example graphical user interfaces (GUIs) are presented herein, some of which provide examples of speaker locations, speaker zones, etc., other implementations are contemplated by the inventors. Furthermore, the described implementations may be implemented in a variety of authoring and / or rendering tools, which may be implemented in a variety of hardware, software, firmware, etc. Accordingly, the teachings of the present disclosure are not intended to be limited to the implementations shown in the drawings and / or described herein, but rather have broad applicability.

[0035] FIG. 1 shows an example of a playback environment with a Dolby Surround 5.1 configuration. Although Dolby Surround 5.1 was developed in the 1990s, this configuration is still widely deployed in movie theater sound system environments. A projector 105 may be configured to project a video image, such as for a movie, onto a screen 150. Audio playback data may be synchronized with the video image and processed by a sound processor 110. A power amplifier 115 may provide speaker feed signals to speakers in the playback environment 100.

[0036] The Dolby Surround 5.1 configuration includes a left surround array 120 and a right surround array 125, each driven collectively by a single channel. The Dolby Surround 5.1 configuration also includes separate channels for the left screen channel 130, the center screen channel 135, and the right screen channel 140. A separate channel for the subwoofer 145 is provided for low-frequency effects (LFE).

[0037] In 2010, Dolby provided an improvement to digital cinema sound by introducing Dolby Surround 7.1. FIG. 2 shows an example of a playback environment with a Dolby Surround 7.1 configuration. Digital projector 205 may be configured to receive digital video data and project a video image onto screen 150. Audio playback data may be processed by sound processor 210. Power amplifier 215 may provide speaker feed signals to speakers in playback environment 200.

[0038] The Dolby Surround 7.1 configuration includes a left lateral surround array 220 and a right lateral surround array 225, each of which may be driven by a single channel. Like Dolby Surround 5.1, the Dolby Surround 7.1 configuration also includes separate channels for left screen channel 230, center screen channel 235, right screen channel 240, and subwoofer 245. However, Dolby Surround 7.1 increases the number of surround channels by dividing the left and right surround channels of Dolby Surround 5.1 into four zones: in addition to left lateral surround array 220 and right lateral surround array 225, separate channels are included for left rear surround speaker 224 and right rear surround speaker 226. Increasing the number of surround zones in playback environment 200 can significantly improve sound localization.

[0039] In an effort to create a more immersive environment, some playback environments may be configured with an increased number of speakers driven by an increased number of channels. Additionally, some playback environments may include speakers that are deployed at various heights, some of which may be above the seating area of ​​the playback environment.

[0040] Figure 3 shows an example of a playback environment with a Hamasaki 22.2 surround sound configuration. Hamasaki 22.2 was developed at the NHK Science and Technology Research Laboratories in Japan as a surround sound component for ultra-high definition television. Hamasaki 22.2 provides 24 speaker channels that can be used to drive speakers arranged in three tiers. In playback environment 300, upper speaker tier 310 can be driven by nine channels. Middle speaker tier 320 can be driven by ten channels. Lower speaker tier 330 can be driven by five channels, two of which are for subwoofers 345a and 345b.

[0041] Thus, the current trend is not only to include more speakers and more channels, but also to include speakers at different heights. As the number of channels increases and speaker layouts transition from 2D to 3D arrays, the task of positioning and rendering sound becomes increasingly difficult.

[0042] This disclosure provides various tools and related user interfaces that enhance functionality and / or reduce authoring complexity for 3D audio sound systems.

[0043] 4A shows an example of a graphical user interface (GUI) depicting speaker zones at various heights in a virtual playback environment. GUI 400 may be displayed on a display device, for example, according to instructions from a logic system, signals received from a user input device, etc. Some such devices are described below with reference to FIG. 21.

[0044] As used herein to refer to a virtual playback environment such as virtual playback environment 404, the term "speaker zone" generally refers to a logical construct that may or may not have a one-to-one correspondence with a playback speaker in a real playback environment. For example, a "speaker zone location" may or may not correspond to a specific playback speaker location in a movie theater playback environment. Instead, the term "speaker zone location" may generally refer to a zone in the virtual playback environment. In some implementations, speaker zones in a virtual playback environment may correspond to virtual speakers through the use of virtualization technology, such as Dolby Headphone™ (sometimes referred to as Mobile Surround™), which generates a virtual surround sound environment in real time using a set of two-channel stereo headphones. GUI 400 has seven speaker zones 402a at a first elevation and two speaker zones 402b at a second elevation, for a total of nine speaker zones in virtual playback environment 404. In this example, speaker zones 1 through 3 are located in the front region 405 of virtual playback environment 404. The front area 405 may correspond, for example, to the area in a cinema playback environment where the screen 150 is located, or the area in a home where a television screen is located, etc.

[0045] Here, speaker zone 4 generally corresponds to speakers in left region 410, and speaker zone 5 corresponds to speakers in right region 415 of virtual playback environment 404. Speaker zone 6 corresponds to left rear region 412, and speaker zone 7 corresponds to right rear region 414 of virtual playback environment 404. Speaker zone 8 corresponds to speakers in upper region 420a, and speaker zone 9 corresponds to speakers in upper region 420b, which may be a virtual ceiling region such as the region of virtual ceiling 520 shown in FIGS. 5D and 5E. Thus, as described in more detail below, the locations of speaker zones 1-9 shown in FIG. 4A may or may not correspond to the locations of playback speakers in a real playback environment. Furthermore, other implementations may include more or fewer speaker zones and / or heights.

[0046] In various implementations described herein, a user interface such as GUI 400 may be used as part of an authoring tool and / or a rendering tool. In some implementations, the authoring tool and / or rendering tool may be implemented via software stored on one or more non-transitory media. The authoring tool and / or rendering tool may be implemented (at least in part) via hardware, firmware, etc., such as the logic system and other devices described below with reference to FIG. 21. In some authoring implementations, the associated authoring tool may be used to generate metadata about the associated audio data. The metadata may include, for example, data indicating the position and / or trajectory of audio objects in three-dimensional space, speaker zone constraint data, etc. The metadata may be generated regarding speaker zones 402 of the virtual playback environment 404 rather than regarding a particular speaker layout of the actual playback environment. The rendering tool may receive the audio data and associated metadata and calculate audio gain and speaker feed signals for the playback environment. Such audio gain and speaker feed signals may be calculated according to an amplitude panning process. The amplitude panning process can create the perception that a sound is coming from a position P in the playback environment. For example, a speaker feed signal can be expressed as x i (t)=g i x(t) i=1,…,N (Equation 1) may be provided to playback speakers 1 to N of the playback environment according to

[0047] In formula (1), x i (t) represents the speaker feed signal applied to speaker i, and g irepresents the gain factor of the corresponding channel, x(t) represents the audio signal, and t represents time. The gain factor may be determined, for example, according to the amplitude panning methods described in Section 2, pp. 3-4, of "Analog Audio Signal Processing," incorporated herein by reference. In some implementations, the gain may be frequency dependent. In some implementations, a time delay may be introduced by replacing x(t) with x(t-Δt).

[0048] In some rendering implementations, audio playback data generated with reference to speaker zones 402 may be mapped to speaker locations in a wide range of playback environments, which may have a Dolby Surround 5.1 configuration, a Dolby Surround 7.1 configuration, a Hamasaki 22.2 configuration, or other configurations. For example, referring to FIG. 2, a rendering tool may map audio playback data for speaker zones 4 and 5 to left lateral surround array 220 and right lateral surround array 225 in a playback environment with a Dolby Surround 7.1 configuration. Audio playback data for speaker zones 1, 2, and 3 may be mapped to left screen channel 230, right screen channel 240, and center screen channel 235, respectively. Audio playback data for speaker zones 6 and 7 may be mapped to left rear surround speaker 224 and right rear surround speaker 226.

[0049] 4B shows another example playback environment. In some implementations, a rendering tool may map audio playback data for speaker zones 1, 2, and 3 to corresponding screen speakers 455 in playback environment 450. The rendering tool may map audio playback data for speaker zones 4 and 5 to left lateral surround array 460 and right lateral surround array 465, and may map audio playback data for speaker zones 8 and 9 to left overhead speaker 470a and right overhead speaker 470b. Audio playback data for speaker zones 6 and 7 may be mapped to left rear surround speaker 480a and right rear surround speaker 480b.

[0050] In some authoring implementations, the authoring tool may be used to generate metadata about audio objects. As used herein, the term "audio object" refers to a stream of audio data and associated metadata. The metadata typically indicates the object's 3D position, rendering constraints, and content type (e.g., dialogue, effects, etc.). Depending on the implementation, the metadata may include other types of data, such as width data, gain data, and trajectory data. Some audio objects may be static, while others may move. Details of an audio object may be authored or rendered according to associated metadata, which may indicate, for example, the audio object's location in three-dimensional space at a given time. When an audio object is monitored or played in a playback environment, it may be rendered according to its position metadata using the playback speakers present in the playback environment, rather than being output to a specific physical channel as in traditional channel-based systems such as Dolby 5.1 or Dolby 7.1.

[0051] Although various authoring and rendering tools are described herein with reference to GUIs that are substantially similar to GUI 400, various other interfaces, including but not limited to GUIs, may be used in conjunction with these authoring and rendering tools. Some such tools can simplify the authoring process by applying various types of constraints. Some implementations are now described with reference to FIG. 5A et seq.

[0052] Figures 5A-5C show examples of speaker responses corresponding to an audio object whose position is constrained to a two-dimensional plane in three-dimensional space. The two-dimensional plane is a hemisphere in this example. In these examples, the speaker responses are calculated by the renderer assuming a nine-speaker configuration, with each speaker corresponding to one of speaker zones 1-9. However, as noted elsewhere in this document, in general, there may not be a one-to-one mapping between speaker zones in the virtual playback environment and playback speakers in the playback environment. Referring first to Figure 5A, audio object 505 is shown in a position at the front left of virtual playback environment 404. Thus, the speaker corresponding to speaker zone 1 exhibits substantial gain, while the speakers corresponding to speaker zones 3 and 4 exhibit moderate gain.

[0053] In this example, the position of the audio object 505 is changed by placing cursor 510 on the audio object 505 and "dragging" the audio object 505 to a desired location within the x-y plane of the virtual playback environment 404. As the object is dragged toward the center of the playback environment, it also maps to the surface of a hemisphere, increasing in height. Here, the increasing height of the audio object 505 is indicated by an increasing diameter of the circle representing the audio object 505. That is, as shown in FIGS. 5B and 5C, the audio object 505 appears larger as it is dragged toward the top center of the virtual playback environment 404. Alternatively or additionally, the height of the audio object 505 may be indicated by a change in color, brightness, a numerical height indication, etc. When the audio object 505 is positioned at the top center of the virtual playback environment 404, as shown in FIG. 5C, the speakers corresponding to speaker zones 8 and 9 exhibit substantial gain, while the other speakers exhibit little or no gain.

[0054] In this implementation, the position of the audio object 505 is constrained to a two-dimensional surface, such as a sphere, ellipsoid, cone, cylinder, wedge, etc. Figures 5D and 5E show examples of two-dimensional surfaces to which the audio object may be constrained. Figures 5D and 5E are cross-sections through the virtual playback environment 404, with the front region 405 shown on the left. In Figures 5D and 5E, the y value of the yz axis increases toward the front region 405 of the virtual playback environment 404 to maintain consistency with the orientation of the x and y axes shown in Figures 5A-5C.

[0055] In the example shown in FIG. 5D, two-dimensional surface 515a is a section of an ellipsoid. In the example shown in FIG. 5E, two-dimensional surface 515b is a section of a wedge. However, the shape, orientation, and position of two-dimensional surface 515 shown in FIGS. 5D and 5E are merely examples. In alternative implementations, at least a portion of two-dimensional surface 515 may extend outside of virtual playback environment 404. In some such implementations, two-dimensional surface 515 may extend above virtual ceiling 520. Thus, the three-dimensional space into which two-dimensional surface 515 extends is not necessarily coextensive with the volume of virtual playback environment 404. In still other implementations, audio objects may be constrained to one-dimensional features such as curves, straight lines, etc.

[0056] FIG. 6A is a flow diagram outlining an example process for constraining the position of an audio object to a two-dimensional plane. As with other flow diagrams provided herein, the operations of process 600 are not necessarily performed in the order shown. Furthermore, process 600 (and other processes provided herein) may include more or fewer operations than those shown and / or described. In this example, blocks 605 through 622 are performed by an authoring tool, and blocks 624 through 630 are performed by a rendering tool. The authoring tool and rendering tool may be implemented on a single device or on two or more devices. While FIG. 6A (and other flow diagrams provided herein) may give the impression that the authoring process and the rendering process are performed sequentially, in many implementations, the authoring process and the rendering process are performed substantially simultaneously. The authoring process and the rendering process may be interactive. For example, the results of the authoring process may be sent to a rendering tool, the corresponding results of the rendering tool may be evaluated by a user, the user may perform further authoring based on these results, etc.

[0057] At block 605, an indication that an audio object position should be constrained to a two-dimensional plane is received. This indication may be received, for example, by a logic system of a device configured to provide an authoring and / or rendering tool. As with other implementations described herein, the logic system may operate according to software instructions, firmware, etc. stored on a non-transitory medium. The indication may be a signal from a user input device (e.g., a touchscreen, a mouse, a trackball, a gesture recognizer, etc.) in response to input from a user.

[0058] At optional block 607, audio data is received. Block 607 is optional in this example because audio data may go to the renderer directly from another source (e.g., a mixing console) that is time-synchronized to the metadata authoring tool. In some such implementations, there may be an implicit mechanism that links each audio stream to a corresponding incoming metadata stream to form an audio object. For example, a metadata stream may contain an identifier for the audio object it represents, e.g., a number from 1 to N. If the rendering device is configured with audio inputs also numbered 1 to N, the rendering tool may automatically assume that an audio object is formed by the metadata stream identified with a certain number (e.g., 1) and the audio data received on the first audio input. Similarly, any metadata stream identified as number 2 may form an object with the audio received on the second audio input channel. In some implementations, audio and metadata may be pre-packaged by an authoring tool to form an audio object, which may be provided to a rendering tool, for example, sent over a network as a TCP / IP packet.

[0059] In an alternative implementation, the authoring tool may only send metadata over a network, and the rendering tool may receive audio from another source (e.g., via a pulse-code modulation (PCM) stream, via analog audio, etc.). In such an implementation, the rendering tool may be configured to group the audio data and metadata to form an audio object. The audio data may be received by the logical system, for example, via an interface. The interface may be, for example, a network interface, an audio interface (e.g., an interface configured for communication via the AES3 standard developed by the Audio Engineering Society and the European Broadcasting Union, also known as AES / EBU, via the Multichannel Audio Digital Interface (MADI) protocol, via an analog signal, etc.), or an interface between the logical system and a memory device. In this example, the data received by the renderer includes at least one audio object.

[0060] In block 610, the (x,y) or (x,y,z) coordinates of the audio object position are received. Block 610 may involve receiving an initial position of the audio object, for example, as described above with reference to FIGS. 5A-5C. Block 610 may also involve receiving an indication that the user has positioned or repositioned the audio object. The audio object coordinates are mapped to a two-dimensional plane in block 615. The two-dimensional plane may be similar to that described above with reference to FIGS. 5D and 5E, or it may be a different two-dimensional plane. In this example, each point in the x-y plane is mapped to a single z value. Thus, block 615 involves mapping the x and y coordinates received in block 610 to z values. In other implementations, different mapping processes and / or coordinate systems may be used. The audio object may be displayed at the (x,y,z) position determined in block 615 (block 620). The audio data and metadata, including the mapped (x, y, z) positions determined in block 615, may be stored in block 621. The audio data and metadata may be sent to a rendering tool (block 622). In some implementations, the metadata may be sent continuously while some authoring processes are being performed, such as while audio objects are being positioned, constrained, displayed in GUI 400, etc.

[0061] In block 623, it is determined whether the authoring process continues. For example, if input is received from the user interface indicating that the user no longer wishes to constrain the audio object positions to a two-dimensional plane, the authoring process may end (block 625). Otherwise, the authoring process may continue, for example, by returning to block 607 or block 610. In some implementations, the rendering process may continue regardless of whether the authoring process continues. In some implementations, the audio objects may be recorded to disk on the authoring platform and then played back for exhibition purposes from a dedicated sound processor or a cinema server connected to a sound processor, such as sound processor 210 of FIG. 2.

[0062] In some implementations, the rendering tool may be software running on a device configured to provide authoring functionality. In other implementations, the rendering tool may be provided on a separate device. The type of communication protocol used for communication between the authoring tool and the rendering tool may vary depending on whether both tools are running on the same device or communicating over a network.

[0063] At block 626, the audio data and metadata (including the (x,y,z) location determined at block 615) are received by the rendering tool. In an alternative implementation, the audio data and metadata may be received separately by the rendering tool and interpreted as audio objects through implicit mechanisms. As noted above, for example, the metadata stream may contain audio object identification codes (e.g., 1, 2, 3, etc.) that may be attached to first, second, and third audio inputs (i.e., digital or analog audio connections) on the rendering system, respectively, to form audio objects that can be rendered to speakers.

[0064] During the rendering process of process 600 (and other rendering processes described herein), panning gain equations may be applied according to the playback speaker layout of a particular playback environment. Thus, the logic system of the rendering tool may receive playback environment data, including an indication of the number of playback speakers in the playback environment and an indication of the position of each playback speaker within the playback environment. This data may be received, for example, by accessing a data structure stored in memory accessible by the logic system or via an interface system.

[0065] In this example, a pan gain equation is applied for the (x, y, z) location to determine (block 628) a gain value to apply (block 630) to the audio data. In some implementations, the audio data, adjusted in level in response to the gain value, may be played back by a playback speaker, such as a headphone speaker (or other speaker) configured to communicate with the logic system of the rendering tool. In some implementations, the playback speaker location may correspond to a speaker zone of a virtual playback environment, such as virtual playback environment 404 described above. The corresponding speaker response may be displayed on a display device, such as those shown in FIGS. 5A-5C.

[0066] At block 635, it is determined whether the process continues. For example, the process may end when input is received from the user interface indicating that the user no longer wishes to continue the rendering process (block 640). Otherwise, the process may continue, for example, by returning to block 626. If the logic system receives an indication that the user wishes to return to the corresponding authoring process, process 600 may return to block 607 or block 610.

[0067] Other implementations may involve imposing various other types of constraints or generating other types of constraint metadata for the audio objects. FIG. 6B is a flow diagram outlining an example process for mapping audio object positions to single speaker positions. This process is sometimes referred to herein as “snapping.” At block 655, an indication that an audio object position may be snapped to a single speaker position or single speaker zone is received. In this example, the indication is that the audio object position is snapped to a single speaker position, as appropriate. The indication may be received by a logic system of a device configured to provide the authoring tool. The indication may correspond to input received from a user input device. However, the indication may also correspond to a category of the audio object (e.g., bullets, speech) and / or a width of the audio object. Information regarding the category and / or width may be received, for example, as metadata about the audio object. In such an implementation, block 657 may occur before block 655.

[0068] In block 656, audio data is received. Coordinates of the audio object position are received in block 657. In this example, the audio object position is displayed according to the coordinates received in block 657 (block 658). Metadata including the audio object coordinates and a snap flag indicating a snap function is saved in block 659. The audio data and metadata are sent by the authoring tool to the rendering tool (block 660).

[0069] In block 662, it is determined whether the authoring process continues. For example, if input is received from the user interface indicating that the user no longer wants audio object positions to snap to speaker positions, the authoring process may end (block 663). Otherwise, the authoring process may continue, for example, by returning to block 665. In some implementations, the rendering process may continue regardless of whether the authoring process continues.

[0070] The audio data and metadata sent by the authoring tool are received by the rendering tool in block 664. A determination is made (e.g., by a logic system) whether to snap the audio object position to a speaker position in block 665. This determination may be based, at least in part, on the distance between the audio object position and the nearest playback speaker position in the playback environment.

[0071] In this example, if it is determined in block 665 that the audio object position should be snapped to a speaker position, then in block 670 the audio object position is mapped to a speaker position, typically the speaker position that is closest to the intended (x,y,z) position received for the audio object. In this case, the gain for audio data played by this speaker position is 1.0, while the gain for audio data played by other speakers is zero. In an alternative implementation, the audio object position may be mapped to a group of speaker positions in block 670.

[0072] For example, referring again to Figure 4B, block 670 may involve snapping the position of an audio object to one of the left overhead speakers 470a. Alternatively, block 670 may involve snapping the position of an audio object to a single speaker and a neighboring speaker, e.g., one or two neighboring speakers. Thus, corresponding metadata may be applied to a small group of playback speakers and / or to individual playback speakers.

[0073] However, if it is determined in block 665 that the audio object position cannot be snapped to the speaker position, for example, if doing so would result in a large position discrepancy compared to the original intended position received for the object, then panning rules are applied (block 675). Panning rules may be applied according to the audio object position and other characteristics of the audio object (width, volume, etc.).

[0074] The determined gain data from block 675 may be applied to the audio data in block 681, and the results may be saved. In some implementations, the resulting audio data may be played through speakers configured for communication with the logic system. If it is determined in block 685 that process 650 continues, process 650 may return to block 664 to continue the rendering process. Alternatively, process 650 may return to block 655 to resume the authoring process.

[0075] Process 650 may involve various types of smoothing operations. For example, the logic system may be configured to smooth the transition in gain applied to audio data when transitioning the mapping of an audio object position from a first single speaker position to a second single speaker position. Referring again to FIG. 4B, if an audio object position is initially mapped to one of the left overhead speakers 470a but later mapped to one of the right rear surround speakers 480b, the logic system may smooth the transition between speakers so that the audio object does not appear to suddenly "jump" from one speaker (or speaker zone) to another. In some implementations, this smoothing may be implemented according to a crossfade rate parameter.

[0076] In some implementations, the logic system may be configured to smooth the transition in gain applied to the audio data when transitioning between mapping an audio object position to a single speaker position and applying panning rules for the audio object position. For example, if it is subsequently determined in block 665 that the audio object position has been moved to a position that is determined to be too far from the nearest speaker, panning rules for the audio object position may be applied in block 675. However, when transitioning from snapping to panning (or vice versa), the logic system may be configured to smooth the transition in gain applied to the audio data. The process may end at block 690, for example, upon receipt of a corresponding input from a user interface.

[0077] Some alternative implementations may involve creating logical constraints. In some cases, for example, a sound mixer may desire more explicit control over the set of speakers used during a particular panning operation. Some implementations allow the user to create a one- or two-dimensional "logical mapping" between a set of speakers and the panning interface.

[0078] FIG. 7 is a flow diagram outlining a process for establishing and using virtual speakers. FIGS. 8A-8C show examples of virtual speakers and corresponding speaker zone responses mapped to line endpoints. Referring first to process 700 of FIG. 7, at block 705, instructions to generate virtual speakers are received. The instructions may, for example, be received by a logic system of an authoring device or may correspond to input received from a user input device.

[0079] At block 710, an indication of a virtual speaker position is received. For example, referring to FIG. 8A, a user may use an input device to position cursor 510 at the location of virtual speaker 805a and select that location, e.g., via a mouse click. At block 715, it is determined (e.g., according to user input) that an additional virtual speaker is to be selected, in this example. The process returns to block 710, and the user selects the location of virtual speaker 805b, shown in FIG. 8A.

[0080] In this example, the user only desires to establish two virtual speaker positions. Thus, in block 715, it is determined (e.g., according to user input) that no additional virtual speakers are selected. As shown in FIG. 8A, a polyline 810 connecting the positions of virtual speakers 805a and 805b may be displayed. In some implementations, the position of audio object 505 is constrained to polyline 810. In some implementations, the position of audio object 505 may be constrained to a parametric curve. For example, a set of control points may be provided according to user input, or a curve-fitting algorithm such as a spline may be used to determine the parametric curve. In block 725, an indication of the audio object position along polyline 810 is received. In some such implementations, the position is represented as a scalar value between 0 and 1. In block 725, the (x, y, z) coordinates of the audio object and the polyline defined by the virtual speakers may be displayed. The audio data and associated metadata, including the resulting scalar position and (x, y, z) coordinates of the virtual speaker, may be displayed (block 727), where the audio data and metadata may be sent to a rendering tool in block 728 via an appropriate communication protocol.

[0081] At block 729, it is determined whether the authoring process continues. If not, process 700 may end (block 730) or may continue with the rendering process, depending on user input. However, as noted above, in many implementations, at least some of the rendering process may be performed in parallel with the authoring process.

[0082] In block 732, audio data and metadata are received by the rendering tool. In block 735, gains to be applied to the audio data are calculated for each virtual speaker position. FIG. 8B shows the speaker response for the position of virtual speaker 805a. FIG. 8C shows the speaker response for the position of virtual speaker 805b. In this example, as with many other examples described herein, the speaker responses shown are for playback speakers whose positions correspond to the positions shown for the speaker zones in GUI 400. Here, virtual speakers 805a and 805b and line 810 are located in a plane that is not close to the playback speakers whose positions correspond to speaker zones 8 and 9. Therefore, gains for these speakers are not shown in FIGS. 8B and 8C.

[0083] As the user moves the audio object 505 to other positions along the line 810, the logic system calculates crossfades corresponding to these positions, for example, according to the audio object scalar position parameters (block 740). In some implementations, a pair-wise panning law (e.g., an energy-conserving sine or power law) may be used to blend between the gain applied to the audio data for the position of virtual speaker 805a and the gain applied to the audio data for the position of virtual speaker 805b.

[0084] At block 742, a determination may be made (e.g., according to user input) whether to continue process 700. The user may be presented (e.g., via a GUI) with options, for example, to continue the rendering process or return to the authoring process. If it is determined that process 700 will not continue, the process terminates (block 745).

[0085] When panning a fast-moving audio object (e.g., an audio object corresponding to a car, jet, etc.), it can be difficult to author a smooth trajectory if the audio object position is selected by the user one point at a time. Lack of smoothness in the audio object trajectory can affect the perceived sound image. Therefore, some authoring implementations provided in this document apply a low-pass filter to the audio object position to smooth the resulting panning gain. Alternative authoring implementations apply a low-pass filter to the gain applied to the audio data.

[0086] Other authoring implementations may allow the user to simulate grabbing, pulling, throwing, or similarly interacting with audio objects. Some such implementations may involve the application of simulated physics laws, such as rule sets used to describe velocity, acceleration, momentum, kinetic energy, application of force, etc.

[0087] 9A-9C show an example of using a virtual tether to drag an audio object. In FIG. 9A, a virtual tether 905 is formed between the audio object 505 and the cursor 510. In this example, the virtual tether 905 has a virtual spring constant. In some such implementations, the virtual spring constant may be selectable according to user input.

[0088] FIG. 9B shows audio object 505 and cursor 510 at a later point in time. The user then moves cursor 510 toward speaker zone 3. The user may move cursor 510 using a mouse, joystick, trackball, gesture detector, or other type of user input device. Virtual string 905 has been stretched, moving audio object 505 closer to speaker zone 8. Audio object 505 is approximately the same size in FIGS. 9A and 9B, indicating that the height of audio object 505 (in this example) has not changed substantially.

[0089] FIG. 9C shows the audio object 505 and cursor 510 at a later point in time. The user then moves the cursor around the speaker zone 9. The virtual string 905 has been stretched further. The audio object 505 has been moved downward, indicated by a decrease in the size of the audio object 505. The audio object 505 has been moved in a smooth arc. This example illustrates one potential benefit of such an implementation: the audio object 505 can be moved in a smoother trajectory than if the user simply selected positions for the audio object 505 point by point.

[0090] FIG. 10A is a flow diagram outlining a process for using a virtual string to move an audio object. Process 1000 begins at block 1005, where audio data is received. At block 1007, an instruction to attach a virtual string between the audio object and a cursor is received. This instruction may be received by the logic system of the authoring device or may correspond to input received from a user input device. Referring to FIG. 9A, a user may position cursor 510 over audio object 505 and then indicate, via a user input device or GUI, that a virtual string 905 should be formed between cursor 510 and audio object 505. Cursor and object position data may be received (block 1010). In this example, as the cursor 510 is moved, cursor velocity and / or acceleration data may be calculated by the logic system according to the cursor position data (block 1015). Position and / or trajectory data for the audio object 505 may be calculated according to the virtual spring constant of the virtual string 905 and the cursor position, velocity, and acceleration data. Some such implementations may involve assigning a virtual mass to the audio object 505 (block 1020). For example, if the cursor 510 is moved at a relatively constant velocity, the virtual string 905 may not stretch, and the audio object 505 may be pulled at a relatively constant velocity. If the cursor 510 accelerates, the virtual string 905 may stretch, and a corresponding force may be applied to the audio object 505 by the virtual string 905. There may be a time delay between the acceleration of the cursor 510 and the force applied by the virtual string 905. In alternative implementations, the position and / or trajectory of the audio object 505 may be determined differently, for example, by applying friction and / or inertia rules to the audio object 505 without assigning a virtual spring constant to the virtual string 905.

[0091] Discrete positions and / or trajectories of the audio object 505 and cursor 510 may be displayed (block 1025). In this example, the logic system samples the audio object position at time intervals (block 1030). In some such implementations, the user may determine the time interval for sampling. Audio object position and / or trajectory metadata, etc. may be saved (block 1034).

[0092] At block 1036, it is determined whether the authoring mode continues. If the user so desires, the process may continue, for example, by returning to block 1005 or block 1010. If not, process 1000 may end (block 1040).

[0093] FIG. 10B is a flow diagram outlining an alternative process for using a virtual string to move an audio object. FIGS. 10C-10E show examples of the process outlined in FIG. 10B. Referring first to FIG. 10B, process 1050 begins at block 1055, where audio data is received. At block 1057, an instruction to attach a virtual string between the audio object and a cursor is received. This instruction may be received by a logic system of the authoring device or may correspond to input received from a user input device. Referring to FIG. 10C, for example, a user may position cursor 510 over audio object 505 and then indicate, via a user input device or GUI, that a virtual string 905 should be formed between cursor 510 and audio object 505.

[0094] In block 1060, cursor and object position data may be received. In block 1062, the logic system may receive an indication (e.g., via a user input device or GUI) that audio object 505 should be held at a designated position, e.g., the position designated by cursor 510. In block 1065, the logic system may receive an indication that cursor 510 has been moved to a new position, which may be displayed along with the position of audio object 505 (block 1067). Referring to FIG. 10D, for example, cursor 510 has moved from the left side to the right side of virtual playback environment 404. However, audio object 510 is still held in the same position shown in FIG. 10C. As a result, virtual string 905 has effectively stretched.

[0095] In block 1069, the logic system receives an indication (e.g., via a user input device or GUI) that the audio object 505 should be released. The logic system may calculate the resulting audio object position and / or trajectory data, which may be displayed (block 1075). The resulting display may be similar to that shown in FIG. 10E, which shows the audio object 505 moving smoothly and rapidly across the virtual playback environment 404. The logic system may save the audio object position and / or trajectory metadata in a memory system (block 1080).

[0096] In block 1085, it is determined whether the authoring process 1050 continues. If the logic system receives an indication that the user so wishes, the process continues. For example, the process 1050 may continue by returning to block 1055 or block 1060. If not, the authoring tool may send the audio data and metadata to the rendering tool (block 1090), after which the process 1050 may end (1095).

[0097] To optimize the realism of the perceived movement of audio objects, it may be desirable to allow a user of an authoring tool (or rendering tool) to select a subset of speakers in the playback environment and limit the set of active speakers to the selected subset. In some implementations, speaker zones and / or groups of speaker zones may be designated as active or inactive during the authoring or rendering process. For example, with reference to FIG. 4A , the speaker zones in front region 405, left region 410, right region 415, and / or top region 420 may be controlled as a group. The speaker zones in the back region, which includes speaker zones 6 and 7 (and, in other implementations, one or more other speaker zones located between speaker zones 6 and 7), may also be controlled as a group. A user interface may be provided for dynamically enabling or disabling specific speaker zones or all of the speakers corresponding to a region containing multiple speaker zones.

[0098] In some implementations, the logic system of the authoring device (or rendering device) may be configured to generate speaker zone constraint metadata according to user input received via a user input system. The speaker zone constraint metadata may include data for disabling selected speaker zones. Some such implementations are now described with reference to FIGS. 11 and 12.

[0099] FIG. 11 shows an example of applying speaker zone constraints in a virtual playback environment. In some such implementations, a user may be able to select speaker zones by clicking on a representation in a GUI such as GUI 400 using a user input device such as a mouse. Here, the user has disabled speaker zones 4 and 5 on the sides of virtual playback environment 404. Speaker zones 4 and 5 may correspond to most (or all) of the speakers in a physical playback environment, such as a movie theater sound system environment. In this example, the user has also constrained the position of audio object 505 to a position along line 1105. With most or all of the speakers along the side walls disabled, panning from screen 150 to the back of virtual playback environment 404 is constrained to avoid using the side speakers. This may produce improved perceived movement from front to back for a wider audience area, especially for audience members sitting near the playback speakers corresponding to speaker zones 4 and 5.

[0100] In some implementations, speaker zone constraints may be enforced throughout all re-rendering modes. For example, speaker zone constraints may be enforced in situations when fewer zones are available for rendering, such as when rendering for a Dolby Surround 7.1 or 5.1 configuration that exhibits only seven or five zones. Speaker zone constraints may also be enforced when a larger number of zones are available for rendering. Thus, speaker zone constraints can be seen as a way to guide re-rendering and provide a non-blind solution to the traditional "upmixing / downmixing" process.

[0101] FIG. 12 is a flow diagram outlining some examples of applying speaker zone constraint rules. Process 1200 begins with block 1205, where one or more instructions to apply speaker zone constraint rules are received. The instructions may be received by a logic system of an authoring or rendering device or may correspond to input received from a user input device. For example, the instructions may correspond to a user selection of one or more speaker zones to be deactivated. In some implementations, block 1205 may involve receiving an instruction of what type of speaker zone constraint rule to apply, for example, as described below.

[0102] In block 1207, audio data is received by the authoring tool. Audio object positions are received (block 1210) and may be displayed (block 1215), for example, according to input from a user of the authoring tool. The position data, in this example, are (x, y, z) coordinates. Here, active and inactive speaker zones for the selected speaker zone constraint rule are also displayed in block 1215. In block 1220, the audio data and associated metadata are saved. In this example, the metadata includes audio object positions and speaker zone constraint metadata, which may include speaker zone identification flags.

[0103] In some implementations, the speaker zone constraint metadata may indicate that the rendering tool should apply panning equations to calculate gains binary, for example, by considering all speakers in the selected (disabled) speaker zone to be "off" and all other speaker zones to be "on." The logic system may be configured to generate speaker zone constraint metadata that includes data for disabling the selected speaker zone.

[0104] In alternative implementations, the speaker zone constraint metadata may instruct the rendering tool to apply a panning equation to calculate gains in a blended manner that includes a certain degree of contribution from speakers in disabled speaker zones. For example, the logic system may be configured to generate speaker zone constraint metadata that instructs the rendering tool to attenuate selected speaker zones by performing the following processes: calculating a first gain that includes a contribution from the selected (disabled) speaker zone; calculating a second gain that does not include a contribution from the selected speaker zone; and blending the first gain with the second gain. In some implementations, a bias may be applied to the first gain and / or the second gain (from a selected minimum to a selected maximum) to allow for a range of potential contributions from the selected speaker zone.

[0105] In this example, in block 1225, the authoring tool sends the audio data and metadata to the rendering tool. The logic system may then determine whether the authoring process continues (block 1227). If the logic system receives an indication that the user wishes to do so, the authoring process may continue. Otherwise, the authoring process may end (block 1229). In some implementations, the rendering process may continue according to user input.

[0106] Audio objects containing audio data and metadata generated by the authoring tool are received by the rendering tool in block 1230. In this example, position data for a particular audio object is received in block 1235. The rendering tool's logic system may apply panning equations to calculate gains for the audio object position data according to speaker zone constraint rules.

[0107] In block 1245, the calculated gain is applied to the audio data. The logic system may store the gain, audio object position, and speaker zone constraint metadata in a memory system. In some implementations, the audio data may be played back by a speaker system. The corresponding speaker responses may be shown on a display in some implementations.

[0108] At block 1248, it is determined whether process 1200 continues. The process may continue if the logic system receives an indication that the user wishes to do so. For example, the rendering process may continue by returning to block 1230 or block 1235. If an indication is received that the user wishes to return to the corresponding authoring process, the process may return to block 1207 or block 1210. Otherwise, process 1200 may end (block 1250).

[0109] The task of positioning and rendering audio objects in a three-dimensional virtual playback environment becomes increasingly difficult. Part of the difficulty relates to the difficulty in representing the virtual playback environment in a GUI. Some authoring and rendering implementations provided in this paper allow users to switch between panning in two-dimensional screen space and panning in three-dimensional room space. Such functionality can help preserve the accuracy of audio object positioning while providing a GUI that is convenient for users.

[0110] 13A and 13B show an example of a GUI that can switch between two-dimensional and three-dimensional views of a virtual playback environment. Referring to FIG. 13A, GUI 400 depicts image 1305 on the screen. In this example, image 1305 is an image of a saber-toothed tiger. In this top view of virtual playback environment 404, the user can easily observe that audio object 505 is near speaker zone 1. Height can be inferred, for example, by the size, color, or some other attribute of audio object 505. However, the relationship of this position to the position of image 1305 can be difficult to determine in this view.

[0111] In this example, GUI 400 can appear to be dynamically rotated around an axis, such as axis 1310. FIG. 13B shows GUI 1300 after the rotation process. In this view, the user can see image 1305 more clearly and use information from image 1305 to more accurately position audio object 505. In this example, the audio object corresponds to the sound that the saber-toothed tiger is looking at. Being able to switch between a top view and a screen view of virtual playback environment 404 allows the user to quickly and accurately select the appropriate height for audio object 505 using information from on-screen material.

[0112] Various other convenient GUIs for authoring and / or rendering are provided herein. Figures 13C-13E show a combination of two-dimensional and three-dimensional representations of a playback environment. Referring first to Figure 13C, a top view of the virtual playback environment 404 is depicted in the left region of the GUI 1310. The GUI 1310 also includes a three-dimensional representation 1345 of the virtual (or real) playback environment. The region 1350 of the three-dimensional representation 1345 corresponds to the screen 150 of the GUI 400. The position of the audio object 505, particularly its height, is clearly visible in the three-dimensional representation 1345. In this example, the width of the audio object 505 is also shown in the three-dimensional representation 1345.

[0113] Speaker layout 1320 depicts speaker positions 1324 through 1340. Each position may indicate a gain corresponding to the location of audio object 505 in virtual playback environment 404. In some implementations, speaker layout 1320 may represent playback speaker positions in a real playback environment, such as a Dolby Surround 5.1 configuration, a Dolby Surround 7.1 configuration, or a Dolby 7.1 configuration with overhead speakers augmented. When the logic system receives an indication of the location of audio object 505 in virtual playback environment 404, the logic system may be configured to map this location to gains for speaker positions 1324 through 1340 in speaker layout 1320, for example, by the amplitude panning process described above. For example, in FIG. 13C , speaker positions 1325, 1335, and 1337 each have a color change that indicates the gain corresponding to the location of audio object 505.

[0114] Referring now to FIG. 13D , the audio object has been moved to a position behind the screen 150. For example, a user may have moved the audio object 505 by placing a cursor over the audio object 505 in the GUI 400 and dragging the object to a new position. This new position is also shown in the three-dimensional rendering 1345, rotated to a new orientation. The response of the speaker layout 1320 may appear substantially the same in FIGS. 13C and 13D . However, in an actual GUI, the speaker positions 1325, 1335, and 1337 may have different appearances (e.g., different brightness or color) to indicate the corresponding gain differences caused by the new position of the audio object 505.

[0115] 13E, audio object 505 may be rapidly moved to a position in the right rear portion of virtual playback environment 404. At the moment depicted in FIG. 13E, speaker position 1326 corresponds to the current position of audio object 505, and speaker positions 1325 and 1337 still correspond to the previous positions of audio object 505.

[0116] FIG. 14A is a flow diagram outlining a process for controlling a device to present a GUI such as those shown in FIGS. 13C-13E. Process 1400 begins at block 1405, in which one or more instructions are received to display audio object positions, speaker zone positions, and playback speaker positions for a playback environment. The speaker zone positions may correspond to a virtual playback environment and / or a real playback environment, for example, as shown in FIGS. 13C-13E. The instructions may be received by a logic system of a rendering and / or authoring device or may correspond to input received from a user input device. For example, the instructions may correspond to a user selection of a playback environment configuration.

[0117] In block 1407, audio data is received. Audio object position data and width are received in block 1410, for example, according to user input. In block 1415, the audio object, speaker zone position, and playback speaker position are displayed. The audio object position may be displayed in two-dimensional and / or three-dimensional views, for example, as shown in FIGS. 13C-13E. The width data may not only be used in audio object rendering, but may also affect how the audio object is displayed (see the rendering of audio object 505 in three-dimensional rendering 1345 of FIGS. 13C-13E).

[0118] The audio data and associated metadata may be recorded (block 1420). In block 1425, the authoring tool sends the audio data and metadata to the rendering tool. The logic system may then determine whether the authoring process continues (block 1427). If the logic system receives an indication that the user wishes to do so, the authoring process may continue (e.g., by returning to block 1405). Otherwise, the authoring process may end (block 1429).

[0119] The audio objects, including audio data and metadata, generated by the authoring tool are received by the rendering tool in block 1430. In this example, position data for a particular audio object is received in block 1435. The rendering tool's logic system may apply panning equations to calculate gains for the audio object position data according to the width metadata.

[0120] In some rendering implementations, a logic system may map speaker zones to playback speakers in a playback environment. For example, the logic system may access a data structure containing speaker zones and corresponding playback speaker locations. Further details and examples are provided below with reference to FIG. 14B.

[0121] In some implementations, a panning equation may be applied (block 1440), e.g., by a logic system, according to the audio object's position, width, and / or other information, such as speaker positions in the playback environment. In block 1445, the audio data is processed according to the gain obtained in block 1440. At least a portion of the resulting audio data may be stored, if desired, along with corresponding audio object position data and other metadata received from the authoring tool. The audio data may be played over speakers.

[0122] The logic system may then determine whether process 1400 continues (block 1448). For example, if the logic system receives an indication that the user wishes to do so, process 1400 may continue. If not, process 1400 may end (block 1449).

[0123] 14B is a flow diagram outlining a process for rendering audio objects for a playback environment. Process 1450 begins at block 1455, where one or more instructions for rendering audio objects for a playback environment are received. The instructions may be received by a logic system of a rendering device or may correspond to input received from a user input device. For example, the instructions may correspond to a user selection of a playback environment configuration.

[0124] At block 1457, audio playback data (including one or more audio objects and associated metadata) is received. At block 1460, playback environment data may be received. The playback environment data may include an indication of the number of playback speakers in the playback environment and an indication of the position of each playback speaker within the playback environment. The playback environment may be a movie theater sound system environment, a home theater environment, etc. In some implementations, the playback environment data may include playback speaker zone layout data that indicates playback speaker zones and playback speaker positions corresponding to the speaker zones.

[0125] The playback environment may be displayed in block 1465. In some implementations, the playback environment may be displayed in a manner similar to the speaker layout 1320 shown in Figures 13C-13E.

[0126] In block 1470, audio objects may be rendered into one or more speaker feed signals for the playback environment. In some implementations, metadata associated with the audio objects may be authored in the manner described above, and the metadata may include gain data corresponding to speaker zones (e.g., corresponding to speaker zones 1 through 9 in GUI 400). A logic system may map speaker zones to playback speakers in the playback environment. For example, the logic system may access a data structure stored in memory that includes speaker zones and corresponding playback speaker positions. A rendering device may have multiple such data structures, each corresponding to a different speaker configuration. In some implementations, a rendering device may have such data structures for multiple standard playback environment configurations, such as Dolby Surround 5.1 configuration, Dolby Surround 7.1 configuration, and / or Hamasaki 22.2 surround sound configuration.

[0127] In some implementations, the metadata about the audio object may include other information from the authoring process. For example, the metadata may include speaker constraint data. The metadata may include information for mapping audio object positions to single playback speaker positions or single playback speaker zones. The metadata may include data that constrains the position of an audio object to a one-dimensional curve or a two-dimensional surface. The metadata may include trajectory data about the audio object. The metadata may include an identifier for the content type (e.g., dialogue, music, or effects).

[0128] Thus, the rendering process may involve using metadata, for example, to impose speaker zone constraints. In some such implementations, the rendering device may provide the user with the option to modify constraints dictated by the metadata, e.g., modify speaker constraints, and re-render accordingly. Rendering may involve generating an overall gain based on one or more of the desired audio object position, the distance from the desired audio object position to a reference position, the velocity of the audio object, or the audio object content type. The corresponding response of the playback speakers may be displayed (block 1475). In some implementations, the logic system may control the speakers to play sounds corresponding to the results of the rendering process.

[0129] In block 1480, the logic system may determine whether process 1450 continues. For example, process 1450 may continue if the logic system receives an indication that the user wishes to do so. For example, process 1450 may continue by returning to block 1457 or block 1460. Otherwise, process 1450 may end (block 1485).

[0130] Control of spread and apparent source width is a feature of some existing surround sound authoring / rendering systems. In this disclosure, the term "spread" refers to distributing the same signal across multiple speakers to blur the sound image. The term "width" refers to decorrelating the output signal to each channel for apparent width control. Width may be an additional scalar value that controls the amount of decorrelation added to each speaker feed signal.

[0131] Some implementations described herein provide 3D axis-oriented spread control. One such implementation is now described with reference to FIGS. 15A and 15B. FIG. 15A illustrates an example of an audio object and associated audio object width in a virtual playback environment. Here, GUI 400 shows an ellipsoid 1505 extending around audio object 505, indicating the audio object width. The audio object width may be indicated by audio object metadata and / or received according to user input. In this example, the x and y dimensions of ellipsoid 1505 are different, although in other implementations these dimensions may be the same. The z dimension of ellipsoid 1505 is not shown in FIG. 15A.

[0132] FIG. 15B shows an example of a diffusion profile corresponding to the audio object width shown in FIG. 15A. Diffusion may be expressed as a three-dimensional vector parameter. In this example, the diffusion profile 1507 can be controlled independently along three dimensions, for example, according to user input. The gain along the x and y axes is indicated by the height of the curves 1510 and 1520, respectively, in FIG. 15B. The gain for each sample 1512 is also indicated by the size of the corresponding circle 1515 in the diffusion profile 1507. The response of the speaker 1510 is indicated by gray shading in FIG. 15B.

[0133] In some implementations, the diffusion profile 1507 may be implemented by separable integrals for each axis. According to some implementations, the minimum diffusion value may be automatically set as a function of speaker placement to avoid tonal discrepancies when panning. Alternatively or additionally, the minimum diffusion value may be automatically set as a function of the velocity of an audio object being panned, so that as the audio object velocity increases, the object becomes increasingly spatially spread out, similar to how fast-moving images in a movie appear blurred.

[0134] When using an audio object-based audio rendering implementation such as that described herein, potentially multiple audio tracks and associated metadata (including, but not limited to, metadata indicating audio object positions in three-dimensional space) may be delivered to the playback environment unmixed. A real-time rendering tool may use such metadata and information about the playback environment to calculate speaker feed signals to optimize the playback of each audio object.

[0135] When many audio objects are mixed to the speaker output, overloads can occur in the digital domain (for example, the digital signal may be clipped before analog conversion) or in the analog domain when the amplified analog signal is played back by the playback speakers. In either case, this leads to audible distortion, which is undesirable. Overloads in the analog domain can even damage the playback speakers.

[0136] Thus, some implementations described in this paper involve "blobbing" of dynamic objects in response to speaker overload. When an audio object is rendered with a given diffusion profile, some implementations may direct energy to an increasing number of neighboring speakers while maintaining a constant overall energy. For example, if the energy for an audio object is uniformly diffused across N speakers, it may contribute a gain of 1 / √N to each speaker output. This approach provides additional mixing "headroom" and can reduce or prevent speaker distortions such as clipping.

[0137] To use a numerical example, suppose a speaker will clip if it receives an input greater than 1.0. Suppose two objects are specified to be blended into Speaker A, one at level 1.0 and the other at level 0.25. If blobbing were not used, the blend levels at Speaker A would sum to 1.25, causing clipping. However, if the first object is blobbed using another Speaker B, then (in some implementations) each speaker will receive that object at 0.707, thereby providing additional "room" at Speaker A for blending additional objects. The second object can then be safely blended into Speaker A without clipping, since the blend level for Speaker A would be 0.707 + 0.25 = 0.957.

[0138] In some implementations, during the authoring stage, each audio object may be mixed into a subset of speaker zones (or into all speaker zones) with a given mix gain. Thus, a dynamic list of all objects contributing to each speaker can be built. In some implementations, this list may be sorted in descending order of energy level, for example, using the product of the signal's original root mean square (RMS) level multiplied by the mix gain. In other implementations, the list may be sorted according to other criteria, such as the relative importance assigned to the audio objects.

[0139] During the rendering process, if an overload is detected for a given playback speaker output, the energy of the audio object may be spread across several playback speakers. For example, the energy of the audio object may be spread using a width or spreading factor proportional to the amount of overload and the relative contribution of each audio object to a given playback speaker. If the same audio object contributes to several overloaded playback speakers, its width or spreading factor may in some implementations be additively increased and applied to the next rendered frame of audio data.

[0140] In general, a hard limiter will clip any value above the threshold to that threshold. As in the example above, if a speaker receives a mixed object at level 1.25 and can only tolerate a maximum level of 1.0, the object will be "hard limited" to 1.0. A soft limiter will begin applying limiting before the absolute threshold is reached, giving a smoother, more perceptually pleasant result. A soft limiter may also use "look ahead" to predict when future clipping may occur, smoothly reducing the gain before clipping occurs, thereby avoiding clipping.

[0141] The various "blobbing" implementations provided herein may be used in conjunction with hard or soft limiters to limit audible distortion while avoiding degradation of spatial accuracy / sharpness. Unlike using only global diffusion or limiters, blobbing implementations can selectively target loud objects or objects of a given content type. Such implementations may be controlled by a mixer. For example, if speaker zone constraint metadata for an audio object indicates that a certain subset of playback speakers should not be used, the rendering device may apply the corresponding speaker zone constraint rules in addition to implementing the blobbing method.

[0142] 16 is a flow diagram outlining a process for blobbing audio objects. Process 1600 begins at block 1605, where one or more instructions to activate audio object blobbing functionality are received. The instructions may be received by a logic system of a rendering device or may correspond to input received from a user input device. In some implementations, the instructions may include a user selection of a playback environment configuration. In alternative implementations, the user may have previously selected a playback environment configuration.

[0143] At block 1607, audio playback data (including one or more audio objects and associated metadata) is received. In some implementations, the metadata may include speaker zone constraint metadata, for example, as described above. In this example, at block 1610, audio object position, time, and diffusion data is parsed from the audio playback data (or otherwise received, for example, via input from a user interface).

[0144] A playback speaker response is determined for the playback environment configuration by applying a panning equation to the audio object data (block 1612), for example, as described above. At block 1615, the audio object position and playback speaker response are displayed (block 1615). The playback speaker response may be played through speakers configured for communication with the logic system.

[0145] In block 1620, the logic system determines whether an overload is detected for any of the playback speakers in the playback environment. If so, the audio object blobbing rules, as described above, are applied until an overload is no longer detected (block 1625). In block 1630, the audio data output may be saved or output to the playback speakers, if desired.

[0146] In block 1635, the logic system may determine whether process 1600 continues. For example, process 1600 may continue if the logic system receives an indication that the user wishes to do so. For example, process 1600 may continue by returning to block 1607 or block 1610. Otherwise, process 1600 may end (block 1640).

[0147] Some implementations provide extended panning gain equations that can be used to image audio object positions in three-dimensional space. Some examples are now described with reference to FIGS. 17A and 17B. FIGS. 17A and 17B show examples of audio objects positioned within a three-dimensional virtual environment. Referring first to FIG. 17A, the position of audio object 505 is seen within virtual playback environment 404. In this example, speaker zones 1 through 7 are positioned in one plane, and speaker zones 8 and 9 are positioned in another plane, as shown in FIG. 17B. However, the number of speaker zones, planes, etc., is merely exemplary, and the concepts described herein may be extended to a different number of speaker zones (or individual speakers) and more than two elevation planes.

[0148] In this example, the height parameter "z", which can range from 0 to 1, maps the position of audio objects to height planes. In this example, the value z=0 corresponds to the base plane containing speaker zones 1-7, and the value z=1 corresponds to the overhead plane containing speaker zones 8 and 9. Values ​​of e between 0 and 1 correspond to a blend between the sound image produced using only the speakers in the base plane and the sound image produced using only the speakers in the overhead plane.

[0149] In the example shown in FIG. 17B, the height parameter for audio object 505 has a value of 0.6. Thus, in one implementation, a first sound image may be generated using a panning equation for the base plane according to the (x,y) coordinates of audio object 505 in the base plane. A second sound image may be generated using a panning equation for the overhead plane according to the (x,y) coordinates of audio object 505 in the overhead plane. The resulting sound image may be generated by combining the first sound image with the second sound image according to the proximity of the audio object 505 to each plane. An energy- or amplitude-conserving function of height z may be applied. For example, where z can vary from 0 to 1, the gain value of the first sound image may be multiplied by cos(z*π / 2), and the gain value of the second sound image may be multiplied by sin(z*π / 2), so that the sum of the squares of both is 1 (energy-conserving).

[0150] Other implementations described herein may involve calculating gains based on two or more panning techniques and generating an overall gain based on one or more parameters, which may include one or more of the following: desired audio object position, distance from the desired audio object position to a reference position, audio object speed or velocity, or audio object content type.

[0151] Some such implementations are now described with reference to Figure 18 et seq. Figure 18 shows example zones corresponding to various panning modes. The size, shape, and extent of these zones are merely exemplary. In this example, near-field panning methods are applied for audio objects located within zone 1805, and far-field panning methods are applied for audio objects located outside zone 1810, within zone 1815.

[0152] Figures 19A-D show examples of applying near-field and far-field panning methods to audio objects at various locations. Referring first to Figure 19A, the audio object is substantially outside the virtual playback environment 1900. This location corresponds to zone 1815 in Figure 18. Therefore, one or more far-field panning methods are applied in this example. In some implementations, the far-field panning method may be based on vector-based amplitude panning (VBAP) formulas known to those skilled in the art. For example, the far-field panning method may be based on the VBAP formulas described in Section 2.3, p. 4, of "VBAP: Virtual Amplitude Panning," in "VBAP," which is incorporated herein by reference. In alternative implementations, other methods may be used to pan audio objects in the far and near fields, such as methods involving the synthesis of corresponding acoustic plane or spherical waves. Related methods are described in "VBAP: Virtual Amplitude Panning," in "VBAP," which is incorporated herein by reference.

[0153] Referring now to FIG. 19B, an audio object is inside virtual playback environment 1900. This location corresponds to zone 1805 in FIG. 18. Therefore, one or more near-field panning methods are applied in this example. Some such near-field panning methods use several speaker zones that surround audio object 505 within virtual playback environment 1900.

[0154] In some implementations, near-field panning methods may involve a "dual-balanced" panning and two sets of gain combinations. In the example depicted in FIG. 19B, the first set of gains corresponds to a front-to-back balance between two sets of speaker zones surrounding the positions of audio object 505 along the y-axis. The corresponding response pertains to all speaker zones in virtual playback environment 1900 except for speaker zones 1915 and 1960.

[0155] In the example depicted in Figure 19C, the second set of gains corresponds to the left-right balance between two sets of speaker zones surrounding the positions of audio object 505 along the x-axis. The corresponding responses pertain to speaker zones 1905 through 1925. Figure 19D shows the result of combining the responses shown in Figures 19B and 19C.

[0156] It may be desirable to blend between different panning modes when an audio object enters or exits the virtual playback environment 1900. Thus, a blend of gains calculated according to the near-field panning method and the far-field panning method is applied to an audio object located within zone 1810 (see FIG. 18). In some implementations, a pair-wise panning law (e.g., an energy-preserving sine or power law) may be used to blend between gains calculated according to the near-field panning method and the far-field panning method. In alternative implementations, the pair-wise panning law may be amplitude-preserving rather than energy-preserving. Thus, rather than the sum of squares equaling one, the sum equals one. It is also possible to process an audio signal using both panning methods independently and then blend the resulting processed signals, e.g., to crossfade the two resulting audio signals.

[0157] It may be desirable to provide a mechanism that allows content creators and / or content players to easily fine-tune various re-renderings for a given authored trajectory. In the context of cinematic mixing, the concept of screen-to-room energy balance is considered important. In some cases, automatic re-rendering of a given sound trajectory (or "pan") leads to a different screen-to-room balance depending on the number of playback speakers in the playback environment. According to some implementations, the screen-to-room bias is controlled according to metadata generated during the authoring process. According to alternative implementations, the screen-to-room bias may be controlled solely on the rendering side (i.e., under the control of the content player) rather than in response to metadata.

[0158] Thus, some implementations described herein provide one or more forms of screen-to-room bias control. In some such implementations, the screen-to-room bias may be implemented as a scaling process. For example, the scaling process may involve scaling speaker positions used in the renderer to determine the original intended trajectory and / or pan gain of an audio object along the front-to-back direction. In some such implementations, the screen-to-room bias control may be a variable value between 0 and a maximum value (e.g., 1). The variation may be controllable, for example, using a GUI, a virtual or physical slider, a knob, etc.

[0159] Alternatively or additionally, screen-to-room bias control may be implemented using some form of speaker area constraint. FIG. 20 shows speaker zones of a playback environment that may be used in the screen-to-room bias control process. In this example, front speaker areas 2005 and rear speaker areas 2010 (or 2015) may be established. The screen-to-room bias may be adjusted as a function of the selected speaker areas. In some such implementations, the screen-to-room bias may be implemented as a scaling process between the front speaker areas 2005 and the rear speaker areas 2010 (or 2015). In alternative implementations, the screen-to-room bias may be implemented binary, for example, by allowing the user to select front bias, rear bias, or no bias. The bias setting for each case may correspond to predetermined (and generally non-zero) bias levels for the front speaker areas 2005 and rear speaker areas 2010 (or 2015). In essence, such an implementation may provide three presets for screen-to-room bias control rather than (or in addition to) a continuous value scaling process.

[0160] According to some such implementations, two additional logical speaker zones may be created in the authoring GUI (e.g., 400) by splitting the side walls into a front wall and a back wall. In some implementations, the two additional logical speaker zones correspond to the left wall / left surround sound and right wall / right surround sound areas of the renderer. Depending on the user's selection of which of these two logical speaker zones is active, the rendering tool may apply preset scaling factors (e.g., as described above) when rendering to a Dolby 5.1 or Dolby 7.1 configuration. The rendering tool may also apply such preset scaling factors when rendering for a playback environment that does not support the definition of these two extra logical zones, for example, because the physical speaker configuration has only one physical speaker on the side wall.

[0161] 21 is a block diagram providing an example of components of an authoring and / or rendering device. In this example, device 2100 includes an interface system 2105. Interface system 2105 may include a network interface, such as a wireless network interface. Alternatively or additionally, interface system 2105 may include a universal serial bus (USB) interface or other such interface.

[0162] Device 2100 includes logic system 2110. Logic system 2110 may include a processor, such as a general-purpose single-chip or multi-chip processor. Logic system 2110 may include a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or discrete hardware components, or a combination thereof. Logic system 2110 may be configured to control other components of device 2100. Although interfaces between components of device 2100 are not shown in FIG. 21 , logic system 2110 may be configured with interfaces for communication with other components. The other components may or may not be configured for communication with each other, as appropriate.

[0163] Logic system 2110 may be configured to perform audio authoring and / or rendering functions, including but not limited to the audio authoring and / or rendering functions described herein. In some such implementations, logic system 2110 may be configured to operate (at least in part) according to software stored on one or more non-transitory media. The non-transitory media may include memory associated with logic system 2110, such as random access memory (RAM) and / or read-only memory (ROM). The non-transitory media may include memory of memory system 2115. Memory system 2115 may include one or more suitable types of non-transitory storage media, such as flash memory, hard drives, etc.

[0164] Display system 2130 may include one or more suitable types of displays, depending on the implementation of device 2100. For example, display system 2130 may include a liquid crystal display, a plasma display, a bi-stable display, etc.

[0165] The user input system 2135 may include one or more devices configured to accept input from a user. In some implementations, the user input system 2135 may include a touchscreen overlaying the display of the display system 2130. The user input system 2135 may include a mouse, a trackball, a gesture detection system, a joystick, one or more GUIs and / or menus presented on the display system 2130, buttons, keyboards, switches, etc. In some implementations, the user input system 2135 may include a microphone 2125; a user may provide voice commands for the device 2100 via the microphone 2125. A logic system may be configured for voice recognition and for controlling at least some operations of the device 2100 according to such voice commands.

[0166] Power system 2140 may include one or more suitable energy storage devices, such as nickel-cadmium batteries or lithium-ion batteries. Power system 2140 may be configured to receive power from an electrical outlet.

[0167] FIG. 22A is a block diagram illustrating several components that may be used for audio content creation. System 2200 may be used for audio content creation in, for example, a mixing studio and / or a dubbing stage. In this example, system 2200 includes an audio and metadata authoring tool 2205 and a rendering tool 2210. In this implementation, audio and metadata authoring tool 2205 and rendering tool 2210 include audio connection interfaces 2207 and 2212, respectively, which may be configured for communication via AES / EBU, MADI, analog, etc. Audio and metadata authoring tool 2205 and rendering tool 2210 include network interfaces 2209 and 2217, respectively, which may be configured to send and receive metadata via TCP / IP or any other suitable protocol. Interface 2220 is configured to output audio data to speakers.

[0168] The system 2200 may include, for example, an existing authoring system, such as a ProTools™ system, that runs a metadata generation tool (i.e., a panner described herein) as a plug-in. The panner may run on a standalone system (e.g., a PC or mixing console) connected to the rendering tool 2210, or it may run on the same physical device as the rendering tool 2210. In the latter case, the panner and renderer may use a local connection, for example, through shared memory. The panner GUI may be remote, such as on a tablet device, laptop, etc. The rendering tool 2210 may have a rendering system including a sound processor configured to run rendering software. The rendering system may include, for example, a personal computer, laptop, etc., including interfaces for audio input and output and an appropriate logic system.

[0169] 22B is a block diagram illustrating some components that may be used for audio playback in a playback environment (e.g., a movie theater). System 2250, in this example, includes a theater server 2255 and a rendering system 2260. Theater server 2255 and rendering system 2260 include network interfaces 2257 and 2262, respectively, which may be configured to send and receive audio objects via TCP / IP or any other suitable protocol. Interface 2264 is configured to output audio data to speakers.

[0170] Various modifications to the implementations described in this disclosure will be readily apparent to those skilled in the art. The general principles defined herein may be applied to other implementations without departing from the spirit or scope of the disclosure. Thus, the scope of the claims is not intended to be limited to the implementations shown herein, but is to be accorded the widest scope consistent with the disclosure, principles, and novel features disclosed herein.

[0171] Several aspects will be described. [Aspect 1] 1. An apparatus having an interface system and a logic system: The logic system: receiving, via the interface system, audio playback data including one or more audio objects and associated metadata; receiving, via the interface system, playback environment data including an indication of a number of playback speakers in the playback environment and an indication of a location of each playback speaker within the playback environment; and rendering the audio objects into one or more speaker feed signals based at least in part on the associated metadata; Each speaker feed signal corresponds to at least one playback speaker in the playback environment. Device. [Aspect 2] 2. The apparatus of claim 1, wherein the playback environment is a movie theater sound system environment. Aspect 3 2. The apparatus of claim 1, wherein the playback environment has a Dolby Surround 5.1 configuration, a Dolby Surround 7.1 configuration, or a Hamasaki 22.2 surround sound configuration. Aspect 4 2. The apparatus of claim 1, wherein the playback environment data includes playback speaker layout data indicating playback speaker positions. Aspect 5 The device of embodiment 1, wherein the playback environment data includes playback speaker zone layout data indicating playback speaker areas and playback speaker positions corresponding to the playback speaker areas. Aspect 6 6. The apparatus of embodiment 5, wherein the metadata includes information for mapping audio object positions to a single playback speaker position. Aspect 7 The apparatus of aspect 1, wherein the rendering includes generating an overall gain based on one or more of a desired audio object position, a distance from the desired audio object position to a reference position, a velocity of the audio object, or an audio object content type. Aspect 8 2. The apparatus of aspect 1, wherein the metadata includes data for constraining the position of the audio object to a one-dimensional curve or a two-dimensional surface. Aspect 9 2. The apparatus of claim 1, wherein the metadata includes trajectory data for an audio object. Aspect 10 10. The apparatus of claim 1, wherein the rendering includes imposing speaker zone constraints. Aspect 11 10. The apparatus of claim 1, further comprising a user input system, wherein the rendering includes applying screen-to-room balance control according to screen-to-room balance control data received from the user input system. Aspect 12 10. The apparatus of claim 1, further comprising a display system, wherein the logic system is configured to control the display system to display a dynamic three-dimensional view of the playback environment. Aspect 13 10. The apparatus of claim 1, wherein the rendering includes controlling audio object diffusion in one or more of three dimensions. Aspect 14 10. The apparatus of claim 1, wherein the rendering includes dynamic object blobbing in response to speaker overload. Aspect 15 2. The apparatus of claim 1, wherein the rendering includes mapping audio object positions onto a plane of a speaker array of the playback environment. Aspect 16 2. The apparatus of claim 1, further comprising a memory device, wherein the interface system comprises an interface between the logic system and the memory device. Aspect 17 The apparatus of embodiment 1, wherein the interface system comprises a network interface. Aspect 18 10. The apparatus of claim 1, wherein the metadata includes speaker zone constraint metadata, and the logic system comprises: Calculating a first gain including the contribution from the selected speaker; Calculate a second gain that does not include the contribution from the selected speaker; by blending the first gain with the second gain; An apparatus configured to attenuate a selected speaker feed signal. Aspect 19 2. The apparatus of claim 1, wherein the metadata includes speaker zone constraint metadata, and the logic system is configured to determine whether to apply panning rules to audio object positions or to map audio object positions to a single speaker position. Aspect 20 20. The apparatus of claim 19, wherein the logic system is configured to smooth the transition in speaker gain when transitioning from a mapping of an audio object position to a first single speaker position to a second single speaker position. Aspect 21 20. The apparatus of claim 19, wherein the logic system is configured to smooth transitions in speaker gain when transitioning between mapping an audio object position to a single speaker position and applying panning rules to the audio object position. Aspect 22 22. The apparatus of any one of aspects 1-21, wherein the logic system is further configured to calculate speaker gains corresponding to virtual speaker positions. Aspect 23 23. The apparatus of claim 22, wherein the logic system is further configured to calculate speaker gains for audio object positions along a one-dimensional curve between virtual speaker positions. Aspect 24 receiving audio playback data including one or more audio objects and associated metadata; receiving playback environment data including an indication of a number of playback speakers in the playback environment and an indication of a location of each playback speaker within the playback environment; and rendering the audio objects into one or more speaker feed signals based at least in part on the associated metadata; Each speaker feed signal corresponds to at least one playback speaker in the playback environment. method. Aspect 25 25. The method of claim 24, wherein the playback environment is a movie theater sound system environment. Aspect 26 25. The method of claim 24, wherein the rendering includes generating an overall gain based on one or more of a desired audio object position, a distance from the desired audio object position to a reference position, a velocity of the audio object, or an audio object content type. Aspect 27 25. The method of embodiment 24, wherein the metadata includes data for constraining the position of the audio object to a one-dimensional curve or a two-dimensional surface. Aspect 28 25. The method of aspect 24, wherein the rendering includes imposing speaker zone constraints. Aspect 29 A non-transitory medium having software stored thereon, said software comprising: receiving audio playback data including one or more audio objects and associated metadata; receiving playback environment data including an indication of a number of playback speakers in the playback environment and an indication of a location of each playback speaker within the playback environment; and rendering the audio objects into one or more speaker feed signals based at least in part on the associated metadata; Each speaker feed signal corresponds to at least one playback speaker in the playback environment. Non-transient medium. Aspect 30 30. The non-transitory medium of claim 29, wherein the playback environment is a movie theater sound system environment. Aspect 31 30. The non-transitory medium of claim 29, wherein the rendering includes generating an overall gain based on one or more of a desired audio object position, a distance from the desired audio object position to a reference position, a velocity of the audio object, or an audio object content type. Aspect 32 30. The non-transitory medium of embodiment 29, wherein the metadata includes data for constraining the position of the audio object to a one-dimensional curve or a two-dimensional surface. Aspect 33 30. The non-transitory media of aspect 29, wherein the rendering includes imposing speaker zone constraints. Aspect 34 30. The non-transitory media of aspect 29, wherein the rendering includes dynamic object blobbing in response to speaker overload. Aspect 35 1. An apparatus having an interface system, a user input system, and a logic system, the logic system comprising: receiving audio data via the interface system; receiving a location of an audio object via the user input system or the interface system; determining a position of the audio object in three-dimensional space, the determining including constraining the position to a one-dimensional curve or a two-dimensional surface within the three-dimensional space; generating metadata associated with the audio object based at least in part on user input received via the user input system, the metadata including data indicative of a position of the audio object in three-dimensional space. Device. Aspect 36 36. The apparatus of claim 35, wherein the metadata includes trajectory data indicating a time-varying position of the audio object within three-dimensional space. Aspect 37 37. The apparatus of embodiment 36, wherein the logic system is configured to calculate the trajectory data according to user input received via the user input system. Aspect 38 37. The apparatus of embodiment 36, wherein the trajectory data comprises a set of positions in three-dimensional space at multiple points in time. Aspect 39 37. The apparatus of embodiment 36, wherein the trajectory data includes an initial position, velocity data, and acceleration data. Aspect 40 37. The apparatus of embodiment 36, wherein the trajectory data includes an initial position and an equation defining positions in three-dimensional space and corresponding times. Aspect 41 37. The apparatus of claim 36, further comprising a display system, wherein the logic system is configured to control the display system to display an audio object trajectory according to the trajectory data. Aspect 42 36. The apparatus of claim 35, wherein the logic system is configured to generate speaker zone constraint metadata according to user input received via the user input system. Aspect 43 43. The apparatus of aspect 42, wherein the speaker zone constraint metadata includes data for disabling selected speakers. Aspect 44 43. The apparatus of aspect 42, wherein the logic system is configured to generate speaker zone constraint metadata by mapping audio object positions to a single speaker. Aspect 45 36. The apparatus of claim 35, further comprising a sound reproduction system, wherein the logic system is configured to control the sound reproduction system at least in part according to the metadata. Aspect 46 36. The apparatus of claim 35, wherein the position of the audio object is constrained to a one-dimensional curve, and the logic system is further configured to generate virtual speaker positions along the one-dimensional curve. Aspect 47 receiving audio data; receiving a position of an audio object; determining a position of the audio object in three-dimensional space, the determining including constraining the position to a one-dimensional curve or a two-dimensional surface within the three-dimensional space; generating metadata associated with the audio object based at least in part on user input, the metadata including data indicative of a position of the audio object within three-dimensional space; method. Aspect 48 48. The method of aspect 47, wherein the metadata includes trajectory data indicating a time-varying position of the audio object within three-dimensional space. Aspect 49 48. The method of aspect 47, wherein generating the metadata includes generating speaker zone constraint metadata according to user input, the speaker zone constraint metadata including data for disabling selected speakers. Aspect 50 48. The method of embodiment 47, wherein the position of the audio object is constrained to a one-dimensional curve, and further comprising generating virtual speaker positions along the one-dimensional curve. Aspect 51 A non-transitory medium having software stored thereon, said software comprising: receiving audio data; receiving a position of an audio object; determining a position of the audio object in three-dimensional space, the determining including constraining the position to a one-dimensional curve or a two-dimensional surface within the three-dimensional space; generating metadata associated with the audio object based at least in part on user input, the metadata including data indicative of a position of the audio object within three-dimensional space; Non-transient medium. Aspect 52 52. The non-transitory medium of claim 51, wherein the metadata includes trajectory data indicating the time-varying position of the audio object within three-dimensional space. Aspect 53 52. The non-transitory medium of claim 51, wherein generating the metadata includes generating speaker zone constraint metadata according to user input, the speaker zone constraint metadata including data for disabling selected speakers. Aspect 54 52. The non-transitory medium of claim 51, wherein the position of the audio object is constrained to a one-dimensional curve, and further comprising generating virtual speaker positions along the one-dimensional curve.

Claims

1. receiving audio playback data including one or more audio objects and metadata associated with each of the one or more audio objects; receiving playback environment data including an indication of a number of playback speakers in a playback environment and an indication of a position of each playback speaker within the playback environment; rendering each audio object into one or more speaker feed signals by applying an amplitude panning process to the audio object, the amplitude panning process being based at least in part on metadata associated with each audio object, the respective positions of one or more virtual speakers and the position of each playback speaker within the playback environment, each speaker feed signal corresponding to at least one of the playback speakers within the playback environment; the metadata associated with each audio object includes audio object coordinates indicating the intended playback position of the audio object within the playback environment and zone constraint metadata, the zone constraint metadata indicating whether rendering of the audio object includes imposing speaker zone constraints; imposing a speaker zone constraint includes disabling one or more playback speakers within a speaker zone indicated by the zone constraint metadata; method.

2. The method of claim 1 , wherein the speaker zones indicated by the zone constraint metadata correspond to one or more of a front region, a left region, a right region, a left rear region, a right rear region, a top region, and a back region.

3. 3. The method of claim 2, wherein the front area corresponds to the area in a cinema reproduction environment where a screen is located, or to the area in a home where a television screen is located.

4. 2. The method of claim 1, wherein disabling one or more playback speakers within a speaker zone indicated by the zone constraint metadata comprises applying a panning equation for calculating gain by considering the one or more playback speakers within the speaker zone indicated by the zone constraint metadata to be off.

5. an interface system; and a logic system, the apparatus being configured to perform the method of any one of claims 1 to 4. Device.

6. A computer program product for carrying out the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Three-dimensional acoustic panning apparatus and program therefor

    JP2010252220A

  • Audio processing method and apparatus

    JP2010511912A

  • Apparatus and Method for Calculating Filter Coefficients for a Predefined Loudspeaker Arrangement

    US20110135124A1

  • Virtual sound source positioning

    US7113610B1

  • Apparatus and method for calculating driving coefficients for loudspeakers of a loudspeaker arrangement for an audio signal associated with a virtual source

    WO2011054876A1