Method and system for generating and interactively rendering object-based audio
By generating object-based audio programs including speaker channel sound beds, replacement speaker channel and object channel groups, the problem that legacy systems cannot render full range audio is solved, and the full range audio experience of legacy systems and personalized audio experience of appropriately configured systems is realized.
Patent Information
- Application Number
- CN202210302370.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2013-06-07
- Filing Date
- 2014-04-03
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2034-04-03
AI Technical Summary
Existing legacy playback systems are unable to effectively parse and render object-based audio programs, resulting in the inability to provide a full range of audio experience.
Object-based audio programs including speaker channel sound beds, replacement speaker channels, and object channel groups are generated and provided, with metadata attached to indicate optional mixing options so that legacy systems can render default mixes while appropriately configured systems can render mixes of extended layers.
Legacy systems can provide a full range of audio experiences, and appropriately configured systems can provide a personalized audio experience, enhancing the flexibility and quality of audio rendering.
Smart Images

Figure CN114708873B_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese invention patent application with application number 201480020223.2 (for which divisional application 201710942931.7 has been submitted), application date April 3, 2014, and titled “Method and system for generating and interactively rendering object-based audio”.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit of the filing date of U.S. Provisional Patent Application No. 61 / 807,922, filed April 3, 2013, and the benefit of the filing date of U.S. Provisional Patent Application No. 61 / 832,397, filed June 7, 2013. Technical Field
[0004] The present invention relates to audio signal processing, and more particularly to encoding, decoding, and interactive rendering of an audio data bitstream comprising audio content (indicative of speaker channels and at least one audio object channel) and metadata supporting interactive rendering of the audio content. Some embodiments of the present invention generate, decode, and / or render audio data in one of the formats known as Dolby Digital (AC-3), Dolby Digital Plus (Enhanced AC-3 or E-AC-3), or Dolby E. Background Art
[0005] Dolby, Dolby Digital, Dolby Digital Plus, and Dolby E are trademarks of Dolby Laboratories Licensing Inc. Dolby Laboratories provides proprietary implementations of AC-3 and E-AC-3 known as Dolby Digital and Dolby Digital Plus, respectively.
[0006] Although the present invention is not limited to encoding audio data according to the E-AC-3 (or AC-3 or Dolby E) format or to delivering, decoding or rendering E-AC-3, AC-3 or Dolby E encoded data, for convenience, the present invention will be described in an embodiment in which it encodes an audio bitstream according to the E-AC-3 or AC-3 or Dolby E format, and delivers, decodes and renders such a bitstream.
[0007] A typical audio data stream includes both audio content (e.g., one or more channels of audio content) and metadata indicating at least one characteristic of the audio content. For example, in an AC-3 bitstream, there are several audio metadata parameters that are specifically intended to alter the sound of a program being delivered to a listening environment.
[0008] An AC-3 or E-AC-3 encoded bitstream includes metadata and may include from one to six channels of audio content. The audio content is audio data that has been compressed using perceptual audio coding. The details of AC-3 encoding are well known and are described in many published references, including:
[0009] ATSC Standard A52 / A: Digital Audio Compression Standard (AC-3), Revision A, Advanced Television Systems Committee, August 20, 2001; and
[0010] U.S. Patents 5,583,962; 5,632,005; 5,633,981; 5,727,119; and 6,021,386.
[0011] Details of Dolby Digital Plus (E-AC-3) encoding are described, for example, in: "Introduction to Dolby Digital Plus, an Enhancement to the Dolby Digital Coding System," AES Conference Paper 6196, 117th AES Convention, October 28, 2004.
[0012] Details of Dolby E encoding are set forth in: “Efficient Bit Allocation, Quantization, and Coding in an Audio Distribution System,” AES Preprint 5068, 107th AES Conference, August 1999, and “Professional Audio Coder Optimized for Use with Video,” AES Preprint 5033, 107th AES Conference, August 1999.
[0013] Each frame of the AC-3 encoded audio bitstream includes metadata and audio content for 1536 samples of digital audio. For a sampling rate of 48kHz, this represents 32 milliseconds of digital audio, or a rate of 31.25 frames per second of audio.
[0014] Each frame of the E-AC-3 coded audio bitstream contains metadata and audio content for 256, 512, 768, or 1536 samples of digital audio, depending on whether the frame contains 1, 2, 3, or 6 audio data blocks, respectively. For a sampling rate of 48kHz, this represents 5.333, 10.667, 16, or 32 milliseconds of digital audio, or audio rates of 189.9, 93.75, 62.5, or 31.25 frames per second, respectively.
[0015] like Figure 1 As shown, each AC-3 frame is divided into multiple parts (segments), including: Figure 2 The audio data file includes: a synchronization word (SW) and a synchronization information (SI) portion of the first of the two error correction words (CRC1); a bit stream information (BSI) portion containing most of the metadata; six audio blocks (AB0 to AB5) containing data-compressed audio content (and may also contain metadata); useless bits (W) containing any unused bits remaining after compressing the audio content; an auxiliary (AUX) information portion that may contain more metadata; and a second of the two error correction words (CRC2).
[0016] like Figure 4 As shown, each E-AC-3 frame is divided into multiple parts (segments), including: Figure 2 The audio data frame comprises a synchronization information (SI) portion containing a synchronization word (SW) (shown in FIG); a bit stream information (BSI) portion containing most of the metadata; one to six audio blocks (AB0 to AB5) containing data-compressed audio content (and may also contain metadata); useless bits (W) containing any unused bits remaining after the compressed audio content; an auxiliary (AUX) information portion that may contain more metadata; and an error correction word (CRC).
[0017] In an AC-3 (or E-AC-3) bitstream, there are several audio metadata parameters that are specifically intended to change the sound of a program delivered to a listening environment. One of the metadata parameters is the DIALNORM parameter, which is included in the BSI segment.
[0018] like Figure 3 As shown, the BSI segment of an AC-3 frame (or E-AC-3 frame) includes a 5-bit parameter ("DIALNORM") indicating the DIALNORM value of the program. If the audio coding mode ("acmod") of the AC-3 frame is "0", a 5-bit parameter ("DIALNORM2") indicating the DIALNORM value of the second audio program carried in the same AC-3 frame is included, indicating that a dual-single or "1+1" channel configuration is used.
[0019] The BSI segment also includes a flag ("addbsie") indicating the presence (or absence) of additional bitstream information following the "addbsie" bit, a parameter ("addbsil") indicating the length of any additional bitstream information following the "addbsil" value, and up to 64 bits of additional bitstream information ("addbsi") following the "addbsil" value.
[0020] The BSI segment is included in Figure 3 Other metadata values not specifically shown in .
[0021] Other types of metadata have been proposed for inclusion in audio bitstreams. For example, PCT International Application Publication No. WO 2012 / 075246 A2, filed December 1, 2011, and assigned to the assignee of the present application, describes methods and systems for generating, decoding, and processing audio bitstreams that include metadata indicating a processing state (e.g., loudness processing state) and characteristics (e.g., loudness) of audio content. This reference also describes the use of metadata for adaptive processing of the audio content of a bitstream and for verifying the validity of the loudness processing state and loudness of the audio content of a bitstream.
[0022] Methods for generating and rendering object-based audio programs are also known. During the generation of such programs, it can be assumed that the loudspeakers to be used for rendering are located at arbitrary locations in the playback environment (or that the loudspeakers are located in a symmetrical configuration within a unit circle). It need not be assumed that the loudspeakers are necessarily located in a (nominal) horizontal plane or in any other predetermined arrangement known at the time of program generation. Typically, metadata included in a program indicates rendering parameters for rendering at least one object of the program at an apparent spatial location or along a trajectory (in a three-dimensional volume), for example, using a three-dimensional loudspeaker array. For example, an object channel of a program may have corresponding metadata indicating a three-dimensional trajectory of the apparent spatial location of the object (indicated by the object channel) to be rendered. The trajectory may include a series of "floor" positions (in the plane of a subgroup of loudspeakers assumed to be located on the floor of the playback environment, or in another horizontal plane of the playback environment), and a series of "above-floor" positions (each position determined by driving a subgroup of loudspeakers assumed to be located in at least one other horizontal plane of the playback environment). For example, examples of rendering object-based audio programs are described in PCT International Application No. PCT / US2001 / 028783, which was published on September 29, 2011 as International Publication No. WO 2011 / 119401 A2 and assigned to the assignee of the present application.
[0023] U.S. Provisional Patent Application No. 61 / 807,922 referenced above and U.S. Provisional Patent Application No. 61 / 832,397 referenced above describe object-based audio programs that are rendered to provide an immersive, personalizable perception of the program's audio content. The content may indicate the atmosphere (sounds produced in or at the viewing event) of a viewing event (e.g., a soccer or football game or another sporting event) and / or a commentary on the viewing event. The program's audio content may indicate a plurality of audio object channels (e.g., indicating user-selectable objects or groups of objects, and typically also a default group of objects to be rendered in the absence of a user's object selection) and at least one speaker channel bed. The speaker channel bed may be a conventional mix of speaker channels of a type that can be included in a conventional broadcast program that does not include object channels (e.g., a 5.1 channel mix).
[0024] U.S. Provisional Patent Applications Nos. 61 / 807,922 and 61 / 832,397, referenced above, describe object-related metadata delivered as part of an object-based audio program that provides mix interactivity (e.g., a substantial degree of mix interactivity) on the playback side, including by allowing an end user to select a mix of the program's audio content for rendering, rather than merely allowing playback of a premixed sound field. For example, a user can select from rendering options provided by the metadata of a typical embodiment of the program of the invention to select a subset of available object channels for rendering and, optionally, also the playback level of at least one audio object (sound source) indicated by the object channel to be rendered. The spatial location at which each selected sound source is rendered can be predetermined by metadata included in the program, but in some embodiments can be selected by the user (e.g., subject to predetermined rules or constraints). In some embodiments, metadata included with the program allows a user to select from a menu of rendering options (e.g., a small number of rendering options, such as a "home team crowd noise" object, a group of "home team crowd noise" and "home team commentary" objects, an "away team crowd noise" object, and a group of "away team crowd noise" and "away team commentary" objects). The menu can be presented to the user via a user interface of a controller, and the controller can be coupled to a set-top device (or other device) configured to decode and render (at least in part) an object-based program. Metadata included with the program can otherwise allow the user to select from a group of options as to which object(s) indicated by the object channel should be rendered, and as to how the rendered objects should be configured.
[0025] U.S. Provisional Patent Applications Nos. 61 / 807,922 and 61 / 832,397 describe object-based audio programs, which are encoded audio bitstreams that indicate at least some of a program's audio content (e.g., speaker channel soundbeds and at least some of the program's object channels) and object-related metadata. At least one additional bitstream or file may indicate some of the program's audio content (e.g., at least some of the object channels) and / or object-related metadata. In some embodiments, the object-related metadata provides a default mix of the object content and soundbed (speaker channel) content with default rendering parameters (e.g., default spatial positions of rendered objects). In some embodiments, the object-related metadata provides a set of selectable "preset" mixes of the object channel and speaker channel content, each preset mix having a predetermined set of rendering parameters (e.g., spatial positions of rendered objects). In some embodiments, the program's object-related metadata (or a preconfiguration of a playback or rendering system not indicated by metadata delivered with the program) provides constraints or conditions on the selectable mixes of the object channel and speaker channel content.
[0026] U.S. Provisional Patent Application Nos. 61 / 807,922 and 61 / 832,397 also describe object-based audio programs that include a set of bitstreams (sometimes referred to as "substreams") that are generated and transmitted in parallel. Multiple decoders can be used to decode them (e.g., if a program includes multiple E-AC-3 substreams, a playback system can utilize multiple E-AC-3 decoders to decode the substreams). Each substream can include a synchronization word (e.g., a time code) to allow the substreams to be synchronized or time-aligned with each other.
[0027] U.S. Provisional Application Nos. 61 / 807,922 and 61 / 832,397 also describe an object-based audio program that is or includes at least one AC-3 (or E-AC-3) bitstream and includes one or more data structures referred to as containers. Each container, including object channel content (and / or object-related metadata), is included in an auxiliary data field (e.g., Figure 1 or Figure 4 ) or in a "Skip Field" segment of the bitstream. Object-based audio programs are also described that are or include a Dolby E bitstream in which object channel content and object-related metadata (e.g., each container of a program that includes object channel content and / or object-related metadata) are included in bit locations of the Dolby E bitstream that typically carry no useful information.
[0028] U.S. Provisional Application No. 61 / 832,397 also describes an object-based audio program that includes metadata for at least one speaker channel group, at least one object channel, and a hierarchical graph (a hierarchical "mixing graph") indicating optional mixes of the speaker channels and object channels (e.g., all optional mixes). The mixing graph can indicate each rule that can be applied to the selection of a subset of speaker and object channels, indicating nodes (each node can indicate an optional channel or channel group, or a classification of optional channels or channel groups) and connections between nodes (e.g., control interfaces for the nodes and / or rules for selecting channels). The mixing graph can indicate necessary data (a "base" layer) and optional data (at least one "extended" layer), and where the mixing graph can be represented as a tree graph, the base layer can be a branch (or two or more branches) of the tree graph, and each extended layer can be an additional branch (or group of branches) of the tree graph.
[0029] U.S. Provisional Application Nos. 61 / 807,922 and 61 / 832,397 also teach that an object-based audio program can be decoded and its speaker channel content can be rendered by a legacy decoder and rendering system that is not configured to parse the program's object channels and object-related metadata. The same program can be rendered by a set-top device (or other decoding and rendering system) that is configured to parse the program's object channels and object-related metadata and render a mix of the speaker channels and object channel content indicated by the program. However, neither U.S. Provisional Application No. 61 / 807,922 nor U.S. Provisional Application No. 61 / 832,397 teaches or suggests how to generate a personalizable object-based audio program that can be rendered by a legacy decoding and rendering system (which is not configured to parse the program's object channels and object-related metadata) to provide a full-range audio experience (e.g., audio intended to be perceived as non-ambient sound from at least one discrete audio object, mixed with the ambient sound), but such that a decoding and rendering system configured to parse the program's object channels and object-related metadata can render a selected mix of the contents of at least one speaker channel and at least one object channel of the program (also providing the full-range audio experience), or such that it would wish to do so. Summary of the Invention
[0030] One class of embodiments of the present invention provides a personalizable object-based program that is compatible with legacy playback systems (which are not configured to parse the program's object channels and object-related metadata), in the sense that the legacy system can render the program's default set of speaker channels to provide a full-range audio experience (wherein "full-range audio experience" in this context means a sound mix represented by the audio content of only the default set of speaker channels, intended to be perceived as a full or complete mix of non-ambient sounds from at least one discrete audio object mixed with other sounds represented by the default set of speaker channels. The other sounds can be ambient sounds.), wherein the same program can be decoded and rendered by a non-legacy playback system (configured to parse the program's object channels and metadata) to render at least one selected preset mix of the content of at least one speaker channel of the program and the non-ambient content of at least one object channel of the program (which can also provide the full-range audio experience). In this document, such a default set of speaker channels (which can be rendered by a legacy system) is sometimes referred to as a speaker channel "sound bed", although this term is not intended to imply that the sound bed must be mixed with additional audio content to provide a full-range audio experience. In fact, in typical embodiments of the present invention, the sound bed does not have to be mixed with additional audio content to provide a full-range audio experience, and the sound bed can be decoded and rendered by a legacy system to provide a full-range audio experience without being mixed with additional audio content. In other embodiments, the object-based audio program of the present invention includes a speaker channel sound bed that represents only non-ambient content (e.g., a mix of different types of non-ambient content) and is capable of being rendered by a legacy system (e.g., to provide a full-range audio experience), and a playback system configured to parse the program's object channels and metadata can render at least one selected preset mix (which can, but does not necessarily, provide a full-range audio experience) of content (e.g., non-ambient and / or ambient content) of at least one speaker channel of the program and at least one object channel of the program.
[0031] Typical embodiments in this class generate, deliver, and / or render an object-based program that includes a base layer (e.g., a 5.1-channel sound bed) that includes a speaker channel sound bed representing all the contents of a default audio program (sometimes referred to as a "default" mix), where the default audio program includes a full set of audio elements (e.g., ambient content mixed with non-ambient content) that provides a full-range audio experience when played. Legacy playback systems (which are not capable of decoding or rendering object-based audio) can decode and render the default mix. An example of ambient content for a default audio program is crowd noise (captured at a sporting event or other viewing event), and an example of non-ambient content for a default audio program includes commentary and / or announcement feeds (related to a sporting event or other viewing event). The program also includes an extension layer (which can be ignored by the legacy playback system) that can be utilized by a suitably configured (non-legacy) playback system to select and render any of a plurality of predetermined mixes of the audio content of the extension layer (or the extension layer and the base layer). Extension layers typically include optional replacement speaker channel groups that allow personalized representation of alternative content (e.g., only the primary ambient content, rather than the mix of ambient and non-ambient content provided by the base layer) and optional object channel groups (e.g., object channels representing the primary non-ambient content and the alternative non-ambient content).
[0032] Providing a base layer and at least one extension layer in a program allows more flexibility to program generation facilities (eg, broadcast head ends) and playback systems (which may be or include set-top boxes or "STBs").
[0033] In some embodiments, the present invention is a method for generating an object-based audio program indicative of audio content (e.g., captured audio content), the audio content including first non-ambient content, second non-ambient content different from the first non-ambient content, and third content different from the first non-ambient content and the second non-ambient content (the third content may be ambient content, but in some cases may also be or include non-ambient content), the method comprising the steps of:
[0034] determining an object channel group comprising N object channels, wherein a first subset of the object channel group indicates the first non-ambient content, the first subset comprising M object channels in the object channel group, each of N and M being an integer greater than zero, and M being equal to or less than N;
[0035] determining a speaker channel bed indicating a default mix of audio content (e.g., a default mix of ambient content and non-ambient content), wherein an object-based speaker channel subset comprising M speaker channels in the bed indicates the second non-ambient content, or a mix of at least some audio content of the default mix and the second non-ambient content;
[0036] determining a set of M replacement speaker channels, wherein each replacement speaker channel in the set of M replacement speaker channels indicates some but not all content of a corresponding speaker channel in the object-based subset of speaker channels;
[0037] generating metadata (sometimes referred to herein as object-related metadata) indicating at least one selectable predetermined alternative mix of content of at least one of the object channels and content of predetermined ones of the speaker channels and / or the replacement speaker channels of the sound bed, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix indicating at least some content of the sound bed and the first non-ambient content but not the second non-ambient content; and
[0038] generating the object-based audio program comprising the speaker channel sound bed, the set of M replacement speaker channels, the object channel groups, and the metadata, such that the speaker channel sound bed is renderable without use of the metadata to provide sound that is perceived as the default mix, and the replacement mix is renderable in response to at least some of the metadata to provide sound that is perceived as a mix that includes at least some of the audio content of the sound bed and the first non-ambient content but not the second non-ambient content.
[0039] Typically, the metadata (object-related metadata) of a program includes (or contains) optional content metadata representing a set of optional experience articulations. Each experience articulation is an optional predetermined ("preset") mix of the audio content of the program (e.g., a mix of the content of at least one object channel and at least one speaker channel in a sound bed, or a mix of the content of at least one object channel and at least one replacement speaker channel, or a mix of at least one object channel and at least one speaker channel in a sound bed and at least one replacement speaker channel). Each preset mix has a predetermined set of rendering parameters (e.g., the spatial position of the rendered object). The user interface of the playback system can present the preset mixes as a limited menu or palette of available mixes.
[0040] In other embodiments, the present invention is a method of rendering audio content determined by an object-based audio program, wherein the program indicates a speaker channel bed, a set of M alternative speaker channels, an object channel group, and metadata, wherein the object channel group includes N object channels, a first subset of the object channel group indicates first non-ambient content, the first subset includes the M object channels in the object channel group, each of N and M is an integer greater than zero, and M is equal to or less than N,
[0041] the speaker channel sound bed indicates a default mix of audio content including second non-ambient content different from the first non-ambient content, wherein an object-based speaker channel subset comprising M speaker channels in the sound bed indicates a mix of the second non-ambient content, or at least some audio content of the default mix, and the second non-ambient content,
[0042] each replacement speaker channel in the set of M replacement speaker channels indicates some, but not all, content of a corresponding speaker channel of the object-based subset of speaker channels, and
[0043] The metadata indicates at least one selectable predetermined alternative mix of the content of at least one of the object channels and the content of predetermined ones of the speaker channels of the sound bed and / or the replacement speaker channels, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix including at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content, the method comprising the steps of:
[0044] (a) providing the object-based audio program to an audio processing unit; and
[0045] (b) parsing the speaker channel sound bed in the audio processing unit and rendering the default mix in response to the speaker channel sound bed without using the metadata.
[0046] In some cases, the audio processing unit is a legacy playback system (or other audio data processing system) that is not configured to parse the program's object channels or metadata. In cases where the audio processing unit is configured to parse the program's object channels, replacement channels, and metadata (and speaker channel sound beds), the method may include the following steps:
[0047] (c) rendering, in the audio processing unit, the alternative mix using at least some of the metadata, including rendering by selecting and mixing content of the first subset of the object channel group and at least one of the alternative speaker channels in response to at least some of the metadata.
[0048] In some embodiments, step (c) comprises driving a speaker to provide sound perceived as a mix comprising at least some of the audio content of the sound bed and the first non-ambient content but not the second non-ambient content.
[0049] Another aspect of the present invention is an audio processing unit (APU) configured to perform any embodiment of the method of the present invention. In another class of embodiments, the present invention is an APU comprising a buffer memory (buffer) that stores (e.g., in a non-transitory manner) at least one frame or other segment of an object-based audio program (including audio content of speaker channel sound beds and object channels and object-related metadata) generated by any embodiment of the method of the present invention. Examples of APUs include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems (pre-processors), post-processing systems (post-processors), audio bitstream processing systems, and combinations of such elements.
[0050] Aspects of the present invention include systems or devices configured to (e.g., programmed to) perform any embodiment of the method of the present invention and computer-readable media (e.g., disks) storing (e.g., in a non-transitory manner) code for implementing any embodiment of the method of the present invention or its steps. For example, the system of the present invention can be or include a programmable general-purpose processor, digital signal processor, or microprocessor that is programmed and / or otherwise configured using software or firmware to perform any of a variety of operations on data, including embodiments of the method of the present invention or its steps. Such a general-purpose processor can be or include a computer system that includes an input device, a memory, and processing circuitry that is programmed (and / or otherwise configured) to perform an embodiment of the method of the present invention (or its steps) in response to data set thereto. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a diagram of an AC-3 frame including the segments into which it is divided.
[0052] Figure 2 is a diagram of a synchronization information (SI) segment of an AC-3 frame including the segments into which it is divided.
[0053] Figure 3 is a diagram of the Bit Stream Information (BSI) segment of an AC-3 frame including the segments into which it is divided.
[0054] Figure 4 is a diagram of an E-AC-3 frame including the segments into which it is divided.
[0055] Figure 5 is a block diagram of an embodiment of a system in which one or more elements of the system may be configured in accordance with embodiments of the present invention.
[0056] Figure 6 is a block diagram of a playback system that may be implemented to perform an embodiment of the method of the present invention.
[0057] Figure 7 is a block diagram of a playback system that may be configured to perform an embodiment of the method of the present invention.
[0058] Figure 8 is a block diagram of a broadcast system configured to generate object-based audio programs (and corresponding video programs) according to an embodiment of the present invention.
[0059] Figure 9 is a diagram of the relationship between object channels of an embodiment of a program of the present invention, indicating which subgroups of the object channels are user-selectable.
[0060] Figure 10 is a block diagram of a system that can be implemented to perform an embodiment of the method of the present invention.
[0061] Figure 11 is a diagram of the content of an object-based audio program generated according to an embodiment of the present invention.
[0062] Figure 12 is a block diagram of an embodiment of a system configured to perform an embodiment of the method of the present invention.
[0063] Symbols and terminology
[0064] Throughout this disclosure, including the claims, the expression "non-ambient sound" means a sound (e.g., a commentary or other monologue, or a conversation) that is perceived or perceptible as emanating from a discrete audio object (or a plurality of audio objects all located therein) located at or within an angular position that is well localizable relative to a listener (i.e., an angular position that subtends a solid angle of no more than approximately 3 steradians relative to the listener, where the entire range centered about the listener's position subtends 4π steradians relative to the listener). Herein, "ambient sound" means a sound that is not a non-ambient sound (e.g., crowd noise perceived by a member of a crowd). Thus, ambient sound herein means a sound that is perceived or perceptible as emanating from a large (or otherwise poorly localizable) angular position relative to the listener.
[0065] Similarly, “non-ambient audio content” (or “non-ambient content”) herein means audio content that is perceived when rendered as sounds emanating from discrete audio objects (or a large number of audio objects all located at or within) that are well-localized relative to a listener (i.e., angular positions subtending a solid angle of no more than approximately 3 steradian degrees relative to the listener), and “ambient audio content” (or “ambient content”) means audio content that is not “non-ambient audio content” (or “non-ambient content”) and that is perceived when rendered as ambient sound.
[0066] Throughout this disclosure, including the claims, expressions that “perform an operation on” a signal or data (e.g., filter, scale, transform, or apply a gain to the signal or data) are used broadly to mean performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., a version of the signal that has undergone preliminary filtering or preprocessing before the operation is performed on the signal).
[0067] Throughout this disclosure, including the claims, the expression "system" is used in a broad sense to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M inputs and the other X-M inputs are received from external sources) may also be referred to as a decoder system.
[0068] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., using software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors that are programmed and / or otherwise configured to perform pipeline processing of audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.
[0069] Throughout this disclosure, including the claims, the expression "audio video receiver" (or "AVR") refers to a receiver in a class of consumer electronic equipment used to control the playback of audio and video content in, for example, a home theater.
[0070] Throughout this disclosure, including the claims, the expression "soundbar" refers to a device that is a piece of consumer electronic equipment (typically installed in a home theater system) and that includes at least one speaker (typically, at least two speakers) and a subsystem for rendering audio for playback for each included speaker (or for playback for each included speaker and at least one additional speaker external to the soundbar).
[0071] Throughout this disclosure, including the claims, the expressions "audio processor" and "audio processing unit" are used interchangeably and broadly refer to a system configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools).
[0072] Throughout this disclosure, including the claims, the expression "metadata" (e.g., as in the expression "processing state metadata") refers to data that is separate and distinct from the corresponding audio data (the audio content of the bitstream that also includes the metadata). The metadata is associated with the audio data and indicates at least one feature or characteristic of the audio data (e.g., which type(s) of processing has been or should be performed on the audio data or the trajectory of an object indicated by the audio data). The association of the metadata with the audio data is time-synchronized. Thus, the current (most recently received or updated) metadata can indicate that the corresponding audio data simultaneously has the indicated feature and / or includes the results of audio data processing of the indicated type.
[0073] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean either a direct or indirect connection. Thus, if a first device couples to a second device, that connection may be through a direct connection or through an indirect connection via other devices and connections.
[0074] Throughout this disclosure, including the claims, the following expressions have the following definitions:
[0075] The terms speaker and loudspeaker are used synonymously to refer to any sound-producing transducer. This definition includes loudspeakers implemented as multiple transducers (e.g., a woofer and a tweeter);
[0076] Loudspeaker feed: An audio signal applied directly to a loudspeaker, or to an amplifier and loudspeaker connected in series;
[0077] Channel (or "audio channel"): a monophonic audio signal. Such a signal can typically be rendered in a manner that is equivalent to applying the signal directly to a loudspeaker at a desired or nominal position. The desired position can be static, as is typically the case with physical loudspeakers, or dynamic;
[0078] Audio program: a set of one or more audio channels (at least one loudspeaker channel and / or at least one object channel) and optionally also associated metadata (e.g. metadata describing the desired spatial audio representation);
[0079] Loudspeaker channel (or "loudspeaker feed channel"): an audio channel associated with a named loudspeaker (at a desired or nominal location), or associated with a named loudspeaker zone within a defined loudspeaker configuration. Loudspeaker channels may be rendered in a manner that is equivalent to applying the audio signal directly to the named loudspeaker (at a desired or nominal location) or loudspeakers within a named loudspeaker zone;
[0080] Object channel: An audio channel that indicates the sound emitted by an audio source (sometimes referred to as an audio "object"). Typically, an object channel determines a parameterized audio source description (e.g., metadata indicating the parameterized audio source description is included in or provided using the object channel). The source description may determine the sound emitted by the source (as a function of time), the apparent position of the source as a function of time (e.g., 3D spatial coordinates), and optionally at least one additional parameter characterizing the source (e.g., apparent source size or width);
[0081] Object-based audio program: an audio program comprising a set of one or more object channels (and optionally also at least one speaker channel) and, optionally, associated metadata (e.g., metadata indicating the trajectory of an audio object emitting the sound indicated by the object channel, or otherwise indicating a desired spatial audio representation of the sound indicated by the object channel, or metadata indicating the identity of at least one audio object that is the source of the sound indicated by the object channel); and
[0082] Rendering: The process of converting an audio program into one or more speaker feeds, or converting an audio program into one or more speaker feeds and converting the speaker feeds into sound using one or more loudspeakers (in the latter case, rendering is sometimes referred to herein as rendering "through" the loudspeakers). An audio channel can be rendered very generally ("at" the desired location) by applying the signal directly to physical loudspeakers at the desired location, or one or more audio channels can be rendered using one of a variety of virtualization techniques designed to be substantially equivalent (to the listener) to such a very general rendering. In the latter case, each audio channel can be converted into one or more speaker feeds to be applied to loudspeakers at known locations, typically different from the desired location, so that the sound emitted by the loudspeakers in response to the feeds will be perceived as emanating from the desired location. Examples of such virtualization techniques include binaural rendering through headphones (e.g., using Dolby headphone processing, which emulates up to 7.1 channels of surround sound for the headphone wearer) and wave field synthesis. DETAILED DESCRIPTION
[0083] Figure 51 is a block diagram of an example of an audio processing chain (audio data processing system), one or more elements of which may be configured according to embodiments of the present invention. The system includes the following elements coupled together as shown: a capture unit 1, a generation unit 3 (which includes an encoding subsystem), a delivery subsystem 5, a decoder 7, an object processing subsystem 9, a controller 10, and a rendering subsystem 11. In variations of the system shown, one or more elements are omitted, or additional audio data processing units are included. Typically, elements 7, 9, 10, and 11 are or are included in a playback system (e.g., an end-user's home theater system).
[0084] Capture unit 1 is typically configured to generate and output PCM (time-domain) samples comprising audio content. The samples may represent multiple audio streams captured by a microphone (e.g., at a sporting or other viewing event). Generation unit 3, typically operated by a broadcaster, is configured to accept the PCM samples as input and output an object-based audio program indicative of the audio content. The program typically comprises or includes an encoded (e.g., compressed) audio bitstream indicative of at least some audio content (sometimes referred to herein as a "main mix") and optionally also comprises or includes at least one additional bitstream or file indicative of some audio content (sometimes referred to herein as a "submix"). Data indicative of the encoded bitstream of audio content (and each generated submix, if any) is sometimes referred to herein as "audio data." If the encoding subsystem of generation unit 3 is configured in accordance with exemplary embodiments of the present invention, the object-based audio program output from unit 3 indicates (i.e., includes) multiple speaker channels of audio data (both the speaker channel "soundbed" and alternate speaker channels), multiple object channels of audio data, and object-related metadata. A program may include a main mix, which in turn includes audio content indicating speaker channel beds, alternate speaker channel audio content, audio content indicating at least one user-selectable object channel (and optionally at least one other object channel), and metadata (including object-related metadata associated with each object channel). A program may also include at least one secondary mix, which includes audio content and / or object-related metadata indicating at least one other object channel (e.g., at least one user-selectable object channel). The program's object-related metadata may include durable metadata (described below). A program (e.g., its main mix) may indicate one or more groups of speaker channels. For example, a main mix may indicate two or more groups of speaker channels (e.g., a 5.1-channel neutral crowd noise bed, a 2.0-channel group indicating alternate speaker channels for home team crowd noise, and a 2.0-channel group indicating alternate speaker channels for away team crowd noise), including at least one user-selectable alternate speaker channel group (which may be selected using the same user interface used for user selection of object channel content or configuration) and a speaker channel bed (which may be rendered in the absence of user selection of other content of the program). The sound bed (which may be referred to as the default sound bed) may be determined by data indicating a configuration (e.g., an initial configuration) of the playback system's speaker set, and optionally, the user may select other audio content of the program to be rendered in place of the default sound bed.
[0085] The program's metadata may indicate at least one (and typically more than one) optional predetermined mix of the content of at least one object channel with the content of predetermined speaker channels and / or alternative speaker channels in the program's sound bed, and may include rendering parameters for each of the mixes. At least one such mix may be an alternative mix indicating that at least some of the sound bed's audio content is mixed with first non-ambient content (indicated by at least one object channel included in the mix) rather than second non-ambient content (indicated by at least one speaker channel in the sound bed).
[0086] Figure 5 The delivery subsystem 5 is configured to store and / or transmit (eg, broadcast) the program generated by the unit 3 (eg, its master mix and each submix, if any submixes are generated).
[0087] In some embodiments, the subsystem 5 implements delivery of an object-based audio program in which the program's audio objects (and at least some corresponding object-related metadata) and speaker channels are transmitted via a broadcast system (in the form of a master mix of the program indicated by the broadcast audio bitstream), and at least some metadata for the program (e.g., object-related metadata indicating constraints on the rendering or mixing of the program's object channels) and / or at least one object channel of the program are delivered in another manner (as a "submix" of the master mix) (e.g., the submix is sent to a particular end user via an Internet Protocol or "IP" network). Alternatively, the end user's decoding and / or rendering system is preconfigured with at least some object-related metadata (e.g., metadata indicating constraints on the rendering or mixing of audio objects of an embodiment of the object-based audio program of the present invention), and such object-related metadata is not broadcast or otherwise delivered (by the subsystem 5) with the corresponding object channels (in the master mix or submix of the object-based audio program).
[0088] In some embodiments, timing and synchronization of portions or elements of an object-based audio program delivered via separate paths (e.g., a main mix broadcast via a broadcast system, and associated metadata sent as submixes over an IP network) is provided via a synchronization word (e.g., a time code) that is sent across all delivery paths (e.g., in the main mix and each corresponding submix).
[0089] Refer again Figure 5, decoder 7 accepts (receives or reads) a program (or at least one bitstream or other element of a program) delivered by delivery subsystem 5 and decodes the program (or each received element thereof). In some embodiments of the invention, the program includes a main mix (an encoded bitstream, such as an AC-3 or E-AC-3 encoded bitstream) and at least one submix of the main mix, and decoder 7 receives and decodes the main mix (and optionally also at least one submix). Optionally, at least one submix of the program that does not need to be decoded (e.g., an object channel) is delivered directly to object processing subsystem 9 by subsystem 5. If decoder 7 is configured according to an exemplary embodiment of the invention, then in exemplary operation the output of decoder 7 includes the following:
[0090] a stream of audio samples indicating the speaker channel soundbed for the program (and typically also indicating alternate speaker channels for the program); and
[0091] A stream of audio samples indicating object channels of a program (e.g., user-selectable audio object channels) and a stream of corresponding object-related metadata.
[0092] The object processing subsystem 9 is coupled to receive (from the decoder 7) the decoded loudspeaker channels, object channels, and object-related metadata of the delivered program, and optionally also receive at least one submix of the program (indicative of at least one other object channel). For example, the subsystem 9 may receive (from the decoder 7) audio samples of the loudspeaker channels of the program and at least one object channel of the program, as well as the object-related metadata of the program, and may also receive (from the delivery subsystem 5) audio samples of at least one other object channel of the program (which has not yet been decoded in the decoder 7).
[0093] Subsystem 9 is coupled and configured to output a selected subset of the full set of object channels indicated by the program, and corresponding object-related metadata, to rendering subsystem 11. Subsystem 9 is also typically configured to pass the decoded speaker channels from decoder 7 unchanged (to subsystem 11), and may be configured to process at least some of the object channels (and / or metadata) asserted thereto to generate its asserted object channels and metadata to subsystem 11.
[0094] The object channel selection performed by subsystem 9 is typically determined by user selection (as indicated by control data set from controller 10 to subsystem 9) and / or rules (e.g., indicating conditions and / or constraints) that subsystem 9 is programmed or otherwise configured to implement. Such rules may be determined by the program's object-related metadata and / or by other data (e.g., data indicating the capabilities and organization of the playback system's speaker array) set to subsystem 9 (e.g., from controller 10 or another external source) and / or by preconfiguration (e.g., programming) of subsystem 9. In some embodiments, controller 10 (via a user interface implemented by controller 10) provides the user (e.g., displays on a touch screen) a menu or palette of selectable "preset" mixes of speaker channel content (i.e., the content of the sound bed speaker channels and / or replacement speaker channels) and object channel content (objects). The selectable preset mixes may be determined by the program's object-related metadata and, typically, also by rules implemented by subsystem 9 (e.g., rules that subsystem 9 has been preconfigured to implement). The user selects from the selectable mixes by inputting commands to the controller 10 (e.g., by activating the touch screen of the controller 10), and in response, the controller 10 sets corresponding control data to the subsystem 9 to enable rendering of the corresponding content according to the present invention.
[0095] Figure 5 Rendering subsystem 11 is configured to render audio content determined by the output of subsystem 9 for playback by speakers (not shown) of the playback system. Subsystem 11 is configured to map audio objects (e.g., default objects and / or user-selected objects that have been selected as a result of user interaction using controller 10) determined by the object channels selected by object processing subsystem 9 to available speaker channels using rendering parameters (e.g., user-selected values and / or default values for spatial position and level) output from subsystem 9 and associated with each selected object. At least some of the rendering parameters are also determined by object-related metadata output from subsystem 9. Rendering system 11 also receives the speaker channels passed by subsystem 9. Typically, subsystem 11 is an intelligent mixer and is configured to determine speaker feeds for available speakers, including by mapping one or more selected (e.g., default-selected) objects to each of a plurality of individual speaker channels and mixing the objects with speaker channel content indicated by each corresponding speaker channel of the program (e.g., each speaker channel in the speaker channel sound bed of the program).
[0096] Figure 12 is a block diagram of an embodiment of another system configured to perform an embodiment of the method of the present invention. Figure 12 The capture unit 1, the generation unit 3 and the delivery subsystem 5 are Figure 5The same numbered components of the system are the same. ( Figure 12 ) units 1 and 3 are operable to generate an object-based audio program according to at least one embodiment of the present invention, and ( Figure 12 The subsystem 5 is configured to deliver such programs to Figure 12 Playback system 111.
[0097] and Figure 5 Unlike the playback system of the decoder 107 (including the decoder 7, object processing subsystem 9, controller 10, and rendering subsystem 11), the playback system 111 is not configured to parse the object channels or object-related metadata of the program. The decoder 107 of the playback subsystem 111 is configured to parse the speaker channel sound bed of the program delivered by the subsystem 5, and the rendering subsystem 109 of the subsystem 111 is coupled and configured to render a default mix (indicated by the speaker channel sound bed) in response to the sound bed (without using the object-related metadata of the program). The decoder 107 may include a buffer 7A that stores (e.g., in a non-transitory manner) at least one frame or other segment of the object-based audio program (including the speaker channel sound bed, audio content of the replacement speaker channels and object channels, and object-related metadata) delivered from the subsystem 5 to the decoder 107.
[0098] In contrast, Figure 5 Typical implementations of the playback system (including decoder 7, object processing subsystem 9, controller 10, and rendering subsystem 11) are configured to parse object channels, object-related metadata, and replacement speaker channels (and speaker channel sound beds indicating a default mix) of object-based programs delivered thereto. In some such implementations, Figure 5 The playback system is configured to render an alternative mix (determined by at least one object channel and at least one alternative speaker channel, and typically also at least one speaker channel sound bed, of the program) in response to at least some object-related metadata, including selecting the alternative mix using at least some of the object-related metadata. In some such implementations, Figure 5 The playback system is capable of operating in a mode in which it renders such an alternative mix in response to the programme's object channel and speaker channel content and metadata, and is also capable of operating in a second mode (which may be triggered by metadata in the programme) in which the decoder 7 parses the programme's speaker channel sound bed, the speaker channel sound bed is set to the rendering subsystem 11, and the rendering subsystem 11 operates (without using the programme's object-related metadata) in response to that sound bed to render a default mix (indicated by that sound bed).
[0099] In one class of embodiments, the present invention is a method for generating an object-based audio program indicative of audio content (e.g., captured audio content), the audio content including first non-ambient content, second non-ambient content different from the first non-ambient content, and third content different from the first non-ambient content and the second non-ambient content, the method comprising the steps of:
[0100] determining an object channel group comprising N object channels, wherein a first subset of the object channel group indicates first non-ambient content, the first subset comprising M object channels in the object channel group, each of N and M being an integer greater than zero, and M being equal to or less than N;
[0101] determining a speaker channel sound bed indicative of a default mix of audio content, wherein an object-based speaker channel subset comprising the M speaker channels in the sound bed indicates second non-ambient content, or a mix of at least some audio content of the default mix and the second non-ambient content;
[0102] determining a set of M replacement speaker channels, wherein each replacement speaker channel in the set of M replacement speaker channels indicates some but not all content of a corresponding speaker channel in the object-based subset of speaker channels;
[0103] generating metadata indicating at least one selectable predetermined alternative mix of content of at least one of the object channels and content of predetermined ones of the speaker channels of the sound bed and / or replacement speaker channels, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix indicating at least some of the audio content of the sound bed and first non-ambient content but not second non-ambient content; and
[0104] An object-based audio program is generated comprising a speaker channel bed, a set of M replacement speaker channels, an object channel group, and metadata such that:
[0105] Speaker channel sound beds are not used with metadata (e.g. Figure 12 The playback system 111 of the system, or the mode of operation in which the decoder 7 analyzes the speaker channel sound bed of the program Figure 5 a playback system in which a speaker channel bed is set to the rendering subsystem 111 and, using the program's object-related metadata, the rendering subsystem 111 operates in response to the bed to render a default mix indicated by the bed to provide sound that is perceptible as the default mix, and
[0106] The alternative mix is responsive to at least some metadata (e.g., via Figure 5A playback system comprising a decoder 7, an object processing subsystem 9, a controller 10 and a rendering subsystem 11, which can be rendered using object-related metadata of a program delivered to the decoder 7 to provide sound that can be perceived as comprising a mix of at least some of the audio content of the sound bed and the first non-ambient content rather than the second non-ambient content.
[0107] In some such embodiments, an object-based audio program is generated such that an alternative mix is renderable (e.g., via Figure 5 ) to provide sound that is perceptible as a mix including the first non-ambient content but not the second non-ambient content, such that the first non-ambient content is perceptible as emanating from a source whose size and position are determined by the subgroup of metadata corresponding to the first subgroup of the object channel group.
[0108] The replacement mix may indicate content of at least one of the speaker channels of the sound bed and the replacement speaker channels and the first subset of the object channel group other than the object-based speaker channel subset of the sound bed. This is accomplished by rendering the replacement mix with the non-ambient content (or a mix of ambient and non-ambient content) of the object-based speaker channel subset of the replacement sound bed, the non-ambient content of the first subset of the object channel group (and the content of the replacement speaker channels, which is typically ambient content or a mix of ambient and non-ambient content).
[0109] In some embodiments, the metadata for a program contains (or includes) optional content metadata indicating a set of optional experience articulations. Each experience articulation is an optional predetermined ("preset") mix of the program's audio content (e.g., a mix of the content of at least one object channel and at least one speaker channel in a sound bed, or a mix of the content of at least one object channel and at least one replacement speaker channel, or a mix of at least one object channel and at least one speaker channel in a sound bed and at least one replacement speaker channel). Each preset mix has a predetermined set of rendering parameters (e.g., spatial positions of rendered objects), which are also typically indicated by the metadata. The preset mixes can be selected by a user interface of the playback system (e.g., by Figure 5 Controller 10 or Figure 6 The user interface implemented by the controller 23 is presented as a limited menu or option panel of available mixes.
[0110] In some embodiments, the program's metadata includes default mix metadata indicating a base layer so that playback by a playback system (e.g., Figure 5 ) can select the default mix (instead of another preset mix) and render the base layer. It is expected that legacy playback systems (e.g., Figure 12The playback system 111) can also render the base layer (and thus the default mix) without using any default mix metadata.
[0111] An object-based audio program (e.g., an encoded bitstream indicative of such a program) generated according to an exemplary embodiment of the present invention is a personalizable object-based audio program (e.g., an encoded bitstream indicative of such a program). According to an exemplary embodiment, audio objects and other audio content are encoded to allow for a selectable full-range audio experience (wherein, in this context, "full-range audio experience" means audio intended to be perceived as ambient sounds mixed with non-ambient sounds (e.g., commentary or dialogue) from at least one discrete audio object. To allow personalization (i.e., selection of a desired mix of audio content), a speaker channel soundbed (e.g., a soundbed indicative of ambient content mixed with non-ambient object channel content) and at least one replacement speaker channel and at least one object channel (typically, multiple object channels) are encoded as distinct elements within the encoded bitstream.
[0112] In some embodiments, a personalizable object-based audio program includes (and allows selection of any of) at least two selectable "preset" mixes of object channel content and / or speaker channel content and a default mix of ambient content and non-ambient content determined by the included speaker channel sound beds. Each selectable mix includes different audio content and thereby provides a different experience to the listener when rendered and reproduced. For example, where the program indicates audio captured at a football game, one preset mix may indicate an atmosphere / effects mix for the home team crowd and another preset mix may indicate an atmosphere / effects mix for the away team crowd. Typically, the default mix and the multiple alternative preset mixes are encoded into a single bitstream. Also optionally, in addition to the speaker channel sound bed that determines the default mix, additional speaker channels (e.g., left and right (stereo) speaker channel pairs) indicating optional audio content (e.g., submixes) are included in the bitstream so that the speaker channels in the additional speaker channels can be selected and mixed with other content of the program (e.g., speaker channel content) in a playback system (e.g., a set-top box, which is sometimes referred to herein as an "STB").
[0113] In other embodiments, the present invention is a method of rendering audio content determined by an object-based audio program, wherein the program indicates a speaker channel sound bed, a set of M replacement speaker channels, an object channel group, and metadata, wherein the object channel group includes N object channels, a first subset of the object channel group indicates first non-ambient content, the first subset includes the M object channels of the object channel group, each of N and M is an integer greater than zero, and M is equal to or less than N,
[0114] the speaker channel sound bed indicates a default mix of audio content including second non-ambient content different from the first non-ambient content, wherein the object-based speaker channel subset comprising the M speaker channels in the sound bed indicates the second non-ambient content, or a mix of at least some audio content of the default mix and the second non-ambient content,
[0115] Each replacement speaker channel in the set of M replacement speaker channels indicates some but not all of the content of a corresponding speaker channel in the object-based subset of speaker channels, and
[0116] metadata indicating at least one selectable predetermined alternative mix of content of at least one object channel with content of predetermined ones of the speaker channels of the sound bed and / or replacement speaker channels, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix comprising at least some of the audio content of the sound bed with first non-ambient content but not second non-ambient content, the method comprising the steps of:
[0117] (a) providing an object-based audio program to an audio processing unit (e.g., operating in a mode where the decoder 7 parses the speaker channel sound bed of the program) Figure 12 Playback system 111 or Figure 5 a playback system of the program, the speaker channel sound bed is set to the rendering subsystem 11, and the rendering subsystem 11 renders a default mix indicated by the sound bed in response to the sound bed without using the program's object-related metadata); and
[0118] (b) In the audio processing unit, the speaker channel sound bed is parsed and a default mix is rendered in response to the speaker channel sound bed without using metadata.
[0119] In some cases, the audio processing unit is a legacy playback system (or other audio data processing system) that is not configured to parse the program's object channels or metadata. Figure 5 In the case of an implementation of a playback system comprising a decoder 7, an object processing subsystem 9, a controller 10 and a rendering subsystem 11 configured to render a selected mix of the program's object channels, sound bed loudspeaker channel content and replacement loudspeaker channel content using object-related metadata of the program delivered to the decoder 7), the method may comprise the steps of:
[0120] (c) in the audio processing unit, rendering the alternative mix using at least some of the metadata, including rendering by selecting and mixing the contents of the first subset of the object channel group and at least one of the alternative speaker channels in response to at least some of the metadata (e.g., this step may be performed by Figure 6 Subsystem 22 and subsystem 24 of the system or by Figure 5 playback system).
[0121] In some embodiments, step (c) comprises the step of driving a speaker to provide sound perceivable as comprising a mix of the at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content.
[0122] In some embodiments, step (c) comprises the following steps:
[0123] (d) selecting a first subset of the object channel group, selecting at least one speaker channel in the speaker channel bed other than the speaker channels in the object-based subset of speaker channels, and selecting the at least one replacement speaker channel in response to the at least some metadata; and
[0124] (e) Mixing the first subset of the group of subject channels and the content of each speaker channel selected in step (d) to determine an alternative mix.
[0125] Step (d) can be carried out by, for example Figure 6 Subsystem 22 of the system or Figure 5 Step (e) may be performed by, for example, a subsystem 9 of the playback system. Figure 6 Subsystem 24 of the system or Figure 5 The playback system subsystem 11 is executed.
[0126] In some embodiments, the method of the present invention generates (or delivers or renders) a personalizable object-based audio program as a bitstream comprising data indicating the following layers:
[0127] A base layer (e.g., a 5.1 channel sound bed) including a speaker channel sound bed indicating all contents of a default audio program (e.g., a default mix of ambient and non-ambient content);
[0128] At least one object channel indicating optional audio content to be rendered (each object channel being an element of an extension layer);
[0129] at least one replacement loudspeaker channel (each replacement loudspeaker channel being an element of an extension layer) which is capable (by means of a suitably configured playback system, e.g. Figure 5 or Figure 6In an embodiment of a playback system of a program, a speaker channel is selected to replace one or more corresponding channels of a base layer, thereby determining a modified base layer including each original (non-replaced) channel of the base layer that is not replaced and each selected replacement speaker channel. The modified base layer may be rendered or may be mixed with the content of at least one of the object channels and then rendered. For example, when the replacement speaker channels include a center channel indicating only ambiance (to replace the center channel of the base layer indicating non-ambient content (e.g., commentary or dialogue) mixed with ambient content), the modified base layer including such replacement speaker channels may be mixed with the non-ambient content of at least one object channel of the program;
[0130] further optionally, at least one alternative speaker channel group (each alternative speaker channel being an element of an extension layer) indicating at least one audio content mix (e.g., each alternative speaker channel group may indicate a different multi-channel ambiance / effects mix), wherein each said alternative speaker channel group is selectable (by a suitably configured playback system) to replace one or more corresponding channels of the base layer; and
[0131] Metadata indicating at least one selectable experience definition (and typically more than one selectable experience definition). Each experience definition is a selectable predetermined ("preset") mix of the audio content of the program (e.g., a mix of the content of at least one object with the content of a speaker channel), each preset mix having a predetermined set of rendering parameters (e.g., spatial positions of rendered objects).
[0132] In some embodiments, the metadata includes default program metadata indicating a base layer (e.g., to enable selection of a default audio program and rendering of the base layer). Typically, the metadata does not include such default program metadata, but includes metadata indicating at least one optional predetermined alternative mix of the content of at least one object channel with predetermined ones of the speaker channels of the sound bed and / or replacement speaker channels (alternative mix metadata), wherein the alternative mix metadata includes rendering parameters for each of the alternative mixes.
[0133] Typically, the metadata indicates selectable preset mixes (e.g., all selectable preset mixes) for the program's speaker channels and object channels. Optionally, the metadata is or includes metadata indicating a layered mix graph indicating selectable mixes (e.g., all selectable mixes) for the program's speaker channels and object channels.
[0134] In one class of embodiments, an encoded bitstream indicating a program of the present invention comprises: a base layer including a speaker channel soundbed indicating a default mix (e.g., a default 5.1 speaker channel mix mixed with ambient content and non-ambient content), metadata, and optional extension channels (at least one object channel and at least one alternate speaker channel). Typically, the base layer includes a center channel indicating non-ambient content (e.g., commentary or dialogue indicated by object channels also included in the program) mixed to ambient sounds (e.g., crowd noise). A decoder (or other element of a playback system) can use the metadata sent with the bitstream to select an alternative "preset" mix (e.g., a mix achieved by discarding (ignoring) the center channel of the default mix and replacing the discarded center channel with an alternate speaker channel to determine a modified speaker channel group (and optionally, mixing the content of at least one object channel with the modified speaker channel group, e.g., by mixing the object channel content indicating an alternate commentary with an alternate center channel of the modified speaker channel group).
[0135] The estimated bitrates of the audio content and associated metadata for an exemplary embodiment of the personalized bitstream of the present invention (encoded as an E-AC-3 bitstream) are indicated in the table below:
[0136]
[0137] In the example illustrated in the table, a 5.1 channel base layer may indicate both ambient and non-ambient content, where the non-ambient content (e.g., commentary on a sporting event) is mixed down to the three front channels ("diverged" non-ambient audio) or mixed down to the center channel only ("non-diverged" non-ambient audio). Alternative channel layers may include a single alternate center channel (e.g., indicating only ambient content for the center channel of the base layer, where the non-ambient content of the base layer is included only in the center channel of the base layer), or three alternate front channels (e.g., indicating only ambient content for the front channels of the base layer, where the non-ambient content of the base layer is interspersed between the front channels of the base layer). Additional sound beds and / or object channels may optionally be included, at the expense of the estimated additional bitrate requirements indicated.
[0138] As noted in the table, a bitrate of 12kbps is typical for metadata indicating each "experience definition", where an experience definition is a specification of an optional "preset" mix of the audio content (e.g., a mix of the content of at least one object channel and a speaker channel sound bed that gives a particular "experience"), including a set of mixing / rendering parameters for that mix (e.g., the spatial positions of rendered objects).
[0139] As indicated in the table, the following bitrates "per object or sound bed" of 2kbps to 5kbps are typical for metadata indicating experience maps. An experience map is a hierarchical mix graph that indicates optional preset mixes of the audio content of a delivered program, and each preset mix includes a certain number (e.g., 0, 1, or 2) of objects and typically at least one speaker channel (e.g., some or all of the speaker channels in a sound bed, and / or at least one replacement speaker channel). In addition to the speaker channels and object channels (objects) that may be included in each mix, the graph typically indicates rules (e.g., grouping and conditional rules). The bitrate requirements of the rules (e.g., grouping and conditional rules) of the experience map are included in a given estimate for each object or sound bed.
[0140] The replacement left, center and right (L / C / R) speaker channels recorded in the table can be selected to replace the left, center and right channels of the base layer, and when the replacement channels are rendered, indicate "separated audio" in the sense that the new content of the replacement speaker channels (i.e., the content that replaces the content of the corresponding channels of the base layer) is spatially separated over the area spanned by the left, center and right speakers (of the playback system).
[0141] In an example of an embodiment of the present invention, an object-based program indicates personalizable audio related to a football game (i.e., a personalizable soundtrack that accompanies a video showing the game). The program's default mix includes ambient content (crowd noise captured at the game) mixed with default commentary (also provided as an optional object channel for the program), two object channels indicating alternative pro-team commentary, and an alternate speaker channel indicating ambient content without the default commentary. The default commentary is unbiased (i.e., not biased in favor of either team). The program offers four experience intelligibility levels: a default mix that includes unbiased commentary, a first alternative mix that includes ambient content and commentary for a first team (e.g., the home team), a second alternative mix that includes ambient content and commentary for a second team (e.g., the away team), and a third alternative mix that includes only ambient content (no commentary). A typical implementation of the delivery of a bitstream including data indicative of a program would have a bitrate requirement of approximately 452 kbps (assuming the base layer is a 5.1 speaker channel soundbed and the default commentary is non-separated and located in only the center channel of that soundbed), allocated as follows: 192 kbps for the 5.1 base layer (indicating the default commentary in the center channel), 48 kbps (indicating an alternative center channel of only non-ambient content, which can be selected to replace the center channel of the base layer and optionally can be mixed with alternative commentary indicated by one of the object channels), 144 kbps for an object layer including three object channels (one channel for "main" or non-biased commentary; one channel for commentary biased toward the first team (e.g., the home team); and one channel for commentary biased toward the second team (e.g., the away team)), 4.5 kbps for object-related metadata (for rendering the object channels), 48 kbps for metadata indicating the four optional experiences, and 1.5 kbps for metadata indicating the experience map (layered mix map).
[0142] In a typical embodiment of a playback system (configured to decode and render a program), metadata included in the program allows a user to select from a menu of rendering options: a default mix including unbiased commentary (which can be rendered by rendering the unmodified sound bed or by replacing the center channel of the sound bed with a replacement center channel and mixing the resulting modified sound bed with the unbiased commentary content of the associated object channel); a first alternate mix (which can be rendered by replacing the center channel of the sound bed with the replacement center channel and mixing the resulting modified sound bed with a commentary mix biased toward the first team); a second alternate mix (which can be rendered by replacing the center channel of the sound bed with the replacement center channel and mixing the resulting modified sound bed with a commentary mix biased toward the second team); and a third alternate mix (which can be rendered by replacing the center channel of the sound bed with the replacement center channel. The menu is typically presented to the user via a user interface coupled (e.g., via a wireless link) to a controller of a set-top device (or other device, such as a TV, AVR, tablet, or phone) that is configured (at least in part) to decode and render object-based programs. In some other implementations, metadata included with the program otherwise allows the user to select from available rendering options.
[0143] In a second example of an embodiment of the present invention, an object-based program indicates personalizable audio related to a football game (i.e., a personalizable soundtrack that accompanies a video showing the game). The program's default mix includes ambient content (crowd noise captured at the game) mixed with a first default, non-biased commentary (the first non-biased commentary is also provided as an optional object channel for the program), five object channels indicating alternative non-ambient content (a second non-biased commentary, commentary biased towards both teams, an announcement feed, and a goal break feed), two alternative speaker channel groups (each of which is a 5.1 speaker channel group indicating a different mix of ambient and non-ambient content, each mix being different from the default mix), and an alternate speaker channel indicating ambient content without the default commentary. The program provides at least 9 experience intelligibility levels: a default mix including first non-biased commentary; a second alternative mix including ambient content and first team commentary; a second alternative mix including ambient content and second team commentary; a third alternative mix including ambient content and second non-biased commentary; a fourth alternative mix including ambient content, first non-biased commentary and an announcement feed; a fifth alternative mix including ambient content, first non-biased commentary and a goal break feed; a sixth alternative mix (determined by the first alternative group of 5.1 speaker channels); a seventh alternative mix (determined by the second alternative group of 5.1 speaker channels); and an eighth alternative mix including only ambient content (no commentary, no announcement feed and no goal break feed). A typically implemented delivery of a bitstream including data indicative of a program would have a bitrate requirement of approximately 987 kbps (assuming the base layer is a 5.1 speaker channel soundbed and that the default commentary is non-separated and presented only in the center channel of the soundbed), allocated as follows: 192 kbps for the 5.1 base layer (indicating the default commentary in the center channel), 48 kbps for an alternative center channel indicating only ambient content, which may be selected to replace the center channel of the base layer and optionally may be mixed with alternative content indicated by one or more object channels ), 384kbps for the object layer comprising 6 object channels (one channel for first non-biased commentary; one channel for second non-biased commentary; one channel for commentary biased towards the first team; one channel for commentary biased towards the second team; one channel for announcement feed; and one channel for goal break feed); 9kbps for object-related metadata (used to render the object channel), 36kbps (for metadata indicating 9 optional experiences), and 30kbps for metadata indicating the experience map (layered mix map).
[0144] In a typical embodiment of a playback system (configured to decode and render the program of the second example), metadata included in the program allows the user to select from a menu of rendering options: a default mix including unbiased commentary (which can be rendered by rendering the unmodified sound bed, or by replacing the center channel of the sound bed with the replacement center channel and mixing the resulting modified sound bed with the first unbiased commentary content of the associated object channels); a first alternate mix (which can be rendered by replacing the center channel of the sound bed with the replacement center channel and mixing the resulting modified sound bed with commentary mixed in favor of the first team); a second alternate mix (which can be rendered by replacing the center channel of the sound bed with the replacement center channel and mixing the resulting modified sound bed with commentary mixed in favor of the second team); a third alternate mix (which can be rendered by replacing the center channel of the sound bed with the replacement center channel and mixing the resulting modified sound bed with commentary mixed in favor of the second team); a fourth alternative mix (which may be rendered by replacing the center channel of the sound bed with the replacement center channel and mixing the resulting modified sound bed with the first non-biased commentary and announcement feed); a fifth alternative mix (which may be rendered by replacing the center channel of the sound bed with the replacement center channel and mixing the resulting modified sound bed with the commentary and goal break feed biased towards the first team); a sixth alternative mix (which may be rendered by rendering the first alternative group of 5.1 speaker channels instead of the sound bed); a seventh alternative mix (which may be rendered by rendering the second alternative group of 5.1 speaker channels instead of the sound bed); and an eighth alternative mix (which may be rendered by replacing the center channel of the sound bed with the replacement center channel). The menu is typically presented to the user through a user interface coupled (e.g., via a wireless link) to a controller of a set-top device (or other device, such as a TV, AVR, tablet, or phone) that is configured (at least in part) to decode and render the object-based program. In some other embodiments, metadata included with the program otherwise allows the user to select from available rendering options.
[0145] In other embodiments, other methods are employed for carrying extension layers of object-based programs (including speaker channels and object channels in addition to those of the base layer). Some of these methods reduce the overall bitrate required to deliver the base layer and the extension layers. For example, joint object coding or receiver-side sound bed mixing can be employed to allow for large bitrate savings in program delivery (at a tradeoff of increased computational complexity and constrained artistic flexibility). For example, joint object coding or receiver-side sound bed mixing can be employed to reduce the bitrate required to deliver the base layer and extension layers of the program of the second example described above from approximately 987 kbps (as noted above) to approximately 750 kbps.
[0146] The examples provided herein indicate the overall bitrate for delivering an entire object-based audio program (including the base layer and the extended layers). In other embodiments, the base layer (soundbed) is delivered in-band (e.g., in the broadcast bitstream), and at least a portion of the extended layers (e.g., object channels, alternative speaker channels, and / or layered mix maps and / or other metadata) are delivered out-of-band (e.g., over an Internet Protocol or "IP" network) to reduce the in-band bitrate. An example of delivering an entire object-based audio program in a manner divided across in-band (broadcast) and out-of-band (Internet) transmission is: a 5.1 base layer, alternative speaker channels, a main commentary object channel, and two alternative 5.1 speaker channel groups (using a total bitstream requirement of approximately 729 kbps) are delivered in-band, and the alternative object channels and metadata (including experience freshness and layered mix maps) (using a total bitrate requirement of approximately 258 kbps) are delivered out-of-band.
[0147] Figure 6 is a block diagram of an embodiment of a playback system that includes a decoder 20, an object processing subsystem 22, a spatial rendering subsystem 25, a controller 23 (which implements a user interface), and optionally digital audio processing subsystems 25, 26, and 27 coupled as shown, and that can be implemented to perform embodiments of the method of the present invention. In some implementations, Figure 6 Elements 20, 22, 24, 25, 26, 27, 29, 31 and 33 of the system are implemented as set-top devices.
[0148] exist Figure 6 In a system, a decoder 20 is configured to receive and decode an encoded signal indicating an object-based audio program (or a master mix of an object-based audio program). Typically, according to embodiments of the present invention, a program (e.g., a master mix of a program) indicates audio content including a sound bed having at least two speaker channels and a set of alternate speaker channels. The program also indicates at least one user-selectable object channel (and optionally at least one other object channel) and object-related metadata corresponding to each object channel. Each object channel indicates an audio object, and thus, for convenience, object channels are sometimes referred to herein as "objects." In one embodiment, the program is an AC-3 or E-AC-3 bitstream (or a master mix thereof) indicating audio objects, object-related metadata, a speaker channel sound bed, and alternate speaker channels. Typically, each audio object is encoded in mono or stereo (i.e., each object channel indicates the left or right channel of an object, or indicates a mono channel of an object), the sound bed is a traditional 5.1 mix, and the decoder 20 can be configured to decode up to 16 channels of audio content simultaneously (including the six speaker channels of the sound bed, as well as the alternate speaker channels and the object channels).
[0149] In some embodiments of the playback system of the present invention, each frame of an input E-AC-3 (or AC-3) encoded bitstream includes one or more metadata "containers." The input bitstream indicates an object-based audio program or a master mix of such a program, and the speaker channels of the program are organized like the audio content of a conventional E-AC-3 (or AC-3) bitstream. One container can be included in the Aux field of the frame, and another container can be included in the addbsi field of the frame. Each container has a core header and includes (or is associated with) one or more payloads. One such payload (of or associated with a container included in the Aux field) can be a set of audio samples for each of one or more object channels of the present invention (associated with a speaker channel soundbed that also has program representation) and object-related metadata associated with each object channel. In such a payload, the samples of some or all of the object channels (and associated metadata) may be organized as standard E-AC-3 (or AC-3) frames, or may be organized in other ways (e.g., they may be included in a submix different from the E-AC-3 or AC-3 bitstream). Another example of such a payload is a set of loudness processing state metadata associated with the audio content of the frame (e.g., in or associated with a container included in the addbsi field or the Aux field).
[0150] In some such embodiments, a decoder (e.g., Figure 6 A decoder (e.g., a decoder of a video frame) will parse the core header of the container in the Aux field and extract the object channels and associated metadata of the present invention from the container (e.g., from the Aux field of an AC-3 or E-AC-3 frame) and / or from the location indicated by the core header (e.g., a submix). After extracting the payload (the object channels and associated metadata), the decoder will perform any necessary decoding on the extracted payload.
[0151] The core header for each container typically includes: at least one ID value indicating the type of payload included in or associated with the container; a substream association indication (indicating which substream the core header is associated with); and protection bits. Such protection bits (which may include or comprise a hash-based message authentication code, or "HMAC") are typically useful for at least one of decrypting, authenticating, or verifying object-related metadata and / or loudness processing state metadata (and optionally other metadata) contained in at least one payload included in or associated with the container and / or corresponding audio data included in the frame. Substreams can be located "in-band" (within the E-AC-3 or AC-3 bitstream) or "out-of-band" (e.g., in a submix bitstream independent of the E-AC-3 or AC-3 bitstream). One type of such payload is a set of audio samples for each of one or more object channels (associated with a speaker channel bed, also indicated by the program) and object-related metadata associated with each object channel. Each object channel is a separate substream and is typically identified in the core header. Another type of payload is loudness processing state metadata.
[0152] Typically, each payload has its own header (or "payload identifier"). Object-level metadata can be carried in each substream that is an object channel. Program-level metadata can be included in the core header of the container and / or in the header of the payload that is a set of audio samples that is one or more object channels (and metadata associated with each object channel).
[0153] In some embodiments, each container in the auxiliary data auxdata (or addbsi) field of a frame has three levels of structure:
[0154] a high-level structure comprising a flag indicating whether the ancillary data (or addbsi) field includes metadata (where, in this context, "metadata" means object channels, object-related metadata, and any other audio content or metadata carried by the bitstream but not typically carried in a conventional E-AC-3 or AC-3 bitstream lacking any container of the described type), at least one ID value indicating what type of metadata is present, and typically also a value indicating how many bits (if any) of the metadata (e.g., of each type of metadata) are present. In this context, an example of one such "type" of metadata is object channel data and associated object-related metadata (i.e., a set of audio samples for each of one or more object channels (also associated with the loudspeaker channel soundbed indicated by the program) and metadata associated with each object channel);
[0155] a mid-level structure including core elements of metadata for each identified type (e.g., a core header, protection values, and payload ID and payload size values, such as those described above, for each identified type of metadata); and
[0156] A low-level structure that includes, for each payload, a core element, where at least one such payload is identified by the core element being presented. An example of such a payload is a set of audio samples for each of one or more object channels (associated with a speaker channel bed also indicated by the program) and metadata associated with each object channel. Another example of such a payload is a payload that includes loudness processing state metadata ("LPSM"), which is sometimes referred to as an LPSM payload.
[0157] Data values in such a three-level structure can be nested. For example, after each payload is identified by the core element (and thus after each payload is identified by the core header of the core element), a protection value for the payload identified by the core element (e.g., an LPSM payload) can be included. In one example, the core header can identify a first payload (e.g., an LPSM payload) and another payload, and after the core header can be a payload ID and payload size value for the first payload, the first payload itself can be after the ID and size value, and the payload ID and payload size value for the second payload can be after the first payload, the second payload itself can be after these ID and size values, and the protection value for one or both of the payloads (or the core element value and one or both of the payloads) can be after the last payload.
[0158] Refer again Figure 6 , the user uses the controller 23 to select the object to be rendered (indicated by the object-based audio program). The controller 23 may be programmed to implement the Figure 6 The controller 23 may be a handheld processing device (e.g., an iPad) that provides a user interface (e.g., an iPad application) that is compatible with the other elements of the system. The user interface may provide the user with (e.g., displayed on a touch screen) a menu or palette of objects, "sound bed" speaker channel content, and optional "preset" mixes that replace the speaker channel content. The optional preset mixes may be determined by the program's object-related metadata and, typically, also by rules implemented by the subsystem 22 (e.g., rules that the subsystem 22 has been preconfigured to implement). The user may select from the optional mixes by inputting commands to the controller 23 (e.g., by activating its touch screen), and in response, the controller 23 sets corresponding control data to the subsystem 22.
[0159] Decoder 20 decodes the speaker channels of the program's speaker channel soundbed (and any replacement speaker channels included in the program) and outputs the decoded speaker channels to subsystem 22. Responsive to the object-based audio program, and responsive to control data from controller 23 indicating a selected subset of the program's entire set of object channels to be rendered, decoder 23 decodes the selected object channels (if necessary) and outputs the selected (e.g., decoded) object channels (each of which may be a pulse code modulated or "PCM" bitstream) and object-related metadata corresponding to the selected object channels to subsystem 22.
[0160] The objects indicated by the decoded object channels are typically or include user-selectable audio objects. For example, the decoder may extract a 5.1 speaker channel sound bed, an alternate speaker channel (indicating ambient content in one of the sound bed speaker channels, rather than non-ambient content in one of the non-sound bed speaker channels), an object channel (e.g., indicating commentary by an announcer from the home team's city), or a Figure 6 ), indicating an object channel for commentary by an announcer from the visiting team's city (e.g., "Commentary-1 Mono" in FIG. Figure 6 ), indicating an object channel consisting of crowd noise from fans of the home team present at a sporting event (e.g., “Commentary-2 Mono” in the example above). Figure 6 ), a left object channel and a right object channel (such as "Fans (Home Team)") indicating the sound produced by the winning goal when the ball is hit by a sports event participant. Figure 6 ), and four object channels indicating special effects (such as Figure 6 Any of the “Commentary-1 Mono” object channel, the “Commentary-2 Mono” object channel, the “Fans (Home Team)” object channel, the “Stereo Ball Sound” object channel, and the “Effect 4x Mono” object channel may be selected (after undergoing any necessary decoding in the decoder 20), and each selected object channel thereof may be passed from the subsystem 22 to the rendering subsystem 24.
[0161] As well as the decoded speaker channels, decoded object channels and decoded object-related metadata from the decoder 20, the input to the object processing subsystem 22 optionally includes external audio object channels set to the system (e.g., one or more submixes of a program set as its main mix to the decoder 20). Examples of objects indicated by such external audio object channels include local commentators (e.g., single-channel audio content delivered by a radio channel), incoming Skype calls, incoming twitter connections (via Figure 6 (converted by a text-to-speech system not shown) and system sounds.
[0162] Subsystem 22 is configured to output a selected subset of the full set of object channels indicated by the program (or a processed version of the selected subset of the full set of object channels) and corresponding object-related metadata for the program, as well as a selected set of speaker channels from the sound bed speaker channels and / or replacement speaker channels. Object channel selection and speaker channel selection may be determined by user selection (as indicated by control data set to subsystem 22 from controller 23) and / or by rules (e.g., indicating conditions and / or constraints) that subsystem 22 has been programmed or otherwise configured to implement. Such rules may be determined by the program's object-related metadata and / or other data (e.g., data indicating the performance and organization of the playback system's speaker array) set to subsystem 22 (e.g., from controller 23 or another external source) and / or by preconfiguration (e.g., programming) of subsystem 22. In some embodiments, the object-related metadata provides a selectable set of "preset" mixes of speaker channel content (of the speaker channel sound bed and / or replacement speaker channels) and objects, and the subsystem 22 uses this metadata to select the object channels that it optionally processes and then sets to the subsystem 24, and the speaker channels that it sets to the subsystem 24. The subsystem 22 typically passes (to the subsystem 24) unchanged the selected subset of the decoded speaker channels (the sound bed speaker channels and typically also the replacement speaker channels) from the decoder 20 (e.g., at least one speaker channel of the sound bed and at least one replacement speaker channel), and processes the selected object channels of the object channels set thereto.
[0163] The object processing (including object selection) performed by subsystem 22 is typically controlled by control data from controller 23 and object-related metadata from decoder 20 (and optionally also object-related metadata of the submix that is set to subsystem 22 rather than from decoder 20), and typically includes the determination of the spatial position and level of each selected object (regardless of whether the object selection is due to user selection or selection made through the application of rules). Typically, default spatial positions and default levels for rendering objects, and optionally user-selected restrictions on objects and their spatial positions and levels, are included in the object-related metadata set to subsystem 20 (e.g., from decoder 20). Such restrictions may indicate prohibited combinations of objects or prohibited spatial positions at which selected objects may be rendered (e.g., to prevent selected objects from being rendered too close to each other). In addition, the loudness of each selected object is typically controlled by object processing subsystem 22 in response to control data input using controller 23 and / or default levels indicated by object-related metadata (e.g., from decoder 20) and / or by pre-configuration of subsystem 22.
[0164] Typically, the decoding performed by decoder 20 includes extracting (from the input program) metadata indicating the type of audio content for each object indicated by the program (e.g., the type of sporting event indicated by the program's audio content, and the names or other identifying indicia (e.g., team logos) of the selectable and default objects indicated by the program). Controller 23 and object processing subsystem 22 receive this metadata or related information indicated by the metadata. Typically, controller 23 also receives (e.g., is programmed with) information related to the playback capabilities of the user's audio system (e.g., the number of speakers, the assumed placement of the speakers, or other assumed organization).
[0165] Figure 6 The spatial rendering subsystem 24 (or the subsystem 24 with at least one downstream device or system) is configured to render the audio content output from the subsystem 22 for playback by the speakers of the user's playback system. One or more of the optionally included digital audio processing subsystems 25, 26, and 27 may implement post-processing of the output of the subsystem 24.
[0166] The spatial rendering subsystem 24 is configured to map the audio object channels selected (or selected and processed) by the object processing subsystem 22 and set to the subsystem 24 (e.g., default selected objects and / or user-selected objects that have been selected due to user interaction using the controller 23) to available speaker channels (e.g., a selected set of sound bed speaker channels and replacement speaker channels determined by the subsystem 22 and passed through the subsystem 22 to the subsystem 24) using the rendering parameters associated with each selected object output from the subsystem 22 (e.g., user-selected and / or default values for spatial position and level). Typically, the subsystem 24 is an intelligent mixer and is configured to determine speaker feeds for the available speakers, including by mapping one, two, or more than two selected object channels to each of a plurality of individual speaker channels, and mixing the selected object channels with the audio content indicated by each corresponding speaker channel.
[0167] Typically, the number of output speaker channels may vary between 2.0 and 7.1, and the speakers to be driven to render the selected audio object channels (to mix with the selected speaker channel content) may be assumed to be located in a (nominal) horizontal plane in the playback environment. In such a case, rendering is performed such that the speakers may be driven to emit sounds mixed with the sounds determined by the speaker channel content, which sounds will be perceived as emanating from different object positions in the plane of the speakers (i.e., one object position or a series of object positions along a trajectory for each selected or default object).
[0168] In some embodiments, the number of full-range speakers to be driven to render audio can be any number within a wide range (not necessarily limited to the range from 2 to 7), and thus the number of output speaker channels is not limited to the range from 2.0 to 7.1.
[0169] In some embodiments, the speakers to be driven to render the audio are assumed to be located at arbitrary locations in the playback environment; not just in a (nominal) horizontal plane. In some such cases, metadata included in the program indicates rendering parameters for rendering at least one object of the program using a three-dimensional speaker array at an arbitrary apparent spatial location (in a three-dimensional volume). For example, an object channel may have corresponding metadata indicating a three-dimensional trajectory of the apparent spatial location of the object (indicated by the object channel) to be rendered. The trajectory may include a series of "floor" positions (in the plane of a subgroup of speakers assumed to be located on the floor of the playback environment, or in another horizontal plane of the playback environment) and a series of "above-floor" positions (each "above-floor" position determined by driving a subgroup of speakers assumed to be located in at least one other horizontal plane of the playback environment). In such a case, rendering may be performed in accordance with the present invention such that the speakers may be driven to emit sound (determined by the associated object channel) mixed with the sound determined by the speaker channel content, which sound will be perceived as emanating from the series of object positions in the three-dimensional space that includes the trajectory. Subsystem 24 may be configured to implement such rendering, or steps thereof, wherein the remaining steps of the rendering are performed by downstream systems or devices (e.g., Figure 6 Rendering subsystem 35) is executed.
[0170] Optionally, a digital audio processing (DAP) stage (e.g., one stage for each of a number of predetermined output speaker channel configurations) is coupled to the output of the spatial rendering subsystem 24 to perform post-processing on the output of the spatial rendering subsystem. Examples of such processing include intelligent equalization or (in the case of stereo output) speaker virtualization processing.
[0171] Figure 6The output of the system (e.g., the output of the spatial rendering subsystem or the DAP stage after the spatial rendering stage) can be a PCM bitstream (which identifies the speaker feeds for the available speakers). For example, where the user's playback system includes a 7.1 speaker array, the system can output a PCM bitstream (generated in subsystem 24) that identifies the speaker feeds for the speakers of such an array, or a post-processed version of such a bitstream (generated in DAP 25). For another example, where the user's playback system includes a 5.1 speaker array, the system can output a PCM bitstream (generated in subsystem 24) that identifies the speaker feeds for the speakers of such an array, or a post-processed version of such a bitstream (generated in DAP 26). For another example, where the user's playback system includes only a left speaker and a right speaker, the system can output a PCM bitstream (generated in subsystem 24) that identifies the speaker feeds for the left and right speakers, or a post-processed version of such a bitstream (generated in DAP 27).
[0172] Figure 6 The system optionally further includes one or both of a re-encoding subsystem 31 and a re-encoding subsystem 33. The re-encoding subsystem 31 is configured to re-encode a PCM bitstream output from the DAP 25 as an E-AC-3 encoded bitstream (indicating a feed for a 7.1 speaker array), and the resulting encoded (compressed) E-AC-3 bitstream can be output from the system. The re-encoding subsystem 33 is configured to re-encode a PCM bitstream output from the DAP 27 as an AC-3 or E-AC-3 encoded bitstream (indicating a feed for a 5.1 speaker array), and the resulting encoded (compressed) AC-3 or E-AC-3 bitstream can be output from the system.
[0173] Figure 6The system optionally also includes a re-encoding (or formatting) subsystem 29 and a downstream rendering subsystem 35 coupled to receive the output of subsystem 29. Subsystem 29 is coupled to receive data (output from subsystem 22) indicating selected audio objects (or default mixes of audio objects), corresponding object-related metadata, and decoded speaker channels (e.g., sound bed speaker channels and replacement speaker channels), and is configured to re-encode (and / or format) such data for rendering by subsystem 35. Subsystem 35, which may be implemented in an AVR or soundbar (or other system or device downstream of subsystem 29), is configured to generate speaker feeds (or bitstreams identifying the speaker feeds) for available playback speakers (speaker array 36) in response to the output of subsystem 29. For example, subsystem 29 may be configured to encode audio into a suitable format for rendering in subsystem 35 by re-encoding data indicating selected (or default) audio objects, corresponding metadata, and speaker channels, and transmit the encoded audio to subsystem 35 (e.g., via an HDMI link). In response to the speaker feeds generated by (or determined by the output of) subsystem 35, available speakers 36 will emit sounds indicative of a mix of the speaker channel contents and the selected (or default) objects, where the objects have apparent source positions determined by the object-related metadata output by subsystem 29. When subsystems 29 and 35 are included, rendering subsystem 24 is optionally omitted from the system.
[0174] In some embodiments, the present invention is a distributed system for rendering object-based audio, wherein a first subsystem (e.g., implemented in a set-top device or a set-top device and a handheld controller) Figure 6 20, 22 and 23) in the components 20, 22 and 23) to implement a part of the rendering (ie, at least one step) (for example, as shown by Figure 6 In one embodiment, the audio rendering system may include selecting audio objects to be rendered and selecting rendering characteristics for each selected object, performed by subsystem 22 and controller 23 of the system, and implementing another portion of the rendering (e.g., generating speaker feeds or immersive rendering of signals that determine speaker feeds in response to outputs of the first subsystem) at a second subsystem (e.g., subsystem 35 implemented in an AVR or soundbar). Some embodiments that provide distributed rendering also implement legacy management to account for multiple portions of audio rendering (any processing of the video corresponding to the rendered audio) being performed at different times and in different subsystems.
[0175] In some embodiments of the playback system of the present invention, each decoder and object processing subsystem (sometimes referred to as a personalization engine) is implemented in a set-top box (STB). For example, Figure 6Elements 20 and 22 and / or Figure 7 All elements of the system. In some embodiments of the playback system of the present invention, multiple renderings are performed on the output of the personalization engine to ensure that all STB outputs (e.g., HDMI output, S / PDIF output, or stereo analog output of the STB) are enabled. Optionally, the selected object channels (and corresponding object-related metadata) and speaker channels (along with the decoded speaker channel soundbed) are passed from the STB to a downstream device (e.g., an AVR or soundbar) that is configured to render a mix of the object channels and speaker channels.
[0176] In one class of embodiments, the object-based audio program of the present invention includes a set of bitstreams (multiple bitstreams, which may be referred to as "substreams") that are generated and transmitted in parallel. In some embodiments within this class, multiple decoders are used to decode the content of the substreams (e.g., the program includes multiple E-AC-3 substreams, and the playback system utilizes multiple E-AC-3 decoders to decode the content of the substreams). Figure 7 is a block diagram of a playback system configured to decode and render an embodiment of the present invention's object-based audio program that includes multiple serial bitstreams delivered in parallel.
[0177] Figure 7 The playback system is Figure 6 A variation of the system in which an object-based audio program comprises a plurality of bit streams (B1, B2, ..., BN, where N is some positive integer) that are delivered in parallel to and received by a playback system. Each bit stream ("substream") B1, B2, ..., and BN is a serial bit stream that includes a time code or other synchronization word (cf. Figure 7 , for convenience, referred to as a "sync word") to enable the substreams to be synchronized or time-aligned with each other. Each substream also includes a different subset of the entire set of object channels and corresponding object-related metadata, and at least one substream includes speaker channels (e.g., a sound bed speaker channel and an alternate speaker channel). For example, in each substream B1, B2, ..., BN, each container including object channel content and object-related metadata includes a unique ID or timestamp.
[0178] Figure 7 The system comprises N deformatters 50 , 51 , . . . , 53 , each coupled and configured to parse a different one of the input substreams and set the metadata (including its sync words) and its audio content to a bitstream synchronization level 59 .
[0179] The deformatter 50 is configured to parse the substream B1 and set its synchronization word (T1), other metadata, and its object channel content (M1) (including the program's object-related metadata and at least one object channel) and its speaker channel audio content (A1) (including at least one speaker channel of the program) to the bitstream synchronization level 59. Similarly, the deformatter 51 is configured to parse the substream B2 and set its synchronization word (T2), other metadata, and its object channel content (M2) (including the program's object-related metadata and at least one object channel) and its speaker channel audio content (A2) (including at least one speaker channel of the program) to the bitstream synchronization level 59. The deformatter 53 is configured to parse the substream BN and set its synchronization word (TN), other metadata, and its object channel content (MN) (including the program's object-related metadata and at least one object channel) and its speaker channel audio content (AN) (including at least one speaker channel of the program) to the bitstream synchronization level 59.
[0180] Figure 7 The system's bitstream synchronization stage 59 typically includes buffers for the audio content and metadata of substreams B1, B2, ..., BN, and a stream offset compensation element coupled and configured to use the synchronization word for each substream to determine any misalignment of data in the input substreams (e.g., because each bitstream is typically carried within the media file via a separate interface and / or trace, misalignment may occur due to the possibility of loss of strict synchronization between them in distribution / contribution). The stream offset compensation element of stage 59 is also typically configured to correct any determined misalignment by setting appropriate control values to the buffers containing the audio data and metadata of the bitstreams so that the time-aligned bits of the speaker channel audio data are read from the buffers to decoders (including decoders 60, 61, and 63), each of which is coupled to a corresponding one of the buffers, and the time-aligned bits of the object channel audio data and metadata are read from the buffers to the object data assembly stage 66.
[0181] The time-aligned bits of the speaker channel audio content A1′ from substream B1 are read from stage 59 to a decoder 60, and the time-aligned bits of the object channel content and metadata M1′ from substream B1 are read from stage 59 to a metadata combiner 66. The decoder 60 is configured to perform decoding on the speaker channel audio data provided thereto and to provide the resulting decoded speaker channel audio to an object processing and rendering subsystem 67.
[0182] Similarly, the time-aligned bits of the speaker channel audio content A2′ from substream B2 are read from stage 59 to a decoder 61, and the time-aligned bits of the object channel content and metadata M2′ from substream B2 are read from stage 59 to a metadata combiner 66. The decoder 61 is configured to perform decoding on the speaker channel audio data provided thereto and to provide the resulting decoded speaker channel audio to an object processing and rendering subsystem 67.
[0183] Similarly, the time-aligned bits of the speaker channel audio content AN′ from substream BN are read from stage 59 to a decoder 63, and the time-aligned bits of the object channel content and metadata MN′ from substream BN are read from stage 59 to a metadata combiner 66. The decoder 63 is configured to perform decoding on the speaker channel audio data provided thereto and to provide the resulting decoded speaker channel audio to an object processing and rendering subsystem 67.
[0184] For example, each substream B1, B2, ..., BN may be an E-AC-3 substream, and each decoder 60, 61, 63, and any other decoders coupled to subsystem 59 in parallel with decoders 60, 61, and 63, may be an E-AC-3 decoder configured to decode the speaker channel content of one of the input E-AC-3 substreams.
[0185] The data object assembler 66 is configured to provide the time-aligned object channel data and metadata for all object channels of a program to the object processing and rendering subsystem 67 in an appropriate format.
[0186] Subsystem 67 is coupled to the output of combiner 66 and the outputs of decoders 60, 61, and 63 (and any other decoders coupled in parallel with decoders 60, 61, and 63 between subsystem 59 and subsystem 67), and controller 68 is coupled to subsystem 67. Subsystem 67 is generally configured to perform object processing (e.g., including processing by) on the outputs of combiner 66 and decoders in an interactive manner in accordance with embodiments of the present invention in response to control data from controller 68. Figure 6 The controller 68 may be configured to perform operations wherein Figure 6 The controller 23 of the system is configured to perform the operations (or variations of such operations) in response to input from the user. The subsystem 67 is also typically configured to perform rendering (e.g., by mixing the speaker channel audio and object channel audio data assigned thereto) in accordance with embodiments of the present invention (e.g., rendering a mix of the sound bed speaker channel content, the replacement speaker channel content, and the object channel content). Figure 6the rendering subsystem 24 or subsystems 24, 25, 26, 31, and 33 of the system, or Figure 6 The operations performed by subsystems 24, 25, 26, 31, 33, 29 and 35 of the system, or variations on such operations).
[0187] exist Figure 7 In one implementation of the system, each substream B1, B2, ..., BN is a Dolby E bitstream. Each such Dolby E bitstream comprises a series of bursts. Each burst may carry speaker channel audio content (content of the sound bed speaker channels and / or replacement speaker channels) and a subgroup of the entire object channel group (which may be a large group) of the object channels of the present invention and object-related metadata (i.e., each burst may indicate some object channels of the entire object channel group and corresponding object-related metadata). Each burst of a Dolby E bitstream typically occupies a time period equivalent to the time period of a corresponding video frame. Each Dolby E bitstream in the group includes a synchronization word (e.g., a time code) so that the bitstreams in the group can be synchronized or time-aligned with each other. For example, in each bitstream, each container including object channel content and object-related metadata may include a unique ID or timestamp so that the bitstreams in the group can be synchronized or time-aligned with each other. In Figure 7 In the noted implementation of the system, each deformatter 50, 51 and 53 (and any other deformatters coupled in parallel with the deformatters 50, 51 and 53) is a SMPTE 337 deformatter, and each decoder 60, 61 and 63 and any other decoders coupled in parallel with the decoders 60, 61 and 63 to the subsystem 59 can be a Dolby E decoder.
[0188] In some embodiments of the present invention, the object-related metadata of the object-based audio program includes persistent metadata. Figure 6The object-related metadata included in the program of the subsystem 20 of the system can include non-persistent metadata (e.g., default levels and / or rendering positions or tracks for user-selectable objects) and persistent metadata. Non-persistent metadata can change at at least one point in the broadcast chain (from the content creation facility that generated the program to the user interface implemented by the controller 23) and persistent metadata, which is not intended to be changeable after the initial generation of the program (typically, in the content creation facility). Examples of persistent metadata include: an object ID for each user-selectable object or other object or group of objects in the program; and a time code or other synchronization word indicating the timing of each user-selectable object or other object relative to the speaker channel content or other elements of the program. Persistent metadata is typically preserved throughout the entire broadcast chain from the content creation facility to the user interface, throughout the entire duration of the program's broadcast, or even during replays of the program. In some embodiments, the audio content (and associated metadata) of at least one user-selectable object is transmitted in the main mix of the object-based audio program, and at least some persistent metadata (e.g., time code) and, optionally, the audio content (and associated metadata) of at least one other object is transmitted in a secondary mix of the program.
[0189] The object-related metadata in some embodiments of the object-based audio program of the present invention is used to save (e.g., even after the broadcast of the program) a user-selected mix of the object content and the speaker channel content. For example, this can provide the selected mix as the default mix whenever the user watches a program of a particular type (e.g., any soccer game) or whenever the user watches any program (of any type) until the user changes his / her selection. For example, during the broadcast of a first program, the user can utilize ( Figure 6 The controller 23 of the system) selects a mix that includes an object with a persistent ID (e.g., an object identified by the user interface of the controller 23 as a "home team crowd noise" object, where the persistent ID indicates "home team crowd noise"). Then, whenever the user watches (and listens to) another program (which includes an object with the same persistent ID), the playback system will automatically render that program using the same mix (i.e., the program's sound bed speaker channels and / or replacement speaker channels mixed with the program's "home team crowd noise" object channels) until the user changes the mix selection. Persistent object-related metadata in some embodiments of the present invention's object-based audio program can make the rendering of some objects mandatory throughout the program (e.g., despite user desire to override such rendering).
[0190] In some embodiments, object-related metadata provides a default mix of object content and speaker channel content with default rendering parameters (e.g., default spatial positions of rendered objects). Figure 6The object-related metadata of a program of subsystem 20 of the system may be a default mix of object content and speaker channel content with default rendering parameters, and subsystems 22 and 24 will cause the program to be rendered using the default mix and using default rendering parameters unless a user utilizes controller 23 to select another mix of object content and speaker channel content and / or another set of rendering parameters.
[0191] In some embodiments, the object-related metadata provides a set of selectable "preset" mixes of the object and speaker channel content, each preset mix having a predetermined set of rendering parameters (e.g., spatial position of the rendered object). The user interface of the playback system can present these as a limited menu or palette of available mixes (e.g., selected by Figure 6 , or a limited menu or options panel displayed by the controller 23 of the system). Each preset mix (and / or each selectable object) may have a persistent ID (e.g., a name, label, or logo). The controller 23 (or the controller of another embodiment of the playback system of the present invention) may be configured to display an indication of such an ID (e.g., on a touch screen of an iPad implementation of the controller 23). For example, there may be a selectable "home team" mix with a persistent ID (e.g., a team logo), regardless of changes (e.g., changes made by the broadcast company) to the details of the audio content or non-persistent metadata of each object of the preset mix.
[0192] In some embodiments, the program's object-related metadata (or reconfiguration of the playback or rendering system that is not indicated by metadata delivered with the program) provides constraints or conditions on the optional mixing of objects and soundbed (speaker channel) content. For example, Figure 6 Implementation of the system may implement digital rights management (DRM), and more specifically, may implement a DRM hierarchy to enable Figure 6 A user of the system can have "tiered" access to the set of audio objects included in an object-based audio program. If a user (e.g., a client associated with the playback system) pays more money (e.g., to the broadcaster), the user may be authorized to decode and select (and listen to) more audio objects of the program.
[0193] For another example, object-related metadata may provide constraints on user selections of objects. An example of such a constraint is if a user utilizes controller 23 to select to render both a "home team crowd noise" object and a "home team announcer" object for a program (i.e., for inclusion in a program generated by Figure 6The constraints may be determined (at least in part) by data related to the playback system (e.g., data input by the user). For example, if the playback system is a stereo system (including only two speakers), then Figure 6 The object processing subsystem 24 (and / or controller 23) of the system may be configured to prevent the user from selecting a mix (identified by the object-related metadata) that cannot be rendered with appropriate spatial resolution through only two speakers. For another example, Figure 6 The object processing subsystem 24 (and / or controller 23) of the system may remove some delivered objects from the assortment of selectable objects indicated by the object-related metadata (and / or other data input to the playback system) for legal (e.g., DRM) reasons or other reasons (e.g., based on the bandwidth of the delivery channel). The user may pay the content creator or broadcaster for more bandwidth, and the system (e.g., Figure 6 The system's object processing subsystem 24 and / or controller 23) may enable the user to select from a large menu of selectable objects and / or object / bed mixes.
[0194] Some embodiments of the present invention (e.g., Figure 6 For example, the default or selected object channels (and corresponding object-related metadata) of a program are transmitted from a set-top device (e.g., from a Figure 6 The system's implemented subsystems 22 and 29) is passed (along with the decoded speaker channels, such as the sound bed speaker channels and the selected group of replacement speaker channels) to a downstream device (e.g., an AVR or sound bar implemented in a set-top device (STB) downstream of the system's implemented subsystems 22 and 29). Figure 6 Subsystem 35 of the present invention. The downstream device is configured to render a mix of the object channels and the speaker channels. The STB may partially render the audio, and the downstream device may complete the rendering (e.g., by generating speaker feeds for driving specific top speakers (e.g., ceiling speakers) to place audio objects at specific apparent source locations, where the output of the STB merely indicates that the objects may be rendered in some non-specific manner in some non-specific top speakers). For example, the STB may not have knowledge of the specific organization of the speakers of the playback system, but the downstream device (e.g., an AVR or sound bar) may have such knowledge.
[0195] In some embodiments, an object-based audio program (e.g., input to Figure 6 Subsystem 20 of the system or Figure 7The programs of elements 50, 51 and 53 of the system) are or include at least one AC-3 (or E-AC-3) bitstream, and each container of the program including object channel content (and / or object related metadata) includes an auxiliary data auxdata field at the end of a frame of the bitstream (e.g., Figure 1 or Figure 4 ). In some such embodiments, each frame of the AC-3 or E-AC-3 bitstream includes one or two metadata containers. One container may be included in the Aux field of the frame, and the other container may be included in the addbsi field of the frame. Each container has a core header and includes (or is associated with) one or more payloads. One such payload (of or associated with a container included in the Aux field) may be a set of audio samples for each of one or more object channels of the present invention (associated with a speaker channel sound bed also indicated by the program) and object-related metadata associated with each object channel. The core header of each container typically includes: at least one ID value indicating the type of payload included in or associated with the container; a substream association indication (indicating which substreams the core header is associated with); and a protection bit. Typically, each payload has its own header (or "payload identifier"). Object-level metadata may be carried in each substream that is an object channel.
[0196] In other embodiments, an object-based audio program (e.g., input to Figure 6 Subsystem 20 of the system or Figure 7The program of elements 50, 51, and 53 of the system is or includes a bitstream that is not an AC-3 bitstream or an E-AC-3 bitstream. In some embodiments, the object-based audio program is or includes at least one Dolby E bitstream, and the program's object channel content and object-related metadata (e.g., each container of the program includes object channel content and / or object-related metadata) are included in bit positions of the Dolby E bitstream that do not normally carry useful information. Each burst of the Dolby E bitstream occupies a time period equivalent to the time period of the corresponding video frame. The object channel (and / or object-related metadata) can be included in guard bands between Dolby E bursts and / or in unused bit positions within each data structure (each having the format of an AES3 frame) within each Dolby E burst. For example, each guard band includes a series of segments (e.g., 100 segments), each of the first X segments of each guard band (e.g., X=20) includes an object channel and object-related metadata, and each of the remaining segments of each guard band can include a guard band symbol. In some embodiments, at least some object channels (and / or object-related metadata) of a program of the present invention are included in the 4 least significant bits (LSBs) of each of two AES3 subframes in each of at least some AES3 frames of a Dolby E bitstream, and data indicating the speaker channels of the program are included in the 20 most significant bits (MSBs) of each of two AES3 subframes in each AES3 frame of the bitstream.
[0197] In some embodiments, the object channels and / or object-related metadata of the program of the present invention are included in metadata containers in the Dolby E bitstream. Each container has a core header and includes (or is associated with) one or more payloads. One such payload (of a container included in the Aux field or associated with a container included in the Aux field) can be a set of audio samples for each of one or more object channels of the present invention (e.g., associated with speaker channels also indicated by the program) and object-related metadata associated with each object. The core header of each container typically includes: at least one ID indicating the type of payload included in or associated with the container; a substream association indication (indicating which substreams the core header is associated with); and protection bits. Typically, each payload has its own header (or "payload identifier"). Object-level metadata can be carried in each substream that is an object channel.
[0198] In some embodiments, an object-based audio program (e.g., input to a Figure 6 Subsystem 20 of the system or Figure 7 The program (of elements 50, 51, and 53 of the system) is decodable and its speaker channel content is renderable. The same program can be rendered according to some embodiments of the present invention by a set-top device (or other decoding and rendering system) configured (according to one embodiment of the present invention) to parse the object channels and object-related metadata of the present invention and render a mix of the speaker channels and object channel content indicated by the program.
[0199] Some embodiments of the present invention are intended to provide end customers with a personalized (and preferably immersive) audio experience in response to a broadcast program, and / or provide new methods for using metadata in a broadcast pipeline. Some embodiments improve microphone capture (e.g., stadium microphone capture) to generate audio programs that provide a more personalized and immersive experience for end users, modify existing production, contribution, and distribution workflows to enable object channels and metadata for the object-based audio programs of the present invention to flow through a professional chain, and create a new playback pipeline (e.g., one implemented in a set-top device) that supports object channels, alternative speaker channels and associated metadata, and conventional broadcast audio (e.g., the speaker channel soundbed included in embodiments of the broadcast audio program of the present invention).
[0200] Figure 8 is a block diagram of a broadcast system configured to generate object-based audio programs (and corresponding video programs) for broadcast, according to an embodiment of the present invention. Figure 8 A set of X microphones of the system (where X is an integer), including microphones 100 , 101 , 102 and 103 , are positioned to capture audio content to be included in the program, and their outputs are coupled to the inputs of an audio console 104 .
[0201] In one class of embodiments, a program includes interactive audio content that indicates the atmosphere of and / or commentary on a viewing event (e.g., a soccer or football game, a car or motorcycle race, or another sporting event). In some embodiments, the program's audio content indicates a plurality of audio objects (including user-selectable objects or groups of objects, and typically also a default group of objects to be rendered if the user does not make an object selection), a speaker channel soundbed (indicating a default mix of captured content), and alternate speaker channels. The speaker channel soundbed can be a conventional mix of speaker channels of the type that would be included in a conventional broadcast program that does not include object channels (e.g., a 5.1 channel mix).
[0202] A subset of microphones (e.g., microphone 100 and microphone 101, and optionally other microphones whose outputs are coupled to audio console 104) is a conventional microphone array that, in operation, captures audio (to be encoded and delivered as a speaker channel soundbed and a set of replacement speaker channels). In operation, another subset of microphones (e.g., microphone 102 and microphone 103, and optionally other microphones whose outputs are coupled to audio console 104) captures audio (e.g., crowd noise and / or other "objects") to be encoded and delivered as object channels of a program. For example, Figure 8 The microphone array of the system may include: at least one microphone (e.g., microphone 100) implemented as a sound field microphone (e.g., a sound field microphone mounted with a heater) and permanently installed in the stadium; at least one stereo microphone (e.g., microphone 102 implemented as a Sennheiser MKH416 microphone or another stereo microphone) directed at the location of spectators supporting one team (e.g., the home team); and at least one other stereo microphone (e.g., microphone 103 implemented as a Sennheiser MKH416 microphone or another stereo microphone) directed at the location of spectators supporting the other team (e.g., the visiting team).
[0203] The broadcast system of the present invention may include a mobile unit (which may be a truck and is sometimes referred to as a "matching truck") located outside the stadium (or other event location) that is the first recipient of audio feeds from microphones in the stadium (or other event location). The matching truck generates an object-based audio program (to be broadcast) by encoding audio content from the microphones for delivery as object channels of the program, generating corresponding object-related metadata (e.g., metadata indicating the spatial location where each object should be rendered) and including such metadata in the program, and encoding audio content from some of the microphones for delivery as a speaker channel soundbed (and a set of replacement speaker channels) for the program.
[0204] For example, in Figure 8 In the system, a console 104, an object processing subsystem 106 (coupled to the output of the console 104), an embedding subsystem 108, and a contribution encoder 110 can be mounted in matching trucks. The object-based audio program generated in subsystem 106 can be combined (e.g., in subsystem 108) with video content (e.g., from cameras placed in a stadium) to generate a combined audio and video signal, which is then encoded (e.g., by encoder 110) to generate an encoded audio / video signal for broadcast (e.g., via Figure 5It should be understood that a playback system for decoding and rendering such an encoded audio / video signal will include a subsystem for parsing the audio content and video content of the delivered audio / video signal (not specifically shown in the figure) and a subsystem for decoding and rendering the audio content according to an embodiment of the present invention (e.g., Figure 6 system), and another subsystem for decoding and rendering video content (not specifically shown in the figure).
[0205] The audio output of the console 104 includes a 5.1 speaker channel sound bed (in the center channel) indicating a default mix of ambient sounds captured at the sporting event and commentary by the announcer (non-ambient content) mixed into its center channel. Figure 8 Indicates the ambient content of the center channel of the sound bed without the replacement speaker channels for comment (in Figure 8 ) (i.e., the captured ambient sound content of the center channel of the sound bed before the commentary is mixed with it to generate the center channel of the sound bed); audio content of stereo object channels indicating crowd noise from fans of the home team present at the event (labeled as "2.0 Home Team"); audio content of stereo object channels indicating crowd noise from fans of the away team present at the event (labeled as "2.0 Away Team"); object channel audio content indicating commentary made by an announcer from the home team's city (labeled as "1.0 Commentary 1"); object channel audio content indicating commentary made by an announcer from the away team's city (labeled as "1.0 Commentary 2"); and object channel audio content indicating the sound produced by the winning goal when it is hit by a participant in a sports event (labeled as "1.0 Kick").
[0206] The object processing subsystem 106 is configured to organize (e.g., group) the audio stream from the console 104 into object channels (e.g., grouping the left and right audio streams labeled "2.0 Away Team" into an Away Team crowd noise object channel) and / or object channel groups, generate object-related metadata indicating the object channels (and / or object channel groups), and encode the object channels (and / or object channel groups), the object-related metadata, the speaker channel soundbed, and each replacement speaker channel (determined based on the audio stream from the console 104) into an object-based audio program (e.g., an object-based audio program encoded as a Dolby E bitstream). Typically, the subsystem 106 is also configured to render (and play on a set of studio monitor speakers) at least a selected subset of the object channels (and / or object channel groups) and the speaker channel soundbed and / or replacement speaker channels (including generating a mix indicating the selected object channels and speaker channels using the object-related metadata) so that the played back sound can be monitored by an operator of the console 104 and the subsystem 106 (e.g., via Figure 8 Monitor Path).
[0207] The interface between the output of subsystem 104 and the input of subsystem 106 may be a multi-channel audio digital interface ("MADT").
[0208] In operation, Figure 8 Subsystem 108 of the system combines the object-based audio program generated in subsystem 106 with video content (e.g., from cameras placed in the stadium) to generate a combined audio and video signal that is provided to encoder 110. The interface between the output of subsystem 108 and the input of subsystem 110 may be a high-definition serial digital interface ("HD-SDI"). In operation, encoder 110 encodes the output of subsystem 108 to generate an encoded audio / video signal for broadcast (e.g., via Figure 5 delivery subsystem 5).
[0209] In some embodiments, a broadcast facility (e.g., Figure 8 The subsystems 106, 108 and 110 of the system are configured to generate a plurality of object-based audio programs indicative of the captured sounds (e.g., by Figure 8The present invention also provides an object-based audio program indicated by the plurality of encoded audio / video signals output by the subsystem 110 of the system. Examples of such object-based audio programs include 5.1 flattened mixes, international mixes, and national mixes. For example, all programs may include a common speaker channel soundbed (and a common set of alternate speaker channels), but the program's object channels (and / or a menu of selectable object channels determined by the program, and / or selectable or non-selectable rendering parameters for rendering and mixing the object channels) may vary from program to program.
[0210] In some embodiments, a broadcaster or other content creator's facility (e.g., Figure 8 The system's subsystems 106, 108, and 110 are configured to generate a single object-based audio program (i.e., a master) that can be rendered in any of a variety of different playback environments (e.g., a 5.1-channel domestic playback system, a 5.1-channel international playback system, and a stereo playback system). The master does not need to be mixed (e.g., downmixed) for playback to clients in any particular environment.
[0211] As noted above, in some embodiments of the invention, the program's object-related metadata (or a preconfiguration of the playback or rendering system that is not indicated by metadata delivered with the program) provides constraints or conditions on the optional mixing of objects and speaker channel content. For example, Figure 6 The implementation of the system can implement DRM levels to enable users to have tiered access to groups of object channels included in an object-based audio program. If the user pays more money (e.g., to the broadcaster), the user can be authorized to decode, select, and render more object channels of the program.
[0212] Will refer to Figure 9 To describe examples of constraints and conditions on user selection of objects (or groups of objects). Figure 9 In the program "P0", program "P0" includes 7 object channels: object channel "N0" indicating neutral crowd noise; object channel "N1" indicating home team crowd noise; object channel "N2" indicating away team crowd noise; object channel "N3" indicating official comments on the event (for example, broadcast comments by commercial radio announcers); object channel "N4" indicating fan comments on the event; object channel "N5" indicating public speech announcements under the event; and object channel "N6" indicating input Twitter connections related to the event (converted via a text-to-speech system).
[0213] The default indication metadata included in program P0 indicates a default object group (one or more "default" objects) and a default rendering parameter group (e.g., the spatial position of each default object in the default object group) that are (by default) included in the rendered mix of the "sound bed" speaker channel content and object channel content indicated by the program. For example, the default object group may be a mix of object channel "N0" (indicating neutral crowd noise) rendered in a diffuse manner (e.g., so as not to be perceived as emanating from any particular source position) and object channel "N3" (indicating official commentary) rendered so as to be perceived as emanating from a source position directly in front of the listener (i.e., at an azimuth angle of 0 degrees relative to the listener).
[0214] ( Figure 9 The program P0 also includes metadata indicating a plurality of sets of user-selectable preset mixes, each preset mix being determined by a subgroup of the program's object channels and a corresponding set of rendering parameters. The user-selectable preset mixes may be presented as a menu on a user interface of a controller of the playback system (e.g., by Figure 6 For example, one such preset mix is Figure 9 A mix of object channel “N0” (indicative of neutral crowd noise), object channel “N1” (indicative of home team crowd noise), and object channel “N4” (indicative of fan commentary) rendered such that channel N0 and N1 content in the mix is perceived as emanating from a source position directly behind the listener (i.e., at an azimuth angle of 180 degrees relative to the listener), wherein a level of channel N1 content in the mix is 3 dB less than a level of channel N0 in the mix, and wherein channel N4 content in the mix is rendered in a diffuse manner (e.g., so as not to be perceived as emanating from any particular source position).
[0215] The playback system can implement the following rules (for example, in Figure 9 The playback system may also implement the following rule (e.g., in Figure 9 Conditional rule “C1” indicated in , which is determined by the program’s metadata): each user-selectable preset mix that includes content of object channel N0 mixed with content of at least one of object channels N1 and N2 must include content of object channel N0 mixed with content of object channel N1, or it must include content of object channel N0 mixed with content of object channel N2.
[0216] The playback system can also implement the following rules (for example, in Figure 9Conditional rule "C2" indicated in , which is determined by the program's metadata): each user-selectable preset mix that includes content of at least one of object channels N3 and N4 must include content of only object channel 3, or it must include content of only object channel N4.
[0217] Some embodiments of the present invention implement conditional decoding (and / or rendering) of object channels of object-based audio programs. For example, a playback system may be configured to enable object channels to be conditionally decoded based on the playback environment or the user's permissions. For example, if a DRM hierarchy is implemented to enable a customer to have "tiered" access to groups of audio object channels included in an object-based audio program, the playback system may be automatically configured (via control bits included in the program's metadata) to prevent the decoding and rendering of some objects unless the playback system is notified that the user has met at least one condition (e.g., a specific amount of money has been paid to the content provider). For example, a user may need to purchase permission to listen to a program that is not subject to the conditions set forth in the program. Figure 9 The "Official Comments" object channel N3 of program P0, and the playback system can achieve Figure 9 The conditional rule "C2" indicated in , makes the object channel N3 unable to be selected unless the playback system is notified that the user of the playback system has purchased the necessary rights.
[0218] For another example, the playback system may be automatically configured (via control bits included in the program's metadata indicating the specific format of the available playback speaker array) to: Figure 9 The conditional rule "C1" indicated in , which makes the preset mix of object channels N0 and N1 not selectable unless the playback system is notified that a 5.1 speaker array is available to render the selected content, but this is not the case if the only available speaker array is a 2.0 speaker array), prevents decoding and selection of some objects.
[0219] In some embodiments, the present invention implements rule-based object channel selection, wherein at least one predetermined rule determines which object channel(s) of an object-based audio program are rendered (e.g., along with the speaker channel sound bed). A user may also specify at least one rule for object channel selection (e.g., by selecting from a menu of available rules presented by a user interface of a playback system controller), and the playback system (e.g., Figure 6 The object processing subsystem 22 of the system may be configured to apply each such rule to determine which object channel(s) of the object-based audio program to be rendered should be included in the audio program to be rendered (e.g., by Figure 6The playback system can determine which object channels of the program meet the predetermined rules based on the object-related metadata in the program.
[0220] For a simple example, consider the following situation: an object-based audio program indicates a sporting event. Instead of manipulating a controller (e.g., Figure 6 The playback system may use a controller 23) to perform static selection of a particular group of objects included in a program (e.g., radio commentary from a particular team or car or bike), and the user manipulates the controller to establish rules (e.g., automatically selecting object channels for rendering that indicate whichever team or car or bike wins or comes first). The playback system applies the rules to implement dynamic selection (during rendering of a single program or a series of different programs) of a series of different subgroups of objects (object channels) included in a program (e.g., a first subgroup of objects indicating one team, automatically followed by a second subgroup of objects indicating the second team in the event that the second team scores and thereby becomes the current winning team). Thus, in some such embodiments, the real-time event steers or influences which object channels are included in the rendered mix. The playback system (e.g., Figure 6 The object processing subsystem 22 of the system can respond to metadata included in the program (e.g., metadata indicating that at least one corresponding object indicates the current winning team, such as crowd noise indicating fans of that team or commentary by a radio announcer associated with the winning team) to select which object channel(s) should be included in the mix of speaker channels and object channels to be rendered. For example, a content creator can include (in an object-based audio program) metadata indicating the placement order (or other hierarchy) of each of at least some of the program's audio object channels (e.g., indicating which object channels correspond to the currently first-place team or car, which object channels correspond to the second-place team or car, etc.). The playback system can be configured to respond to such metadata by selecting and rendering only the object channels that meet user-specified rules (e.g., object channels associated with the "n"-th place team, as indicated by the program's object-related metadata).
[0221] Examples of object-related metadata associated with object channels of the object-based audio program of the present invention include (but are not limited to): metadata indicating detailed information about how to render the object channel; dynamic temporal metadata (e.g., indicating the object's pan trajectory, object size, gain, etc.); and metadata used by the AVR (or other devices or systems downstream of the decoding and object processing subsystems of some implementations of the system of the present invention) to render the object channel (e.g., using organizational knowledge of the available playback speaker array). Such metadata may specify constraints on object position, gain, muting, or other rendering parameters and / or constraints on how an object interacts with other objects (e.g., constraints on which additional objects can be selected if a particular object is selected), and / or may specify default objects and / or default rendering parameters (to be used if the user does not select other objects and / or rendering parameters).
[0222] In some embodiments, at least some of the object-related metadata (and optionally also at least some of the object channels) of the object-based audio program of the present invention is sent in a separate bitstream or other container (e.g., a submix that a user may need to pay extra to receive and / or use) from the program's speaker channel soundbed and conventional metadata. Without access to such object-related metadata (or both), a user can decode and render the speaker channel soundbed, but cannot select the program's audio objects and cannot render the program's audio objects in a mix with the audio indicated by the speaker channel soundbed. Each frame of the object-based audio program of the present invention may include audio content for multiple object channels and corresponding object-related metadata.
[0223] An object-based audio program generated (or transmitted, stored, buffered, decoded, rendered, or otherwise processed) according to some embodiments of the present invention includes a speaker channel soundbed, at least one alternative speaker channel, at least one object channel, and metadata indicating a layered graph (sometimes referred to as a layered "mixing graph") that indicates optional mixes of the speaker channels and object channels (e.g., all optional mixes). For example, the mixing graph indicates each rule that can be applied to the selection of a subset of speaker channels and object channels. Typically, an encoded audio bitstream indicates at least some (i.e., at least a portion) of the program's audio content (e.g., the speaker channel soundbed and at least some of the program's object channels) and object-related metadata (including metadata indicating the mixing graph), and optionally at least one additional encoded audio bitstream or file indicates some of the program's audio content and / or object-related metadata.
[0224] The hierarchical audio mix graph indicates nodes (each node may indicate an optional channel or channel group, or a classification of optional channels or channel groups) and connections between nodes (e.g., control interfaces to the nodes and / or rules for selecting channels), and includes essential data (a "base" layer) as well as optional (i.e., optionally omitted) data (at least one "extension" layer). Typically, the hierarchical audio mix graph is included in one of the coded audio bitstreams indicating a program, and may be evaluated by a graph traversal (implemented by a playback system, such as an end-user's playback system) to determine a default mix of channels and options for modifying the default mix.
[0225] In the case where the mix graph can be represented as a tree graph, the base layer can be one branch (or two or more branches) of the tree graph, and each extended layer can be another branch (or another set of two or more branches) of the tree graph. For example, one branch of the tree graph (indicated by the base layer) can indicate optional channels and channel groups available to all end users, and another branch of the tree graph (indicated by the extended layer) can indicate additional optional channels and / or channel groups available to only some end users (e.g., such an extended layer can be provided only to end users who are authorized to use it). Figure 9 is an example of a tree graph including object channel nodes (eg, nodes indicating object channels N0, N1, N2, N3, N4, N5, and N6) of a mixing graph and other elements.
[0226] Typically, the base layer contains (indicates) the graph structure and control interfaces to the nodes of the graph (eg, pan control interface and gain control interface).The base layer is necessary to map any user interaction to the decoding / rendering process.
[0227] Each extension layer contains (indicates) an extension to the base layer.The extension is not immediately necessary for mapping user interaction to the decoding process and therefore it may be transmitted at a slower rate and / or delayed, or omitted.
[0228] In some implementations, the base layer is included as metadata for a separate substream of the program (eg, transmitted as metadata for a separate substream).
[0229] According to some embodiments of the present invention, an object-based audio program generated (or transmitted, stored, buffered, decoded, rendered, or otherwise processed) includes a speaker channel soundbed, at least one alternative speaker channel, at least one object channel, and metadata indicating a mix map (which may or may not be a layered mix map) indicating optional mixes of the speaker channels and object channels (e.g., all optional mixes). An encoded audio bitstream (e.g., a Dolby E or E-AC-3 bitstream) indicates at least a portion of the program, and metadata indicating the mix map (and typically also optional object channels and / or speaker channels) is included in each frame of the bitstream (or in each frame of a subset of frames of the bitstream). For example, each frame may include at least one metadata segment and at least one audio data segment, and the mix map may be included in at least one metadata segment of each frame. Each metadata segment (which may be referred to as a "container") may have a format that includes a metadata segment header (and optionally other elements) and one or more payloads following the metadata segment header. Each metadata payload is itself identified by a payload header. If the mixing map exists in the metadata segment, the mixing map is included in one of the metadata payloads of the metadata segment.
[0230] In some embodiments, an object-based audio program generated (or transmitted, stored, buffered, decoded, rendered, or otherwise processed) according to the present invention includes at least two speaker channel beds, at least one object channel, and metadata indicating a mixing graph (which may or may not be a layered mixing graph). The mixing graph indicates optional mixes of the speaker channels and the object channels and includes at least one "bedmix" node. Each "bedmix" node defines a predetermined mix of the speaker channel beds and thereby indicates or implements a predetermined set of mixing rules (optionally with user-selectable parameters) for mixing the speaker channels of two or more speaker beds of the program.
[0231] Consider the following example: the audio program is associated with a soccer (football) game between Team A (the home team) and Team B in a stadium, and includes a 5.1 speaker channel soundbed for the entire crowd in the stadium (determined by microphone feeds), a stereo feed for the portion of the crowd biased towards Team A (i.e., audio captured from spectators located in the portion of the stadium occupied primarily by fans of Team A), and another stereo feed for the portion of the crowd biased towards Team B (i.e., audio captured from spectators located in the portion of the stadium occupied primarily by fans of Team B). These three feeds (a 5.1 channel neutral bed, a 2.0 channel "Team A" bed, and a 2.0 channel "Team B" bed) can be mixed on a mixing console to generate four 5.1 speaker channel beds (which can be referred to as "fan zone" beds): unbiased, home team biased (a mix of the neutral and Team A beds), away team biased (a mix of the neutral and Team B beds), and opposite (the neutral bed mixed with the Team A bed panned to one side of the space and the Team B bed panned to the opposite side of the space). However, transmitting four mixed 5.1 channel beds is expensive in terms of bit rate. Thus, bitstream embodiments of the present invention include metadata that specifies bed mixing rules (for mixing of speaker channel beds, e.g., to generate four specified mixes of a 5.1 channel bed) to be implemented by a playback system (e.g., in an end user's home) based on user mix selections, and the speaker channel beds that may be mixed according to those rules (e.g., the original 5.1 channel bed and two biased stereo speaker channel beds). In response to the bed mix node of the mix graph, the playback system may present options to the user (e.g., via a user interface defined by Figure 6 In response to the user selection of the 5.1 channel sound bed of the mix, the playback system (e.g., Figure 6 The subsystem 22 of the system will use the (unmixed) speaker channel sound bed transmitted in the bitstream to generate the selected mix.
[0232] In some embodiments, the bed mixing rules contemplate the following operations (which may have predetermined parameters or user selectable parameters):
[0233] Sound bed "rotation" (i.e., panning the speaker channel sound beds to the left, right, front, or back). For example, to create the "opposite mix" mentioned above, the stereo Team A sound bed will be rotated to the left of the playback speaker array (with the L and R channels of the Team A sound bed mapped to the L and Ls channels of the playback system), and the stereo Team B sound bed will be rotated to the right of the playback speaker array (with the L and R channels of the Team B sound bed mapped to the R and Rs channels of the playback system). Thus, the user interface of the playback system will present the end user with a choice of one of the four aforementioned "unbiased" sound bed mixes, "home team biased" sound bed mixes, "away team biased" sound bed mixes, and "opposite" sound bed mixes, and upon user selection of the "opposite" sound bed mix, the playback system will implement the appropriate sound bed rotation during rendering of the "opposite" sound bed mix; and
[0234] A sudden duck (i.e., attenuation) of a specific speaker channel (target channel) in a sound bed mix (typically, to create headroom). For example, in the soccer game example mentioned above, the user interface of the playback system may present the end user with a choice of one of the four aforementioned "unbiased" sound bed mixes, "home team biased" sound bed mixes, "away team biased" sound bed mixes, and "relative" sound bed mixes, and in response to the user selecting the "relative" sound bed mix, the playback system may achieve the target sudden duck during rendering of the "relative" sound bed mix by causing each of the L, Ls and R, Rs channels of the neutral 5.1 channel sound bed to be ducked (attenuated) by a predetermined amount (specified by metadata in the bitstream) before mixing the attenuated 5.1 channel sound bed with the stereo "Team A" and "Team B" sound beds to generate the "relative" sound bed mix.
[0235] In another class of embodiments, an object-based audio program generated (or transmitted, stored, buffered, decoded, rendered, or otherwise processed) according to the present invention includes substreams, and the substreams indicate at least one speaker channel soundbed, at least one object channel, and object-related metadata. The object-related metadata includes "substream" metadata (indicating the substream structure of the program and / or the manner in which the substreams should be decoded) and typically also includes a mix map indicating optional mixes (e.g., all optional mixes) of the speaker channels and object channels. The substream metadata can indicate which substreams of the program should be decoded independently of other substreams of the program and which substreams of the program should be decoded in association with at least one other substream of the program.
[0236] For example, in some embodiments, an encoded audio bitstream indicates the audio content of at least some (i.e., at least a portion) of a program (e.g., at least one speaker channel soundbed, at least one alternate speaker channel, and at least some of the program's object channels) and metadata (e.g., a mix map and substream metadata, and optionally other metadata), and at least one additional encoded audio bitstream (or file) indicates the audio content and / or metadata of some programs. In the case where each bitstream is a Dolby E bitstream (or is encoded in a manner consistent with the SMPTE 337 format for carrying non-PCM data in an AES3 serial digital audio bitstream), the bitstreams can collectively indicate multiple, up to 8 channels of audio content, with each bitstream carrying up to 8 channels of audio data and typically also including metadata. Each bitstream can be viewed as a substream of a combined bitstream indicating all audio data and metadata carried by all bitstreams.
[0237] For another example, in some embodiments, an encoded audio bitstream indicates multiple metadata substreams (e.g., mix maps and substream metadata, and optionally other object-related metadata) and the audio content of at least one audio program. Typically, each substream indicates one or more of the program's channels (and typically metadata). In some cases, the multiple substreams of the encoded audio bitstream indicate the audio content of several audio programs, such as a "main" audio program (which can be a multi-channel program) and at least one other audio program (e.g., a program that is a commentary on the main audio program).
[0238] An encoded audio bitstream indicating at least one audio program necessarily includes at least one "independent" substream of audio content. The independent substream indicates at least one channel of the audio program (e.g., the independent substream may indicate the five full-range channels of a conventional 5.1-channel audio program). This audio program is referred to herein as a "main" program.
[0239] In some cases, an encoded audio bitstream indicates two or more audio programs (a "main" program and at least one other audio program). In such cases, the bitstream includes two or more independent substreams: a first independent substream indicating at least one channel of the main program; and at least one other independent substream indicating at least one channel of another audio program (a program different from the main program). Each independent bitstream can be independently encoded, and the decoder can be operated to decode only a subset (rather than all) of the independent substreams of the encoded bitstream.
[0240] Optionally, the coded audio bitstream indicating the main program (and optionally at least one other audio program) includes at least one "slave" substream of audio content. Each slave substream is associated with an independent substream of the bitstream and indicates at least one additional channel of the program (e.g., the main program) whose content is indicated by the associated independent substream (i.e., the slave substream indicates at least one channel of the program not indicated by the associated independent substream, while the associated independent substream indicates at least one channel of the program).
[0241] In the example of an encoded bitstream including an independent substream (indicating at least one channel of a main program), the bitstream also includes a dependent substream (associated with the independent substream) indicating one or more additional speaker channels of the main program. Such additional speaker channels are in addition to the main program channels indicated by the independent substream. For example, if the independent substream indicates the standard format left, right, center, left surround, and right surround full-range speaker channels of a 7.1-channel main program, the dependent substream may indicate two additional full-range speaker channels of the main program.
[0242] According to the E-AC-3 standard, a conventional E-AC-3 bitstream must indicate at least one independent substream (e.g., a single AC-3 bitstream) and can indicate up to 8 independent substreams. Each independent substream of an E-AC-3 bitstream can be associated with up to 8 dependent substreams.
[0243] In (refer to Figure 11 In an exemplary embodiment to be described, an object-based audio program includes at least one speaker channel soundbed, at least one object channel, and metadata. The metadata includes "substream" metadata (indicating the substream structure of the program's audio content and / or the manner in which the substreams of the program's audio content should be decoded), and typically also includes a mix map indicating an optional mix of the speaker channels and the object channels. The audio program is associated with an English soccer game. The encoded audio bitstream (e.g., an E-AC-3 bitstream) indicates the program's audio content and metadata. Figure 11 As shown, the audio content of the program (and thus the audio content of the bitstream) comprises four independent substreams. Figure 11 The other independent substream (labeled as substream "I0" in the Figure 11The third independent substream (labeled as substream "I1") in FIG11A ) indicates a 2.0 channel "Team A" sound bed ("M Crowd") indicating the sound of the portion of the match crowd that is biased towards one team ("Team A"), a 2.0 channel "Team B" sound bed ("LivP Crowd") indicating the sound of the portion of the match crowd that is biased towards the other team ("Team B"), and a mono object channel ("Sky Commentary 1") indicating the commentary on the match. Figure 11 The fourth independent sub-stream (labeled as “I2” in the figure) indicates an object channel audio content (labeled as “2 / 0” goal) indicating a sound generated by a winning goal when the ball is hit by a participant in a soccer event, and three object channels (“Sky Commentary 2”, “Man Commentary”, and “Liv Commentary”), the object channel audio content (labeled as “2 / 0” goal) indicating a sound generated by a winning goal when the ball is hit by a participant in a soccer event, and each object channel (“Sky Commentary 2”, “Man Commentary”, “Liv Commentary”) indicating a different commentary on the soccer match. Figure 11 ) indicates an object channel (labeled as "PA"), an object channel (labeled as "Radio"), and an object channel (labeled as "Goal Break"), where the object channel (labeled as "PA") indicates the sound produced by the stadium public address system at a soccer game, the object channel (labeled as "Radio") indicates the radio broadcast of the soccer game, and the object channel (labeled as "Goal Break") indicates the scoring scored during the soccer game.
[0244] exist Figure 11 In this example, substream I0 includes a mix map and metadata ("obj md") for a program, where the metadata ("obj md") includes at least some substream metadata and at least some object channel-related metadata. Each of substreams I1, I2, and I3 includes metadata ("obj md") including at least some object channel-related metadata and, optionally, at least some substream metadata.
[0245] exist Figure 11In an example, the substream metadata of the bitstream indicates that coupling should be "off" (such that each independent substream is decoded independently of the other independent substreams) between each pair of independent substreams during decoding, and the substream metadata of the bitstream indicates that coupling should be "on" (such that these channels are not decoded independently of each other) or "off" (such that these channels are decoded independently of each other) for program channels within each substream. For example, the substream metadata indicates that coupling should be "on" within each of the two stereo speaker channel sound beds of substream I1 (a 2.0-channel "Team A" sound bed and a 2.0-channel "Team B" sound bed), but coupling should be disabled across the speaker channel sound bed of substream I1 and between the mono object channel and each of the speaker channel sound beds of substream I1 (such that the mono object channel and the speaker channel sound bed are decoded independently of each other). Similarly, the substream metadata indicates that coupling should be "on" (such that the speaker channels of that sound bed are decoded in association with each other) within the 5.1 speaker channel sound bed of substream I0.
[0246] In some embodiments, speaker channels and object channels are included ("packaged") in substreams of an audio program in a manner appropriate to the mixing graph of the program. For example, if the mixing graph is a tree graph, all channels of one branch of the graph may be included in one substream, and all channels of another branch of the graph may be included in another substream.
[0247] Figure 10 It is a block diagram of a system that implements an embodiment of the present invention.
[0248] Figure 10 The object processing system (object processor) 200 of the system includes a metadata generation subsystem 210, an intermediate (mezzanine) encoder 212, and a simulation subsystem 211 coupled as shown. The metadata generation subsystem 210 is coupled to receive a captured audio stream (e.g., a stream indicating sound captured by a microphone located at the viewing event, and optionally other audio streams) and is configured to organize (e.g., group) the audio stream from the console 104 into speaker channel beds, replacement speaker channel groups, and a plurality of object channels and / or object channel groups. The subsystem 210 is also configured to generate object-related metadata indicating the object channels (and / or object channel groups). The encoder 212 is configured to encode the object channels (and / or object channel groups), the object-related metadata, and the speaker channels into an intermediate type of object-based audio program (e.g., an object-based audio program encoded as a Dolby E bitstream).
[0249] The simulation subsystem 211 of the object processor 200 is configured to render at least a selected subset of object channels (and / or object channel groups) and speaker channels (and play them on a set of studio monitor speakers) (including generating a mix indicating the selected object channels and speaker channels by using object-related metadata) so that the played back sound can be monitored by an operator of the subsystem 200.
[0250] Figure 10 The transcoder 202 of the system includes an intermediate decoder subsystem (intermediate decoder) 213 and an encoder 214 coupled as shown. The intermediate decoder 213 is coupled and configured to receive and decode the intermediate type of object-based audio program output from the object processor 200. The decoded output of the decoder 213 is re-encoded by the encoder 214 into a format suitable for broadcast. In one embodiment, the encoded object-based audio program output from the encoder 214 is an E-AC-3 bitstream (and thus the encoder 214 is Figure 10 2. In other embodiments, the encoded object-based audio program output from encoder 214 is an AC-3 bitstream or has some other format. The object-based audio program output by transcoder 202 is broadcast (or otherwise delivered) to a large number of end users.
[0251] The decoder 204 is included in one such end-user's playback system. The decoder 204 includes a decoder 215 and a rendering subsystem (renderer) 216 coupled as shown. The decoder 215 accepts (receives or reads) and decodes the object-based audio program delivered from the transcoder 202. If the decoder 215 is configured according to a typical embodiment of the present invention, the output of the decoder 215 in typical operation includes: an audio sample stream indicating the speaker channel soundbed of the program; and an audio sample stream indicating the object channels of the program (e.g., user-selectable audio object channels) and a corresponding object-based metadata stream. In one embodiment, the encoded object-based audio program input to the decoder 215 is an E-AC-3 bitstream, and thus the decoder 215 Figure 10 It is marked as "DD+decoder".
[0252] The renderer 216 of the decoder 204 includes an object processing subsystem that is coupled to receive the decoded speaker channels, object channels, and object-related metadata of the delivered program (from the decoder 215). The renderer 216 also includes a rendering subsystem configured to render the audio content determined by the object processing subsystem for playback by the speakers (not shown) of the playback system.
[0253] Typically, the object processing subsystem of the renderer 216 is configured to output a selected subset of the full set of object channels indicated by the program, along with corresponding object-related metadata, to the rendering subsystem of the renderer 216. The object processing subsystem of the renderer 216 is also typically configured to pass the decoded speaker channels from the decoder 215 unchanged (to the rendering subsystem). According to embodiments of the present invention, the object channel selection performed by the object processing subsystem is determined, for example, by rules (e.g., indicating conditions and / or constraints) that the renderer 216 has been programmed or otherwise configured to implement and / or by user selection.
[0254] Figure 10 Each element 200, 202 and 204 (and Figure 8 Each element 104, 106, 108 and 110) can be implemented as a hardware system. The input to such a hardware implementation of processor 200 (or processor 106) is typically a Multi-Channel Audio Digital Interface ("MADI") input. Typically, Figure 8 The processor 106 and Figure 10 Each encoder 212, 214 includes a frame buffer. Typically, a frame buffer is a buffer memory coupled to receive an encoded input audio bitstream, and in operation the buffer memory stores (e.g., in a non-transitory manner) at least one frame of the encoded audio bitstream, and a series of frames of the encoded audio bitstream are set from the buffer memory to a downstream device or system. Furthermore, typically, Figure 10 Each decoder 213, 215 includes a frame buffer. Typically, the frame buffer is a buffer memory coupled to receive an encoded input audio bitstream, and in operation, the buffer memory stores (e.g., in a non-transitory manner) at least one frame of the encoded audio bitstream to be decoded by the decoder 213 or 215.
[0255] Figure 8 Processor 106 (or Figure 10 Any component or element of the subsystems 200, 202 and / or 204) may be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASICs, FPGAs, or other integrated circuits).
[0256] It should be understood that in some embodiments, the object-based audio program of the present invention is generated and / or delivered as an unencoded (e.g., baseband) representation indicating program content (including metadata). For example, such a representation may include PCM audio samples and associated metadata. The unencoded (uncompressed) representation may be delivered in any of a variety of ways, including as at least one data file (e.g., stored in a non-transitory manner in a memory such as on a computer-readable medium) or as a bitstream in AES-3 format or a serial digital interface (SDI) format (or another format).
[0257] One aspect of the present invention is an audio processing unit (APU) configured to perform any embodiment of the method of the present invention. Examples of APUs include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems (pre-processors), post-processing systems (post-processors), audio bitstream processing systems, and combinations of these elements.
[0258] In one class of embodiments, the present invention is an APU comprising a buffer memory (buffer) that stores (e.g., in a non-transitory manner) at least one frame or other segment of an object-based audio program (including the audio content of the speaker channel sound bed and the object channels, and object-related metadata) that has been generated by any embodiment of the method of the present invention. For example, Figure 5 The generating unit 3 may include a buffer 3A that stores (e.g., in a non-transitory manner) at least one frame or other segment of the object-based audio program generated by the unit 3 (including the audio content of the speaker channel sound bed and the object channels, and object-related metadata). For another example, Figure 5 The decoder 7 may include a buffer 7A that stores (e.g., in a non-transitory manner) at least one frame or other segment of an object-based audio program (including audio content of speaker channel sound beds and object channels, and object-related metadata) delivered from the subsystem 5 to the decoder 7.
[0259] Embodiments of the present invention may be implemented in hardware, firmware, or software, or a combination thereof (e.g., as a programmable logic array). For example, the present invention may be implemented in hardware, firmware, or software, or a combination thereof (e.g., as a programmable logic array). Figure 8 or Figure 7 subsystem 106 of the system, or Figure 6 all or some of the elements 20, 22, 24, 25, 26, 29, 35, 31 and 35 of the system, or Figure 10All or some of the elements 200, 202, and 204 are implemented as, for example, a programmed general-purpose processor, digital signal processor, or microprocessor. Unless otherwise indicated, the algorithms or processes included as part of the present invention are not inherently related to any particular computer or other device. In particular, various general-purpose machines may be used with programs written according to the teachings herein, or it may be more convenient to construct more specialized devices (e.g., integrated circuits) to perform the required method steps. Thus, the present invention may be implemented on one or more programmable computer systems (e.g., Figure 6 Each programmable computer system includes at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. The program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.
[0260] Each such program can be implemented in any desired computer language (including machine, assembly or high-level procedural, logical or object-oriented programming languages) to communicate with a computer system. In any case, the language can be a compiled language or an interpreted language.
[0261] For example, when implemented by a sequence of computer software instructions, the various functions and steps of the embodiments of the present invention may be implemented by a multi-threaded software instruction sequence running in appropriate digital signal processing hardware, in which case the various devices, steps and functions of the embodiments may correspond to portions of the software instructions.
[0262] Each such computer program is preferably stored on or downloaded to a storage medium or device (e.g., solid-state memory or media, or magnetic or optical media) readable by a general or special purpose programmable computer for configuring and operating the computer when the storage medium or device is read by a computer system to perform the processes described herein. The system of the present invention may also be implemented as a computer-readable storage medium configured with (i.e., storing) a computer program, wherein the storage medium so configured causes the computer system to operate in a specific and predefined manner to perform the functions described herein.
[0263] A number of embodiments of the present invention have been described. It will be appreciated that various modifications may be made without departing from the spirit and scope of the present invention. In light of the above teachings, numerous modifications and variations may be made to the present invention. It will be appreciated that, within the scope of the appended claims, the present invention may be practiced otherwise than as specifically described herein.
[0264] The present invention includes the following technical solutions.
[0265] Solution 1. A method for generating an object-based audio program indicating audio content, the audio content including first non-ambient content, second non-ambient content different from the first non-ambient content, and third content different from the first non-ambient content and the second non-ambient content, the method comprising the steps of:
[0266] determining an object channel group comprising N object channels, wherein a first subset of the object channel group indicates the first non-ambient content, the first subset comprising M object channels in the object channel group, each of N and M being an integer greater than zero, and M being equal to or less than N;
[0267] determining a speaker channel bed indicative of a default mix of audio content, wherein an object-based speaker channel subset comprising M speaker channels in the sound bed indicates the second non-ambient content, or a mix of at least some audio content of the default mix and the second non-ambient content;
[0268] determining a set of M replacement speaker channels, wherein each replacement speaker channel in the set of M replacement speaker channels indicates some but not all content of a corresponding speaker channel in the object-based subset of speaker channels;
[0269] generating metadata indicating at least one selectable predetermined alternative mix of content of at least one of the object channels and content of predetermined ones of the speaker channels and / or the replacement speaker channels of the sound bed, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix indicating at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content; and
[0270] generating the object-based audio program comprising the speaker channel sound bed, the set of M replacement speaker channels, the object channel groups, and the metadata, such that the speaker channel sound bed is renderable without use of the metadata to provide sound that is perceived as the default mix, and the replacement mix is renderable in response to at least some of the metadata to provide sound that is perceived as a mix that includes at least some of the audio content of the sound bed and the first non-ambient content but not the second non-ambient content.
[0271] Option 2. A method according to Option 1, wherein at least some of the metadata is optional content metadata, which indicates a set of optional predetermined mixes of the audio content of the program and includes a predetermined rendering parameter set for each of the predetermined mixes.
[0272] Option 3. A method according to any one of Options 1 to 2, wherein the object-based audio program is a coded bitstream comprising frames, the coded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, each of the frames of the coded bitstream indicates at least one data structure, the at least one data structure is a container comprising some content of the object channel and some of the metadata, and at least one of the containers is included in an auxiliary data auxdata field or an additional bitstream information addbsi field of each frame.
[0273] Solution 4. The method according to any one of Solutions 1 to 2, wherein the object-based audio program is a Dolby E bitstream comprising a series of bursts and guard bands between burst pairs.
[0274] Option 5. A method according to any one of Options 1 to 2, wherein the object-based audio program is an unencoded representation of the audio content and the metadata indicating the program, and the unencoded representation is a bitstream or at least one data file stored in a memory in a non-transitory manner.
[0275] Option 6. A method according to any one of Options 1 to 5, wherein at least some of the metadata indicates a layered mixing graph, the layered mixing graph indicates an optional mix of the speaker channels, the replacement speaker channels and the object channels of the sound bed, and the layered mixing graph includes a basic layer of metadata and at least one extended layer of metadata.
[0276] Option 7. A method according to any one of Options 1 to 6, wherein at least some of the metadata indicates a mixing graph, the mixing graph indicating an optional mix of the speaker channels, the replacement speaker channels and the object channels of the sound bed, the object-based audio program is an encoded bitstream comprising frames, and each of the frames of the encoded bitstream includes metadata indicating the mixing graph.
[0277] Option 8. The method according to any one of options 1 to 7, wherein the object-based audio program indicates captured audio content.
[0278] Option 9. The method according to any one of Options 1 to 8, wherein the default mix is a mix of ambient content and non-ambient content.
[0279] Option 10. The method according to any one of Options 1 to 9, wherein the third content is environmental content.
[0280] Option 11. A method according to Option 10, wherein the ambient content indicates ambient sounds when viewing an event, the first non-ambient content indicates a commentary on the viewing event, and the second non-ambient content indicates an alternative commentary on the viewing event.
[0281] Embodiment 12. A method of rendering audio content determined by an object-based audio program, wherein the program indicates a speaker channel sound bed, a set of M replacement speaker channels, an object channel group, and metadata, wherein the object channel group includes N object channels, a first subset of the object channel group indicates first non-ambient content, the first subset includes the M object channels in the object channel group, each of N and M is an integer greater than zero, and M is equal to or less than N,
[0282] the speaker channel sound bed indicates a default mix of audio content including second non-ambient content different from the first non-ambient content, wherein an object-based speaker channel subset comprising M speaker channels in the sound bed indicates a mix of the second non-ambient content, or at least some audio content of the default mix, and the second non-ambient content,
[0283] each replacement speaker channel in the set of M replacement speaker channels indicates some, but not all, content of a corresponding speaker channel of the object-based subset of speaker channels, and
[0284] The metadata indicates at least one selectable predetermined alternative mix of the content of at least one of the object channels and the content of predetermined ones of the speaker channels of the sound bed and / or the replacement speaker channels, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix including at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content, the method comprising the steps of:
[0285] (a) providing the object-based audio program to an audio processing unit; and
[0286] (b) parsing the speaker channel sound bed in the audio processing unit and rendering the default mix in response to the speaker channel sound bed without using the metadata.
[0287] Solution 13. The method according to Solution 12, wherein the audio processing unit is configured to parse the object channel and the metadata of the program, and the method further comprises the steps of:
[0288] (c) rendering, in the audio processing unit, the alternative mix using at least some of the metadata, including rendering by selecting and mixing content of the first subset of the object channel group and at least one of the alternative speaker channels in response to at least some of the metadata.
[0289] Option 14. The method according to Option 13, wherein step (c) comprises the steps of:
[0290] (d) responsive to the at least some metadata, selecting the first subset of the group of object channels, selecting at least one speaker channel in the speaker channel bed other than speaker channels in the subset of object-based speaker channels, and selecting the at least one replacement speaker channel; and
[0291] (e) mixing the content of each speaker channel and the first subset of the subject channel group selected in step (d) to determine the replacement mix.
[0292] Option 15. A method according to any one of Options 13 to 14, wherein step (c) includes the step of: driving a speaker to provide sound that can be perceived as a mix of at least some of the audio content of the sound bed and the first non-ambient content rather than the second non-ambient content.
[0293] Option 16. The method according to any one of options 13 to 15, wherein step (c) comprises the steps of:
[0294] In response to the replacement mix, a speaker feed is generated for driving a speaker to emit sound, wherein the sound includes object channel sounds indicative of the first non-ambient content and the object channel sounds are perceptible as emanating from at least one apparent source location determined by the first subset of the object channel groups.
[0295] 17. The method according to any one of 13 to 16, wherein step (c) comprises the steps of:
[0296] providing a menu of mixes available for selection, each mix of at least one subgroup of the mixes comprising contents of the subgroup of subject channels and the subgroup of replacement speaker channels; and
[0297] The alternative mix is selected by selecting one of the mixes indicated by the menu.
[0298] Option 18. A method according to any one of Options 13 to 17, wherein the menu is presented through a user interface of a controller, the controller is coupled to a set-top device, and the set-top device is coupled to receive the object-based audio program and is configured to perform step (c).
[0299] Option 19. A method according to any one of Options 12 to 18, wherein the object-based audio program includes a set of bitstreams, and wherein step (a) includes the step of: sending the bitstream of the object-based audio program to the audio processing unit.
[0300] Option 20. The method according to any one of Option 12 to Option 19, wherein the default mix is a mix of ambient content and non-ambient content.
[0301] Option 21. A method according to Option 20, wherein the ambient content indicates ambient sounds when viewing an event, the first non-ambient content indicates a commentary on the viewing event, and the second non-ambient content indicates an alternative commentary on the viewing event.
[0302] Scheme 22. A method according to any one of Schemes 12 to 21, wherein the object-based audio program is a coded bitstream comprising frames, the coded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, each of the frames of the coded bitstream indicates at least one data structure, the at least one data structure is a container comprising some content of the object channel and some of the metadata, and at least one of the containers is included in the auxiliary data auxdata field or the additional bitstream information addbsi field of each of the frames.
[0303] Embodiment 23. The method of any one of embodiments 12 to 21, wherein the object-based audio program is a Dolby E bitstream comprising a series of bursts and guard bands between pairs of bursts.
[0304] Option 24. A method according to any one of Options 12 to 21, wherein the object-based audio program is an unencoded representation of the audio content and the metadata indicating the program, and the unencoded representation is a bitstream or at least one data file stored in a memory in a non-transitory manner.
[0305] Option 25. A method according to any one of Options 12 to 24, wherein at least some of the metadata indicates a layered mixing graph, the layered mixing graph indicates an optional mix of the speaker channels, the replacement speaker channels and the object channels of the sound bed, and the layered mixing graph includes a basic layer of metadata and at least one extended layer of metadata.
[0306] Option 26. A method according to any one of Options 12 to 25, wherein at least some of the metadata indicates a mixing graph, the mixing graph indicating an optional mix of the speaker channels, the replacement speaker channels and the object channels of the sound bed, the object-based audio program is an encoded bitstream comprising frames, and each of the frames of the encoded bitstream includes metadata indicating the mixing graph.
[0307] Embodiment 27. A system for generating an object-based audio program indicative of audio content, the audio content comprising first non-ambient content, second non-ambient content different from the first non-ambient content, and third content different from the first non-ambient content and the second non-ambient content, the system comprising:
[0308] A first subsystem is configured to determine:
[0309] an object channel group comprising N object channels, wherein a first subset of the object channel group indicates the first non-ambient content, the first subset comprising M object channels in the object channel group, each of N and M being an integer greater than zero, and M being equal to or less than N,
[0310] a speaker channel bed indicating a default mix of audio content, wherein an object-based speaker channel subset comprising M speaker channels in the bed indicates the second non-ambient content, or a mix of at least some audio content of the default mix and the second non-ambient content, and
[0311] a set of M replacement speaker channels, wherein each replacement speaker channel in the set of M replacement speaker channels indicates some but not all content of a corresponding speaker channel in the object-based speaker channel subset,
[0312] wherein the first subsystem is further configured to generate metadata indicating at least one optional predetermined alternative mix of content of at least one of the object channels and content of predetermined ones of the speaker channels and / or the replacement speaker channels of the sound bed, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix indicating at least some audio content of the sound bed and the first non-ambient content instead of the second non-ambient content; and
[0313] an encoding subsystem coupled to the first subsystem and configured to generate the object-based audio program such that the object-based audio program includes the speaker channel sound bed, the set of M replacement speaker channels, the object channel groups, and the metadata, and such that the speaker channel sound bed is renderable without use of the metadata to provide sound that is perceived as the default mix, and the replacement mix is renderable in response to at least some of the metadata to provide sound that is perceived as a mix that includes at least some of the audio content of the sound bed and the first non-ambient content but not the second non-ambient content.
[0314] Option 28. A system according to Option 27, wherein at least some of the metadata is optional content metadata, which indicates a set of optional predetermined mixes of the audio content of the program and includes a predetermined set of rendering parameters for each of the predetermined mixes.
[0315] Option 29. A system according to any one of Option 27 to Option 28, wherein the default mix is a mix of ambient content and non-ambient content.
[0316] Option 30. A system according to any one of Option 27 to Option 28, wherein the third content is environmental content.
[0317] Option 31. A system according to Option 30, wherein the ambient content indicates ambient sounds when viewing an event, the first non-ambient content indicates a commentary on the viewing event, and the second non-ambient content indicates an alternative commentary on the viewing event.
[0318] Scheme 32. A system according to any one of Schemes 27 to 31, wherein the encoding subsystem is configured to generate the object-based audio program so that the object-based audio program is an encoded bit stream including frames, the encoded bit stream is an AC-3 bit stream or an E-AC-3 bit stream, each of the frames of the encoded bit stream indicates at least one data structure, the at least one data structure is a container including some content of the object channel and some of the metadata, and at least one of the containers is included in the auxiliary data auxdata field or the additional bit stream information addbsi field of each of the frames.
[0319] Option 33. A system according to any one of Options 27 to 32, wherein the encoding subsystem is configured to generate the object-based audio program so that the object-based audio program is a DolbyE bitstream comprising a series of burst pulses and a guard band between burst pulse pairs.
[0320] Option 34. A system according to any one of Options 27 to 33, wherein at least some of the metadata indicates a layered mixing graph, the layered mixing graph indicating an optional mix of the speaker channels, the replacement speaker channels and the object channels of the sound bed, and the layered mixing graph includes a basic layer of metadata and at least one extended layer of metadata.
[0321] Option 35. A system according to any one of Options 27 to 34, wherein at least some of the metadata indicates a mixing graph, the mixing graph indicating an optional mix of the speaker channels, the replacement speaker channels and the object channels of the sound bed, the object-based audio program is an encoded bitstream comprising frames, and each of the frames of the encoded bitstream includes metadata indicating the mixing graph.
[0322] Embodiment 36. An audio processing unit configured to render audio content determined by an object-based audio program, wherein the program indicates a speaker channel bed, a set of M replacement speaker channels, an object channel group, and metadata, wherein the object channel group includes N object channels, a first subset of the object channel group indicates first non-ambient content, the first subset includes the M object channels in the object channel group, each of N and M is an integer greater than zero, and M is equal to or less than N,
[0323] the speaker channel sound bed indicates a default mix of audio content including second non-ambient content different from the first non-ambient content, wherein an object-based speaker channel subset comprising M speaker channels in the sound bed indicates a mix of the second non-ambient content, or at least some audio content of the default mix, and the second non-ambient content,
[0324] each replacement speaker channel in the set of M replacement speaker channels indicates some, but not all, content of a corresponding speaker channel of the object-based subset of speaker channels, and
[0325] The metadata indicates at least one optional predetermined alternative mix of the content of at least one of the object channels and the content of predetermined ones of the speaker channels of the sound bed and / or the replacement speaker channels, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix including at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content, and the audio processing unit comprises:
[0326] a first subsystem coupled to receive the object-based audio program and configured to parse the speaker channel soundbed, the replacement speaker channels, the object channels, and the metadata of the program; and
[0327] a rendering subsystem coupled to the first subsystem and operable in a first mode to render the default mix in response to the speaker channel sound bed without using the metadata, wherein the rendering subsystem is further operable in a second mode to render the alternative mix using at least some of the metadata, including by selecting and mixing contents of the first subset of the object channel groups and at least one of the alternative speaker channels in response to at least some of the metadata.
[0328] Solution 37. The audio processing unit according to Solution 36, wherein the rendering subsystem comprises:
[0329] a first subsystem operable in the second mode to select, in response to the at least some metadata, the first subset of the group of object channels, at least one speaker channel in the speaker channel bed other than speaker channels in the subset of object-based speaker channels, and the at least one replacement speaker channel; and
[0330] A second subsystem coupled to the first subsystem and operable in the second mode to mix content of the first subset of the object channel groups selected by the first subsystem with content of each speaker channel to thereby determine the alternative mix.
[0331] Option 38. An audio processing unit according to any one of Options 36 to 37, wherein the rendering subsystem is configured to generate a speaker feed for driving a speaker to emit sound in response to the replacement mix, and the sound can be perceived as a mix of at least some of the audio content of the sound bed and the first non-ambient content rather than the second non-ambient content.
[0332] Option 39. An audio processing unit according to any one of Options 36 to 38, wherein the rendering subsystem is configured to generate a speaker feed for driving a speaker to emit sound in response to the replacement mix, wherein the sound includes object channel sounds indicative of the first non-ambient content, and the object channel sounds can be perceived as emanating from at least one apparent source position determined by the first subset of the object channel group.
[0333] Option 40. An audio processing unit according to any one of Options 36 to 39, further comprising a controller coupled to the rendering subsystem, wherein the controller is configured to provide a menu of mixes that can be selected, each mix in at least one subgroup of the mixes including the contents of the subgroup of the object channels and the subgroup of the replacement speaker channels.
[0334] Option 41. An audio processing unit according to Option 40, wherein the controller is configured to implement a user interface for displaying the menu.
[0335] Embodiment 42: The audio processing unit according to any one of embodiments 40 to 41, wherein the first subsystem and the rendering subsystem are implemented in a set-top device, and the controller is coupled to the set-top device.
[0336] Embodiment 43. An audio processing unit according to any one of Embodiments 36 to 42, wherein the default mix is a mix of ambient content and non-ambient content.
[0337] Option 44. An audio processing unit according to Option 43, wherein the ambient content indicates ambient sounds when viewing an event, the first non-ambient content indicates a commentary on the viewing event, and the second non-ambient content indicates an alternative commentary on the viewing event.
[0338] Scheme 45. An audio processing unit according to any one of Schemes 36 to 44, wherein the object-based audio program is a coded bit stream comprising frames, the coded bit stream is an AC-3 bit stream or an E-AC-3 bit stream, each of the frames of the coded bit stream indicates at least one data structure, the at least one data structure is a container comprising some content of the object channel and some of the metadata, and at least one of the containers is included in the auxiliary data auxdata field or the additional bit stream information addbsi field of each of the frames.
[0339] Embodiment 46. The audio processing unit according to any one of embodiments 36 to 44, wherein the object-based audio program is a DolbyE bitstream comprising a series of burst pulses and a guard band between burst pulse pairs.
[0340] Solution 47. An audio processing unit comprising:
[0341] buffer memory; and
[0342] at least one audio processing subsystem coupled to the buffer memory, wherein the buffer memory stores at least one segment of an object-based audio program, wherein the program indicates a speaker channel bed, a set of M replacement speaker channels, an object channel group, and metadata, wherein the object channel group includes N object channels, a first subset of the object channel group indicates first non-ambient content, the first subset includes the M object channels in the object channel group, each of N and M is an integer greater than zero, and M is equal to or less than N,
[0343] the speaker channel sound bed indicates a default mix of audio content including second non-ambient content different from the first non-ambient content, wherein an object-based speaker channel subset comprising M speaker channels in the sound bed indicates a mix of the second non-ambient content, or at least some audio content of the default mix, and the second non-ambient content,
[0344] each replacement speaker channel in the set of M replacement speaker channels indicates some, but not all, content of a corresponding speaker channel of the object-based subset of speaker channels, and
[0345] the metadata indicating at least one selectable predetermined alternative mix of the content of at least one of the object channels and the content of predetermined ones of the speaker channels of the sound bed and / or the replacement speaker channels, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix that includes at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content,
[0346] And wherein each of the segments comprises: data indicative of audio content of the speaker channel sound bed, data indicative of audio content of the replacement speaker channel, data indicative of audio content of the object channel, and at least a portion of the metadata.
[0347] Embodiment 48. An audio processing unit according to embodiment 47, wherein the object-based audio program is an encoded bitstream comprising frames, each of the segments being one of the frames.
[0348] Scheme 49. An audio processing unit according to any one of Schemes 47 to 48, wherein the encoded bit stream is an AC-3 bit stream or an E-AC-3 bit stream, each of the frames indicates at least one data structure, the at least one data structure is a container including some content of at least one of the object channels and some of the metadata, and at least one of the containers is included in the auxiliary data auxdata field or the additional bit stream information addbsi field of each of the frames.
[0349] Embodiment 50: The audio processing unit according to any one of embodiments 47 to 48, wherein the object-based audio program is a DolbyE bitstream comprising a series of burst pulses and a guard band between burst pulse pairs.
[0350] Scheme 51. An audio processing unit according to any one of Schemes 47 to 48, wherein the object-based audio program is an unencoded representation of the audio content and the metadata indicating the program, and the unencoded representation is a bitstream or at least one data file stored in a memory in a non-transitory manner.
[0351] Embodiment 52: The audio processing unit according to any one of embodiments 47 to 51, wherein the buffer memory stores the segments in a non-transitory manner.
[0352] Option 53. An audio processing unit according to any one of Option 47 to Option 52, wherein the audio processing subsystem is an encoder.
[0353] Embodiment 54. The audio processing unit according to any one of embodiments 47 to 53, wherein the audio processing subsystem is configured to: parse the speaker channel sound bed, the replacement speaker channel, the object channel and the metadata.
[0354] Embodiment 55. An audio processing unit according to any one of embodiments 47 to 54, wherein the audio processing subsystem is configured to: render the default mix in response to the speaker channel sound bed without using the metadata.
[0355] Option 56. An audio processing unit according to any one of Options 47 to 55, wherein the audio processing subsystem is configured to render the replacement mix using at least some of the metadata, including performing the rendering by selecting and mixing the contents of the first subgroup of the object channel group and at least one of the replacement speaker channels in response to at least some of the metadata.
Claims
1. A method for rendering audio content of an audio program, wherein: The audio program includes at least one loudspeaker channel, at least one object channel, and object-related metadata, the object-related metadata including rendering parameters for spatially rendering the at least one object channel as a spatial audio object, and metadata indicating selectable predetermined mixes corresponding to subgroups of the at least one loudspeaker channel and the object channel, the method comprising the following steps: a) receiving the audio program; b) providing a menu of selectable predetermined mixes to a user via a controller; c) receiving from the controller a selection by the user of one of the selectable predetermined mixes; d) rendering the selected mix using the rendering parameters of the object channel for the selected mix.
2. An audio processing unit configured to render audio content of an audio program, wherein The program includes at least one loudspeaker channel, at least one object channel, and object-related metadata, the object-related metadata including rendering parameters for spatially rendering the at least one object channel as a spatial audio object, and metadata indicating a selectable predetermined mix corresponding to a subgroup of the at least one loudspeaker channel and the object channel, the audio processing unit comprising: a first subsystem configured to receive the audio program; a controller coupled to the first subsystem and configured to provide a menu of selectable predetermined mixes to a user; a second subsystem configured to receive from the controller a selection by the user of one of the selectable predetermined mixes; A rendering subsystem is coupled to the first subsystem and the second subsystem and is configured to render the selected mix using rendering parameters for an object channel of the selected mix.
3. A playback system comprising: a plurality of deformatters, each of the plurality of deformatters configured to receive a corresponding bitstream of a plurality of serial bitstreams of a parallel delivery of an object-based audio program, and determine, by parsing the corresponding bitstream, a first output corresponding to speaker channel audio content of the corresponding bitstream, a second output corresponding to a synchronization word of the corresponding bitstream, and a third output including other metadata of the corresponding bitstream and object channel content; a bitstream synchronizer coupled to the plurality of deformatters and configured to use a synchronization word for each of the plurality of serial bitstreams to determine any misalignment of data in the bitstream, correct the determined misalignment, and output time-aligned speaker channel audio content for the bitstream; a plurality of decoders, each of the plurality of decoders configured to decode time-aligned speaker channel audio content of a corresponding bit stream of the plurality of serial bit streams; an object data combiner configured to provide all time-aligned third outputs of the plurality of serial bit streams to an object processing and rendering subsystem; Controller; as well as The object processing and rendering subsystem is configured to perform object processing on outputs of the object data combiner and the plurality of decoders in response to control data from a controller.
Citation Information
Patent Citations
Methods and systems for generating and rendering object based audio with conditional rendering metadata
CN107731239A
Encoder / decoder for multidimensional sound fields
US5583962A
Encoder / decoder for multidimensional sound fields
US5632005A
Method and apparatus for adjusting dynamic range and gain in an encoder / decoder for multidimensional sound fields
US5633981A
Method and apparatus for efficient implementation of single-sideband filter banks providing accurate measures of spectral magnitude and phase
US5727119A