Method and system for generating and interactively rendering object-based audio
By generating object-based audio programs that include speaker channel soundbeds, replacement speaker channels, and object channels, and attaching metadata to indicate optional mixing parameters, the system addresses the legacy system's inability to render both non-ambient and ambient sounds, enabling a full range of audio experiences and personalized rendering.
Patent Information
- Application Number
- CN202210300855.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2013-06-07
- Filing Date
- 2014-04-03
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2034-04-03
AI Technical Summary
Existing legacy playback systems cannot effectively parse and render object-based audio programs, and cannot provide a full range of audio experiences, especially in rendering personalized mixes of non-ambient sound and ambient sound in legacy systems.
Generate object-based audio programs that include speaker channel soundbeds, replacement speaker channels, and object channels, with attached metadata to indicate optional mixing parameters. This allows legacy systems to render default mixes, while properly configured systems can render mixes from extended layers to provide a personalized experience.
Legacy systems can render default mixes, while properly configured systems can render personalized mixes, providing a full-range audio experience and enhancing the flexibility and personalization of audio content rendering.
Smart Images

Figure CN114613373B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese invention patent application No. 201480020223.2 (which has been filed as a divisional application 201710942931.7), filed on April 3, 2014, entitled "Method and System for Generating and Interactively Rendering Object-Based Audio".
[0002] Cross-references to related applications
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 61 / 807,922, filed April 3, 2013, and the benefit of U.S. Provisional Patent Application No. 61 / 832,397, filed June 7, 2013. Technical Field
[0004] This invention relates to audio signal processing, and more specifically, to the encoding, decoding, and interactive rendering of audio data bitstreams including audio content (indicating speaker channels and at least one audio object channel) and metadata supporting interactive rendering of the audio content. Some embodiments of the invention generate, decode, and / or render audio data in one of the formats referred to as Dolby Digital (AC-3), Dolby Digital+ (Enhanced AC-3 or E-AC-3), or Dolby E. Background Technology
[0005] Dolby, Dolby Digital, Dolby Digital+, and Dolby E are trademarks of Dolby Laboratories, Inc. Dolby Laboratories provides proprietary implementations of AC-3 and E-AC-3, respectively known as Dolby Digital and Dolby Digital+.
[0006] While the present invention is not limited to encoding audio data according to E-AC-3 (or AC-3 or Dolby E) format or to delivering, decoding or rendering E-AC-3, AC-3 or Dolby E encoded data, for convenience, the invention will be described in embodiments in which it encodes audio bitstreams according to E-AC-3 or AC-3 or Dolby E format and delivers, decodes and renders such bitstreams.
[0007] A typical audio data stream includes both audio content (e.g., one or more channels of audio content) and metadata indicating at least one characteristic of the audio content. For example, in an AC-3 bitstream, there are several audio metadata parameters that are specifically intended to alter the sound of the program delivered to the listening environment.
[0008] AC-3 or E-AC-3 encoded bitstreams include metadata and may include 1 to 6 channels of audio content. The audio content is audio data that has been compressed using perceptual audio coding. The details of AC-3 encoding are well-known and are described in numerous publicly available references, including:
[0009] ATSC Standard A52 / A: Digital Audio Compression Standard (AC-3), Revision A, Advanced Television Systems Committee, August 20, 2001; and
[0010] U.S. Patents 5,583,962; 5,632,005; 5,633,981; 5,727,119; and 6,021,386.
[0011] The details of Dolby Digital Plus (E-AC-3) encoding are illustrated in, for example, in the following paper: “Introduction to Dolby Digital Plus, an Enhancement to the Dolby Digital Coding System”, AES Conference Paper 6196, 117th AES Conference, October 28, 2004.
[0012] The details of Dolby E encoding are described in the following: “Efficient Bit Allocation, Quantization, and Coding in an Audio Distribution System”, AES Preprint 5068, 107th AES Conference, August 1999; and “Professional Audio Coder Optimized for Use with Video”, AES Preprint 5033, 107th AES Conference, August 1999.
[0013] Each frame of an AC-3 encoded audio bitstream includes metadata and audio content for 1536 samples of digital audio. At a sampling rate of 48 kHz, this translates to 32 milliseconds of digital audio, or 31.25 frames per second.
[0014] Depending on whether the frames contain 1, 2, 3, or 6 audio data blocks, each frame of an E-AC-3 encoded audio bitstream contains metadata and audio content of 256, 512, 768, or 1536 samples for digital audio. For a sampling rate of 48 kHz, this represents 5.333, 10.667, 16, or 32 milliseconds of digital audio, or 189.9, 93.75, 62.5, or 31.25 frames per second, respectively.
[0015] like Figure 1 As shown, each AC-3 frame is divided into multiple parts (segments), including: containing (such as Figure 2 The synchronization information (SI) portion of the synchronization word (SW) and the first error correction word (CRC1) of the two error correction words; the bit stream information (BSI) portion containing most of the metadata; the six audio blocks (AB0 to AB5) containing the compressed audio content (and may also contain metadata); the useless bits (W) containing any unused bits remaining after the compressed audio content; the auxiliary (AUX) information portion that may contain more metadata; and the second error correction word (CRC2) of the two error correction words.
[0016] like Figure 4 As shown, each E-AC-3 frame is divided into multiple parts (segments), including: containing (such as Figure 2 The synchronization information (SI) part of the synchronization word (SW) shown; the bit stream information (BSI) part containing most of the metadata; 1 to 6 audio blocks (AB0 to AB5) containing the compressed audio content (and may also contain metadata); the unused bits (W) containing any unused bits remaining after the compressed audio content; the auxiliary (AUX) information part which may contain more metadata; and the error correction word (CRC).
[0017] In an AC-3 (or E-AC-3) bitstream, there are several audio metadata parameters specifically designed to alter the sound of the program delivered to the listening environment. One of these metadata parameters is the DIALNORM (Dialogue Normalization) parameter, which is included in the BSI segment.
[0018] like Figure 3 As shown, the BSI segment of an AC-3 frame (or E-AC-3 frame) includes a 5-bit parameter (“DIALNORM”) indicating the DIALNORM value of the program. If the audio coding mode (“acmod”) of the AC-3 frame is “0”, it includes a 5-bit parameter (“DIALNORM2”) indicating the DIALNORM value of a second audio program carried in the same AC-3 frame, indicating the use of a dual-single or “1+1” channel configuration.
[0019] The BSI segment also includes a flag (“addbsie”) indicating the presence (or absence) of additional bitstream information following the “addbsie” bit, a parameter (“addbsil”) indicating the length of any additional bitstream information following the “addbsil” value, and up to 64 bits of additional bitstream information (“addbsi”) following the “addbsil” value.
[0020] BSI segmentation includes Figure 3 Other metadata values not specifically shown in the document.
[0021] Including other types of metadata in audio bitstreams has been proposed. For example, PCT International Application Publication No. WO 2012 / 075246 A2, filed on December 1, 2011 and assigned to the assignee of this application, describes a method and system for generating, decoding, and processing audio bitstreams that include metadata indicating the processing state (e.g., loudness processing state) and characteristics (e.g., loudness) of the audio content. This reference also describes adaptive processing of the audio content of the bitstream using metadata and verification of the validity of the loudness processing state and loudness of the audio content of the bitstream using metadata.
[0022] Methods for generating and rendering object-based audio programs are also known. During the generation of such a program, it can be assumed that the loudspeakers to be used for rendering are located at any position in the playback environment (or that the speakers are located in a unit circle in a symmetrical configuration). It is not necessary to assume that the speakers must be located in a (nominal) horizontal plane or in any other predetermined arrangement known at the time of program generation. Typically, the metadata included in the program indicates rendering parameters for at least one object of the program, for example, using a three-dimensional loudspeaker array at an apparent spatial location or along a trajectory (in a three-dimensional volume). For example, the object channel of the program may have corresponding metadata indicating a three-dimensional trajectory (indicated by the object channel) of the apparent spatial location of the object to be rendered. The trajectory may include a series of “floor” positions (in a plane of a subgroup of loudspeakers assumed to be located on the floor of the playback environment, or in another horizontal plane of the playback environment) and a series of “above-floor” positions (each position is determined by driving a subgroup of loudspeakers assumed to be located in at least one other horizontal plane of the playback environment). For example, an example of rendering an object-based audio program is described in PCT International Application No. PCT / US2001 / 028783, which was published on September 29, 2011, with International Publication No. WO 2011 / 119401 A2 and assigned to the assignee of this application.
[0023] U.S. Provisional Patent Application No. 61 / 807,922 and No. 61 / 832,397, cited above, describe immersive, personalized, and perceptually object-based audio programs rendered to provide audio content for a program. The content may indicate the atmosphere (sounds emanating in or at the event being watched) and / or commentary on the event being watched, such as a football or rugby match or other sporting event. The program's audio content may indicate multiple audio object channels (e.g., user-selectable objects or groups of objects, and typically a default group of objects to be rendered in the absence of user object selection) and at least one speaker channel bed. The speaker channel bed may be a conventional mix (e.g., a 5.1 channel mix) of a type of speaker channel that can be included in a regular broadcast program that does not include object channels.
[0024] The U.S. Provisional Patent Applications Nos. 61 / 807,922 and 61 / 832,397 cited above describe object-related metadata delivered as part of an object-based audio program that provides mixing interactivity (e.g., a high degree of mixing interactivity) on the playback side. This includes allowing end users to select the mixing of the program's audio content for rendering, rather than simply allowing playback of the pre-mixed sound field. For example, a user could select from rendering options provided by metadata in a typical implementation of the program of this invention to select a subgroup of available object channels for rendering, and optionally also the playback level of at least one audio object (sound source) indicated by the object channel to be rendered. The spatial location of each selected sound source being rendered can be predetermined by the metadata included in the program, but in some implementations, it can be selected by the user (e.g., subject to predetermined rules or constraints). In some implementations, the metadata included in the program allows the user to select from a menu of rendering options (e.g., a small number of rendering options, such as a "Home Team Crowd Noise" object, a group of "Home Team Crowd Noise" and "Home Team Commentary" objects, a "Away Team Crowd Noise" object, and a group of "Away Team Crowd Noise" and "Away Team Commentary" objects). The menu can be presented to the user through a controller's user interface, and the controller can be coupled to a set-top device (or other device) configured to decode and render (at least partially) object-based programming. The metadata included in the program may otherwise allow the user to select from the option groups as to which object(s) indicated by the object channel should be rendered, and how the objects to be rendered should be configured.
[0025] U.S. Provisional Patent Applications Nos. 61 / 807,922 and 61 / 832,397 describe object-based audio programs, which are encoded audio bitstreams indicating the audio content of at least some programs (e.g., speaker channel soundbeds and object channels of at least some programs) and object-related metadata. At least one additional bitstream or file may indicate the audio content of some programs (e.g., at least some object channels) and / or object-related metadata. In some embodiments, the object-related metadata provides a default mix of object content and soundbed (speaker channel) content with default rendering parameters (e.g., default spatial location of rendered objects). In some embodiments, the object-related metadata provides a set of optional “preset” mixes of object channel and speaker channel content, each preset mix having a predetermined set of rendering parameters (e.g., spatial location of rendered objects). In some embodiments, the object-related metadata of the program (or a pre-configuration of the playback or rendering system, not indicated by metadata delivered with the program) provides constraints or conditions regarding optional mixes of object channel and speaker channel content.
[0026] U.S. Provisional Patent Applications Nos. 61 / 807,922 and 61 / 832,397 also describe object-based audio programs that include a set of bitstreams (sometimes referred to as “substreams”) generated and transmitted in parallel. Multiple decoders can be used to decode them (e.g., if the program includes multiple E-AC-3 substreams, the playback system can utilize multiple E-AC-3 decoders to decode the substreams). Each substream can include a synchronization word (e.g., a timecode) to allow the substreams to be synchronized or time-aligned with each other.
[0027] U.S. Provisional Applications Nos. 61 / 807,922 and 61 / 832,397 also describe an object-based audio program that is or includes at least one AC-3 (or E-AC-3) bitstream and includes one or more data structures referred to as containers. Each container, including object channel content (and / or object-related metadata), is included in an auxiliary data field at the end of a frame of the bitstream (e.g., Figure 1 or Figure 4 In the AUX segment shown, or in the "skip field" segment of the bitstream. Also described are object-based audio programs that are or include Dolby E bitstreams, in which object channel content and object-related metadata (e.g., each container of the program that includes object channel content and / or object-related metadata) are included in the bit positions of the Dolby E bitstream that do not normally carry useful information.
[0028] U.S. Provisional Application No. 61 / 832,397 also describes an object-based audio program that includes metadata of at least one speaker channel group, at least one object channel, and a hierarchical graph (a hierarchical “mix graph”) indicating optional mixes (e.g., all optional mixes) for the speaker channels and object channels. The mix graph may indicate each rule applicable to the selection of subgroups of speaker and object channels, indicating nodes (each node may indicate an optional channel or channel group, or a classification of optional channels or channel groups), and connections between nodes (e.g., control interfaces to nodes and / or rules for selecting channels). The mix graph may indicate necessary data (“basic” layers) and optional data (at least one “extended” layer), and where the mix graph can be represented as a tree graph, the basic layer may be one branch (or two or more branches) of the tree graph, and each extended layer may be another branch (or group of branches) of the tree graph.
[0029] U.S. Provisional Applications Nos. 61 / 807,922 and 61 / 832,397 also teach that object-based audio programs can be decoded, and that their speaker channel content can be rendered using legacy decoders and rendering systems (which are not configured to parse the program's object channels and object-related metadata). The same program can be rendered by a set-top device (or other decoding and rendering system) configured to parse the program's object channels and object-related metadata and to render a mix of speaker channel and object channel content as indicated by the program. However, neither U.S. Provisional Application No. 61 / 807,922 nor U.S. Provisional Application No. 61 / 832,397 teaches or implies how to generate a personalized object-based audio program that can be rendered by a legacy decoding and rendering system (not configured to parse the program's object channels and object-related metadata) to provide a full-range audio experience (e.g., audio intended to be perceived as non-ambient sound from at least one discrete audio object mixed with ambient sound), but makes a decoding and rendering system configured to parse the program's object channels and object-related metadata capable of rendering a selected mix of the content of at least one speaker channel and at least one object channel of the program (also providing a full-range audio experience), or makes it desirable to do so. Summary of the Invention
[0030] One embodiment of the invention provides personalized, object-based programming compatible with legacy playback systems (which are not configured to parse object channels and object-related metadata of a program), whereby the legacy system can render a program's default speaker channel group to provide a full-range audio experience (wherein, in this context, "full-range audio experience" means a sound mix represented only by the audio content of the default speaker channel group, intended to be perceived as a full or complete mix of non-ambient sounds from at least one discrete audio object mixed with other sounds represented by the default speaker channel group. These other sounds can be ambient sounds.), wherein the same program can be decoded and rendered by a non-legacy playback system (configured to parse object channels and metadata of a program) to render at least one selected preset mix of the content of at least one speaker channel of the program and the non-ambient content of at least one object channel of the program (which can also provide a full-range audio experience). In this document, such a default speaker channel group (which can be rendered by a legacy system) is sometimes referred to as a speaker channel “sound bed,” although this term is not intended to imply that the sound bed must be mixed with additional audio content to provide a full-range audio experience. In fact, in typical embodiments of the invention, the sound bed does not necessarily need to be mixed with additional audio content to provide a full-range audio experience, and the sound bed can be decoded and rendered by a legacy system to provide a full-range audio experience without mixing with additional audio content. In other embodiments, the object-based audio program of the invention includes a speaker channel sound bed, which represents only non-ambient content (e.g., a mix of different types of non-ambient content) and is capable of being rendered by a legacy system (e.g., to provide a full-range audio experience), and a playback system configured to parse the object channels and metadata of the program can render at least one selected preset mix of the content (e.g., non-ambient and / or ambient content) of at least one speaker channel and at least one object channel of the program (which may, but must, provide a full-range audio experience).
[0031] Typical implementations of this type of approach generate, deliver, and / or render object-based programs comprising a base layer (e.g., a 5.1-channel sound bed) including speaker channel sound beds representing all content of a default audio program (sometimes referred to as the "default" mix). The default audio program includes a complete set of audio elements (e.g., ambient content mixed with non-ambient content) that provide a full-range audio experience when played. Legacy playback systems (which cannot decode or render object-based audio) can decode and render the default mix. Examples of ambient content for the default audio program are crowd noise (captured at sporting events or other viewing events), and examples of non-ambient content for the default audio program include commentary and / or announcement feeds (related to sporting events or other viewing events). The program also includes an extension layer (which can be ignored by legacy playback systems) that can be utilized by appropriately configured (non-legacy) playback systems to select and render any of a plurality of predetermined mixes of audio content for the extension layer (or the extension layer and the base layer). Extension layers typically include optional alternative speaker channel groups that allow for personalized representation of alternative content (e.g., only the main ambient content, rather than a mix of ambient and non-ambient content provided by the base layer) and optional object channel groups (e.g., object channels representing the main non-ambient content and alternative non-ambient content).
[0032] Providing a basic layer and at least one extended layer in a program gives program generation facilities (e.g., broadcast front-ends) and playback systems (which may be or include set-top boxes or “STBs”) greater flexibility.
[0033] In some embodiments, the present invention is a method for generating an object-based audio program that indicates audio content (e.g., captured audio content), said audio content including first non-environmental content, second non-environmental content different from the first non-environmental content, and third content different from the first and second non-environmental content (the third content may be ambient content, but in some cases may also be or include non-environmental content), said method comprising the steps of:
[0034] Determine an object channel group comprising N object channels, wherein a first subgroup of the object channel group indicates the first non-environmental content, the first subgroup comprising M object channels in the object channel group, each of N and M being a positive integer, and M being equal to or less than N;
[0035] Determine a speaker channel sound bed that indicates the default mix of audio content (e.g., a default mix of ambient content and non-ambient content), including an object-based speaker channel subgroup of M speaker channels in the sound bed that indicates the second non-ambient content, or at least some audio content of the default mix mixed with the second non-ambient content.
[0036] A set of M alternative speaker channels is determined, wherein each alternative speaker channel in the set of M alternative speaker channels indicates some, but not all, of the contents of the corresponding speaker channel in the object-based speaker channel subgroup;
[0037] Generate metadata (sometimes referred to herein as object-related metadata) indicating at least one of the contents of the object channel and at least one of the contents of the speaker channels and / or predetermined speaker channels of the sound bed, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is an alternative mix indicating at least some contents of the sound bed and the first non-ambient content instead of the second non-ambient content; and
[0038] Generate the object-based audio program comprising the speaker channel sound bed, the set of M alternative speaker channels, the object channel group, and the metadata, such that the speaker channel sound bed is renderable without using the metadata to provide sound that is perceived as the default mix, and the alternative mix is renderable in response to at least some of the metadata to provide sound that is perceived as a mix comprising the at least some audio content of the sound bed and the first non-ambient content instead of the second non-ambient content.
[0039] Typically, program metadata (object-related metadata) includes (or contains) optional content metadata representing a set of optional experience clarity. Each experience clarity is an optional pre-defined (“preset”) mix of the program’s audio content (e.g., a mix of content from at least one object channel and at least one speaker channel in the soundbed, or a mix of content from at least one object channel and at least one alternative speaker channel, or a mix of content from at least one object channel and at least one speaker channel in the soundbed and at least one alternative speaker channel). Each preset mix has a pre-defined set of rendering parameters (e.g., the spatial location of the rendered objects). The playback system’s user interface can present preset mixes as a limited menu or palette of available mixes.
[0040] In other embodiments, the present invention is a method for rendering audio content determined by an object-based audio program, wherein the program indicates a speaker channel soundbed, a set of M alternative speaker channels, an object channel group, and metadata, wherein the object channel group includes N object channels, and a first subgroup of the object channel group indicates a first non-ambient content, the first subgroup including M object channels in the object channel group, where each of N and M is a positive integer, and M is equal to or less than N.
[0041] The speaker channel sound bed indicates a default mix of audio content including a second non-ambient content that is different from the first non-ambient content. This includes an object-based speaker channel subgroup of M speaker channels in the sound bed indicating a mix of the second non-ambient content, or at least some audio content of the default mix, with the second non-ambient content.
[0042] Each of the M alternative speaker channels in the set indicates some, but not all, of the contents of the corresponding speaker channel in the object-based speaker channel subgroup, and
[0043] The metadata indicates at least one of the contents of the object channel and at least one of the contents of the speaker channel of the sound bed and / or the contents of a predetermined speaker channel in the replacement speaker channel, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is a replacement mix that includes at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content, the method comprising the steps of:
[0044] (a) providing the object-based audio program to the audio processing unit; and
[0045] (b) In the audio processing unit, the speaker channel sound bed is parsed, and the default mix is rendered in response to the speaker channel sound bed without using the metadata.
[0046] In some cases, the audio processing unit is a legacy playback system (or other audio data processing system) that is not configured to parse the object channel or metadata of a program. When the audio processing unit is configured to parse the object channel, replacement channel, and metadata (as well as the speaker channel sound bed) of a program, the method may include the following steps:
[0047] (c) In the audio processing unit, the replacement mix is rendered using at least some of the metadata, including rendering by selecting and mixing the contents of the first subgroup of the object channel group and at least one of the replacement speaker channels in response to at least some of the metadata.
[0048] In some implementations, step (c) includes the following steps: driving a loudspeaker to provide a sound that can be perceived as a mixture of at least some of the audio content of the sound bed and the first non-ambient content but not the second non-ambient content.
[0049] Another aspect of the invention is an audio processing unit (APU) configured to perform any embodiment of the method of the invention. In another type of embodiment, the invention is an APU including a buffer memory (buffer) that stores (e.g., in a non-transient manner) at least one frame or other segment of an object-based audio program generated by any embodiment of the method of the invention (including audio content of speaker channel sound beds and object channels, as well as object-related metadata). Examples of APUs include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, preprocessing systems (preprocessors), postprocessing systems (postprocessors), audio bitstream processing systems, and combinations of such elements.
[0050] The aspects of this invention include systems or apparatus configured (e.g., programmed) to perform any embodiment of the methods of the invention, and computer-readable media (e.g., disks) storing (e.g., in a non-transitory manner) code for implementing any embodiment of the methods or steps of the invention. For example, the system of the invention may be or includes a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the methods or steps of the invention. Such a general-purpose processor may be or includes a computer system comprising input devices, memory, and processing circuitry programmed (and / or otherwise configured) to perform embodiments of the methods (or steps of the invention) in response to data set thereto. Attached Figure Description
[0051] Figure 1 It is a diagram that includes the AC-3 frames into which they are divided.
[0052] Figure 2 It is a diagram that includes the synchronization information (SI) segments of the AC-3 frames into which they are divided.
[0053] Figure 3 It is a diagram of the bit stream information (BSI) segments of the AC-3 frames, which are divided into segments.
[0054] Figure 4 It is a diagram of E-AC-3 frames, which are divided into segments.
[0055] Figure 5 This is a block diagram of an implementation of the system, wherein one or more elements of the system can be configured according to an embodiment of the invention.
[0056] Figure 6 This is a block diagram of a playback system that can be implemented as an embodiment of the method of the present invention.
[0057] Figure 7 This is a block diagram of a playback system that can be configured to perform an embodiment of the method of the present invention.
[0058] Figure 8 This is a block diagram of a broadcasting system configured to generate object-based audio programs (and corresponding video programs) according to an embodiment of the present invention.
[0059] Figure 9 This is a diagram showing the relationship between object channels in an embodiment of the program of the present invention, indicating which subgroups of the object channels are user-selectable.
[0060] Figure 10 This is a block diagram of a system that can be implemented as an embodiment of the method of the present invention.
[0061] Figure 11 This is a diagram of the content of an object-based audio program generated according to an embodiment of the present invention.
[0062] Figure 12 This is a block diagram of an implementation of a system configured to carry out an embodiment of the method of the present invention.
[0063] Symbols and terms
[0064] Throughout this disclosure, including the claims, the expression "non-ambient sound" means sound (e.g., commentary or other monologue, or dialogue) emanating from or within a discrete audio object (or a large number of audio objects located here or within) situated at or within an angular position well-localized relative to the listener (i.e., the angular position subtends a solid angle no greater than approximately 3 solid radians relative to the listener, wherein the entire range centered on the listener's position subtends a solid angle of 4π solid radians relative to the listener). Here, "ambient sound" means sound that is not non-ambient sound (e.g., crowd noise perceived by a member of a crowd). Therefore, ambient sound herein means sound emanating from a large (or otherwise poorly localized) angular position relative to the listener.
[0065] Similarly, “non-ambient audio content” (or “non-ambient content”) in this document refers to audio content perceived when rendered as sound emanating from a discrete audio object (or a large number of audio objects located here or within) at or within an angular position that is well-positioned relative to the listener (i.e., the angular position is relative to a solid angle of no more than approximately 3 solid radians relative to the listener), and “ambient audio content” (or “ambient content”) refers to audio content that is not “non-ambient audio content” (or “non-ambient content”) and is perceived when rendered as ambient sound.
[0066] Throughout this disclosure, including the claims, the expression “to” a signal or data (e.g., to filter, scale, transform, or apply gain to the signal or data) is broadly used to mean directly performing an operation on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing before the operation is performed on the signal).
[0067] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to an apparatus, system, or subsystem. For example, a subsystem implementing a decoder may be called a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, wherein the subsystem generates M inputs and the other X-M inputs are received from an external source) may also be called a decoder system.
[0068] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or apparatus that is programmable or otherwise configurable (e.g., using software or firmware) to perform operations on data (e.g., audio, video, or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing of audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.
[0069] Throughout this disclosure, including the claims, the term "audio-video receiver" (or "AVR") refers to a receiver in a class of consumer electronic devices used for controlling the playback of audio and video content, such as in a home theater.
[0070] Throughout this disclosure, including the claims, the term "soundbar" refers to a device that is a consumer electronic device (typically installed in a home theater system) and includes at least one speaker (typically at least two speakers) and a subsystem for rendering audio for playback by each of the included speakers (or for playback by each of the included speakers and at least one additional speaker outside the soundbar).
[0071] Throughout this disclosure, including the claims, the terms "audio processor" and "audio processing unit" are used interchangeably and, more broadly, refer to a system configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, preprocessing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools).
[0072] Throughout this disclosure, including the claims, the expression "metadata" (e.g., as in the expression "processing status metadata") refers to data that is separate from and distinct from the corresponding audio data (audio content of the bitstream that also includes metadata). Metadata is associated with audio data and indicates at least one characteristic or property of the audio data (e.g., what type of processing has been performed or should be performed on the audio data or the trajectory of an object indicated by the audio data). The association between metadata and audio data is time-synchronized. Thus, current (latest received or updated) metadata can indicate that the corresponding audio data simultaneously possesses the indicated characteristics and / or includes the result of audio data processing of the indicated type.
[0073] Throughout this disclosure, including the claims, the terms "coupled" or "coupled to" are used to indicate a direct or indirect connection. Thus, if a first device is coupled to a second device, the connection can be either a direct connection or an indirect connection via other devices and connections.
[0074] Throughout this disclosure, including the claims, the following expressions have the following definitions:
[0075] The terms "speaker" and "loudspeaker" are used synonymously to refer to any sound transducer. This definition includes loudspeakers that are implemented as multiple transducers (e.g., low-frequency speakers and high-frequency speakers);
[0076] Speaker feed: Audio signals applied directly to the loudspeaker, or audio signals applied to an amplifier and loudspeaker connected in series;
[0077] Channel (or “audio channel”): A mono audio signal. Such a signal can typically be rendered in a manner that is equivalent to applying the signal directly to a loudspeaker at a desired or nominal location. The desired location can be static, as is typical for physical loudspeakers, or it can be dynamic;
[0078] Audio program: a set of one or more audio channels (at least one speaker channel and / or at least one object channel) and optionally associated metadata (e.g., metadata describing the desired spatial audio representation);
[0079] Speaker channel (or "speaker feed channel"): An audio channel associated with a named loudspeaker (at the desired or nominal location) or with a named speaker area within a defined speaker configuration. A speaker channel can be rendered in a manner that is equivalent to applying an audio signal directly to the named loudspeaker (at the desired or nominal location) or the speaker within the named speaker area;
[0080] Object channel: An audio channel that indicates the sound emitted by an audio source (sometimes referred to as an audio "object"). Typically, an object channel defines a parameterized description of the audio source (e.g., metadata indicating the parameterized audio source description is included in or provided using the object channel). The source description may define the sound emitted by the source (as a function of time), the apparent location of the source as a function of time (e.g., 3D spatial coordinates), and optionally at least one additional parameter characterizing the source (e.g., apparent source size or width);
[0081] Object-based audio programs: comprising an audio program with one or more object channels (and optionally at least one speaker channel) and optionally associated metadata (e.g., metadata indicating the trajectory of an audio object emitting a sound indicated by an object channel, or metadata otherwise indicating the desired spatial audio representation of the sound indicated by an object channel, or metadata indicating the identifier of at least one audio object as the source of the sound indicated by an object channel); and
[0082] Rendering: The process of converting an audio program into a feed for one or more speakers, or the process of converting an audio program into a feed for one or more speakers and converting the speaker feed into sound using one or more amplifiers (in the latter case, rendering is sometimes referred to as rendering "through" an amplifier). Audio channels can be rendered very generally ("at" the desired location) by directly applying a signal to a physical amplifier at the desired location, or one or more audio channels can be rendered using one of a variety of virtualization techniques designed to be substantially equivalent (to the listener) to this very general rendering. In the latter case, each audio channel can be converted into a feed for one or more speakers to be applied to an amplifier at a known location that is typically different from the desired location, such that the sound emitted by the amplifier in response to the feed will be perceived as emanating from the desired location. Examples of such virtualization techniques include binaural rendering through headphones (e.g., using Dolby headphone processing, which simulates up to 7.1 channels of surround sound for the headphone wearer) and wavefield synthesis. Detailed Implementation
[0083] Figure 5This is a block diagram of an example audio processing chain (audio data processing system), wherein one or more elements of the system can be configured according to embodiments of the invention. The system includes the following elements coupled together as shown: a capture unit 1, a generation unit 3 (which includes an encoding subsystem), a delivery subsystem 5, a decoder 7, an object processing subsystem 9, a controller 10, and a rendering subsystem 11. In variations of the illustrated system, one or more elements are omitted, or additional audio data processing units are included. Typically, elements 7, 9, 10, and 11 are a playback system (e.g., a home theater system for an end user) or are included in such a playback system.
[0084] Capture unit 1 is typically configured to generate and output PCM (temporal) samples that include audio content. The samples may indicate multiple audio streams captured by a microphone (e.g., during a sporting event or other viewing event). Production unit 3, typically operated by a broadcaster, is configured to accept PCM samples as input and output an object-based audio program indicating audio content. The program typically includes or comprises an encoded (e.g., compressed) audio bitstream indicating at least some audio content (sometimes referred to herein as the “master mix”) and optionally also includes at least one additional bitstream or file indicating some audio content (sometimes referred to herein as a “submix”). The data of the encoded bitstream indicating the audio content (and each generated submix, if any) is sometimes referred to herein as “audio data”. If the encoding subsystem of production unit 3 is configured according to a typical embodiment of the invention, the object-based audio program output from unit 3 indicates (i.e., includes) multiple speaker channels (speaker channel “sound bed” and alternative speaker channels), multiple object channels of audio data, and object-related metadata. A program may include a master mix, which in turn includes audio content indicating speaker channel sound beds, audio content indicating alternative speaker channels, audio content indicating at least one user-selectable object channel (and optionally at least one other object channel), and metadata (including object-related metadata associated with each object channel). A program may also include at least one submix, which includes audio content indicating at least one other object channel (e.g., at least one user-selectable object channel) and / or object-related metadata. The object-related metadata of the program may include durable metadata (described below). A program (e.g., its master mix) may indicate one or more sets of speaker channels. For example, the master mix may indicate two or more sets of speaker channels (e.g., a 5.1 channel neutral crowd noise sound bed, a 2.0 channel group indicating alternative speaker channels for the home crowd noise, and a 2.0 channel group indicating alternative speaker channels for the away crowd noise), including at least one user-selectable alternative speaker channel group (which can be selected using the same user interface used for user selection of object channel content or configuration) and a speaker channel sound bed (which may be rendered in the absence of user selection of other content for the program). The sound bed (which may be referred to as the default sound bed) can be determined by data indicating the configuration (e.g., initial configuration) of the speaker group of the playback system, and optionally, the user can select other audio content of the program to be rendered instead of the default sound bed.
[0085] The program's metadata may indicate at least one (and typically more than one) optional predetermined mix of the content of at least one object channel and the content of a predetermined speaker channel and / or a replacement speaker channel in the program's sound bed, and may include rendering parameters for each said mix. At least one such mix may be a replacement mix indicating at least some audio content of the sound bed with a first non-ambient content (indicated by at least one object channel included in the mix) rather than a second non-ambient content (indicated by at least one speaker channel in the sound bed).
[0086] Figure 5 The delivery subsystem 5 is configured to store and / or transmit (e.g., broadcast) the programs generated by unit 3 (e.g., its main mix and each submix, if any submix is generated).
[0087] In some implementations, subsystem 5 delivers object-based audio programs, wherein the audio objects of the program (and at least some corresponding object-related metadata) and speaker channels are transmitted via a broadcast system (in the form of a main mix of the program indicated by the broadcast audio bitstream), and at least some metadata of the program (e.g., object-related metadata indicating constraints on the rendering or mixing of object channels of the program) and / or at least one object channel of the program is delivered in another manner (as a “secondary mix” of the main mix) (e.g., the secondary mix is sent to a specific end user via Internet Protocol or an “IP” network). Alternatively, the end user’s decoding and / or rendering system is pre-configured with at least some object-related metadata (e.g., metadata indicating constraints on the rendering or mixing of audio objects of an embodiment of the object-based audio program of the present invention), and such object-related metadata is not broadcast or otherwise delivered (by subsystem 5) without the corresponding object channel (in the main mix or secondary mix of the object-based audio program).
[0088] In some implementations, the timing and synchronization of portions or elements of an object-based audio program delivered via separate paths (e.g., the main mix broadcast via a broadcast system, and associated metadata sent as a submix via an IP network) are provided via synchronization words (e.g., timecodes), which are sent via all delivery paths (e.g., the main mix and each corresponding submix).
[0089] Refer again Figure 5Decoder 7 receives (receives or reads) the program (or at least one bitstream or other element of the program) delivered by delivery subsystem 5, and decodes the program (or each of its received elements). In some embodiments of the invention, the program includes a main mix (encoded bitstream, e.g., AC-3 or E-AC-3 encoded bitstream) and at least one submix of the main mix, and decoder 7 receives and decodes the main mix (and optionally at least one submix). Optionally, at least one submix of the program that does not need to be decoded (e.g., object channel) is delivered directly by subsystem 5 to object processing subsystem 9. If decoder 7 is configured according to a typical embodiment of the invention, the output of decoder 7 in typical operation includes the following:
[0090] The audio sample stream indicating the speaker channel soundbed of the program (and usually also indicating the program's alternative speaker channels); and
[0091] The stream of audio samples for the program's object channel (e.g., a user-selectable audio object channel) and the corresponding object-related metadata.
[0092] The object processing subsystem 9 is coupled (from decoder 7) to receive the decoded speaker channels, object channels, and object-related metadata of the delivered program, and optionally also to receive at least one submix of the program (indicating at least one other object channel). For example, subsystem 9 may receive (from decoder 7) audio samples of the program's speaker channels and at least one object channel of the program, as well as the program's object-related metadata, and may also receive (from delivery subsystem 5) audio samples of at least one other object channel of the program (which has not yet undergone decoding in decoder 7).
[0093] Subsystem 9 is coupled and configured to output the selected subgroup of all object channels indicated by the program, along with the corresponding object-related metadata, to rendering subsystem 11. Subsystem 9 is also typically configured to pass the decoded speaker channels from decoder 7 through (to subsystem 11) unchanged, and may be configured to process at least some of the object channels (and / or metadata) claimed thereto to generate their claimed object channels and metadata to subsystem 11.
[0094] The selection of object channels performed by subsystem 9 is typically determined by user selection (such as control data instructions set from controller 10 to subsystem 9) and / or by rules (e.g., indicating conditions and / or constraints) programmed or otherwise configured to be implemented by subsystem 9. Such rules may be determined by object-related metadata of the program and / or by other data set to subsystem 9 (e.g., data indicating the performance and organization of the speaker array of the playback system) and / or by pre-configuring (e.g., programming) subsystem 9. In some implementations, controller 10 (via a user interface implemented by controller 10) provides the user with (e.g., displayed on a touchscreen) a menu or palette of optional “preset” mixes of speaker channel content (i.e., the content of the bed speaker channels and / or replacement speaker channels) and object channel content (objects). Optional preset mixes may be determined by the program’s object-related metadata and also typically by rules implemented by subsystem 9 (e.g., rules pre-configured to be implemented by subsystem 9). The user selects from the optional mixes by inputting a command to the controller 10 (e.g., by activating the touchscreen of the controller 10), and in response, the controller 10 sets the corresponding control data to the subsystem 9 to render the corresponding content according to the invention.
[0095] Figure 5 The rendering subsystem 11 is configured to render audio content determined by the output of subsystem 9 for playback by speakers (not shown) of the playback system. Subsystem 11 is configured to map audio objects (e.g., default objects and / or user-selected objects selected as a result of user interaction using controller 10) determined by the object channels selected by object processing subsystem 9 to available speaker channels using rendering parameters (e.g., user-selected values and / or default values for spatial location and level) associated with each selected object output from subsystem 9. At least some rendering parameters are also determined by object-related metadata output from subsystem 9. Rendering system 11 also receives speaker channels passed through subsystem 9. Typically, subsystem 11 is a smart mixer and is configured to determine the speaker feed for available speakers, including by mapping one or more selected (e.g., default selected) objects to each of a large number of individual speaker channels, and mixing the objects with speaker channel content indicated by each corresponding speaker channel of the program (e.g., each speaker channel in the program's speaker channel bed).
[0096] Figure 12 This is a block diagram of another system configured to carry out an embodiment of the method of the present invention. Figure 12 The capture unit 1, the generation unit 3, and the delivery subsystem 5 are together with Figure 5The components with the same number in the system are identical. Figure 12 Units 1 and 3 are operable to generate object-based audio programs according to at least one embodiment of the invention, and ( Figure 12 Subsystem 5 is configured to deliver such programs to Figure 12 The playback system 111.
[0097] and Figure 5 Unlike the playback system (including decoder 7, object processing subsystem 9, controller 10, and rendering subsystem 11), playback system 111 is not configured to parse the object channels or object-related metadata of the program. Decoder 107 of playback subsystem 111 is configured to parse the speaker channel sound bed of the program delivered by subsystem 5, and rendering subsystem 109 of subsystem 111 is coupled and configured to render the default mix (indicated by the speaker channel sound bed) in response to the sound bed (without using the program's object-related metadata). Decoder 107 may include buffer 7A, which (e.g., in a non-transient manner) stores at least one frame or other segment of the object-based audio program delivered from subsystem 5 to decoder 107 (including the speaker channel sound bed, audio content of the alternative speaker channels and object channels, and object-related metadata).
[0098] In comparison, Figure 5 A typical implementation of the playback system (including decoder 7, object processing subsystem 9, controller 10, and rendering subsystem 11) is configured to parse the object channels, object-related metadata, and alternative speaker channels (and speaker channel sound beds indicating the default mix) of the object-based program delivered there. In some such implementations, Figure 5 The playback system is configured to render a replacement mix (determined by at least one object channel and at least one replacement speaker channel, and typically at least one additional speaker channel sound bed) in response to at least some object-related metadata, including selecting the replacement mix using at least some object-related metadata. In some such implementations, Figure 5 The playback system can operate in a mode in which it renders such an alternative mix in response to the object channel and speaker channel content and metadata of the program, and can also operate in a second mode in which the decoder 7 parses the speaker channel sound bed of the program (which can be triggered by metadata in the program), the speaker channel sound bed is set to the rendering subsystem 11, and the rendering subsystem 11 (without using the object-related metadata of the program) operates in response to the sound bed to render the default mix (indicated by the sound bed).
[0099] In one embodiment, the present invention is a method for generating an object-based audio program that indicates audio content (e.g., captured audio content), the audio content including first non-contextual content, second non-contextual content different from the first non-contextual content, and third content different from the first and second non-contextual content, the method comprising the steps of:
[0100] Determine an object channel group comprising N object channels, wherein a first subgroup of the object channel group indicates a first non-contextual content, the first subgroup comprising M object channels in the object channel group, where each of N and M is a positive integer, and M is equal to or less than N;
[0101] Determine the speaker channel sound bed that indicates the default mix of audio content, including an object-based speaker channel subgroup of M speaker channels in the sound bed that indicates a second non-ambient content, or a mix of at least some audio content of the default mix with the second non-ambient content.
[0102] A set of M alternative speaker channels is determined, wherein each alternative speaker channel in the set of M alternative speaker channels indicates some, but not all, of the contents of the corresponding speaker channel in an object-based speaker channel subgroup;
[0103] Generate metadata indicating at least one of the contents of an object channel and at least one optional predetermined alternative mix of the contents of a predetermined speaker channel and / or a replacement speaker channel in the speaker channel of the sound bed, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is a replacement mix indicating at least some audio content of the sound bed and a first non-ambient content rather than a second non-ambient content; and
[0104] Generate an object-based audio program that includes a speaker channel soundbed, a set of M replacement speaker channels, an object channel group, and metadata, such that:
[0105] The speaker channel sound bed can be used without metadata (e.g., via...). Figure 12 The system's playback system 111, or operates in a mode by resolving the speaker channel sound bed of the program using decoder 7. Figure 5 In the playback system, the speaker channel sound bed is set to the rendering subsystem 11, and, in response to the sound bed operation (using the program's object-related metadata) the rendering subsystem 111 can render the default mix indicated by the sound bed, to provide a sound perceptible as the default mix, and
[0106] The mix replacement responds to at least some metadata (e.g., via...) Figure 5The playback system, comprising decoder 7, object processing subsystem 9, controller 10, and rendering subsystem 11, is renderable using object-related metadata of the program delivered to decoder 7 to provide sound that can be perceived as a mixture of at least some audio content including the sound bed with a first non-ambient content rather than a second non-ambient content.
[0107] In some such implementations, object-based audio programs are generated such that replacement mixes are renderable in response to at least some metadata (e.g., via...). Figure 5 The playback system provides sound that is perceptible as a mix including a first non-environmental content but not a second non-environmental content, such that the first non-environmental content is perceptible as emanating from a source whose size and position are determined by a subgroup of metadata corresponding to a first subgroup of the object channel group.
[0108] The substitution mix can specify at least one of the speaker channels of the sound bed and the substitution speaker channel, as well as the content of the first subgroup of the object channel group, rather than the content of the object-based speaker channel subgroup of the sound bed. This is accomplished by rendering a substitution mix of the non-ambient content (or a mix of ambient and non-ambient content) of the object-based speaker channel subgroup of the substitution sound bed, the non-ambient content of the first subgroup of the object channel group (and the content of the substitution speaker channel, which is typically ambient content or a mix of ambient and non-ambient content).
[0109] In some implementations, the program's metadata includes (or includes) optional content metadata indicating a set of optional experience clarity. Each experience clarity is an optional, predetermined ("preset") mix of the program's audio content (e.g., a mix of content from at least one object channel and at least one speaker channel in the soundbed, or a mix of content from at least one object channel and at least one alternative speaker channel, or a mix of content from at least one object channel and at least one speaker channel in the soundbed and at least one alternative speaker channel). Each preset mix has a predetermined set of rendering parameters (e.g., the spatial location of the rendered objects), which is also typically indicated by the metadata. The preset mixes may be provided by the user interface of the playback system (e.g., by...). Figure 5 Controller 10 or Figure 6 The user interface implemented by controller 23 is presented as a limited menu or palette available for mixing.
[0110] In some implementations, the program's metadata includes default mix metadata indicating the base layer, so that playback systems configured to recognize and use the default mix metadata (e.g., Figure 5 The implementation of the playback system (e.g., the default mix) is able to select the default mix (instead of another preset mix) and render the base layer. Legacy playback systems not configured to recognize or use default mix metadata are not expected to be suitable for this purpose. Figure 12The playback system 111 can also render the base layer (and thus the default mix) without using any default mix metadata.
[0111] The object-based audio program generated according to a typical embodiment of the invention (e.g., an encoded bitstream indicating such a program) is a personalizable object-based audio program (e.g., an encoded bitstream indicating such a program). According to a typical embodiment, audio objects and other audio content are encoded to allow for an optional full-range audio experience (wherein, in this context, "full-range audio experience" means audio intended to be perceived as mixed with ambient sounds (e.g., commentary or dialogue) from at least one discrete audio object). To allow for personalization (i.e., selecting the desired mix of audio content), speaker channel sound beds (e.g., sound beds indicating ambient content mixed with non-ambient object channel content) and at least one alternative speaker channel, as well as at least one object channel (typically, multiple object channels), are encoded as distinct elements within the encoded bitstream.
[0112] In some implementations, a personalizable object-based audio program includes (and allows selection of any of) at least two optional “preset” mixes of object channel content and / or speaker channel content, as well as a default mix of ambient and non-ambient content determined by the included speaker channel soundbed. Each optional mix includes different audio content and thus provides a different experience to the listener when rendered and reproduced. For example, in the case of a program indicating audio captured at a football match, one preset mix could indicate the atmosphere / effects mix for the home team crowd, and another preset mix could indicate the atmosphere / effects mix for the away team crowd. Typically, the default mix and multiple alternative preset mixes are encoded into a single bitstream. Alternatively, in addition to determining the speaker channel sound bed of the default mix, additional speaker channels (e.g., left and right (stereo) speaker channel pairs) that indicate optional audio content (e.g., submixes) are included in the bitstream, such that the speaker channels in the additional speaker channels can be selected and mixed with other content of the program (e.g., speaker channel content) in a playback system (e.g., a set-top box, which is sometimes referred to herein as "STB").
[0113] In other embodiments, the present invention is a method for rendering audio content determined from an object-based audio program, wherein the program indicates a speaker channel sound bed, a set of M replacement speaker channels, an object channel group, and metadata, wherein the object channel group includes N object channels, a first subgroup in the object channel group indicates a first non-ambient content, the first subgroup including M object channels in the object channel group, where each of N and M is a positive integer, and M is equal to or less than N.
[0114] The speaker channel sound bed indicates the default mix of audio content including a second non-ambient content that is different from the first non-ambient content, wherein an object-based speaker channel subgroup including M speaker channels in the sound bed indicates the mix of the second non-ambient content, or at least some audio content of the default mix, with the second non-ambient content.
[0115] Each of the M alternative speaker channels in this group indicates some, but not all, of the contents of the corresponding speaker channel in the object-based speaker channel subgroup, and
[0116] The metadata indicates at least one optional predetermined alternative mix of the content of at least one object channel and the content of a predetermined speaker channel and / or a replacement speaker channel in the speaker channels of the sound bed, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is a replacement mix comprising at least some audio content of the sound bed and a first non-ambient content rather than a second non-ambient content, the method comprising the steps of:
[0117] (a) Providing object-based audio programs to an audio processing unit (e.g., in a mode operation where decoder 7 resolves the speaker channel sound bed of the program). Figure 12 The playback system 111 or Figure 5 The playback system, in which the speaker channel sound bed is set to the rendering subsystem 11, and the rendering subsystem 11 renders the default mix indicated by the sound bed in response to the sound bed without using the program's object-related metadata; and
[0118] (b) In the audio processing unit, the speaker channel sound bed is parsed, and the default mix is rendered in response to the speaker channel sound bed without using metadata.
[0119] In some cases, the audio processing unit is a legacy playback system (or other audio data processing system) that is not configured to parse the program's object channel or metadata. When the audio processing unit is configured to parse the program's object channel, replacement channel, and metadata (e.g., Figure 5 The implementation of a playback system including decoder 7, object processing subsystem 9, controller 10, and rendering subsystem 11 is such that, when the rendering subsystem 11 is configured to render a selected mix of the program's object channel, sound bed speaker channel content, and replacement speaker channel content using object-related metadata of the program delivered to decoder 7, the method may include the steps of:
[0120] (c) In the audio processing unit, the replacement mix is rendered using at least some metadata, including rendering by selecting and mixing the contents of a first subgroup of the object channel group and at least one of the replacement speaker channels in response to at least some metadata (e.g., this step can be performed by...). Figure 6 Subsystems 22 and 24 of the system or by Figure 5 (to be executed by the playback system).
[0121] In some implementations, step (c) includes the following steps: driving a loudspeaker to provide sound that is perceived as a mixture of at least some of the audio content, including a sound bed, with a first non-ambient content but not a second non-ambient content.
[0122] In some implementations, step (c) includes the following steps:
[0123] (d) In response to the at least some metadata, select a first subgroup of the object channel group, select at least one speaker channel in the speaker channel bed other than the speaker channels in the object-based speaker channel subgroup, and select the at least one alternative speaker channel; and
[0124] (e) Mix the contents of the first subgroup of the object channel group and each speaker channel selected in step (d) to determine the replacement mix.
[0125] Step (d) can be, for example, Figure 6 Subsystem 22 of the system or Figure 5 The playback system subsystem 9 performs the operation. Step (e) can be performed by, for example, Figure 6 Subsystem 24 of the system Figure 5 The playback system is executed by subsystem 11.
[0126] In some embodiments, the method of the present invention generates (or delivers or renders) a personalized, object-based audio program as a bitstream comprising data indicating several layers:
[0127] A base layer (e.g., a 5.1 channel sound bed) that includes speaker channel sound beds that indicate all content of the default audio program (e.g., the default mix of ambient and non-ambient content);
[0128] Indicates at least one object channel of the optional audio content to be rendered (each object channel is an element of the extension layer);
[0129] At least one replacement speaker channel (each replacement speaker channel is an element of the expansion layer), which is capable of (through a properly configured playback system, e.g., Figure 5 or Figure 6(In an implementation of the playback system) one or more corresponding channels of the base layer are selected to replace, thereby determining a modified base layer comprising each original (non-replaced) channel of the base layer that has not been replaced and each selected replacement speaker channel. The modified base layer may be rendered or may be mixed with the content of at least one of the object channels and then rendered. For example, when the replacement speaker channel includes a central channel that only indicates atmosphere (to replace the central channel of the base layer that indicates non-ambient content (e.g., commentary or dialogue) mixed with ambient content), the modified base layer including such a replacement speaker channel may be mixed with the non-ambient content of at least one object channel of the program;
[0130] Optionally, at least one alternative speaker channel group (each alternative speaker channel is an element of an extended layer) is indicated for at least one audio content mix (e.g., each alternative speaker channel group may indicate a different multi-channel atmosphere / effect mix), wherein each of the alternative speaker channel groups can be selected (via a suitably configured playback system) to replace one or more corresponding channels of the base layer; and
[0131] Metadata indicating at least one optional experience resolution (typically, more than one optional experience resolution). Each experience resolution is an optional, pre-defined (“preset”) mix of the program’s audio content (e.g., a mix of the content of at least one object with the content of a speaker channel), and each preset mix has a pre-defined set of rendering parameters (e.g., the spatial location of the rendered object).
[0132] In some implementations, the metadata includes default program metadata indicating the base layer (e.g., to enable the selection of a default audio program and rendering of the base layer). Typically, the metadata does not include such default program metadata, but includes metadata indicating at least one object channel's content relative to the content of a predetermined speaker channel and / or a replacement speaker channel in the sound bed (alternate mix metadata), wherein the alternative mix metadata includes rendering parameters for each of the alternative mixes.
[0133] Typically, metadata indicates the optional preset mixes (e.g., all optional preset mixes) for the speaker channels and object channels of a program. Optionally, metadata is or includes metadata indicating a layered mix diagram that indicates the optional mixes (e.g., all optional mixes) for the speaker channels and object channels of a program.
[0134] In one implementation, the encoded bitstream indicating a program of the present invention includes: a base layer comprising a speaker channel soundbed indicating a default mix (e.g., a default 5.1 speaker channel mix with both ambient and non-ambient content), metadata, and optional extended channels (at least one object channel and at least one alternative speaker channel). Typically, the base layer includes a central channel indicating non-ambient content (e.g., comments or dialogue indicated by object channels also included in the program) that indicates mixing to ambient sounds (e.g., crowd noise). A decoder (or other element of the playback system) can use the metadata sent with the bitstream to select alternative "preset" mixes (e.g., by discarding (ignoring) the central channel of the default mix and replacing the discarded central channel with an alternative speaker channel to determine a modified speaker channel group (and optionally, by mixing the content of at least one object channel with the modified speaker channel group, e.g., by mixing the content of an object channel indicating alternative comments with an alternative central channel of the modified speaker channel group)).
[0135] The table below indicates the estimated bitrate of the audio content and related metadata for an exemplary embodiment of the personalized bitstream (encoded as an E-AC-3 bitstream) of the present invention:
[0136]
[0137] In the examples illustrated in the table, the 5.1 channel base layer can indicate both ambient and non-ambient content, where non-ambient content (e.g., commentary on a sporting event) is mixed to the three front channels (“diverged” non-ambient audio) or to the center channel only (“non-diverged” non-ambient audio). Alternate channel layers can include a single alternate center channel (e.g., indicating ambient-only content in the center channel of the base layer if the non-ambient content of the base layer is only included in the center channel), or three alternate front channels (e.g., indicating ambient-only content in the front channels of the base layer if the non-ambient content of the base layer is distributed among the front channels). Additional bed and / or object channels may optionally be included at the expense of additional estimated bitrate requirements.
[0138] As indicated in the table, a bitrate of 12kbps is typical for the metadata indicating each “experience clarity”, where experience clarity is a specification of an optional “preset” mix of audio content (e.g., a mix of the content of at least one object channel and speaker channel soundbed that gives a particular “experience”), including a set of mix / render parameters for that mix (e.g., the spatial location of the rendered object).
[0139] As indicated in the table, the bitrates of 2kbps to 5kbps "per object or soundbed" are typical for metadata indicating experience maps. An experience map is a layered mix diagram of optional preset mixes indicating the audio content of a delivered program, and each preset mix includes a certain number (e.g., 0, 1, or 2) of objects and typically at least one speaker channel (e.g., some or all of the speaker channels in a soundbed, and / or at least one alternative speaker channel). In addition to the speaker channels and object channels (objects) that may be included in each mix, the diagram typically indicates rules (e.g., grouping and conditional rules). The bitrate requirements for the rules (e.g., grouping and conditional rules) of the experience map are included in a given estimate for each object or soundbed.
[0140] The table records the replacement left, center, and right (L / C / R) speaker channels that can be selected to replace the left, center, and right channels of the base layer, and when the replacement channels are rendered, they are separated in the space of the new content of the replacement speaker channel (i.e., the content of the content of the corresponding channel of the base layer that it replaces) in the sense that it is separated in the area spanned by the left, center, and right speakers (of the playback system).
[0141] In an example of an embodiment of the invention, the object-based programming indicates personalized audio related to a football match (i.e., personalized soundtracks accompanying the video of the match). The default mix of the program includes ambient content (crowd noise captured at the match) mixed with default commentary (also provided as an optional object channel for the program), two object channels indicating alternative team-biased commentary, and an alternative speaker channel indicating ambient content without default commentary. The default commentary is unbiased (i.e., does not favor any team). The program offers four levels of clarity: a default mix including unbiased commentary, a first alternative mix including ambient content and commentary for a first team (e.g., the home team), a second alternative mix including ambient content and commentary for a second team (e.g., the away team), and a third alternative mix including only ambient content (no commentary). A typical implementation of the bitstream delivery, including data indicating the program, would have a bitrate requirement of approximately 452 kbps (assuming the base layer is a 5.1 speaker channel sound bed, and the default commentary is non-separate and located in the center channel only of that sound bed), allocated as follows: 192 kbps for the 5.1 base layer (indicating the default commentary in the center channel), 48 kbps (indicating an alternative center channel for non-ambient content only, which can be selected to replace the center channel of the base layer and optionally mixed with an alternative commentary indicated by one of the object channels), 144 kbps for the object layer including three object channels (one channel for "main" or unbiased commentary; one channel for commentary biased towards the first team (e.g., the home team); and one channel for commentary biased towards the second team (e.g., the away team)), 4.5 kbps for object-related metadata (for rendering the object channels), 48 kbps for metadata indicating four optional experiences, and 1.5 kbps for metadata indicating the experience map (layered mix map).
[0142] In a typical implementation of a playback system (configured to decode and render programs), metadata included in the program allows the user to select from a menu of rendering options: a default mix with unbiased commentary (which can be rendered either by rendering an unmodified soundbed or by replacing the central channel of the soundbed with a replacement central channel and mixing the resulting modified soundbed with unbiased commentary content from the relevant object channel); a first alternative mix (which can be rendered by replacing the central channel of the soundbed with a replacement central channel and mixing the resulting modified soundbed with commentary biased towards a first team); a second alternative mix (which can be rendered by replacing the central channel of the soundbed with a replacement central channel and mixing the resulting modified soundbed with commentary biased towards a second team); and a third alternative mix (which can be rendered by replacing the central channel of the soundbed with a replacement central channel). The menu is typically presented to the user via a user interface (e.g., via a wireless link) coupled to a controller of a set-top device (or other device, such as a TV, AVR, tablet, or telephone) configured to (at least partially) decode and render object-based programs. In some other implementations, the metadata included in the program allows the user to select from available rendering options in other ways.
[0143] In a second example of an embodiment of the invention, the object-based program indicates personalized audio related to a football match (i.e., personalized soundtracks accompanying the video of the match). The default mix of the program includes ambient content (crowd noise captured at the match) mixed with a first default unbiased commentary (the first unbiased commentary is also provided as an optional object channel for the program), five object channels indicating alternative non-ambient content (a second unbiased commentary, commentary biased towards both teams, announcement feed, and goal report feed), two alternative speaker channel groups (each of which is a 5.1 speaker channel group indicating different mixes of ambient and non-ambient content, each mix being different from the default mix), and alternative speaker channels indicating ambient content without the default commentary. The program offers at least nine levels of clarity: a default mix including first unbiased commentary; a second alternative mix including ambient content and first team commentary; a second alternative mix including ambient content and second team commentary; a third alternative mix including ambient content and second unbiased commentary; a fourth alternative mix including ambient content, first unbiased commentary, and announcement feed; a fifth alternative mix including ambient content, first unbiased commentary, and goal report feed; a sixth alternative mix (determined by the first alternative group of 5.1 speaker channels); a seventh alternative mix (determined by the second alternative group of 5.1 speaker channels); and an eighth alternative mix including ambient content only (no commentary, no announcement feed, and no goal report feed). A typical implementation of the bitstream delivery, including data indicating the program, would have a bitrate requirement of approximately 987 kbps (assuming the base layer is a 5.1 speaker channel sound bed, and the default commentary is non-separate and is presented only in the center channel of the sound bed), allocated as follows: 192 kbps for the 5.1 base layer (indicating the default commentary in the center channel), and 48 kbps (for indicating an alternative center channel for ambient-only content, which can be selected to replace the center channel of the base layer and optionally mixed with alternative content indicated by one or more object channels). 384kbps is used for the object layer, which includes six object channels (one channel for first unbiased comments; one channel for second unbiased comments; one channel for comments biased towards the first team; one channel for comments biased towards the second team; one channel for announcement feeds; and one channel for goal report feeds); 9kbps is used for object-related metadata (for rendering object channels); 36kbps is used for metadata indicating nine optional experiences; and 30kbps is used for metadata indicating the experience map (layered mix map).
[0144] In a typical implementation of the playback system (configured to decode and render the program of the second example), the metadata included in the program allows the user to select from a menu of rendering options: a default mix of unbiased commentary (which can be rendered either by rendering an unmodified soundbed or by replacing the central channel of the soundbed with a replacement central channel and mixing the resulting modified soundbed with the first unbiased commentary content of the relevant object channel); a first alternative mix (which can be rendered by replacing the central channel of the soundbed with a replacement central channel and mixing the resulting modified soundbed with commentary biased towards the first team); a second alternative mix (which can be rendered by replacing the central channel of the soundbed with a replacement central channel and mixing the resulting modified soundbed with commentary biased towards the second team); a third alternative mix (which can be rendered by replacing the central channel of the soundbed with a replacement central channel and mixing the resulting modified soundbed with commentary biased towards the second team); and a third alternative mix (which can be rendered by replacing the central channel with a replacement central channel). The following are alternative mixes: a fourth alternative mix (which can be rendered by replacing the central channel of the sound bed with the central channel replacement and mixing the resulting modified sound bed with the first unbiased comment and announcement feeds); a fifth alternative mix (which can be rendered by replacing the central channel of the sound bed with the central channel replacement and mixing the resulting modified sound bed with the first unbiased comment and announcement feeds); a sixth alternative mix (which can be rendered by rendering the first alternative group of 5.1 speaker channels instead of the sound bed); a seventh alternative mix (which can be rendered by rendering the second alternative group of 5.1 speaker channels instead of the sound bed); and an eighth alternative mix (which can be rendered by replacing the central channel of the sound bed with the central channel replacement). Menus are typically presented to the user via a user interface (e.g., via a wireless link) coupled to a controller of a set-top device (or other device, such as a TV, AVR, tablet, or telephone) configured to (at least partially) decode and render object-based programming. In some other implementations, metadata included in the program allows the user to select from available rendering options in other ways.
[0145] In other implementations, other methods are employed for carrying the extended layer of object-based programming, which includes speaker channels and object channels in addition to those of the base layer. Some of these methods reduce the overall bit rate required to deliver the base and extended layers. For example, joint object coding or receiver-side bed mixing can be used to allow for large bit rate savings in program delivery (at the cost of increased computational complexity and constrained artistic flexibility). For example, joint object coding or receiver-side bed mixing can be used to reduce the bit rate required to deliver the base and extended layers of the program in the second example described above from approximately 987 kbps (as noted above) to approximately 750 kbps.
[0146] The examples provided herein indicate the overall bitrate used to deliver the entire object-based audio program, including the base layer and extension layers. In other implementations, the base layer (sound bed) is delivered in-band (e.g., in the broadcast bitstream), and at least a portion of the extension layers (e.g., object channels, alternative speaker channels, and / or layered mix maps and / or other metadata) is delivered out-of-band (e.g., via Internet Protocol or an “IP” network) to reduce the in-band bitrate. An example of delivering the entire object-based audio program in a manner that divides transmission across in-band (broadcast) and out-of-band (Internet) is: the 5.1 base layer, alternative speaker channels, main comment object channel, and two alternative 5.1 speaker channel groups (using a total bitstream requirement of approximately 729 kbps) are delivered in-band, and alternative object channels and metadata (including experience clarity and layered mix maps) (using a total bitrate requirement of approximately 258 kbps) are delivered out-of-band.
[0147] Figure 6 This is a block diagram of an implementation of a playback system, which includes a decoder 20, an object processing subsystem 22, a spatial rendering subsystem 25, a controller 23 (which implements a user interface), and optionally digital audio processing subsystems 25, 26, and 27, coupled as shown, and can be implemented to perform the methods of the present invention. In some implementations, Figure 6 The system components 20, 22, 24, 25, 26, 27, 29, 31 and 33 are implemented as top-mounted devices.
[0148] exist Figure 6 In the system, decoder 20 is configured to receive and decode encoded signals indicating an object-based audio program (or the main mix of an object-based audio program). Typically, according to embodiments of the invention, the program (e.g., the main mix of the program) indicates audio content comprising a sound bed having at least two speaker channels and a group of alternative speaker channels. The program also indicates at least one user-selectable object channel (and optionally at least one other object channel) and object-related metadata corresponding to each object channel. Each object channel indicates an audio object, and thus, for convenience, the object channel is sometimes referred to herein as an "object". In one embodiment, the program is an AC-3 or E-AC-3 bitstream (or including its main mix) indicating the audio object, object-related metadata, speaker channel sound bed, and alternative speaker channels. Typically, the individual audio objects are encoded in mono or stereo (i.e., each object channel indicates the left or right channel of the object, or a mono channel indicating the object), the sound bed is a conventional 5.1 mix, and decoder 20 can be configured to simultaneously decode up to 16 channels of the audio content (including the 6 speaker channels of the sound bed, and the alternative speaker channels and object channels).
[0149] In some embodiments of the playback system of the present invention, each frame of the input E-AC-3 (or AC-3) encoded bitstream includes one or more metadata "containers". The input bitstream indicates an object-based audio program or the main mix of such a program, and the speaker channels of the program are organized in the same way as the audio content of a regular E-AC-3 (or AC-3) bitstream. One container may be included in the Aux field of the frame, and another container may be included in the addbsi field of the frame. Each container has a core header and includes one or more payloads (or is associated with one or more payloads). One such payload (of or associated with a container included in the Aux field) may be a set of audio samples of each of the object channels of the present invention (as opposed to a speaker channel soundbed also representing a program) and object-related metadata associated with each object channel. In such payloads, some or all of the samples in the object channel (and associated metadata) may be organized as standard E-AC-3 (or AC-3) frames, or may be organized in other ways (e.g., they may be included in a submix different from the E-AC-3 or AC-3 bitstream). Another example of such payloads is a group of loudness processing status metadata associated with the audio content of the frame.
[0150] In some such implementations, the decoder (e.g., Figure 6 The decoder will parse the core header of the container in the Aux field and extract the object channel and associated metadata of the invention from that container (e.g., from the Aux field of an AC-3 or E-AC-3 frame) and / or from the location indicated by the core header (e.g., the submix). After extracting the payload (object channel and associated metadata), the decoder will perform any necessary decoding on the extracted payload.
[0151] Each container's core header typically includes: at least one ID value indicating the type of payload included in or associated with the container; a substream association indicator (indicating which substreams the core header is associated with); and guard bits. Such guard bits (which may contain or include a hash-based message authentication code or "HMAC") are generally useful for at least one of the following: decryption, authentication, or verification of object-related metadata and / or loudness processing status metadata (and optionally other metadata) contained in at least one payload included in or associated with the container, and / or the corresponding audio data included in the frame. Substreams can be located "in-band" (in an E-AC-3 or AC-3 bitstream) or "out-of-band" (e.g., in a submix bitstream independent of the E-AC-3 or AC-3 bitstream). One type of such payload is a set of audio samples for each object channel in one or more object channels (associated with speaker channel sound beds also indicated by the program) and object-related metadata associated with each object channel. Each object channel is a separate substream and is typically identified in the core header. Another type of payload is loudness processing status metadata.
[0152] Typically, each payload has its own header (or "payload identifier"). Object-level metadata can be carried in each substream for an object channel. Program-level metadata can be included in the core header of the container and / or in the header of a payload that is a set of audio samples as one or more object channels (and the metadata associated with each object channel).
[0153] In some implementations, each container in the auxiliary data auxdata (or addbsi) field of the frame has a three-level structure:
[0154] The high-level structure includes a flag indicating whether the auxiliary data (or addbsi) field includes metadata (where, in this context, "metadata" means object channel, object-related metadata, and any other audio content or metadata carried by the bitstream but not typically in any container lacking the described type in a regular E-AC-3 or AC-3 bitstream), at least one ID value indicating what type of metadata is presented, and usually a value indicating how many bits (if metadata) are presented. In this context, an example of such a "type" of metadata is object channel data and associated object-related metadata (i.e., a set of audio samples for each object channel in one or more object channels (also related to the speaker channel soundbed indicated by the program) and the metadata associated with each object channel).
[0155] The intermediate structure includes core elements of the metadata for each type of identification (e.g., for each identification type, the core header, protection value, payload ID, and payload size value for the types mentioned above); and
[0156] The low-level structure, where at least one such payload is identified by the presented core element, includes a payload for each core element. An example of such a payload is a set of audio samples for each of one or more object channels (as opposed to speaker channel sound beds also indicated by the program), along with metadata associated with each object channel. Another example of such a payload is a payload that includes loudness processing state metadata (“LPSM”), sometimes referred to as an LPSM payload.
[0157] Data values in such a three-tier structure can be nested. For example, after each payload is identified by the core element (and thus after each payload is identified in the core header of the core element), a protection value for the payload identified by the core element (e.g., the LPSM payload) can be included. In one example, the core header may identify a first payload (e.g., the LPSM payload) and another payload. After the core header may be the payload ID and payload size value of the first payload, the first payload itself may follow the ID and size value, and the payload ID and payload size value of the second payload may follow the first payload, the second payload itself may follow these IDs and size values, and the protection value for one or two payloads (or the core element value and one or two payloads) may follow the last payload.
[0158] Refer again Figure 6 The user uses controller 23 to select the object to be rendered (indicated by the object-based audio program). Controller 23 can be programmed to implement... Figure 6 The system is a handheld processing device (e.g., an iPad) with a user interface (e.g., an iPad application) compatible with other components of the system. The user interface may provide the user with (e.g., displayed on a touchscreen) menus or palettes of objects, “soundbed” speaker channel content, and optional “preset” mixes that can replace the speaker channel content. Optional preset mixes may be determined by object-related metadata of the program and, generally also by rules implemented by subsystem 22 (e.g., rules that subsystem 22 has been pre-configured to implement). The user can select from the optional mixes by inputting commands to controller 23 (e.g., by activating their touchscreen), and in response, controller 23 sets the corresponding control data to subsystem 22.
[0159] Decoder 20 decodes the speaker channels of the program's speaker channel sound bed (as well as any alternative speaker channels included in the program) and outputs the decoded speaker channels to subsystem 22. In response to an object-based audio program, and in response to control data from controller 23 instructing the selected subgroup of the entire set of object channels to be rendered, decoder 23 decodes the selected object channels (if necessary) and outputs the selected (e.g., decoded) object channels (each object channel may be pulse code modulated or a "PCM" bitstream) and the object-related metadata corresponding to the selected object channels to subsystem 22.
[0160] The objects indicated by the decoded object channels are typically or include user-selectable audio objects. For example, the decoder can extract the 5.1 speaker channel soundbed, replace speaker channels (indicating ambient content of one of the soundbed speaker channels, rather than non-ambient content of one of the soundbed speaker channels), and indicate object channels that indicate commentary from a broadcaster from the home team's city (such as...). Figure 6 The “Comment-1 Mono” shown indicates the object channel for commentary (such as) given by the announcer from the away team's city. Figure 6 The “Comment-2 Mono” shown indicates the object channel of the noise from the crowd of fans appearing at the home team's stadium at the sporting event (such as...). Figure 6 The “fans (home team)” shown indicates the left and right object channels (e.g., the left and right object channels for the sound produced by the deciding ball when the ball is played by a participant in a sporting event). Figure 6 The “stereo spherical sound” shown), and the four object channels indicating special effects (such as...) Figure 6 (As shown in "Effect 4x Mono"). Any object channel among the "Comment-1 Mono" object channel, "Comment-2 Mono" object channel, "Fan (Home Team)" object channel, "Stereo Ball Sound" object channel, and "Effect 4x Mono" object channel (after undergoing any necessary decoding in decoder 20) can be selected, and each of these selected object channels will be passed from subsystem 22 to rendering subsystem 24.
[0161] Similar to the decoded speaker channels, decoded object channels, and decoded object-related metadata from decoder 20, input to object processing subsystem 22 may optionally include external audio object channels set to the system (e.g., one or more submixes of a program set to decoder 20 as its main mix). Examples of objects indicated by such external audio object channels include local commentators (e.g., single-channel audio content delivered via a radio channel), incoming Skype calls, and incoming Twitter connections (via...). Figure 6 (The text is not shown in the image; it is converted to speech by a system) and the system sound.
[0162] Subsystem 22 is configured to output selected subgroups of all object channels (or processed versions of selected subgroups of all object channels) and corresponding object-related metadata of the program, as well as selected speaker channel groups in the bed speaker channels and / or replacement speaker channels. Object channel selection and speaker channel selection can be determined by user selection (as indicated by control data set from controller 23 to subsystem 22) and / or by rules (e.g., indication conditions and / or constraints) that subsystem 22 has been programmed into or otherwise configured to implement. Such rules can be determined by object-related metadata of the program and / or other data (e.g., data indicating the performance and organization of the speaker array of the playback system) set to subsystem 22 from controller 23 or other external sources, and / or by pre-configuring (e.g., programming) subsystem 22. In some implementations, object-related metadata provides a set of optional "preset" mixes of speaker channel content and objects (speaker channel bed and / or alternative speaker channels), and subsystem 22 uses this metadata to select the object channels it optionally processes and then sets to subsystem 24, as well as the speaker channels it sets to subsystem 24. Subsystem 22 typically processes the selected subgroup of decoded speaker channels (bed speaker channels and typically alternative speaker channels) from decoder 20 (e.g., at least one speaker channel of the bed and at least one alternative speaker channel) and the selected object channels set thereto.
[0163] Object processing (including object selection) performed by subsystem 22 is typically controlled by control data from controller 23 and object-related metadata from decoder 20 (and optionally, object-related metadata for submixes set to subsystem 22 rather than decoder 20), and typically includes determining the spatial location and level of each selected object (regardless of whether the object selection is due to user selection or selection via rule application). Typically, default spatial locations and default levels for rendering objects, and optionally, restrictions on user selection of objects and their spatial locations and levels, are included in the object-related metadata set to subsystem 20 (e.g., from decoder 20). Such restrictions may indicate prohibited combinations of objects or prohibited spatial locations that selected objects can use to render (e.g., to prevent selected objects from being rendered too close to each other). Additionally, the loudness of each selected object is typically controlled by object processing subsystem 22 in response to control data input from controller 23 and / or by default levels indicated by object-related metadata (e.g., from decoder 20) and / or by pre-configuration of subsystem 22.
[0164] Typically, the decoding performed by decoder 20 includes extracting metadata (from the input program) indicating the type of audio content for each object indicated by the program (e.g., the type of sporting event indicated by the program's audio content, and the names or other identifying markers (e.g., team logos) of optional and default objects indicated by the program). Controller 23 and object processing subsystem 22 receive this metadata or related information indicated by it. Additionally, controller 23 typically receives (e.g., is programmed to) information related to the playback performance of the user's audio system (e.g., the number of speakers, assumed speaker placement, or other assumed organization).
[0165] Figure 6 The spatial rendering subsystem 24 (or subsystem 24 having at least one downstream device or system) is configured to render audio content output from subsystem 22 for playback by the speakers of a user's playback system. Optionally included, one or more of the digital audio processing subsystems 25, 26, and 27 can perform post-processing of the output from subsystem 24.
[0166] The spatial rendering subsystem 24 is configured to use rendering parameters (e.g., user-selected and / or default values for spatial location and level) output from subsystem 22 associated with each selected object to map audio object channels (e.g., default-selected objects, and / or user-selected objects selected due to user interaction using controller 23) selected (or selected and processed by subsystem 22) and set to subsystem 24 to available speaker channels (e.g., a set of bed speaker channels determined by subsystem 22 and passed through subsystems 22 to subsystem 24, as well as alternative speaker channels). Typically, subsystem 24 is a smart mixer and is configured to determine speaker feeds for available speakers, including by mapping one, two, or more selected object channels to each of a large number of individual speaker channels, and by mixing the selected object channels with audio content indicated by each corresponding speaker channel.
[0167] Typically, the number of output speaker channels can vary between 2.0 and 7.1, and the speakers to be driven to render the selected audio object channels (to mix with the content of the selected speaker channels) can be assumed to be located in the (nominal) horizontal plane of the playback environment. In such a case, rendering is performed such that the speakers can be driven to emit sound mixed with the sound determined by the content of the speaker channels, which will be perceived as emanating from different object locations in the plane of the speakers (i.e., one object location or a series of object locations along a trajectory for each selected or default object).
[0168] In some implementations, the number of full-range speakers to be driven to render audio can be any number in a wide range (not necessarily limited to the range of 2 to 7), and thus the number of output speaker channels is not limited to the range of 2.0 to 7.1.
[0169] In some implementations, the speakers to be driven to render audio are assumed to be located at arbitrary locations within the playback environment; not just in the (nominal) horizontal plane. In some such cases, metadata included in the program indicates rendering parameters for using a three-dimensional speaker array to render at least one object of the program at an arbitrary apparent spatial location (in the three-dimensional volume). For example, an object channel may have corresponding metadata indicating a three-dimensional trajectory of the apparent spatial location of the object to be rendered (indicated by the object channel). The trajectory may include a series of “floor” locations (in the plane of a speaker subgroup assumed to be located on the floor of the playback environment, or in another horizontal plane of the playback environment) and a series of “above-floor” locations (each “above-floor” location is determined by driving a speaker subgroup assumed to be located in at least one other horizontal plane of the playback environment). In such cases, rendering can be performed according to the invention such that the speakers can be driven to emit sound mixed with sound determined by the speaker channel content (determined by the associated object channel), which will be perceived as emanating from a series of object locations in the three-dimensional space including the trajectory. Subsystem 24 may be configured to implement such rendering or its steps, wherein the remaining steps of rendering are handled by downstream systems or devices (e.g., Figure 6 The rendering subsystem 35) is used to execute this.
[0170] Optionally, a digital audio processing (DAP) level (e.g., one level in each of a large number of predetermined output speaker channel configurations) is coupled to the output of the spatial rendering subsystem 24 to perform post-processing on the output of the spatial rendering subsystem. Examples of such processing include intelligent equalization or (in the case of stereo output) speaker virtualization processing.
[0171] Figure 6The system output (e.g., the output of the spatial rendering subsystem or a DAP level following the spatial rendering level) can be a PCM bitstream (which determines the speaker feed for the available speakers). For example, in the case where the user's playback system includes a 7.1 speaker array, the system can output a PCM bitstream that determines the speaker feed for such an array (generated in subsystem 24), or a post-processed version of such a bitstream (generated in DAP 25). For another example, in the case where the user's playback system includes a 5.1 speaker array, the system can output a PCM bitstream that determines the speaker feed for such an array (generated in subsystem 24), or a post-processed version of such a bitstream (generated in DAP 26). For another example, in the case where the user's playback system includes only left and right speakers, the system can output a PCM bitstream that determines the speaker feed for both the left and right speakers (generated in subsystem 24), or a post-processed version of such a bitstream (generated in DAP 27).
[0172] Figure 6 The system may optionally include one or both of re-encoding subsystems 31 and 33. Re-encoding subsystem 31 is configured to re-encode a PCM bitstream (indicating a feed for a 7.1 speaker array) output from DAP 25 as an E-AC-3 encoded bitstream, and may output the resulting encoded (compressed) E-AC-3 bitstream from the system. Re-encoding subsystem 33 is configured to re-encode a PCM bitstream (indicating a feed for a 5.1 speaker array) output from DAP 27 as an AC-3 or E-AC-3 encoded bitstream, and may output the resulting encoded (compressed) AC-3 or E-AC-3 bitstream from the system.
[0173] Figure 6The system may optionally include a re-encoding (or formatting) subsystem 29 and a downstream rendering subsystem 35 coupled to receive the output of subsystem 29. Subsystem 29 is coupled to receive data (output from subsystem 22) indicating the selected audio object (or the default mix of the audio object), the corresponding object-related metadata, and the decoded speaker channels (e.g., bed speaker channels and replacement speaker channels), and is configured to re-encode (and / or format) such data for rendering by subsystem 35. Subsystem 35, which may be implemented in an AVR or soundbar (or other systems or devices downstream of subsystem 29), is configured to generate a speaker feed (or determine a bitstream of the speaker feed) for use with the playback speakers (speaker array 36) in response to the output of subsystem 29. For example, subsystem 29 may be configured to encode audio into an appropriate format for rendering in subsystem 35 by re-encoding data indicating the selected (or default) audio object, the corresponding metadata, and the speaker channels, and to transmit the encoded audio (e.g., via an HDMI link) to subsystem 35. In response to a speaker feed generated by (or determined by) the output of subsystem 35, available speaker 36 will emit a mixed sound indicating the speaker channel content and a selected (or default) object, wherein the object has a apparent source location determined by object-related metadata output by subsystem 29. When subsystems 29 and 35 are included, rendering subsystem 24 may optionally be omitted from this system.
[0174] In some embodiments, the present invention is a distributed system for rendering object-based audio, wherein, in a first subsystem (e.g., implemented in a set-top device or a set-top device and a handheld controller), Figure 6 A portion of the rendering (i.e., at least one step) is implemented in components 20, 22, and 23 (e.g., as by...). Figure 6 Subsystem 22 and controller 23 of the system perform the selection of audio objects to be rendered and the selection of rendering features for each selected object, and another part of the rendering is performed in a second subsystem (e.g., subsystem 35 implemented in an AVR or soundbar) (e.g., immersive rendering that generates speaker feeds or determines the signal of the speaker feeds in response to the output of the first subsystem). Some implementations of distributed rendering also implement legacy management to account for different times and different subsystems, performing multiple parts of audio rendering (any processing of video corresponding to the rendered audio) at different times and in different subsystems.
[0175] In some embodiments of the playback system of the present invention, each decoder and object processing subsystem (sometimes referred to as a personalization engine) is implemented in the set-top unit (STB). For example, it can be implemented in the STB. Figure 6Components 20 and 22 and / or Figure 7 All components of the system. In some embodiments of the playback system of the present invention, the output of the personalization engine is rendered multiple times to ensure that all STB outputs (e.g., STB HDMI output, S / PDIF output, or stereo analog output) are made available. Optionally, the selected object channel (and corresponding object-related metadata) and the speaker channel (along with the decoded speaker channel sound bed) are passed from the STB to a downstream device (e.g., an AVR or soundbar) configured to render the mix of the object channel and the speaker channel.
[0176] In one type of implementation, the object-based audio program of the present invention comprises a set of bitstreams (multiple bitstreams, which may be referred to as "substreams") generated and transmitted in parallel. In some implementations of this type, multiple decoders are used to decode the content of the substreams (e.g., the program comprises multiple E-AC-3 substreams, and the playback system utilizes multiple E-AC-3 decoders to decode the content of the substreams). Figure 7 This is a block diagram of a playback system configured to decode and render an implementation of the present invention of an object-based audio program comprising multiple serial bit streams delivered in parallel.
[0177] Figure 7 The playback system is Figure 6 A variant of the system, in which the object-based audio program comprises multiple bitstreams (B1, B2, ..., BN, where N is a positive integer), which are delivered in parallel to and received by a playback system. Each bitstream (“substream”) B1, B2, ..., and BN is a serial bitstream that includes timecodes or other synchronization words (see reference). Figure 7 For convenience, these are referred to as "synchronization words" to enable substreams to synchronize or time-align with each other. Each substream also includes a distinct subgroup of object channels and corresponding object-related metadata, and at least one substream includes a speaker channel (e.g., a bed speaker channel and an alternate speaker channel). For example, in each substream B1, B2, ..., BN, each container that includes the object channel content and object-related metadata includes a unique ID or timestamp.
[0178] Figure 7 The system includes N deformatters 50, 51, ..., 53, each coupled and configured to parse different substreams in the input substream and set the metadata (including its synchronization word) and its audio content to bitstream synchronization level 59.
[0179] Deformatter 50 is configured to parse substream B1 and set its synchronization word (T1), other metadata, and its object channel content (M1) (including object-related metadata of the program and at least one object channel) and its speaker channel audio content (A1) (including at least one speaker channel of the program) to bitstream synchronization level 59. Similarly, deformatter 51 is configured to parse substream B2 and set its synchronization word (T2), other metadata, and its object channel content (M2) (including object-related metadata of the program and at least one object channel) and its speaker channel audio content (A2) (including at least one speaker channel of the program) to bitstream synchronization level 59. Deformatter 53 is configured to parse substream BN and set its synchronization word (TN), other metadata, and its object channel content (MN) (including object-related metadata of the program and at least one object channel) and its speaker channel audio content (AN) (including at least one speaker channel of the program) to bitstream synchronization level 59.
[0180] Figure 7 The system's bitstream synchronization stage 59 typically includes buffers for the audio content and metadata of substreams B1, B2, ..., BN, and a stream bias compensation element coupled and configured to use the synchronization word of each substream to determine any misalignment in the data of the input substream (e.g., misalignment may occur due to the possibility of losing strict synchronicity in distribution / contribution, since each bitstream is typically carried within the media file via separate interfaces and / or traces). The stream bias compensation element of stage 59 is also typically configured to correct any determined misalignment by setting appropriate control values to the buffers containing audio data and metadata of the bitstreams, such that the time-aligned bits of the speaker channel audio data are read from the buffers to the decoders (including decoders 60, 61, and 63), each decoder being coupled to a corresponding buffer in the buffers, and the time-aligned bits of the object channel audio data and metadata are read from the buffers to the object data combination stage 66.
[0181] The time-aligned bit slave level 59 of the speaker channel audio content A1' from substream B1 is read to decoder 60, and the time-aligned bit slave level 59 of the object channel content and metadata M1' from substream B1 is read to metadata combiner 66. Decoder 60 is configured to decode the speaker channel audio data set thereto and set the resulting decoded speaker channel audio to object processing and rendering subsystem 67.
[0182] Similarly, the time-aligned bit slave level 59 of the speaker channel audio content A2' from substream B2 is read to decoder 61, and the time-aligned bit slave level 59 of the object channel content and metadata M2' from substream B2 is read to metadata combiner 66. Decoder 61 is configured to perform decoding on the speaker channel audio data set thereto, and set the resulting decoded speaker channel audio to object processing and rendering subsystem 67.
[0183] Similarly, the time-aligned bit slave level 59 of the speaker channel audio content AN' from substream BN is read to decoder 63, and the time-aligned bit slave level 59 of the object channel content and metadata MN' from substream BN is read to metadata combiner 66. Decoder 63 is configured to perform decoding on the speaker channel audio data set thereto, and set the resulting decoded speaker channel audio to object processing and rendering subsystem 67.
[0184] For example, each substream B1, B2, ..., BN can be an E-AC-3 substream, and each decoder 60, 61, 63, as well as any other decoder coupled in parallel with decoders 60, 61, and 63 to subsystem 59, can be an E-AC-3 decoder configured to decode the speaker channel content of one of the input E-AC-3 substreams.
[0185] The data object combiner 66 is configured to set the time-aligned object channel data and metadata of all object channels of the program to the object processing and rendering subsystem 67 in an appropriate format.
[0186] Subsystem 67 is coupled to the output of combiner 66 and the outputs of decoders 60, 61, and 63 (and any other decoders coupled in parallel with decoders 60, 61, and 63 between subsystem 59 and subsystem 67), and controller 68 is coupled to subsystem 67. Subsystem 67 is typically configured to perform object processing (e.g., including by...) on the outputs of combiner 66 and decoders in an interactive manner according to embodiments of the invention, in response to control data from controller 68. Figure 6 The steps performed by subsystem 22 of the system, or variations thereof. Controller 68 can be configured to perform operations, wherein... Figure 6 The system controller 23 is configured to perform the operation (or a variation thereof) in response to input from the user. Subsystem 67 is also typically configured to perform rendering (e.g., by...) on the speaker channel audio and object channel audio data set thereto, according to embodiments of the invention (e.g., rendering a mix of the bed speaker channel content, the replacement speaker channel content, and the object channel content). Figure 6The system's rendering subsystem 24 or subsystems 24, 25, 26, 31, and 33, or Figure 6 The operations performed by subsystems 24, 25, 26, 31, 33, 29, and 35 of the system, or variations thereof.
[0187] exist Figure 7 In one implementation of the system, each substream B1, B2, ..., BN is a Dolby E bitstream. Each such Dolby E bitstream comprises a series of bursts. Each burst may carry speaker channel audio content (content of the bed speaker channel and / or replacement speaker channels) as well as a subgroup and object-related metadata of the entire object channel group (which may be a large group) of the object channels of the present invention (i.e., each burst may indicate some object channels of the entire object channel group and the corresponding object-related metadata). Each burst of the Dolby E bitstream typically occupies a time period equivalent to the time period of the corresponding video frame. Each Dolby E bitstream in the group includes a synchronization word (e.g., a timecode) to enable the bitstreams in the group to be synchronized or time-aligned with each other. For example, in each bitstream, each container including object channel content and object-related metadata may include a unique ID or timestamp to enable the bitstreams in the group to be synchronized or time-aligned with each other. Figure 7 In the specified implementation of the system, each deformatter 50, 51, and 53 (and any other deformatter coupled in parallel with deformatters 50, 51, and 53) is an SMPTE 337 deformatter, and each decoder 60, 61, and 63, and any other decoder coupled in parallel with decoders 60, 61, and 63 to subsystem 59, can be a Dolby E decoder.
[0188] In some embodiments of the present invention, object-related metadata for object-based audio programs includes persistent metadata. For example, metadata input to... Figure 6The object-related metadata included in the program of subsystem 20 of the system may include non-persistent metadata (e.g., for user-selectable objects, default rank and / or rendering position or trajectory) and persistent metadata. Non-persistent metadata may vary at at least one point in the broadcast chain (from the content creation facility that generated the program to the user interface implemented by controller 23), while persistent metadata is not intended to be variable after the initial program generation (typically, in the content creation facility). Examples of persistent metadata include: object IDs for each user-selectable object or other object or group of objects in the program; and timecodes or other synchronization words indicating the timing of each user-selectable object or other object relative to the speaker channel content or other elements of the program. Persistent metadata is typically maintained throughout the entire broadcast chain from the content creation facility to the user interface, throughout the entire duration of the program's broadcast, or even during program replays. In some implementations, the audio content (and associated metadata) of at least one user-selectable object is transmitted in the main mix of the object-based audio program, and at least some persistent metadata (e.g., timecode) and optionally the audio content (and associated metadata) of at least one other object are transmitted in the secondary mix of the program.
[0189] In some embodiments of the object-based audio program of the present invention, object-related metadata is used to store (e.g., even after the program is broadcast) the user-selected mix of object content and speaker channel content. For example, this can provide the selected mix as the default mix whenever a user watches a specific type of program (e.g., any football match) or whenever a user watches any program (of any type), until the user changes his / her selection. For example, during the broadcast of the first program, the user can utilize ( Figure 6 The system's controller 23 selects a mix that includes objects with persistent IDs (e.g., objects identified by the controller 23's user interface as "main crowd noise" objects, where the persistent ID indicates "main crowd noise"). Thus, whenever a user watches (and listens to) another program that includes objects with the same persistent ID, the playback system automatically renders the program using the same mix (i.e., the program's bed speaker channel and / or alternative speaker channel mixed with the program's "main crowd noise" object channel) until the user changes the mix selection. In some embodiments of the object-based audio program of the present invention, persistent object-related metadata can make the rendering of some objects mandatory throughout the program (e.g., even if the user expects to defeat such rendering).
[0190] In some implementations, object-related metadata provides a default mix of object content and speaker channel content with default rendering parameters (e.g., the default spatial location of the rendered object). For example, it is input to... Figure 6The program's object-related metadata in subsystem 20 can be a default mix with default rendering parameters for object content and speaker channel content, and subsystems 22 and 24 will render the program using the default mix and default rendering parameters, unless the user uses controller 23 to select another mix and / or another set of rendering parameters for object content and speaker channel content.
[0191] In some implementations, object-related metadata provides a set of optional "preset" mixes for the object and speaker channel content, each preset mix having a predetermined set of rendering parameters (e.g., the spatial location of the rendered object). The playback system's user interface can present these as a limited menu or palette of available mixes (e.g., by...). Figure 6 The system's controller 23 displays a limited menu or palette. Each preset mix (and / or each optional object) may have a persistent ID (e.g., name, tag, or logo). The controller 23 (or the controller of another embodiment of the playback system of the present invention) may be configured (e.g., on a touchscreen implemented on an iPad) to display an indication of such IDs. For example, there may be an optional "main" mix with a persistent ID (e.g., a team logo), regardless of changes to the details of the audio content or non-persistent metadata of each object in the preset mix (e.g., changes made by the broadcasting company).
[0192] In some implementations, the program's object-related metadata (or the reconfiguration of the playback or rendering system, which is not indicated by metadata delivered with the program) provides constraints or conditions for optional mixing of the object and sound bed (speaker channel) content. For example, Figure 6 The system implementation can achieve Digital Rights Management (DRM), and more specifically, it can implement DRM hierarchies to enable... Figure 6 Users of the system are able to have "hierarchical" access to groups of audio objects included in object-based audio programs. If a user (e.g., a customer associated with the playback system) (e.g., paying a broadcaster) pays more money, the user can be authorized to decode and select (and listen to) more audio objects in the program.
[0193] For another example, object-related metadata can provide constraints on user selections of objects. An example of such constraints is: if a user uses controller 23 to select whether to render both the "Home Team Crowd Noise" object and the "Home Team Announcer" object of a program (i.e., for inclusion in the program by...). Figure 6In the mix determined by subsystem 24, the metadata included in the program can ensure that subsystem 24 renders the two selected objects by a predetermined relative spatial position. Constraints can be (at least partially) determined by data related to the playback system (e.g., user input data). For example, if the playback system is a stereo system (including only two speakers), then... Figure 6 The system's object processing subsystem 24 (and / or controller 23) can be configured to prevent the user from selecting mixes that cannot be rendered at the appropriate spatial resolution using only two speakers (identified by object-related metadata). For another example, Figure 6 The system's object processing subsystem 24 (and / or controller 23) may remove some delivered objects from the optional object categories indicated by object-related metadata (and / or other data input to the playback system) for legitimate (e.g., DRM) or other reasons (e.g., based on delivery channel bandwidth). Users can pay content creators or broadcasters for more bandwidth, and therefore the system (e.g., ...) Figure 6 The system's object processing subsystem 24 and / or controller 23) allows the user to select from a larger menu of selectable objects and / or object / sound bed mixes.
[0194] Some embodiments of the present invention (e.g., Figure 6 The playback system, including the aforementioned components 29 and 35, implements distributed rendering. For example, the program's default or selected object channel (and corresponding object-related metadata) is received from the set-top device (e.g., from...). Figure 6 The implementation subsystems 22 and 29 of the system are passed (along with the selected group of decoded speaker channels, such as bed speaker channels and replacement speaker channels) to downstream devices (e.g., AVRs or soundbars implemented downstream of the set-top box (STB) implementing subsystems 22 and 29). Figure 6 (Subsystem 35). Downstream devices are configured to render the mix of object channels and speaker channels. The STB can render the audio in part, and the downstream device can perform the rendering (e.g., by generating speaker feeds to drive specific top-level speakers (e.g., ceiling speakers) to place audio objects at specific apparent source locations, where the STB's output only indicates that the object can be rendered in some non-specific way in some non-specific top-level speakers). For example, the STB may not have knowledge of the specific organization of the speakers in the playback system, but the downstream device (e.g., an AVR or soundbar) may have such knowledge.
[0195] In some implementations, object-based audio programs (e.g., input to...) Figure 6 Subsystem 20 of the system or Figure 7The programs of system elements 50, 51, and 53 are or include at least one AC-3 (or E-AC-3) bitstream, and each container of a program that includes object channel content (and / or object-related metadata) is included in the auxiliary data auxdata field at the end of the frame of the bitstream (e.g., Figure 1 or Figure 4 In some such implementations, each frame of an AC-3 or E-AC-3 bitstream includes one or two metadata containers. One container may be included in the Aux field of the frame, and another container may be included in the addbsi field of the frame. Each container has a core header and includes one or more payloads (or is associated with one or more payloads). One such payload (of or associated with a container included in the Aux field) may be a set of audio samples of each of one or more object channels in the object channel of the present invention (associated with the speaker channel sound bed also indicated by the program) and object-related metadata associated with each object channel. The core header of each container typically includes: at least one ID value indicating the type of payload included in or associated with the container; a substream association indicator (indicating which substreams the core header is associated with); and guard bits. Typically, each payload has its own header (or "payload identifier"). Object-level metadata may be carried in each substream that is an object channel.
[0196] In other implementations, object-based audio programs (e.g., input to...) Figure 6 Subsystem 20 of the system or Figure 7The programs of system elements 50, 51, and 53 are or include bitstreams that are not AC-3 or E-AC-3 bitstreams. In some embodiments, the object-based audio program is or includes at least one Dolby E bitstream, and the object channel content and object-related metadata of the program (e.g., each container of the program including object channel content and / or object-related metadata) are included in bit positions of the Dolby E bitstream that do not typically carry useful information. Each burst of the Dolby E bitstream occupies a time period equivalent to the time period of the corresponding video frame. The object channel (and / or object-related metadata) may be included in guard bands between Dolby E bursts and / or in unused bit positions within each data structure (each in a format with AES3 frames) within each Dolby E burst. For example, each guard band includes a series of segments (e.g., 100 segments), each of the first X segments of each guard band (e.g., X = 20) includes the object channel and object-related metadata, and each of the remaining segments of each guard band may include a guard band symbol. In some embodiments, at least some object channels (and / or object-related metadata) of the program of the present invention are included in the four least significant bits (LSBs) of each of the two AES3 subframes of each frame in at least some AES3 frames of the Dolby E bitstream, and data indicating the speaker channels of the program are included in the twenty most significant bits (MSBs) of each of the two AES3 subframes of each AES3 frame of the bitstream.
[0197] In some implementations, the object channels and / or object-related metadata of the program of the present invention are included in metadata containers within a Dolby E bitstream. Each container has a core header and includes one or more payloads (or is associated with one or more payloads). One such payload (of or associated with a container included in the Aux field) may be a set of audio samples of each object channel in one or more object channels of the present invention (e.g., associated with a speaker channel also indicated by the program) and object-related metadata associated with each object. The core header of each container typically includes: at least one ID indicating the type of payload included in or associated with the container; a substream association indicator (indicating which substreams the core header is associated with); and guard bits. Typically, each payload has its own header (or "payload identifier"). Object-level metadata may be carried in each substream that is an object channel.
[0198] In some implementations, object-based audio programs (e.g., input to...) are rendered using legacy decoders and legacy rendering systems (which are not configured to parse the object channels and object-related metadata of the present invention). Figure 6 Subsystem 20 of the system or Figure 7 The programs of system elements 50, 51, and 53 are decodeable, and their speaker channel content is renderable. The same program can be rendered using a set-top device (or other decoding and rendering system) configured (according to one embodiment of the invention) to parse the object channel and object-related metadata of the invention and to render the mix of speaker channel and object channel content indicated by the program.
[0199] Some embodiments of the present invention are intended to provide end-users with a personalized (and preferably immersive) audio experience in response to broadcast programming, and / or to provide new methods for using metadata in broadcast pipelines. Some embodiments improve microphone capture (e.g., stadium microphone capture) to generate audio programs that provide end-users with a more personalized and immersive experience, modifying existing generation, contribution, and distributed workflows to enable object channels and metadata of the object-based audio programs of the present invention to flow through a professional chain, and creating new playback pipelines (e.g., a playback pipeline implemented in a set-top device) that support object channels, replace speaker channels and associated metadata, and conventionally broadcast audio (e.g., speaker channel sound beds included in embodiments of the broadcast audio programs of the present invention).
[0200] Figure 8 This is a block diagram of a broadcasting system configured to generate object-based audio programs (and corresponding video programs) for broadcasting, according to an embodiment of the present invention. Figure 8 A set of X microphones (where X is an integer), including microphones 100, 101, 102 and 103, are positioned to capture audio content to be included in the program, and their outputs are coupled to the input of audio console 104.
[0201] In one implementation, the program includes interactive audio content indicating the atmosphere and / or commentary on the event being watched (e.g., a football or rugby match, a car or motorcycle race, or other sporting event). In some implementations, the program's audio content indicates multiple audio objects (including user-selectable objects or groups of objects, and typically a default group of objects to be rendered if the user does not select any objects), a speaker channel sound bed (indicating a default mix of the captured content), and alternative speaker channels. The speaker channel sound bed can be a conventional mix (e.g., a 5.1 channel mix) of speaker channels that can be included in a regular broadcast program that does not include object channels.
[0202] Microphone subgroups (e.g., microphones 100 and 101, and optionally other microphones whose outputs are coupled to audio console 104) are conventional microphone arrays that capture audio in operation (to be encoded and delivered as a speaker channel sound bed and a set of alternative speaker channels). In operation, another subgroup of microphones (e.g., microphones 102 and 103, and optionally other microphones whose outputs are coupled to audio console 104) captures audio (e.g., crowd noise and / or other "objects") to be encoded and delivered as object channels of the program. For example, Figure 8 The system's microphone array may include: at least one microphone (e.g., microphone 100) implemented as a sound field microphone (e.g., a sound field microphone with a heater) and permanently installed in the stadium; at least one stereo microphone (e.g., microphone 102 implemented as a Sennheiser MKH416 microphone or another stereo microphone) pointed at a location supporting the spectators of one team (e.g., the home team); and at least one other stereo microphone (e.g., microphone 103 implemented as a Sennheiser MKH416 microphone or another stereo microphone) pointed at a location supporting the spectators of another team (e.g., the away team).
[0203] The broadcasting system of the present invention may include a mobile unit (which may be a truck and is sometimes referred to as a “matching truck”) located outside a stadium (or other event location) as the first receiver of audio feeds from microphones in the stadium (or other event location). The matching truck generates object-based audio programs (to be broadcast) by encoding audio content from the microphones for delivery as object channels of the program, generating corresponding object-related metadata (e.g., metadata indicating the spatial location where each object should be rendered) and including such metadata in the program, and encoding audio content from some microphones for delivery as speaker channel sound beds (and a set of alternative speaker channels) of the program.
[0204] For example, in Figure 8 In the system, console 104, object processing subsystem 106 (coupled to the output of console 104), embedding subsystem 108, and contribution encoder 110 can be mounted in a matching bogie. Object-based audio programs generated in subsystem 106 can be combined (e.g., from cameras placed in the stadium) with video content (e.g., in subsystem 108) to generate combined audio and video signals, which are then (e.g., encoded by encoder 110) to generate encoded audio / video signals for broadcast (e.g., via...). Figure 5The delivery subsystem 5). It should be understood that the playback system for decoding and rendering such encoded audio / video signals will include a subsystem (not specifically shown in the figure) for parsing the audio and video content of the delivered audio / video signals, and a subsystem for decoding and rendering the audio content according to embodiments of the present invention (e.g., with...). Figure 6 The system includes similar or identical subsystems, as well as another subsystem for decoding and rendering video content (not specifically shown in the diagram).
[0205] The audio outputs of console 104 include: a default mix of ambient sounds captured during sporting events and commentary by the announcer (non-ambient content) mixed to its center channel, and a 5.1 speaker channel sound bed (in... Figure 8 The middle channel is marked "5.1 Neutral"; it indicates the ambient content of the central channel of the sound bed without commenting on the alternative speaker channels (in...). Figure 8 The following are audio contents: (1.0 Replacement) (i.e., the ambient sound content captured in the central channel of the sound bed before the commentary is mixed with it to generate the central channel of the sound bed); (2.0 Home Team) audio content of the stereo object channel indicating the crowd noise from the fans of the home team appearing at the event; (2.0 Away Team) audio content of the stereo object channel indicating the crowd noise from the fans of the away team appearing at the event; (1.0 Comment 1) audio content of the object channel indicating the commentary made by the announcer from the home team's city; (1.0 Comment 2) audio content of the object channel indicating the commentary made by the announcer from the away team's city; and (1.0 Kick) audio content of the object channel indicating the sound produced by the decisive ball when the ball is played by the participants in the sports event.
[0206] The object processing subsystem 106 is configured to organize (e.g., group) audio streams from console 104 into object channels (e.g., grouping the left and right audio streams labeled “2.0 Away” into an away crowd noise object channel) and / or object channel groups to generate object-related metadata indicating the object channels (and / or object channel groups), and to encode the object channels (and / or object channel groups), object-related metadata, speaker channel sound beds, and each replacement speaker channel (determined based on the audio streams from console 104) into object-based audio programs (e.g., object-based audio programs encoded as Dolby E bitstreams). Typically, subsystem 106 is also configured to render (and play on a set of studio monitor speakers) at least selected subgroups of object channels (and / or object channel groups) and speaker channel sound beds and / or replacement speaker channels (including generating mixes indicating the selected object channels and speaker channels using object-related metadata) such that the played-back sound can be monitored by operators of console 104 and subsystem 106 (e.g., via...). Figure 8 (As indicated by the "monitor path").
[0207] The interface between the output of subsystem 104 and the input of subsystem 106 can be a multi-channel audio digital interface (“MADT”).
[0208] During operation, Figure 8 Subsystem 108 combines the object-based audio program generated in subsystem 106 with video content (e.g., from cameras placed in a stadium) to generate a combined audio and video signal, which is then fed to encoder 110. The interface between the output of subsystem 108 and the input of subsystem 110 can be a High Definition Serial Digital Interface (“HD-SDI”). In operation, encoder 110 encodes the output of subsystem 108 to generate encoded audio / video signals for broadcast (e.g., via...). Figure 5 Delivery subsystem 5).
[0209] In some implementations, broadcasting facilities (e.g., Figure 8 Subsystems 106, 108, and 110 of the system are configured to generate multiple object-based audio programs (e.g., from...) that indicate the captured sound. Figure 8The subsystem 110 outputs multiple encoded audio / video signals indicating object-based audio programs. Examples of such object-based audio programs include 5.1 flattened mixes, international mixes, and domestic mixes. For example, all programs may include a common speaker channel soundbed (and a common set of alternative speaker channels), but the program's object channels (and / or a menu of optional object channels determined by the program, and / or optional or non-optional rendering parameters for rendering and mixing object channels) may vary from program to program.
[0210] In some implementations, the facilities of the broadcaster or other content creator (e.g., Figure 8 Subsystems 106, 108, and 110 of the system are configured to generate a single object-based audio program (i.e., the original) that can be rendered in any playback environment, such as a 5.1-channel domestic playback system, a 5.1-channel international playback system, and a stereo playback system. The original does not need to be mixed (e.g., downmixed) to be played to a client in any particular environment.
[0211] As noted above, in some embodiments of the invention, the program's object-related metadata (or a pre-configured playback or rendering system not indicated by metadata delivered with the program) provides constraints or conditions for optional mixing of the object and speaker channel content. For example, Figure 6 The system implementation can establish DRM hierarchies, enabling users to have tiered access to groups of object channels included in object-based audio programs. If a user (e.g., pays a broadcaster) more money, they can be authorized to decode, select, and render more object channels of a program.
[0212] Reference Figure 9 Examples of constraints and conditions for user selection of objects (or groups of objects). Figure 9 In the program "P0", there are 7 target channels: "N0" indicating neutral crowd noise; "N1" indicating home team crowd noise; "N2" indicating away team crowd noise; "N3" indicating official comments on the event (e.g., broadcast comments by commercial radio announcers); "N4" indicating fan comments on the event; "N5" indicating public announcements related to the event; and "N6" indicating Twitter links related to the event (converted via a text-to-speech system).
[0213] The default indicator metadata included in program P0 indicates that the default object group (one or more "default" objects) and the default rendering parameter group (e.g., the spatial location of each default object in the default object group) are (by default) included in the mix rendered by the program-indicated "soundbed" speaker channel content and object channel content. For example, the default object group could be a mix of object channel "N0" (indicating neutral crowd noise) rendered in a diffuse manner (e.g., not perceived as emanating from any particular source location) and object channel "N3" (indicating official commentary) rendered as if emanating from a source location directly in front of the listener (i.e., at an azimuth angle of 0 degrees relative to the listener).
[0214] ( Figure 9 The program P0 also includes metadata indicating multiple sets of user-selectable preset mixes, each preset mix determined by a subgroup of the program's object channels and a corresponding set of rendering parameters. User-selectable preset mixes can be presented as menus on the playback system's controller's user interface (e.g., via...). Figure 6 The system controller 23 displays a menu. For example, one such preset mix is... Figure 9 The mix of object channel "N0" (indicating neutral crowd noise), object channel "N1" (indicating main crowd noise), and object channel "N4" (indicating fan comments) is rendered such that the content of channels N0 and N1 in the mix is perceived as emanating from a source location directly behind the listener (i.e., at an azimuth angle of 180 degrees relative to the listener), wherein the level of the content of channel N1 in the mix is 3 dB lower than the level of the content of channel N0 in the mix, and wherein the content of channel N4 in the mix is rendered in a diffuse manner (e.g., so as not to be perceived as emanating from any particular source location).
[0215] The playback system can implement the following rules (e.g., in Figure 9 The grouping rule "G" (determined by the program's metadata) states that each user-selectable preset mix, including at least one of object channels N0, N1, and N2, must include content from object channel N0 only, or content from object channel N0 mixed with content from at least one of object channels N1 and N2. The playback system can also implement the following rules (e.g., in...). Figure 9 The condition rule “C1” indicated in the program (which is determined by the program’s metadata) states that each user-selectable preset mix that includes content of object channel N0 mixed with content of at least one of object channels N1 and N2 must include content of object channel N0 mixed with content of object channel N1, or it must include content of object channel N0 mixed with content of object channel N2.
[0216] The playback system can also implement the following rules (for example, in Figure 9The condition rule “C2” indicated in the program (which is determined by the program’s metadata) states that each user-selectable preset mix that includes content from at least one of object channels N3 and N4 must include content from object channel 3 only, or it must include content from object channel N4 only.
[0217] Some embodiments of the present invention implement conditional decoding (and / or rendering) of object channels in object-based audio programs. For example, a playback system can be configured such that object channels can be conditionally decoded based on the playback environment or the user's permissions. For instance, if a DRM hierarchy is implemented to give clients "tiered" access to groups of audio object channels included in an object-based audio program, the playback system can be automatically configured (via control bits included in the program's metadata) to prevent the selection of decoding and rendering of certain objects unless the playback system is notified to the user that at least one condition has been met (e.g., paying a specific amount of money to the content provider). For example, a user might need to purchase permissions to listen. Figure 9 The program P0's "official commentary" object channel N3, and the replay system can achieve this. Figure 9 The condition rule "C2" specified in the document prevents object channel N3 from being selected unless the playback system is notified that the playback system user has purchased the necessary permissions.
[0218] For another example, the playback system can be automatically configured (via control bits included in the program's metadata, indicating the specific format of the available playback speaker array): if the playback speaker array does not meet certain conditions (e.g., the playback system can implement...). Figure 9 The conditional rule "C1" indicates that the preset mixes for object channels N0 and N1 cannot be selected unless the playback system is notified that a 5.1 speaker array is available to render the selected content (but not if the only available speaker array is a 2.0 speaker array), preventing the decoding and selection of some objects.
[0219] In some embodiments, the invention implements rule-based object channel selection, wherein at least one predetermined rule determines which(s) of an object-based audio program is rendered (e.g., along with the speaker channel soundbed). The user can also specify at least one rule for object channel selection (e.g., by selecting from a menu of available rules presented by the user interface of the playback system controller), and the playback system (e.g., Figure 6 The system's object processing subsystem 22) can be configured to apply each such rule to determine which(s) of the object channels(s) of the object-based audio program to be rendered should be included in the rendering (e.g., via...). Figure 6In the mixing of subsystem 20 or subsystems 24 and 35 of the system. The playback system can determine which object(s) of the program meet the predetermined rules based on the object-related metadata in the program.
[0220] For a simple example, consider the following scenario: an object-based audio program indicating a sporting event. Instead of manipulating a controller (e.g., ...), Figure 6 The controller 23) performs static selection of specific object groups included in a program (e.g., radio commentary from a particular team or car or bicycle), and the user manipulates the controller to establish rules (e.g., automatically selecting object channels indicating whether a team or car or bicycle wins or is first for rendering). The playback system applies these rules to achieve dynamic selection of a series of different subgroups of objects (object channels) included in a program (e.g., indicating the first subgroup of objects for one team, and automatically following the second subgroup of objects for the second team when the second team scores and thus becomes the current winning team). Thus, in some such implementations, real-time event control or influence is applied to determine which object channels are included in the rendered mix. The playback system (e.g., Figure 6 The system's object processing subsystem 22 can respond to metadata included in the program (e.g., metadata indicating at least one corresponding object refers to the currently winning team, such as crowd noise indicating the team's fans or commentary from a radio announcer associated with the winning team) to select which(s) object channels should be included in the mix of speaker channels and object channels to be rendered. For example, content creators may include (in an object-based audio program) metadata indicating the placement order (or other hierarchy) of each of at least some audio object channels in the program (e.g., indicating which object channels correspond to the currently first-ranked team or car, which object channels correspond to the second-ranked team or car, etc.). The playback system can be configured to respond to such metadata by selecting and rendering only object channels that satisfy user-specified rules (e.g., object channels associated with the "n"th" ranked team, as indicated by the program's object-related metadata).
[0221] Examples of object-related metadata relating to object channels in the object-based audio program of the present invention include (but are not limited to): metadata indicating detailed information about how the object channel is rendered; dynamic temporary metadata (e.g., indicating the object's translation trajectory, object size, gain, etc.); and metadata for use by an AVR (or other devices or systems downstream of the decoding and object processing subsystem of some implementations of the system of the present invention) to render the object channel (e.g., using organizational knowledge of the available playback speaker array). Such metadata may specify constraints on object position, gain, mute, or other rendering parameters and / or constraints on how an object interacts with other objects (e.g., constraints on which other objects can be selected when a particular object is selected), and / or may specify default objects and / or default rendering parameters (to be used when the user does not select other objects and / or rendering parameters).
[0222] In some implementations, at least some object-related metadata (and optionally at least some object channels) of the object-based audio program of the present invention are sent in separate bitstreams or other containers (e.g., submixes that the user may need to pay extra to receive and / or use) from the speaker channel sound bed of the program. Without access to such object-related metadata (or object-related metadata and object channels), the user can decode and render the speaker channel sound bed, but cannot select the audio objects of the program and cannot render the audio objects of the program with a mix of audio indicated by the speaker channel sound bed. Each frame of the object-based audio program of the present invention may include audio content from multiple object channels and corresponding object-related metadata.
[0223] An object-based audio program generated (or transmitted, stored, buffered, decoded, rendered, or otherwise processed) according to some embodiments of the invention includes a speaker channel soundbed, at least one alternative speaker channel, at least one object channel, and metadata indicating a layered diagram (sometimes referred to as a layered "mix diagram"), which indicates optional mixes (e.g., all optional mixes) for the speaker channels and object channels. For example, the mix diagram indicates each rule that can be applied to the selection of subgroups of speaker channels and object channels. Typically, an encoded audio bitstream indicates at least some (i.e., at least a portion) of the program's audio content (e.g., at least some object channels among the speaker channel soundbed and the program's object channels) and object-related metadata (including metadata indicating the mix diagram), and optionally at least one additional encoded audio bitstream or file indicating some of the program's audio content and / or object-related metadata.
[0224] A layered mix diagram indicates nodes (each node may indicate an optional channel or channel group, or a classification of optional channels or channel groups) and connections between nodes (e.g., control interfaces for nodes and / or rules for selecting channels), and includes necessary data (“basic” layers) as well as optional (i.e., optionally omitted) data (at least one “extended” layer). Typically, a layered mix diagram is included in one audio bitstream of the encoded audio bitstream indicating the program, and can be evaluated (by a playback system, such as the end user's playback system) through graph traversal to determine the default mix for a channel and options for modifying the default mix.
[0225] In the case where the mix diagram can be represented as a tree diagram, the base layer can be one branch (or two or more branches) of the tree diagram, and each extension layer can be another branch (or another set of two or more branches) of the tree diagram. For example, one branch of the tree diagram (indicated by the base layer) can indicate the optional channels and channel groups available to all end users, and another branch of the tree diagram (indicated by the extension layer) can indicate additional optional channels and / or channel groups available only to some end users (e.g., such extension layers can be provided only to end users authorized to use them). Figure 9 This is an example of a tree graph that includes object channel nodes of the mix diagram (e.g., nodes indicating object channels N0, N1, N2, N3, N4, N5, and N6) and other elements.
[0226] Typically, the base layer contains the graph structure and control interfaces for the graph's nodes (e.g., translation and gain control interfaces). The base layer is necessary for mapping any user interactions to the decoding / rendering process.
[0227] Each extension layer contains (indicating) an extension to the base layer. Extensions are not immediately necessary for mapping user interactions to decoding processing, and therefore can be transmitted at a slower rate and / or with delays, or omitted altogether.
[0228] In some implementations, the base layer is included as metadata of a separate substream of the program (e.g., it is transmitted as metadata of a separate substream).
[0229] An object-based audio program generated (or transmitted, stored, buffered, decoded, rendered, or otherwise processed) according to some embodiments of the invention includes a speaker channel soundbed, at least one alternative speaker channel, at least one object channel, and metadata indicating a mix diagram (which may or may not be a layered mix diagram) indicating optional mixes (e.g., all optional mixes) for the speaker channel and object channel. An encoded audio bitstream (e.g., a Dolby E or E-AC-3 bitstream) indicates at least a portion of the program, and the metadata indicating the mix diagram (and typically also optional object channels and / or speaker channels) is included in each frame of the bitstream (or in each frame of a subgroup of frames of the bitstream). For example, each frame may include at least one metadata segment and at least one audio data segment, and the mix diagram may be included in at least one metadata segment in each frame. Each metadata segment (which may be referred to as a "container") may have a format including a metadata segment header (and optionally other elements) and one or more payloads following the metadata segment header. Each metadata payload is identified by a payload header itself. If the mix diagram exists in the metadata segment, the mix diagram is included in one of the metadata payloads of the metadata segment.
[0230] In some embodiments, an object-based audio program generated (or transmitted, stored, buffered, decoded, rendered, or otherwise processed) according to the present invention includes at least two speaker channel sound beds, at least one object channel, and metadata indicating a mix diagram (which may or may not be a hierarchical mix diagram). The mix diagram indicates optional mixing of the speaker channels and object channels and includes at least one “speaker bed mix” node. Each “speaker bed mix” node defines a predetermined mix for the speaker channel sound bed and thereby indicates or implements a predetermined set of mixing rules (optionally, with user-selectable parameters) for mixing the speaker channels of two or more speaker sound beds of the program.
[0231] Consider the following example: an audio program is associated with a football (soccer) match between Team A (home team) and Team B in a stadium, and includes a 5.1 speaker channel sound bed for the entire crowd in the stadium (determined by microphone feeds), a stereo feed biased towards the crowd of Team A (i.e., audio captured from spectators located in the section of the stadium mainly occupied by fans of Team A), and another stereo feed biased towards the crowd of Team B (i.e., audio captured from spectators located in the section of the stadium mainly occupied by fans of Team B). The three feeds (a 5.1-channel neutral bed, a 2.0-channel "Team A" bed, and a 2.0-channel "Team B" bed) can be mixed on the mixing console to generate four 5.1-speaker channel beds (which can be referred to as "fan zone" beds): unbiased, home-team biased (a mix of the neutral and Team A beds), guest-team biased (a mix of the neutral and Team B beds), and relative (the neutral bed mixed with the Team A bed shifted to one side of the space and the Team B bed shifted to the opposite side of the space). However, transmitting four 5.1-channel beds of mixes is expensive in terms of bit rate. Thus, the implementation of the bitstream of the present invention includes metadata specifying the bed mixing rules (for mixing speaker channel beds, e.g., a 5.1-channel bed to generate four specified mixes) to be implemented by a playback system (e.g., in the end user's home) based on the user's mixing selections, and speaker channel beds that can be mixed according to these rules (e.g., the original 5.1-channel bed and two biased stereo speaker channel beds). In response to the bed mixing nodes of the mix diagram, the playback system can present options to the user (e.g., via...) Figure 6 The system controller 23 implements a user interface (displayed) to select one of four specified 5.1-channel sound beds for the mix. In response to the user selection of the 5.1-channel sound bed for the mix, the playback system (e.g., Figure 6 Subsystem 22) of the system will use the (unmixed) speaker channel sound bed transmitted in the bitstream to generate the selected mix.
[0232] In some implementations, the sound bed mixing rule is conceived to operate as follows (which may have predetermined parameters or user-selectable parameters):
[0233] The sound bed is "rotated" (i.e., the speaker channel sound bed is shifted to the left, right, front, or rear). For example, to create the aforementioned "opposite" mix, the stereo team A sound bed will be rotated to the left of the playback speaker array (the L and R channels of team A's sound bed are mapped to the L and Ls channels of the playback system), and the stereo team B sound bed will be rotated to the right of the playback speaker array (the L and R channels of team B's sound bed are mapped to the R and Rs channels of the playback system). Thus, the playback system's user interface will present the end user with one of four options: "unbiased" sound bed mix, "home team biased" sound bed mix, "away team biased" sound bed mix, and "opposite" sound bed mix. When the user selects the "opposite" sound bed mix, the playback system will implement the appropriate sound bed rotation during the rendering of the "opposite" sound bed mix; and
[0234] A sudden drop (duck) (i.e., attenuation) (typically to create headroom) in a specific speaker channel (target channel) of a bed mix. For example, in the football example mentioned above, the playback system's user interface could present the end user with a choice of one of the four aforementioned "unbiased" bed mixes, "home team biased" bed mixes, "away team biased" bed mixes, and "relative" bed mixes. In response to the user selecting the "relative" bed mix, the playback system could achieve a target drop during the rendering of the "relative" bed mix by pre-setting each drop (attenuation) (specified by metadata in the bitstream) in the L, Ls, and R, Rs channels of the neutral 5.1 channel bed before mixing the attenuated 5.1 channel bed with the stereo "Team A" and "Team B" bed mixes to generate the "relative" bed mix.
[0235] In another embodiment, an object-based audio program generated (or transmitted, stored, buffered, decoded, rendered, or otherwise processed) according to the invention includes substreams, and each substream indicates at least one speaker channel soundbed, at least one object channel, and object-related metadata. The object-related metadata includes "substream" metadata (indicating the substream structure of the program and / or how the substreams should be decoded), and typically also includes a mix diagram indicating optional mixes (e.g., all optional mixes) for the speaker channel and object channel. The substream metadata can indicate which substreams of the program should be decoded independently of other substreams of the program and which substreams of the program should be decoded in association with at least one other substream of the program.
[0236] For example, in some implementations, the encoded audio bitstream indicates at least some (i.e., at least a portion) of the audio content of a program (e.g., at least one speaker channel bed, at least one alternative speaker channel, and at least some program object channels) and metadata (e.g., mix diagrams and substream metadata, and optionally other metadata), and at least one additional encoded audio bitstream (or file) indicates some of the program's audio content and / or metadata. Where each bitstream is a Dolby E bitstream (or encoded in a manner consistent with the SMPTE 337 format for carrying non-PCM data in AES3 serial digital audio bitstreams), the bitstreams may collectively indicate multiple audio contents up to eight channels, where each bitstream carries audio data for up to eight channels and typically also includes metadata. Each bitstream can be considered as a substream indicating a combined bitstream of all audio data and metadata carried by all the bitstreams.
[0237] In another example, in some implementations, the encoded audio bitstream indicates multiple metadata substreams (e.g., mix diagrams and substream metadata, and optionally other object-related metadata) and the audio content of at least one audio program. Typically, each substream indicates one or more channels of the program (and usually also metadata). In some cases, the multiple substreams of the encoded audio bitstream indicate the audio content of several audio programs, such as a “main” audio program (which may be a multi-channel program) and at least one other audio program (e.g., a program that provides commentary on the main audio program).
[0238] An encoded audio bitstream indicating at least one audio program must include at least one “independent” substream of audio content. An independent substream indicates at least one channel of the audio program (e.g., an independent substream could indicate all five full-range channels of a typical 5.1-channel audio program). In this document, this audio program is referred to as the “main” program.
[0239] In some cases, an encoded audio bitstream indicates two or more audio programs (a "main" program and at least one other audio program). In such cases, the bitstream comprises two or more independent substreams: a first independent substream indicating at least one channel of the main program; and at least one other independent substream indicating at least one channel of another audio program (a program different from the main program). Each independent bitstream can be encoded independently, and the decoder can operate to decode only a subgroup (not all) of the independent substreams of the encoded bitstream.
[0240] Optionally, the encoded audio bitstream indicating the main program (and optionally at least one other audio program) includes at least one “subordinate” substream of audio content. Each subordinate substream is associated with an independent substream of the bitstream and indicates at least one additional channel of the program (e.g., the main program) whose content is indicated by the associated independent substream (i.e., the subordinate substream indicates at least one channel of the program not indicated by the associated independent substream, while the associated independent substream indicates at least one channel of the program).
[0241] In an example of an coded bitstream that includes an independent substream (indicating at least one channel of the main program), the bitstream also includes subordinate substreams (associated with the independent substream) indicating one or more additional speaker channels of the main program. These additional speaker channels are supplementary to the main program channel indicated by the independent substream. For example, if the independent substream indicates the standard format left, right, center, left surround, and right surround full-range speaker channels of a 7.1 channel main program, then the subordinate substreams could indicate two other full-range speaker channels of the main program.
[0242] According to the E-AC-3 standard, a regular E-AC-3 bitstream must indicate at least one independent substream (e.g., a single AC-3 bitstream), and can indicate up to eight independent substreams. Each independent substream of an E-AC-3 bitstream can be associated with up to eight subordinate substreams.
[0243] In (refer to) Figure 11 In the exemplary embodiment described, an object-based audio program includes at least one speaker channel soundbed, at least one object channel, and metadata. The metadata includes "substream" metadata (indicating the substream structure of the program's audio content and / or how the substreams of the program's audio content should be decoded), and typically also includes a mix diagram indicating optional mixing for the speaker channel and object channel. The audio program is associated with an English football match. Encoded audio bitstreams (e.g., E-AC-3 bitstreams) indicate the program's audio content and metadata. Figure 11 As shown, the program's audio content (and thus the bitstream audio content) comprises four independent substreams. One independent substream (in...) Figure 11 The sub-stream labeled "I0" indicates the 5.1 speaker channel sound bed, which indicates neutral crowd noise at a football match. Another independent sub-stream (in...) Figure 11The sub-stream labeled "I1" indicates the 2.0 channel "Team A" soundbed ("M Crowd"), the 2.0 channel "Team B" soundbed ("LivP Crowd"), and the mono object channel ("Sky Commentary 1"). The 2.0 channel "Team A" soundbed ("M Crowd") indicates the portion of the sound from the match crowd biased towards one team ("Team A"), the 2.0 channel "Team B" soundbed ("LivP Crowd") indicates the portion of the sound from the match crowd biased towards the other team ("Team B"), and the mono object channel ("Sky Commentary 1") indicates the commentary on the match. A third independent sub-stream (in...) Figure 11 The first object channel (labeled "I2") indicates the audio content of the target channel (labeled "2 / 0" goal) and three target channels ("Sky Commentary 2", "Man Commentary", and "Liv Commentary"), the audio content of which (labeled "2 / 0" goal) indicates the sound produced by the winning goal when the ball is hit by a participant in a football event, and each target channel ("Sky Commentary 2", "Man Commentary", "Liv Commentary") indicates a different commentary on a football match. The fourth independent substream (in...) Figure 11 The sub-stream “I3” indicates the object channel (“PA”), the object channel (“radio”), and the object channel (“goal report”). The object channel (“PA”) indicates the sound generated by the stadium public address system at the football match venue, the object channel (“radio”) indicates the radio broadcast of the football match, and the object channel (“goal report”) indicates the scoring during the football match.
[0244] exist Figure 11 In the example, substream I0 includes the program's mix diagram and metadata ("obj md"), which includes at least some substream metadata and at least some object channel-related metadata. Each of substreams I1, I2, and I3 includes metadata ("obj md"), which includes at least some object channel-related metadata and optionally at least some substream metadata.
[0245] exist Figure 11In the example, the substream metadata of the bitstream indicates that coupling between each pair of independent substreams should be "off" during decoding (so that each independent substream is decoded independently of the other independent substreams), and the substream metadata of the bitstream indicates that coupling within each substream should be either "on" (so that these channels are not decoded independently of each other) or "off" (so that these channels are decoded independently of each other). For example, the substream metadata indicates that each internal coupling in the two stereo speaker channel sound beds (2.0 channel "Team A" sound bed and 2.0 channel "Team B" sound bed) of substream I1 should be "on," but coupling across the speaker channel sound bed of substream I1 and between each sound bed in the mono object channel and the speaker channel sound bed of substream I1 should be disabled (so that the mono object channel and the speaker channel sound bed are decoded independently of each other). Similarly, the substream metadata indicates that the internal coupling in the 5.1 speaker channel sound bed of substream I0 should be "on" (so that the speaker channels of that sound bed are decoded correlatedly with each other).
[0246] In some implementations, speaker channels and object channels are included (“encapsulated”) in substreams of the audio program in a manner suitable for a mix diagram of the program. For example, if the mix diagram is a tree diagram, all channels of one branch of the diagram can be included in one substream, and all channels of another branch of the diagram can be included in another substream.
[0247] Figure 10 This is a block diagram of a system for implementing embodiments of the present invention.
[0248] Figure 10 The system's object processing system (object processor) 200 includes a metadata generation subsystem 210, an intermediate (mezzanine) encoder 212, and a simulation subsystem 211, coupled as shown. The metadata generation subsystem 210 is coupled to receive captured audio streams (e.g., streams indicating sound captured by microphones located at the viewing event, and optionally other audio streams) and is configured to organize (e.g., group) the audio streams from console 104 into speaker channel sound beds, replacement speaker channel groups, and a plurality of object channels and / or object channel groups. Subsystem 210 is also configured to generate object-related metadata indicating object channels (and / or object channel groups). Encoder 212 is configured to encode the object channels (and / or object channel groups), object-related metadata, and speaker channels into an intermediate type of object-based audio program (e.g., an object-based audio program encoded as a Dolby E bitstream).
[0249] The simulation subsystem 211 of the object processor 200 is configured to render (and play on a set of studio monitor speakers) at least a selected subgroup of object channels (and / or object channel groups) and speaker channels (including generating a mix that indicates the selected object channels and speaker channels by using object-related metadata), so that the played-back sound can be monitored by the operator of the subsystem 200.
[0250] Figure 10 The system's transcoder 202 includes an intermediate decoder subsystem (intermediate decoder) 213 and an encoder 214 coupled as shown. The intermediate decoder 213 is coupled and configured to receive and decode intermediate-type object-based audio programs output from the object processor 200. The decoded output of decoder 213 is re-encoded by encoder 214 into a broadcast-suitable format. In one embodiment, the encoded object-based audio program output from encoder 214 is an E-AC-3 bitstream (and thus encoder 214...). Figure 10 (This is labeled "DD+ encoder"). In other embodiments, the encoded object-based audio program output from encoder 214 is an AC-3 bitstream or has some other format. The object-based audio program output from transcoder 202 is broadcast (or otherwise delivered) to a large number of end users.
[0251] Decoder 204 is included in such an end-user playback system. Decoder 204 includes a decoder 215 and a rendering subsystem (renderer) 216 coupled as shown. Decoder 215 accepts (receives or reads) and decodes object-based audio programs delivered from transcoder 202. If decoder 215 is configured according to a typical embodiment of the invention, in typical operation, the output of decoder 215 includes: an audio sample stream indicating the speaker channel sound bed of the program; and an audio sample stream indicating the object channel of the program (e.g., a user-selectable audio object channel) and the corresponding object-based metadata stream. In one embodiment, the encoded object-based audio program input to decoder 215 is an E-AC-3 bitstream, and thus decoder 215... Figure 10 It is marked as "DD+ decoder".
[0252] The renderer 216 of the decoder 204 includes an object processing subsystem coupled to receive (from the decoder 215) decoded speaker channels, object channels, and object-related metadata of the delivered program. The renderer 216 also includes a rendering subsystem configured to render audio content determined by the object processing subsystem for playback by speakers (not shown) of the playback system.
[0253] Typically, the object processing subsystem of renderer 216 is configured to output the selected subgroup of the full group of object channels indicated by the program, along with the corresponding object-related metadata, to the rendering subsystem of renderer 216. The object processing subsystem of renderer 216 is also typically configured to allow the decoded speaker channels from decoder 215 to pass unchanged (to the rendering subsystem). According to embodiments of the invention, the object channel selection performed by the object processing subsystem is determined, for example, by rules (e.g., indication conditions and / or constraints) and / or user selections that renderer 216 has been programmed to or otherwise configured to implement.
[0254] Figure 10 Each of the components 200, 202, and 204 (and) Figure 8 Each of the components 104, 106, 108, and 110 can be implemented as a hardware system. The inputs of such a hardware implementation of processor 200 (or processor 106) are typically multi-channel audio digital interface (“MADI”) inputs. Figure 8 processor 106 and Figure 10 Each encoder 212, 214 includes a frame buffer. Typically, the frame buffer is a buffer memory coupled to receive an encoded input audio bitstream, and in operation, the buffer memory (e.g., in a non-transient manner) stores at least one frame of the encoded audio bitstream, and a series of frames of the encoded audio bitstream are set from the buffer memory to downstream devices or systems. Furthermore, typically... Figure 10 Each decoder 213, 215 includes a frame buffer. Typically, the frame buffer is a buffer memory coupled to receive an encoded input audio bitstream, and in operation, the buffer memory stores at least one frame of the encoded audio bitstream to be decoded by decoder 213 or 215 (e.g., in a non-transient manner).
[0255] Figure 8 Processor 106 (or Figure 10 Any component or element of the subsystems 200, 202 and / or 204 may be implemented in hardware, software or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA or other integrated circuits).
[0256] It should be understood that, in some embodiments, the object-based audio program of the present invention is generated and / or delivered as an unencoded (e.g., baseband) representation indicating program content (including metadata). For example, such a representation may include PCM audio samples and associated metadata. The unencoded (uncompressed) representation may be delivered in any of a variety of ways, including as at least one data file (e.g., stored in a non-transitory manner in memory, such as on a computer-readable medium) or as a bitstream in AES-3 format or Serial Digital Interface (SDI) format (or another format).
[0257] One aspect of the present invention is an audio processing unit (APU) configured to perform any embodiment of the methods of the present invention. Examples of APUs include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, preprocessing systems (preprocessors), postprocessing systems (postprocessors), audio bitstream processing systems, and combinations of these elements.
[0258] In one embodiment, the invention is an APU including a buffer memory (buffer) that stores (e.g., in a non-transient manner) at least one frame or other segment (including audio content of speaker channel soundbeds and object channels, and object-related metadata) of an object-based audio program generated by any embodiment of the method of the invention. For example, Figure 5 The generation unit 3 may include a buffer 3A, which (e.g., in a non-transient manner) stores at least one frame or other segment of the object-based audio program generated by unit 3 (including audio content of the speaker channel sound bed and the object channel, as well as object-related metadata). For another example, Figure 5 The decoder 7 may include a buffer 7A, which (e.g., in a non-transient manner) stores at least one frame or other segment of an object-based audio program delivered from the subsystem 5 to the decoder 7 (including audio content of the speaker channel sound bed and the object channel, as well as object-related metadata).
[0259] Embodiments of the present invention can be implemented in hardware, firmware, or software, or a combination thereof (e.g., as a programmable logic array). For example, hardware or firmware that is appropriately programmed (or otherwise configured) can... Figure 8 or Figure 7 Subsystem 106 of the system, or Figure 6 All or some of the system's components 20, 22, 24, 25, 26, 29, 35, 31, and 35, or Figure 10All or some of the elements 200, 202, and 204 are implemented as, for example, programmable general-purpose processors, digital signal processors, or microprocessors. Unless otherwise specified, the algorithms or processes included as part of this invention are not inherently related to any particular computer or other device. Specifically, various general-purpose machines can be used with programs written according to the teachings herein, or they may be more readily adapted to construct more specialized devices (e.g., integrated circuits) to perform the desired method steps. Thus, the invention can be implemented in one or more programmable computer systems (e.g., Figure 6 The programmable computer system is implemented by executing one or more computer programs on all or some of the elements 20, 22, 24, 25, 26, 29, 35, 31, and 35. Each programmable computer system includes at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.
[0260] Each such program can be implemented in any desired computer language (including machine, assembly, or high-level procedural, logical, or object-oriented programming languages) to communicate with the computer system. In any case, the language can be a compiled language or an interpreted language.
[0261] For example, when implemented by a sequence of computer software instructions, the various functions and steps of embodiments of the present invention can be implemented by a sequence of multi-threaded software instructions running in appropriate digital signal processing hardware. In this case, the various means, steps and functions of the embodiments can correspond to portions of the software instructions.
[0262] Each such computer program is preferably stored in or downloaded to a storage medium or device readable by a general-purpose or special-purpose programmable computer (e.g., solid-state memory or medium, or magnetic or optical medium) for configuring and operating the computer when the storage medium or device is read by a computer system to perform the processes described herein. The system of the present invention can also be implemented as a computer-readable storage medium configured (i.e., storing) a computer program, wherein such a storage medium causes the computer system to operate in a specific and predefined manner to perform the functions described herein.
[0263] Several embodiments of the invention have been described. It should be understood that various modifications can be made without departing from the spirit and scope of the invention. Numerous modifications and variations can be made to the invention in light of the foregoing teachings. It should be understood that the invention can be practiced differently than those specifically described herein, within the scope of the appended claims.
[0264] The present invention includes the following technical solutions.
[0265] Solution 1. A method for generating an object-based audio program that indicates audio content, the audio content including first non-contextual content, second non-contextual content different from the first non-contextual content, and third content different from the first non-contextual content and the second non-contextual content, the method comprising the steps of:
[0266] Determine an object channel group comprising N object channels, wherein a first subgroup of the object channel group indicates the first non-environmental content, the first subgroup comprising M object channels in the object channel group, each of N and M being a positive integer, and M being equal to or less than N;
[0267] Determine a speaker channel sound bed that indicates the default mix of audio content, including an object-based speaker channel subgroup of M speaker channels in the sound bed that indicates the second non-ambient content, or a mix of at least some audio content of the default mix with the second non-ambient content;
[0268] A set of M alternative speaker channels is determined, wherein each alternative speaker channel in the set of M alternative speaker channels indicates some, but not all, of the contents of the corresponding speaker channel in the object-based speaker channel subgroup;
[0269] Generate metadata indicating at least one of the contents of the object channel and at least one of the contents of the speaker channel and / or the contents of a predetermined speaker channel in the replacement speaker channel of the sound bed, wherein the metadata includes rendering parameters for each of the replacement mixes, and at least one of the replacement mixes is a replacement mix indicating at least some audio content of the sound bed and the first non-ambient content instead of the second non-ambient content; and
[0270] Generate the object-based audio program comprising the speaker channel sound bed, the set of M alternative speaker channels, the object channel group, and the metadata, such that the speaker channel sound bed is renderable without using the metadata to provide sound that is perceived as the default mix, and the alternative mix is renderable in response to at least some of the metadata to provide sound that is perceived as a mix comprising the at least some audio content of the sound bed and the first non-ambient content instead of the second non-ambient content.
[0271] Option 2. According to the method of Option 1, wherein at least some of the metadata is optional content metadata, the optional content metadata indicating a set of optional predetermined mixes of the audio content of the program and including a predetermined set of rendering parameters for each predetermined mix.
[0272] Option 3. The method according to any one of Options 1 to 2, wherein the object-based audio program is an encoded bitstream comprising frames, the encoded bitstream being an AC-3 bitstream or an E-AC-3 bitstream, each frame of the encoded bitstream indicating at least one data structure, the at least one data structure being a container comprising some content of the object channel and some of the metadata, and at least one of the containers being included in an auxiliary data auxdata field or an additional bitstream information addbsi field of each frame.
[0273] Option 4. The method according to any one of Options 1 to 2, wherein the object-based audio program is a Dolby E bitstream comprising a series of bursts and guard bands between burst pairs.
[0274] Option 5. The method according to any one of Options 1 to 2, wherein the object-based audio program is an unencoded representation of the audio content and metadata of the program, and the unencoded representation is a bitstream or at least one data file stored in memory in a non-transitory manner.
[0275] Option 6. The method according to any one of Options 1 to 5, wherein at least some of the metadata indicates a layered mixing diagram, the layered mixing diagram indicating optional mixing of the speaker channels, the alternative speaker channels and the object channels of the sound bed, and the layered mixing diagram includes a base layer of metadata and at least one extended layer of metadata.
[0276] Option 7. The method according to any one of Options 1 to 6, wherein at least some of the metadata indicates a mixing diagram, the mixing diagram indicating optional mixing of the speaker channels, the alternative speaker channels and the object channels of the sound bed, the object-based audio program being an encoded bitstream comprising frames, and each frame of the encoded bitstream including metadata indicating the mixing diagram.
[0277] Option 8. The method according to any one of Options 1 to 7, wherein the object-based audio program indicates the captured audio content.
[0278] Option 9. The method according to any one of Options 1 to 8, wherein the default mix is a mix of ambient content and non-ambient content.
[0279] Option 10. The method according to any one of Options 1 to 9, wherein the third content is environmental content.
[0280] Option 11. The method according to Option 10, wherein the environmental content indicates ambient sound during the viewing event, the first non-environmental content indicates commentary on the viewing event, and the second non-environmental content indicates alternative commentary on the viewing event.
[0281] Solution 12. A method for rendering audio content determined from an object-based audio program, wherein the program indicates a speaker channel soundbed, a set of M alternative speaker channels, an object channel group, and metadata, wherein the object channel group includes N object channels, a first subgroup of the object channel group indicates first non-ambient content, the first subgroup including M object channels in the object channel group, where each of N and M is a positive integer, and M is equal to or less than N.
[0282] The speaker channel sound bed indicates a default mix of audio content including a second non-ambient content that is different from the first non-ambient content. This includes an object-based speaker channel subgroup of M speaker channels in the sound bed indicating a mix of the second non-ambient content, or at least some audio content of the default mix, with the second non-ambient content.
[0283] Each of the M alternative speaker channels in the set indicates some, but not all, of the contents of the corresponding speaker channel in the object-based speaker channel subgroup, and
[0284] The metadata indicates at least one of the contents of the object channel and at least one of the contents of the speaker channel of the sound bed and / or the contents of a predetermined speaker channel in the replacement speaker channel, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is a replacement mix that includes at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content, the method comprising the steps of:
[0285] (a) providing the object-based audio program to the audio processing unit; and
[0286] (b) In the audio processing unit, the speaker channel sound bed is parsed, and the default mix is rendered in response to the speaker channel sound bed without using the metadata.
[0287] Solution 13. The method according to Solution 12, wherein the audio processing unit is configured to parse the object channel and the metadata of the program, the method further comprising the step of:
[0288] (c) In the audio processing unit, the replacement mix is rendered using at least some of the metadata, including rendering by selecting and mixing the contents of the first subgroup of the object channel group and at least one of the replacement speaker channels in response to at least some of the metadata.
[0289] Option 14. The method according to Option 13, wherein step (c) includes the step:
[0290] (d) In response to the at least some metadata, select the first subgroup of the object channel group, select at least one speaker channel in the speaker channel bed other than the speaker channels in the object-based speaker channel subgroup, and select the at least one alternative speaker channel; and
[0291] (e) Mix the contents of the first subgroup and each speaker channel of the object channel group selected in step (d) to determine the replacement mix.
[0292] Option 15. The method according to any one of Options 13 to 14, wherein step (c) includes the step of: driving a loudspeaker to provide a sound that can be perceived as a mixture of the at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content.
[0293] Option 16. The method according to any one of Options 13 to 15, wherein step (c) comprises the step:
[0294] In response to the replacement mix, a speaker feed is generated to drive a speaker to emit sound, wherein the sound includes an object channel sound indicating the first non-ambient content, and the object channel sound can be perceived as emanating from at least one apparent source location determined by the first subgroup of the object channel group.
[0295] Option 17. The method according to any one of Options 13 to 16, wherein step (c) comprises the step:
[0296] A menu is provided for selecting mixes, each mix in at least one subgroup comprising the contents of the object channel subgroup and the replacement speaker channel subgroup; and
[0297] The replacement mix is selected by choosing one of the mixes indicated by the menu.
[0298] Option 18. The method according to any one of Options 13 to 17, wherein the menu is presented through a user interface of a controller, the controller being coupled to a set-top device, and the set-top device being coupled to receive the object-based audio program and configured to perform step (c).
[0299] Scheme 19. The method according to any one of Schemes 12 to 18, wherein the object-based audio program comprises a set of bitstreams, wherein step (a) comprises the step of: sending the bitstreams of the object-based audio program to the audio processing unit.
[0300] Option 20. The method according to any one of Options 12 to 19, wherein the default mix is a mix of ambient content and non-ambient content.
[0301] Option 21. The method according to Option 20, wherein the environmental content indicates ambient sound during the viewing event, the first non-environmental content indicates commentary on the viewing event, and the second non-environmental content indicates alternative commentary on the viewing event.
[0302] Option 22. The method according to any one of Options 12 to 21, wherein the object-based audio program is an encoded bitstream comprising frames, the encoded bitstream being an AC-3 bitstream or an E-AC-3 bitstream, each frame of the encoded bitstream indicating at least one data structure, the at least one data structure being a container comprising some content of the object channel and some of the metadata, and at least one of the containers being included in an auxiliary data auxdata field or an additional bitstream information addbsi field of each frame.
[0303] Option 23. The method according to any one of Options 12 to 21, wherein the object-based audio program is a Dolby E bitstream comprising a series of bursts and guard bands between burst pairs.
[0304] Option 24. The method according to any one of Options 12 to 21, wherein the object-based audio program is an unencoded representation of the audio content and metadata of the program, and the unencoded representation is a bitstream or at least one data file stored in memory in a non-transitory manner.
[0305] Option 25. The method according to any one of Options 12 to 24, wherein at least some of the metadata indicates a layered mix diagram indicating optional mixing of the speaker channels, the alternative speaker channels, and the object channels of the sound bed, and the layered mix diagram includes a base layer of metadata and at least one extended layer of metadata.
[0306] Option 26. The method according to any one of Options 12 to 25, wherein at least some of the metadata indicates a mixing diagram indicating optional mixing of the speaker channels, the alternative speaker channels, and the object channels of the sound bed, the object-based audio program being an encoded bitstream comprising frames, and each frame of the encoded bitstream including metadata indicating the mixing diagram.
[0307] Solution 27. A system for generating object-based audio programs that indicate audio content, said audio content including first non-contextual content, second non-contextual content different from the first non-contextual content, and third content different from the first non-contextual content and the second non-contextual content, said system comprising:
[0308] The first subsystem is configured to determine:
[0309] An object channel group comprising N object channels, wherein a first subgroup of the object channel group indicates the first non-environmental content, the first subgroup comprising M object channels from the object channel group, where each of N and M is a positive integer, and M is equal to or less than N.
[0310] A speaker channel sound bed that indicates the default mix of audio content, including an object-based speaker channel subgroup of M speaker channels in the sound bed indicating the second non-ambient content, or a mix of at least some audio content of the default mix with the second non-ambient content, and
[0311] A set of M alternative speaker channels, wherein each alternative speaker channel in the set of M alternative speaker channels indicates some, but not all, of the contents of a corresponding speaker channel in the object-based speaker channel subgroup.
[0312] The first subsystem is further configured to generate metadata indicating at least one of the contents of the object channel and at least one of the contents of the speaker channel and / or the contents of a predetermined speaker channel in the replacement speaker channel of the sound bed, wherein the metadata includes rendering parameters for each of the replacement mixes, and at least one of the replacement mixes is a replacement mix indicating at least some audio content of the sound bed and the first non-ambient content instead of the second non-ambient content; and
[0313] An encoding subsystem, coupled to the first subsystem, is configured to generate the object-based audio program, such that the object-based audio program includes the speaker channel sound bed, the set of M alternative speaker channels, the object channel group, and the metadata, and such that the speaker channel sound bed is renderable without using the metadata to provide sound that is perceived as the default mix, and the alternative mix is renderable in response to at least some of the metadata to provide sound that is perceived as a mix including the at least some audio content of the sound bed and the first non-ambient content instead of the second non-ambient content.
[0314] Option 28. The system according to Option 27, wherein at least some of the metadata is optional content metadata, the optional content metadata indicating a set of optional predetermined mixes of the audio content of the program and including a predetermined set of rendering parameters for each predetermined mix.
[0315] Option 29. The system according to any one of Options 27 to 28, wherein the default mix is a mix of ambient content and non-ambient content.
[0316] Option 30. The system according to any one of Options 27 to 28, wherein the third content is environmental content.
[0317] Solution 31. The system according to Solution 30, wherein the environmental content indicates ambient sound during the viewing event, the first non-environmental content indicates commentary on the viewing event, and the second non-environmental content indicates alternative commentary on the viewing event.
[0318] Option 32. The system according to any one of Options 27 to 31, wherein the encoding subsystem is configured to generate the object-based audio program such that the object-based audio program is an encoded bitstream comprising frames, the encoded bitstream being an AC-3 bitstream or an E-AC-3 bitstream, each frame of the encoded bitstream indicating at least one data structure, the at least one data structure being a container comprising some content of the object channel and some of the metadata, and at least one of the containers being included in an auxiliary data auxdata field or an additional bitstream information addbsi field of each frame.
[0319] Option 33. The system according to any one of Options 27 to 32, wherein the encoding subsystem is configured to generate the object-based audio program such that the object-based audio program is a Dolby E bitstream comprising a series of bursts and guard bands between burst pairs.
[0320] Option 34. The system according to any one of Options 27 to 33, wherein at least some of the metadata indicates a layered mixing diagram indicating optional mixing of the speaker channels, the alternative speaker channels, and the object channels of the sound bed, and the layered mixing diagram includes a base layer of metadata and at least one extended layer of metadata.
[0321] Option 35. The system according to any one of Options 27 to 34, wherein at least some of the metadata indicates a mixing diagram indicating optional mixing of the speaker channels, the alternative speaker channels, and the object channels of the sound bed, the object-based audio program being an encoded bitstream comprising frames, and each frame of the encoded bitstream including metadata indicating the mixing diagram.
[0322] Solution 36. An audio processing unit configured to render audio content determined by an object-based audio program, wherein the program indicates a speaker channel sound bed, a set of M alternative speaker channels, an object channel group, and metadata, wherein the object channel group includes N object channels, a first subgroup of the object channel group indicates first non-contextual content, the first subgroup including M object channels in the object channel group, where each of N and M is a positive integer, and M is equal to or less than N.
[0323] The speaker channel sound bed indicates a default mix of audio content including a second non-ambient content that is different from the first non-ambient content. This includes an object-based speaker channel subgroup of M speaker channels in the sound bed indicating a mix of the second non-ambient content, or at least some audio content of the default mix, with the second non-ambient content.
[0324] Each of the M alternative speaker channels in the set indicates some, but not all, of the contents of the corresponding speaker channel in the object-based speaker channel subgroup, and
[0325] The metadata indicates at least one of the contents of the object channel and at least one of the contents of the speaker channel and / or the predetermined speaker channel in the replacement speaker channel of the sound bed, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is a replacement mix that includes at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content, the audio processing unit comprising:
[0326] A first subsystem, coupled to receive the object-based audio program, and configured to parse the speaker channel soundbed, the alternative speaker channel, the object channel, and the metadata of the program; and
[0327] A rendering subsystem, coupled to the first subsystem, is capable of operating in a first mode to render the default mix in response to the speaker channel sound bed without using the metadata. The rendering subsystem is also capable of operating in a second mode to render the alternative mix using at least some of the metadata, including rendering by selecting and mixing the contents of the first subgroup of the object channel group and at least one of the alternative speaker channels in response to at least some of the metadata.
[0328] Solution 37. The audio processing unit according to Solution 36, wherein the rendering subsystem includes:
[0329] A first subsystem, capable of operating in the second mode, thereby selecting, in response to the at least some metadata, a first subgroup of the object channel group, at least one speaker channel in the speaker channel bed other than the speaker channels in the object-based speaker channel subgroup, and the at least one alternative speaker channel; and
[0330] A second subsystem, coupled to the first subsystem and capable of operating in the second mode, mixes the contents of the first subgroup of the object channel group selected by the first subsystem with the contents of each speaker channel, thereby determining the replacement mix.
[0331] Option 38. An audio processing unit according to any one of Options 36 to 37, wherein the rendering subsystem is configured to generate a speaker feed for driving a speaker to emit sound in response to the replacement mix, the sound being perceptible as a mix including the at least some audio content of the sound bed and the first non-ambient content but not the second non-ambient content.
[0332] Solution 39. An audio processing unit according to any one of Solutions 36 to 38, wherein the rendering subsystem is configured to generate a speaker feed for driving a speaker to emit sound in response to the replacement mix, wherein the sound includes an object channel sound indicating the first non-ambient content, the object channel sound being perceptible as emanating from at least one apparent source location determined by the first subgroup of the object channel group.
[0333] Option 40. The audio processing unit according to any one of Options 36 to 39 further includes a controller coupled to the rendering subsystem, wherein the controller is configured to provide a menu of mixes that can be selected, each mix in at least one subgroup of the mixes including the contents of the subgroup of the object channel and the subgroup of the alternative speaker channel.
[0334] Option 41. The audio processing unit according to Option 40, wherein the controller is configured to implement a user interface for displaying the menu.
[0335] Option 42. The audio processing unit according to any one of Options 40 to 41, wherein the first subsystem and the rendering subsystem are implemented in a set-top device, and the controller is coupled to the set-top device.
[0336] Solution 43. The audio processing unit according to any one of Solutions 36 to 42, wherein the default mix is a mix of ambient content and non-ambient content.
[0337] Solution 44. The audio processing unit according to Solution 43, wherein the ambient content indicates ambient sound during the viewing event, the first non-ambient content indicates a comment on the viewing event, and the second non-ambient content indicates alternative comments on the viewing event.
[0338] Option 45. The audio processing unit according to any one of Options 36 to 44, wherein the object-based audio program is an encoded bitstream comprising frames, the encoded bitstream being an AC-3 bitstream or an E-AC-3 bitstream, each frame of the encoded bitstream indicating at least one data structure, the at least one data structure being a container comprising some content of the object channel and some of the metadata, and at least one of the containers being included in an auxiliary data auxdata field or an additional bitstream information addbsi field of each frame.
[0339] Scheme 46. The audio processing unit according to any one of Schemes 36 to 44, wherein the object-based audio program is a Dolby E bitstream comprising a series of bursts and guard bands between burst pairs.
[0340] Solution 47. An audio processing unit, comprising:
[0341] Buffer memory; and
[0342] At least one audio processing subsystem is coupled to the buffer memory, wherein the buffer memory stores at least one segment of an object-based audio program, wherein the program indicates a speaker channel soundbed, a set of M alternative speaker channels, an object channel group, and metadata, wherein the object channel group comprises N object channels, and a first subgroup of the object channel group indicates a first non-ambient content, the first subgroup comprising M object channels in the object channel group, where each of N and M is a positive integer, and M is equal to or less than N.
[0343] The speaker channel sound bed indicates a default mix of audio content including a second non-ambient content that is different from the first non-ambient content. This includes an object-based speaker channel subgroup of M speaker channels in the sound bed indicating a mix of the second non-ambient content, or at least some audio content of the default mix, with the second non-ambient content.
[0344] Each of the M alternative speaker channels in the set indicates some, but not all, of the contents of the corresponding speaker channel in the object-based speaker channel subgroup, and
[0345] The metadata indicates at least one of the contents of the object channel and at least one of the contents of the speaker channel and / or the predetermined speaker channel in the replacement speaker channel of the sound bed, wherein the metadata includes rendering parameters for each of the alternative mixes, and at least one of the alternative mixes is a replacement mix that includes at least some audio content from the sound bed and the first non-ambient content but not the second non-ambient content.
[0346] Furthermore, each of the segments includes: data indicating the audio content of the speaker channel sound bed, data indicating the audio content of the replacement speaker channel, data indicating the audio content of the object channel, and at least a portion of the metadata.
[0347] Option 48. The audio processing unit according to Option 47, wherein the object-based audio program is an encoded bitstream comprising frames, and each segment is one of the frames.
[0348] Scheme 49. The audio processing unit according to any one of Schemes 47 to 48, wherein the encoded bitstream is an AC-3 bitstream or an E-AC-3 bitstream, each frame indicates at least one data structure, the at least one data structure being a container including some content of at least one of the object channels and some of the metadata, and at least one of the containers being included in the auxiliary data auxdata field or the additional bitstream information addbsi field of each frame.
[0349] Option 50. The audio processing unit according to any one of Options 47 to 48, wherein the object-based audio program is a Dolby E bitstream comprising a series of bursts and guard bands between burst pairs.
[0350] Option 51. The audio processing unit according to any one of Options 47 to 48, wherein the object-based audio program is an uncoded representation of the audio content and metadata indicating the program, and the uncoded representation is a bitstream or at least one data file stored in memory in a non-transitory manner.
[0351] Option 52. The audio processing unit according to any one of Options 47 to 51, wherein the buffer memory stores the segments in a non-transitory manner.
[0352] Option 53. An audio processing unit according to any one of Options 47 to 52, wherein the audio processing subsystem is an encoder.
[0353] Option 54. An audio processing unit according to any one of Options 47 to 53, wherein the audio processing subsystem is configured to: parse the speaker channel sound bed, the replacement speaker channel, the object channel, and the metadata.
[0354] Option 55. An audio processing unit according to any one of Options 47 to 54, wherein the audio processing subsystem is configured to render the default mix in response to the speaker channel sound bed without using the metadata.
[0355] Option 56. An audio processing unit according to any one of Options 47 to 55, wherein the audio processing subsystem is configured to render the replacement mix using at least some of the metadata, including rendering the render by selecting and mixing the contents of the first subgroup of the object channel group and at least one of the replacement speaker channels in response to at least some of the metadata.
Claims
1. A method for rendering the audio content of an audio program, wherein, The audio program includes at least one object channel and metadata, wherein the metadata indicates at least one optional predetermined content mix including the at least one object channel, and wherein the metadata includes rendering parameters for each optional predetermined content mix. The method includes the following steps: (a) Receive the at least one object channel of the audio program and the metadata; (b) Providing the audio program to a controller with a set of optional predetermined audio content mixes including the at least one optional predetermined content mix, wherein the controller is configured to provide an interface relating to the selectable mixes, wherein the metadata includes data for determining the set of optional predetermined audio content mixes; (c) Receiving from the controller a selection of the set of optional predetermined audio content mixes, wherein the selection indicates a selected subgroup of the set of optional predetermined audio content mixes of the audio program; and (d) Rendering the at least one object channel based on at least some metadata in the metadata indicating the selected subgroup of the set of optional predetermined audio content mixing for the audio program, wherein the rendering includes selecting and mixing the content of the at least one object channel in response to the at least some metadata in the metadata indicating the selected subgroup of the set of optional predetermined audio content mixing for the audio program.
2. An audio processing unit configured to render the audio content of an audio program, wherein, The audio program includes at least one object channel and metadata, wherein the metadata indicates at least one optional predetermined content mix including the at least one object channel, and the metadata includes rendering parameters for each optional predetermined content mix. The audio processing unit includes: A first subsystem is configured to receive the at least one object channel; A second subsystem, coupled to the first subsystem, is configured to provide the controller with a set of optional predetermined audio content mixes, including the at least one optional predetermined content mix, of the audio program, wherein the controller is configured to provide an interface relating to the selectable mixes, wherein the metadata includes data for determining the set of optional predetermined audio content mixes. A third subsystem, coupled to the second subsystem, is configured to receive from the controller a selection of the set of optional predetermined audio content mixes, wherein the selection indicates a selected subgroup of the set of optional predetermined audio content mixes for the audio program; and A rendering subsystem, coupled to the first subsystem and the third subsystem, is configured to render the at least one object channel based on at least some metadata in the metadata indicating the selected subgroup of the set of optional predetermined audio content mixing for the audio program, wherein the rendering includes selecting and mixing the content of the at least one object channel in response to the at least some metadata in the metadata indicating the selected subgroup of the set of optional predetermined audio content mixing for the audio program.
3. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, perform the method according to claim 1.
Citation Information
Patent Citations
Methods and systems for generating and rendering object based audio with conditional rendering metadata
CN107731239A
Encoder / decoder for multidimensional sound fields
US5583962A
Encoder / decoder for multidimensional sound fields
US5632005A
Method and apparatus for adjusting dynamic range and gain in an encoder / decoder for multidimensional sound fields
US5633981A
Method and apparatus for efficient implementation of single-sideband filter banks providing accurate measures of spectral magnitude and phase
US5727119A
Cited By
Method and system for generating and interactively rendering object-based audio
CN121438845A
Method and system for generating and interactively rendering object-based audio
CN121438846A
Method and system for generating and interactively rendering object-based audio
CN121438847A