Smart access to personalized audio
The method generates a bitstream with containers for object channels and metadata, enabling efficient extraction and rendering of personalized audio programs by selecting and decoding only required substreams, addressing inefficiencies in existing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-29
AI Technical Summary
Existing audio encoding and decoding systems require all audio object and speaker channels for personalized audio experiences, making it inefficient to remove unnecessary data without significant computational effort, and unable to generate sub-bitstreams containing only needed data.
A method for generating a bitstream with containers that include object channels, relation metadata, and presentation data, allowing decoders to efficiently extract and render personalized audio programs by selecting and decoding only required substreams.
Enables resource-efficient generation of personalized audio programs by allowing decoders to select and decode only necessary substreams, reducing computational complexity and bitrate.
Smart Images

Figure 2026123000000001_ABST
Abstract
Description
Technical Field
[0001] This document relates to audio signal processing, and more particularly, to the encoding, decoding, and interactive rendering of an audio data bitstream that includes audio content and metadata that supports the interactive rendering of the audio content.
Background Art
[0002] Audio encoding and decoding that enables a personalized audio experience typically needs to carry all of the audio object channels and / or audio speaker channels that may potentially be required for a personalized audio experience. In particular, audio data / metadata is typically such that portions that are not required for a personalized audio program cannot be easily removed from the bitstream that includes such a personalized audio program.
[0003] Typically, all the data for an audio program (audio data and metadata) is stored together within a bitstream. The receiver / decoder needs to parse at least the complete metadata to understand which parts of the bitstream (e.g., which speaker channels and / or object channels) are needed for the personalized audio program. In addition, stripping the parts of the bitstream that are not needed for the personalized audio program is typically not possible without considerable computational effort. In particular, it may be required that parts of the bitstream that are not needed for a given playback scenario / a given personalized audio program need to be decoded. Then, in order to generate the personalized audio program, it may be necessary to mute these parts of the bitstream during playback. Furthermore, it may not be possible to efficiently generate sub-bitstreams from the bitstream, where the sub-bitstream contains only the data needed for the personalized audio program. [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] This paper addresses the technical problem of providing a bitstream for an audio program that enables a decoder of the bitstream to derive a personalized audio program from the bitstream in a resource-efficient manner. [Means for solving the problem]
[0005] In one aspect, a method for generating a bitstream representing an object-based audio program is described. The bitstream includes a sequence of containers for corresponding sequences of audio program frames of the object-based audio program. The first container in the sequence of containers includes multiple substream entities for multiple substreams of the object-based audio program. Furthermore, the first container includes a presentation section. The method includes determining a set of object channels representing the audio content of at least some audio signals from a set of audio signals, where the set of object channels includes a sequence of object channel frames. The method also includes providing or determining a set of object relation metadata for the set of object channels, where the set of object relation metadata includes a sequence of object relation metadata frames. The first audio program frame of the object-based audio program includes a first set of object channel frames from the set of object channel frames and a corresponding first set of object relation metadata frames. Furthermore, the method includes inserting the first set of object channel frames and the first set of object relation metadata frames into each set of object channel substream entities of the plurality of substream entities in the first container. In addition, the method includes inserting presentation data into the presentation section, where the presentation data represents at least one presentation, which includes a set of substream entities from the plurality of substream entities that are presented simultaneously.
[0006] In another aspect, a bitstream representing an object-based audio program is described. The bitstream includes a sequence of containers for corresponding sequences of audio program frames of the object-based audio program. The first container in the sequence of containers includes the first audio program frame of the object-based audio program. The first audio program frame includes a first set of object channel frames and a corresponding first set of object relation metadata frames. The set of object channels represents the audio content of at least some audio signals from a set of audio signals. The first container includes a plurality of substream entities for a plurality of substreams of the object-based audio program. Each of the plurality of substream entities includes a set of object channel substream entities for the first set of object channel frames. The first container further includes a presentation section having presentation data, where the presentation data represents at least one presentation of the object-based audio program. The presentation includes a set of substream entities from the plurality of substream entities to be presented simultaneously.
[0007] In another aspect, this paper describes a method for generating a personalized audio program from the bitstream outlined herein. The method includes extracting presentation data from the presentation section, where the presentation data represents a presentation for the personalized audio program, and the presentation includes a set of substream entities from the multiple substream entities to be presented simultaneously. Furthermore, the method includes extracting one or more object channel frames and corresponding one or more object relation metadata frames from the set of object channel substream entities in the first container based on the presentation data.
[0008] In a further aspect, a system (e.g., an encoder) that generates a bitstream representing an object-based audio program is described. The bitstream includes a sequence of containers for corresponding sequences of audio program frames of the object-based audio program. The first container in the sequence of containers includes multiple substream entities for multiple substreams of the object-based audio program. The first container further includes a presentation section. The system is configured to determine a set of object channels representing the audio content of at least some audio signals from a set of audio signals, where the set of object channels includes a sequence of object channel frames. Furthermore, the system is configured to determine a set of object relation metadata for the set of object channels, where the set of object relation metadata includes a sequence of object relation metadata frames. The first audio program frame of the object-based audio program includes a first set of object channel frames from the set of object channel frames and a corresponding first set of object relation metadata frames. In addition, the system is configured to insert the first set of object channel frames and the first set of object relation metadata frames into each set of object channel substream entities of the plurality of substream entities in the first container. Furthermore, the system is configured to insert presentation data into the presentation section, where the presentation data represents at least one presentation, which includes a set of substream entities from the plurality of substream entities to be presented simultaneously.
[0009] In another aspect, a system is described for generating a personalized audio program from a bitstream containing an object-based audio program, such as the bitstream described herein. The system includes extracting presentation data from the presentation section, where the presentation data represents a presentation for the personalized audio program, and the presentation includes a set of substream entities from the plurality of substream entities to be presented simultaneously. Furthermore, the system is configured to extract one or more object channel frames and corresponding one or more object relation metadata frames from the set of object channel substream entities in the first container based on the presentation data.
[0010] In a further aspect, a software program is described. This software program may be adapted for execution on a processor, so that when executed on the processor, it performs the method steps outlined in this paper.
[0011] From another perspective, a storage medium is described. This storage medium may have a software program adapted to perform the method steps outlined in this paper when executed on a processor.
[0012] In a further aspect, a computer program product is described. This computer program may contain executable instructions for performing the method steps outlined in this paper when executed on a computer.
[0013] It should be noted that the methods and systems outlined in this patent application, including their preferred embodiments, may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application may be combined in any way. In particular, the features of the claims may be combined with each other in any way. [Brief explanation of the drawing]
[0014] The present invention is described below in an illustrative manner with reference to the accompanying drawings. [Figure 1] This is a block diagram of an exemplary audio processing chain. [Figure 2] This is a block diagram of an exemplary audio encoder. [Figure 3] This is an illustrative block diagram of an audio decoder. [Figure 4] This figure shows exemplary presentation data and exemplary subsystems for an audio program. [Figure 5] This figure shows an exemplary structure of a bitstream containing presentation data. [Figure 6] This is a flowchart illustrating an exemplary method for generating a bitstream containing presentation data. [Modes for carrying out the invention]
[0015] As described above, this paper addresses the technical problem of providing a bitstream for a general audio program that enables a decoder of the bitstream to generate a personalized audio program from the bitstream in a resource-efficient manner. In particular, the generation of the personalized audio program should be performed with relatively low computational complexity. Furthermore, the bitstream containing the general audio program should exhibit a relatively low bitrate.
[0016] Figure 1 shows a block diagram of an exemplary audio processing chain (also referred to as an audio data processing system). This system includes the following elements, combined as shown in the figure: capture unit 1, production unit 3 (which includes an encoding subsystem), delivery subsystem 5, decoder 7, object processing subsystem 9, controller 10, and rendering subsystem 11. Variations of the illustrated system may omit one or more of these elements, or include additional audio data processing units. Typically, elements 7, 9, 10, and 11 are included in a playback and / or decoding system (e.g., an end-user home theater system).
[0017] Capture unit 1 is typically configured to generate a PCM (time-domain) sample containing audio content and to output the PCM sample. The sample may represent multiple streams of audio captured by a microphone (e.g., at a sporting event or other spectator event). Production unit 3, typically operated by a broadcaster, accepts the PCM sample as input and is configured to output an object-based audio program representing the audio content. The program is typically an encoded (e.g., compressed) audio bitstream representing the audio content and presentation data that allows various personalized audio programs to be derived from the bitstream, or includes such a bitstream. The data of the encoded bitstream representing the audio content is sometimes referred to in this paper as "audio data". The object-based audio program output from unit 3 may represent (i.e., include) multiple speaker channels ("beds" of speaker channels), multiple object channels of audio data, and object-relation metadata. An audio program may include presentation data that can be used to select various combinations of speaker channels and / or object channels to generate various personalized audio programs (which may also be referred to as various experiences). For example, an object-based audio program may include a main mix, which includes audio content indicating speaker channel beds, audio content indicating at least one user-selectable object channel (and at least one other optional object channel), and object relationship metadata associated with each object channel.The program may also include at least one side mix containing audio content and / or object relation metadata that indicates at least one other object channel (e.g., at least one user-selectable object channel). The audio program may indicate one or more beds of speaker channels, or it may not indicate any beds. For example, the audio program (or a particular mix / presentation) may indicate two or more beds of speaker channels (e.g., a 5.1 channel neutral crowd noise bed, a 2.0 channel home team crowd noise bed, and a 2.0 away team crowd noise bed), which include at least one user-selectable bed (which can be selected using a user interface used for user selection of object channel content or configuration) and a default bed (which is rendered if no other bed is selected by the user). The default bed may be determined by data indicating the configuration of the speaker set of the playback system (e.g., initial configuration), and optionally, the user may select another bed to be rendered instead of the default bed.
[0018] The delivery subsystem 5 in Figure 1 is configured to store and / or transmit (e.g., broadcast) the audio program generated by unit 3. Decoder 7 accepts (receives or reads) the audio program delivered by delivery subsystem 5 and decodes the program (or one or more accepted elements thereof). Object processing subsystem 9 is coupled to receive the decoded speaker channels, object channels, and object relation metadata of the audio program delivered (from decoder 7). Subsystem 9 is coupled to and configured to output a selected subset of the entire set of object channels represented by the audio program, along with the corresponding object relation metadata, to rendering subsystem 11. Subsystem 9 is typically configured to pass the decoded speaker channels from decoder 7 directly to subsystem 11, immutably.
[0019] The object channel selection performed by subsystem 9 may be determined by user selection (one or more) (indicated by control data presented to subsystem 9 from controller 10) and / or by rules programmed or otherwise configured for subsystem 9 to implement (e.g., indicating conditions and / or constraints). Such rules may be determined by object relation metadata of the audio program and / or by other data presented to subsystem 9 (e.g., from controller 10 or another external source) (e.g., data indicating the functionality and organization of the speaker array of the playback system) and / or by pre-configuring subsystem 9 (e.g., by programming). Controller 10 may provide the user with a menu or palette of selectable “preset” mixes or presentations of object and “bed” speaker channel content (e.g., displayed on a touchscreen) via a user interface implemented by controller 10. The selectable preset mixes or presentations may also be determined by presentation data contained within the audio program and, possibly, by rules implemented by subsystem 9 (e.g., rules pre-configured for subsystem 9 to implement). The user selects from the available mixes / presentations by entering a command into the controller 10 (for example, by operating its touchscreen), and in response, the controller 10 presents the corresponding control data to the subsystem 9.
[0020] The rendering subsystem 11 in Figure 1 is configured to render audio content determined by the output of subsystem 9 for playback by the speakers (not shown) of the playback system. Subsystem 11 is configured to map audio content determined by object channels selected by object processing subsystem 9 (e.g., default objects and / or user-selected objects selected as a result of user interaction using controller 10) to available speaker channels using rendering parameters (e.g., user-selected and / or default values for spatial position and level) output from subsystem 9 that are associated with each selected object. At least some of the rendering parameters may be determined by object relation metadata output from subsystem 9. The rendering system 11 may also receive a bed of speaker channels that has been passed through by subsystem 9. Typically, subsystem 11 is an intelligent mixer configured to determine speaker feeds for available speakers. This involves mapping one or more selected (e.g., default selected) objects to each of several separate speaker channels and mixing those objects with “bed” audio content indicated by the corresponding speaker channels in the program’s speaker channel bed.
[0021] FIG. 2 is a block diagram of a broadcast system configured to generate an object-based audio program (and corresponding video program) for broadcast. A set of X microphones (X being an integer greater than 0, 1 or 2), including microphones 100, 101, 102, 103 of the system of FIG. 2, are positioned to capture audio content to be included in the audio program, and their outputs are coupled to the inputs of an audio console 104. The audio program may include interactive audio content that indicates the atmosphere within or about a viewer event (such as a soccer or rugby match, an automobile or motorcycle race or another sports event) and / or commentary about the viewer event. The audio program may include a plurality of audio objects (including user-selectable objects or object sets and also typically a default set of objects that are rendered when there is no user object selection), and a mixing (or "bed") of speaker channels of the audio program. The bed of speaker channels may be a normal mixing of speaker channels (such as a 5.1 channel mixing) of a type that may be included in a normal broadcast program that does not include object channels.
[0022] A subset of microphones (e.g., microphones 100 and 101, and optionally other microphones whose outputs are coupled to audio console 104) may, in operation, be a regular microphone array capturing audio (to be encoded and delivered as speaker channel beds). Another subset of microphones (e.g., microphones 102 and 103, and optionally other microphones whose outputs are coupled to audio console 104) may, in operation, capture audio (e.g., crowd noise and / or other "objects") to be encoded and delivered as object channels of the program. For example, the microphone array of the system in Figure 2 may include at least one microphone (e.g., microphone 100) implemented as a sound field microphone and permanently installed in the stadium; at least one stereo microphone (e.g., microphone 102) directed towards the positions of spectators supporting one team (e.g., the home team) and at least one other stereo microphone (e.g., microphone 103) directed towards the positions of spectators supporting the other team (e.g., the away team).
[0023] The broadcast system of FIG. 2 may include a mobile unit located outside the stadium (or other event location), which may be a truck and is sometimes referred to as a "game truck". This mobile unit is the first recipient of the audio feed from the microphones within the stadium (or other event location). The game truck generates an object-based audio program (to be broadcast). This includes encoding the audio content from the microphones for delivery as object channels of the audio program, generating corresponding object-related metadata (e.g., metadata indicating the spatial position where each object should be rendered), including such metadata in the audio program, and / or encoding the audio content from several microphones for delivery as the bed of the speaker channels of the audio program.
[0024] For example, in the system of Figure 2, a console 104, an object processing subsystem 106 (coupled to the output of console 104), an embedded subsystem 108, and a contributing encoder 110 may be installed within the match track. The object-based audio program generated in subsystem 106 may be combined with video content (for example, from cameras located within the stadium) (for example, within subsystem 108) to generate a combined audio and video signal. This combined signal is then encoded (for example, by encoder 110) to generate an encoded audio / video signal for broadcast (for example, by delivery subsystem 5 in Figure 1). It should be understood that a playback system that decodes and renders such an encoded audio / video signal would include subsystems (not shown individually) for parsing the audio and video content of the delivered audio / video signal, a subsystem for decoding and rendering the audio content, and another subsystem (not shown individually) for decoding and rendering the video content.
[0025] The audio output of console 104 may include, for example, a 5.1 speaker channel bed (labeled "5.1 neutral" in Figure 2) showing sounds captured at a sporting event; stereo object channel audio content (labeled "2.0 home") showing crowd noise from home team fans present at the event; stereo object channel audio content (labeled "2.0 away") showing crowd noise from visiting team fans present at the event; object channel audio content (labeled "1.0 cmm1") showing commentary from an announcer from the home team's city; object channel audio content (labeled "1.0 cmm2") showing commentary from an announcer from the visiting team's city; and object channel audio content (labeled "1.0 ball kick") showing sounds generated by a match ball when it is struck by a participant in a sporting event.
[0026] The object processing subsystem 106 is configured to organize the audio stream from console 104 into object channels (for example, grouping the left and right audio streams labeled "2.0 Away" into an expeditionary crowd noise object channel) and / or into sets of object channels (for example, grouping them), generate object relation metadata indicating those object channels (and / or sets of object channels), and encode those object channels (and / or sets of object channels), object relation metadata, and speaker channel beds (determined from the audio stream from console 104) as an object-based audio program (for example, an object-based audio program encoded as an AC-4 bitstream). Alternatively, encoder 110 may be configured to generate an object-based audio program, which may be encoded as an AC-4 bitstream, for example. In such a case, object processing subsystem 106 may focus on generating audio content (for example, using the Dolby E+ format), while encoder 110 may focus on generating a bitstream for transmission or distribution.
[0027] Subsystem 106 may further be configured to render (and play back on a set of studio monitor speakers) at least a selected subset of speaker channel beds and object channels (and / or object channel sets) (including by using object relation metadata to generate a mix / presentation indicating selected object channels (one or more) and speaker channels), the sound thereby played back can be monitored by the console 104 and the operator(s) of subsystem 106 (as shown by the “monitoring path” in Figure 2).
[0028] The interface between the output of subsystem 104 and the input of subsystem 106 may be a multi-channel audio digital interface (MADI).
[0029] In operation, subsystem 108 of the system in Figure 2 combines the object-based audio program generated in subsystem 106 with video content (e.g., from cameras located within the stadium) to generate a combined audio and video signal, which is then presented to encoder 110. The interface between the output of subsystem 108 and the input of subsystem 110 may be a high-definition serial digital interface (HD-SDI). In operation, encoder 110 encodes the output of subsystem 108, thereby generating an encoded audio / video signal for broadcast (e.g., by transmission subsystem 5 in Figure 1).
[0030] Broadcast facilities (for example, subsystems 106, 108, and 110 of the system in Figure 2) may be configured to generate various presentations of elements of an object-based audio program. Examples of such presentations include the flattened mix, international mix, and domestic mix of 5.1. For example, all presentations may include a common bed of speaker channels, but the object channels of a presentation (and / or the selectable object channels and / or the menu of selectable or non-selectable rendering parameters for rendering and mixing the object channels, as determined by the presentation) may differ from presentation to presentation.
[0031] Object-related metadata of an audio program (or pre-configured settings for a playback or rendering system that are not indicated by metadata delivered with the audio program) may impose constraints or conditions on the selectable mixing / presentation of object and bed (speaker channel) content. For example, a DRM hierarchy may be implemented to allow a user to have tiered access to a set of audio channels contained within an object-based audio program. A user may be permitted to decode, select, and render more object channels in the audio program if they pay a higher amount (e.g., to a broadcaster).
[0032] Figure 3 is a block diagram of an exemplary playback system that includes a decoder 20, an object processing subsystem 22, a spatial rendering subsystem 24, a controller 23 (which implements the user interface), and optionally digital audio processing subsystems 25, 26, and 27, combined as shown in the figure. In some implementations, elements 20, 22, 24, 25, 26, 27, 29, 31, and 33 of the system in Figure 3 are implemented as set-top devices.
[0033] In the system shown in Figure 3, the decoder 20 is configured to receive and decode an encoded signal representing an object-based audio program. The audio program represents audio content, for example, two speaker channels (i.e., "beds" of at least two speaker channels). The audio program also represents at least one user-selectable object channel (and optionally at least one other object channel) and object relation metadata corresponding to each object channel. Each object channel represents an audio object, and therefore, object channels are sometimes referred to as "objects" in this paper for convenience. The audio program may be contained within an AC-4 bitstream representing the audio objects, object relation metadata, and / or beds of speaker channels. Typically, individual audio objects are mono or stereo encoded (i.e., each object channel is a monophonic channel representing the left or right channel of an object or an object), the bed may be a traditional 5.1 mix, and the decoder 20 may be configured to simultaneously decode the audio content of a predetermined number (e.g., 16 or more) channels of audio content (e.g., including the six speaker channels of the bed and e.g., 10 or more object channels). The incoming bitstream may represent a number (e.g., more than 10) of audio objects, and it may not be necessary to decode all of them in order to achieve a particular mix / presentation.
[0034] As described above, the audio program may include zero speaker channels and one or more beds in addition to one or more object channels. The speaker channel beds and / or object channels may form substreams of the bitstream containing the audio program. Thus, the bitstream may contain multiple substreams, where a substream represents a speaker channel bed or one or more object channels. Furthermore, the bitstream may include presentation data (for example, contained within a presentation section of the bitstream), where the presentation data represents one or more different presentations. A presentation may define a particular mixture of substreams. In other words, a presentation may define speaker channel beds and / or one or more object channels that should be mixed together to provide a personalized audio program.
[0035] Figure 4 shows several substreams 411, 412, 413, and 414. Each substream 411, 412, 413, and 414 contains audio data 421 and 424, where the audio data 421 and 424 may correspond to speaker channel beds or to audio data (i.e., audio channels) of audio objects. For example, substream 411 may contain speaker channel beds 421, and substream 414 may contain object channels 424. Furthermore, each substream 411, 412, 413, and 414 may contain metadata 431 and 434 (e.g., default metadata) associated with the audio data 421 and 424 that can be used to render the associated audio data 421 and 424. For example, substream 411 may contain speaker relation metadata (for the bed of speaker channel 421), and substream 414 may contain object relation metadata (for object channel 424). In addition, substreams 411, 412, 413, and 414 may contain alternative metadata 441 and 444 to provide one or more alternative ways of rendering the associated audio data 421 and 424.
[0036] Furthermore, Figure 4 shows different presentations 401, 402, and 403. Presentation 401 shows a selection of substreams 411, 412, 413, and 414 to be used for presentation 401, thereby defining a personalized audio program. Furthermore, presentation 401 may show metadata 431, 441 (for example, one of the default metadata 431 or alternative metadata 441) to be used for the substream 411 selected for presentation 401. In the illustrated example, presentation 401 describes a personalized audio program that includes substreams 411, 412, and 414.
[0037] Therefore, the use of presentations 401, 402, and 403 provides an efficient means of signaling various personalized audio programs within a general object-based audio program. In particular, presentations 401, 402, and 403 may be such that decoders 7 and 20 can easily select one or more substreams 411, 412, 413, and 414 required for a particular presentation 401 without having to decode the complete bitstream of a general object-based audio program. For example, a re-multiplexer (not shown in Figure 3) may be configured to easily extract one or more substreams 411, 412, 413, and 414 from the complete bitstream to generate a new bitstream for a personalized audio program of a particular presentation 401. In other words, a new bitstream carrying a reduced number of presentations may be efficiently generated from a bitstream with a relatively large number of presentations 401, 402, and 403. A possible scenario is a relatively large bitstream with a relatively large number of presentations reaching the STB. The STB may be configured to focus on personalization (i.e., selecting presentations) and may be configured to repackage a single presentation bitstream (without decoding the audio data). The single presentation bitstream (and audio data) may then be decoded in a suitable remote decoder, for example, in an AVR (Audio / Video Receiver) or in a mobile home device such as a tablet PC.
[0038] The decoder (for example, decoder 20 in Figure 3) may parse the presentation data to identify presentation 401 for rendering. Furthermore, decoder 200 may extract the substreams 411, 412, and 414 required for presentation 401 from the locations indicated by the presentation data. After extracting substreams 411, 412, and 414 (speaker channel, object channel, and associated metadata), the decoder may perform any necessary decoding on the extracted substreams 411, 412, and 414 (for example, on them alone).
[0039] The bitstream may be an AC-4 bitstream, and presentations 401, 402, and 403 may be AC-4 presentations. These presentations allow for easy access to the parts of the bitstream required for a particular presentation (audio data 421 and metadata 431). In this way, the decoder or receiver system 20 can easily access the required parts of the bitstream without having to parse deep into other parts of the bitstream. This also makes it possible, for example, to transfer only the required parts of the bitstream to another device without having to reconstruct the entire structure or even decode and encode the bitstream substreams 411, 412, 413, and 414. In particular, a reduced structure derived from the bitstream may be extracted.
[0040] Referring again to Figure 3, the user may use Controller 23 to select an object (indicated by an object-based audio program) to be rendered. For example, the user may select a specific presentation 401. Controller 23 may be a handheld processing device (e.g., iPad®) programmed to implement a user interface (e.g., an iPad® app) compatible with other elements of the system in Figure 3. The user interface may provide the user with a menu or palette (e.g., displayed on a touchscreen) of selectable presentations 401, 402, 403 (e.g., a "preset" mix) of objects and / or "bed" speaker channel content. Presentations 401, 402, 403 may be provided with name tags within the menu or palette. The selectable presentations 401, 402, 403 may be determined by the presentation data of the bitstream and, possibly, by rules implemented by subsystem 22 (e.g., rules pre-configured to be implemented by subsystem 22). The user may select from the available presentations by entering a command into the controller 23 (for example, by activating the touchscreen of the controller 23), and in response, the controller 23 may present the corresponding control data to the subsystem 22.
[0041] In response to an object-based audio program and control data from controller 23 indicating a selected presentation 401, decoder 20 (if necessary) decodes the speaker channels of the speaker channel bed of the selected presentation 401 and outputs the decoded speaker channels to subsystem 22. In response to an object-based audio program and control data from controller 23 indicating a selected presentation 401, decoder 20 (if necessary) decodes a selected object channel and outputs the selected (e.g., decoded) object channels (each of which may be pulse code modulated or a "PCM" bitstream) and object relation metadata corresponding to the selected object channel to subsystem 22.
[0042] The objects indicated by the decoded object channels are typically user-selectable audio objects or contain user-selectable audio objects. For example, as shown in Figure 3, the decoder 20 may include a speaker channel bed, an object channel showing commentary by an announcer from the home team's city ("Comment 1 Mono"), an object channel showing commentary by an announcer from the visiting team's city ("Comment 2 Mono"), a stereo object channel showing crowd noise from home team fans at a sporting event ("Fans (Home)"), left and right object channels showing sounds produced by the game ball when it is hit by participants in the sporting event ("Ball Sound Stereo"), and four object channels showing special effects ("Effect 4x Mono"). Any of the "Comment 1 Mono", "Comment 2 Mono", "Fan (Home)", "Ball Sound Stereo", and "Effect 4x Mono" object channels may be selected as part of presentation 401, and each selected is passed from subsystem 22 to rendering subsystem 24 (after receiving any necessary decoding in decoder 20).
[0043] Subsystem 22 is configured to output a selected subset of the full set of object channels presented by the audio program and the corresponding object relation metadata of the audio program. Object selection may be determined by user selection (indicated by control data presented to subsystem 22 from controller 23) and / or rules programmed or otherwise configured for subsystem 22 to implement (e.g., indicating conditions and / or constraints). Such rules may be determined by the object relation metadata of the program and / or by other data presented to subsystem 22 (e.g., from controller 23 or another external source) (data indicating the functionality and organization of the speaker array of the playback system) and / or by pre-configuring subsystem 22 (e.g., programming). As described above, the bitstream may include presentation data that provides a set of selectable “preset” mixtures (i.e., presentations 401, 402, 403) of object and “bed” speaker channel content. Subsystem 22 passes the decoded speaker channels from decoder 20 typically immutably (to subsystem 24) and processes the selected object channels presented thereto.
[0044] The spatial rendering subsystem 24 (or subsystem 24 together with at least one downstream device or system) in Figure 3 is configured to render the audio content output from subsystem 22 for playback by the user's playback system speakers. Optionally, one or more of the included digital audio processing subsystems 25, 26, and 27 may implement post-processing for the output of subsystem 24.
[0045] The spatial rendering subsystem 24 is configured to map audio object channels selected by the object processing subsystem 22 to available speaker channels using rendering parameters output from subsystem 22 (e.g., user-selected and / or default values for spatial position and level) associated with each selected object. The spatial rendering system 24 also receives the decoded speaker channel bed passed through subsystem 22. Typically, subsystem 24 is an intelligent mixer, configured to determine the speaker feed for available speakers, including mapping one, two, or more selected object channels to each of several individual speaker channels and mixing the selected object channels(s) with the “bed” audio content represented by the corresponding speaker channels in the program’s speaker channel bed.
[0046] Speakers driven to render audio can be located at any position in the playback environment, not merely in a (nominal) horizontal plane. In some such cases, metadata included in the program indicates rendering parameters for rendering at least one object of the program at any apparent spatial position (in a three-dimensional volume) using a three-dimensional array of speakers. For example, an object channel may have corresponding metadata indicating a three-dimensional trajectory of the apparent spatial position where the object (indicated by the object channel) should be rendered. The trajectory may include a sequence of "floor" positions (in the plane of a subset of speakers assumed to be located on the floor or another horizontal plane of the playback environment) and a sequence of "above-floor" positions (determined by driving a subset of speakers assumed to be located in at least one other horizontal plane of the playback environment). In such cases, rendering can be performed according to the present invention so that the speaker is driven to emit a sound (determined by the associated object channel) that is perceived as originating from a sequence of object positions in three-dimensional space including the trajectory, mixed with a sound determined by the "bed" audio content. Subsystem 24 may be configured to implement such rendering or a step thereof, and the remaining steps of rendering may be performed by a downstream system or device (for example, rendering subsystem 35 in Figure 3).
[0047] Optionally, a digital audio processing (DAP) stage (for example, one for each of several predetermined output speaker channel configurations) is coupled to the output of the spatial rendering subsystem 24 to perform post-processing on the output of the spatial rendering subsystem. Examples of such processing include intelligent equalization or (in the case of stereo output) speaker virtualization processing.
[0048] The output of the system in Figure 3 (for example, the output of the spatial rendering subsystem or the DAP stage following the spatial rendering stage) may be a PCM bitstream (which determines the speaker feed for the available speakers). For example, if the user's playback system includes a 7.1 array of speakers, the system may output a PCM bitstream (generated in subsystem 24) or a post-processed version of such a bitstream (generated in DAP 25) that determines the speaker feed for the speakers in such an array. As another example, if the user's playback system includes a 5.1 array of speakers, the system may output a PCM bitstream (generated in subsystem 24) or a post-processed version of such a bitstream (generated in DAP 26) that determines the speaker feed for the speakers in such an array. As yet another example, if the user's playback system includes only left and right speakers, the system may output a PCM bitstream (generated in subsystem 24) or a post-processed version of such a bitstream (generated in DAP 27) that determines the speaker feed for the left and right speakers.
[0049] The system in Figure 3 optionally also includes one or both of the re-encoding subsystems 31 and 33. Re-encoding subsystem 31 is configured to re-encode the PCM bitstream (showing a feed for a 7.1 speaker array) output from DAP 25 as an encoded bitstream (e.g., AC-4 or AC-3 bitstream), and the resulting encoded (compressed) AC-3 bitstream may be output from the system. Re-encoding subsystem 33 is configured to re-encode the PCM bitstream (showing a feed for a 5.1 speaker array) output from DAP 27 as an encoded bitstream (e.g., AC-4 or AC-3 bitstream), and the resulting encoded (compressed) bitstream may be output from the system.
[0050] The system in Figure 3 optionally also includes a re-encoding (or formatting) subsystem 29 and a downstream rendering subsystem 35 coupled to receive the output of subsystem 29. Subsystem 29 is coupled to receive data (output from subsystem 22) indicating selected audio objects (or a default mixture of audio objects), corresponding object-relation metadata, and speaker channel beds, and is configured to re-encode (and / or format) such data for rendering by subsystem 35. Subsystem 35 may be implemented in an AVR or soundbar (or other system or device downstream of subsystem 29) and is configured to generate a speaker feed (or bitstream determining the speaker feed) for available playback speakers (speaker array 36) in response to the output of subsystem 29. For example, subsystem 29 may generate encoded audio by re-encoding the data indicating selected (or default) audio objects, corresponding metadata, and speaker channel beds into a format suitable for rendering in subsystem 35, and is configured to transmit the encoded audio to subsystem 35 (for example, via an HDMI® link). In response to the speaker feed generated by subsystem 35 (or determined by its output), the available speakers 36 emit a sound that represents a mixture of the speaker channel bed and the selected (or default) object(s), with the object(s) having an apparent source location determined by the object relation metadata of subsystem 29's output. When subsystems 29 and 35 are included, the rendering subsystem 24 is optionally omitted from the system.
[0051] As described above, the use of presentation data is beneficial because it allows the decoder 20 to efficiently select one or more substreams 411, 412, 413, 414 required for a particular presentation 401. In view of this, the decoder 20 may be configured to extract one or more substreams 411, 412, 413, 414 of a particular presentation 401 and reconstruct a new bitstream containing (typically only) one or more substreams 411, 412, 413, 414 of the particular presentation 401. This extraction and reconstruction of the new bitstream can be performed without actually decoding and re-encoding the one or more substreams 411, 412, 413, 414. Thus, the generation of a new bitstream for a particular presentation 401 can be performed in a resource-efficient manner.
[0052] The system in Figure 3 may also be a distributed system for rendering object-based audio, in which part of the rendering (i.e., at least one step) (for example, the selection of audio objects to be rendered and the selection of rendering characteristics for each selected object, as performed by subsystems 22 and controller 23 of the system in Figure 3) is implemented in a first subsystem (for example, elements 20, 22 and 23 of Figure 3, implemented in a set-top device or a set-top device and handheld controller), and another part of the rendering (for example, immersive rendering, in which a speaker feed or a signal determining the speaker feed is generated in response to the output of the first subsystem) is implemented in a second subsystem (for example, subsystem 35, implemented in an AVR or soundbar). Latency management may be implemented to take into account different times and different subsystems in which the parts of the audio rendering (and any processing of the video corresponding to the rendered audio) are performed.
[0053] As shown in Figure 5, a typical audio program may be transmitted in a bitstream 500 containing a sequence of containers 501. Each container 501 may contain data for a specific frame of the audio program. A specific frame of the audio program may correspond to a specific temporal segment of the audio program (for example, 20 milliseconds of the audio program). Thus, each container 501 in the sequence of containers 501 may carry data for a specific frame in the sequence of frames of a typical audio program. The data for a frame may be contained within a frame entity 502 of the container 501. The frame entity may be identified using a syntax element of the bitstream 500.
[0054] As described above, bitstream 500 may carry multiple substreams 411, 412, 413, and 414, where each substream 411 contains a bed of speaker channels 421 or an object channel 424. Thus, frame entity 502 may contain multiple corresponding substream entities 520. Furthermore, frame entity 502 may contain a presentation section 510 (also referred to as a Table of Contents (TOC) section). Presentation section 510 may contain TOC data 511 which may show, for example, several presentations 401, 402, and 403 contained within presentation section 510. Furthermore, presentation section 510 may contain one or more presentation entities 512, each carrying data for defining one or more presentations 401, 402, and 403. Substream entity 520 may contain content subentities 521 for carrying audio data 421 and 424 of the frames of substream 411. Furthermore, the substream entity 520 may include a metadata subentity 522 for carrying the corresponding metadata 431, 441 of the frames of the substream 411.
[0055] Figure 6 shows a flowchart of an exemplary method 600 for generating a bitstream 500 representing an object-based audio program (i.e., a general audio program). The bitstream 500 represents a bitstream format such that the bitstream 500 contains a sequence of containers 501 for corresponding sequences of audio program frames of the object-based audio program. In other words, each frame (i.e., each temporal segment) of the object-based audio program may be inserted into a container of a sequence of containers that may be defined by the bitstream format. The containers may be defined using specific container syntax elements of the bitstream format. As an example, the bitstream format may correspond to the AC-4 bitstream format. In other words, the bitstream 500 to be generated may be an AC-4 bitstream.
[0056] Furthermore, the bitstream format may be such that the first container 501 in the sequence of containers 501 (i.e., at least one of the containers 501 in the sequence of containers 501) contains multiple substream entities 520 for multiple substreams 411, 412, 413, 414 of the object-based audio program. As outlined above, the audio program may contain multiple substreams 411, 412, 413, 414, each substream 411, 412, 413, 414 may contain speaker channel beds 421 or object channels 424 or both. The bitstream format may be such that each container 501 in the sequence of containers 501 provides a dedicated substream entity 520 for the corresponding substreams 411, 412, 413, 414. In particular, each substream entity 520 may contain data relating to the frames of the corresponding substreams 411, 412, 413, 414. The frames of substreams 411, 412, 413, and 414 may also be the frames of speaker channel bed 421, which are referred to here as speaker channel frames. Alternatively, the frames of substreams 411, 412, 413, and 414 may also be the frames of object channel, which are referred to here as object channel frames. Substream entities 520 may be defined by the corresponding syntax elements of the bitstream format.
[0057] Furthermore, the first container 501 may include a presentation section 510. In other words, the bitstream format may allow the definition of a presentation section 510 for all of the containers 501 in a sequence of containers 501 (for example, using appropriate syntax elements). The presentation section 510 may be used to define different presentations 401, 402, 403 for different personalized audio programs that can be generated from a (general) object-based audio program.
[0058] Method 600 includes determining a set of object channels 424 that represent the audio content of at least some of the audio signals from a set of audio signals 601. The set of audio signals may represent captured audio content, for example, audio content captured using the system described in the context of Figure 2. The set of object channels 424 may include multiple object channels 424. Furthermore, the set of object channels 424 includes a sequence of object channel frames. In other words, each object channel includes a sequence of object channel frames. As a result, the set of object channels includes a sequence of object channel frames, and the set of object channel frames at a particular point in time includes the object channel frames of the set of object channels at that particular point in time.
[0059] Furthermore, method 600 includes providing or determining a set of object relation metadata 434, 444 for a set of object channels 424, where the set of object relation metadata 434, 444 includes a sequence of set of object relation metadata frames. In other words, the object relation metadata for a given object channel is segmented into a sequence of object relation metadata frames. As a result, the set of object relation metadata for a corresponding set of object channels includes a sequence of set of object relation metadata frames.
[0060] Therefore, object relation metadata frames may be provided for the corresponding object channel frames (for example, using the object processor 106 described in the context of Figure 2). As described above, object channel 424 may provide various variations of object relation metadata 434, 444. For example, a default variation 434 of object relation metadata and one or more alternative variations 444 of object relation metadata may be provided. This allows various perspectives (e.g., various positions within a stadium) to be simulated. Alternatively or additionally, speaker channel bed 421 may provide various variations of speaker relation metadata 431, 441. For example, a default variation 431 of speaker relation metadata and one or more alternative variations 441 of speaker relation metadata may be provided. This allows various rotations of speaker channel bed 421 to be defined. Similar to object relation metadata, speaker relation metadata may also change over time.
[0061] Therefore, an audio program may have a set of object channels. As a result, the first audio program frame of an object-based audio program includes a first set of object channel frames from a sequence of sets of object channel frames and a corresponding first set of object relation metadata frames from a sequence of sets of object relation metadata frames.
[0062] Method 600 further includes inserting the first set of object channel frames and the first set of object relation metadata frames into each set of object channel substream entities 520 of the plurality of substream entities 520 of the first container 501 603. Thus, for each object channel 421 of the object-based audio program, substreams 411, 412, 413, and 414 can be generated. Each substream 411, 412, 413, and 414 may be identified within the bitstream 500 via the respective substream entity 520 that carries the substreams 411, 412, 413, and 414. As a result, various substreams 411, 412, 413, and 414 can be identified and potentially extracted by the decoders 7, 20 in a resource-efficient manner without the need to decode the complete bitstream 500 and / or substreams 411, 412, 413, and 414.
[0063] Furthermore, method 600 includes inserting presentation data into the presentation section 510 of the bitstream 500 604. The presentation data may indicate at least one presentation 401, which may define a personalized audio program. In particular, the at least one presentation 401 may include or indicate a set of substream entities 520 from the plurality of substream entities 520 to be presented simultaneously. Thus, presentation 401 may indicate which one or more of the substreams 411, 412, 413, 414 of the object-based audio program are selected to generate the personalized audio program. As outlined above, presentation 401 may identify a subset of the complete set of substreams 411, 412, 413, 414 (i.e., less than the total number of substreams 411, 412, 413, 414).
[0064] Insertion of presentation data allows the corresponding decoders 7, 20 to identify and extract one or more substreams 411, 412, 413, 414 from the bitstream 500 in order to generate a personalized audio program without having to decode or parse the complete bitstream 500.
[0065] Method 600 may include determining a speaker channel bed 421 that represents the audio content of one or more audio signals from the set of audio signals. The speaker channel bed 421 may include one or more of the following: 2.0 channels, 5.1 channels, 5.1.2 channels, 7.1 channels, and / or 7.1.4 channels. The speaker channel bed 421 may be used to provide a basis for a personalized audio program. In addition, one or more object channels 424 may be used to provide a personalized variation of the personalized audio program.
[0066] The speaker channel bed 421 may contain a sequence of speaker channel frames, and the first audio program frame of an object-based audio program may contain the first speaker channel frame of the sequence of speaker channel frames. Method 600 may further include inserting the first speaker channel frame into a speaker channel substream entity 520 of the plurality of substream entities 520 of the first container 501. In this case, presentation 401 of presentation section 510 may include or show the speaker channel substream entity 520. Alternatively or additionally, presentation 401 may include or show one or more object channel substream entities 520 from a set of object channel substream entities.
[0067] Method 600 may further include providing speaker relation metadata 431, 441 for speaker channel beds 421. The speaker relation metadata 431, 441 may include a sequence of speaker relation metadata frames. The first speaker relation metadata frame from the sequence of speaker relation metadata frames may be inserted into the speaker channel substream entity 520. It should be noted that multiple speaker channel beds 421 may be inserted into the corresponding multiple speaker channel substream entities 520.
[0068] As outlined in the context of Figure 4, the presentation data may represent multiple presentations 401, 402, 403, each containing different sets of substream entities 520 for different personalized audio programs. The different sets of substream entities 520 may include different combinations of the one or more speaker channel substream entities 520, the one or more object channel substream entities 520, and / or different combinations of metadata variations 434, 444 (e.g., default metadata 434 or alternative metadata 444).
[0069] The presentation data within the presentation section 510 may be segmented into different presentation data entities 512 for different presentations 401, 402, and 403 (for example, using appropriate syntax elements of a bitstream format). Method 600 may further include inserting table of contents (TOC) data into the presentation section 510. The TOC data may indicate the location of the various presentation data entities 512 within the presentation section 510 and / or identifiers for the various presentations 401, 402, and 403 contained within the presentation section 510. Thus, the TOC data may be used by the corresponding decoders 7, 20 to efficiently identify and extract the various presentations 401, 402, and 403. Alternatively or additionally, the presentation data entities 512 for the various presentations 401, 402, and 403 may be included sequentially within the presentation section 510. If the TOC does not indicate the location of the various presentation data entities 512, the corresponding decoders 7 and 20 may identify and extract the various presentations 401, 402, and 403 by sequentially parsing through the various presentation data entities 512. This may be a bitrate-efficient method for signaling the various presentations 401, 402, and 403.
[0070] The substream entity 520 may include a content subentity 521 for audio content or audio data 424 and a metadata subentity 522 for related metadata 434, 444. The subentities 521, 522 may be identified by appropriate syntax elements of the bitstream format. In this way, the corresponding decoders 7, 20 can identify the audio data and corresponding metadata for the object channels or speaker channel beds in a resource-efficient manner.
[0071] As already mentioned above, the metadata frame for a corresponding channel frame may contain multiple different variations or groups 434, 444 of the metadata. Presentation 401 may indicate which variation or group 434 of the metadata should be used to render the corresponding channel frame. This can increase the degree of personalization of the audio program (e.g., listening / browsing perspective).
[0072] A speaker channel bed 421 typically includes one or more speaker channels, each to be presented by one or more speakers 36 of the presentation environment. On the other hand, an object channel 424 is typically presented by a combination of speakers 36 of the presentation environment. Object relation metadata 434, 444 of the object channel 424 may indicate the position from which the object channel 424 should be rendered within the presentation environment. The position of the object channel 424 may change over time. As a result, the combination of speakers 36 for rendering the object channel 424 may change along the sequence of object channel frames of the object channel 424, and / or the pan of the speaker 36 of the speaker combination may change along the sequence of object channel frames of the object channel 424.
[0073] Presentations 401, 402, and 403 may include target device configuration data for the target device configuration. In other words, presentations 401, 402, and 403 may depend on the target device configuration used for rendering presentations 401, 402, and 403. The target device configuration may differ with respect to the number of speakers, the speaker positions, and / or the number of audio channels that can be processed and rendered. Exemplary target device configurations include a 2.0 (stereo) target device configuration or a 5.1 target device configuration with left and right speakers, etc. The target device configuration typically includes a spatial rendering subsystem 24 described in the context of Figure 3.
[0074] Therefore, presentations 401, 402, and 403 may show different audio resources to be used for different target device configurations. Target device configuration data may show a set of substream entities 520 from the plurality of substream entities 520 and / or a variation of metadata 434 to be used to render presentation 401 for a particular target device configuration. In particular, target device configuration data may show such information for multiple different target device configurations. For example, presentation 401 may include various sections having target device configuration data for various target device configurations.
[0075] By doing so, the corresponding decoder or demultiplexer can efficiently identify the audio resources (one or more substreams 411, 412, 413, 414, one or more variations of metadata 441) that should be used for a particular target device configuration.
[0076] The bitstream format may allow for further (intermediate) layers for defining a personalized audio program. In particular, the bitstream format may allow for the definition of substream groups containing one, two, or more of the substreams 411, 412, 413, 414. Substream groups may be used to group various audio content, such as ambient content, dialogue, and / or effects. Presentation 401 may indicate a substream group. In other words, Presentation 401 may identify one, two, or more substreams to be rendered simultaneously by referring to a substream group containing the one, two, or more substreams. Thus, substream groups provide an efficient means for identifying two or more substreams (possibly related to one another).
[0077] Presentation section 510 may include one or more substream group entities (not shown in Figure 5) for defining one or more corresponding substream groups. Substream group entities may be located after or downstream of presentation data entities 512. Substream group entities may indicate one or more substreams 411, 412, 413, 414 contained within the corresponding substream group. Presentation 401 (defined within the corresponding presentation data entity 512) may indicate a substream group entity in order to include a corresponding substream group in presentation 401. Decoders 7, 20 may parse through presentation data entities 512 to identify a particular presentation 401. If presentation 401 refers to a substream group or substream group entity, decoders 7, 20 may continue to parse through presentation section 510 to identify the definition of the substream group contained within the substream group entity in presentation section 510. Therefore, decoders 7 and 20 may determine substreams 411, 412, 413, and 414 for a particular presentation 401 by parsing through the presentation data entities 512 and through the substream group entities of the presentation section 510.
[0078] Therefore, the method 600 for generating the bitstream 500 may include inserting data for identifying one, two, or more of the plurality of substreams into the substream group entity of the presentation section 510. As a result, the substream group entity includes data for defining the substream group.
[0079] Defining substream groups can be beneficial in light of bitrate reduction. In particular, multiple substreams 411, 412, 413, and 414 used collectively within multiple presentations 401, 402, and 403 may be grouped within a substream group. As a result, the multiple substreams 411, 412, 413, and 414 can be efficiently identified within presentations 401, 402, and 403 by referring to the substream group. Furthermore, defining substream groups can provide a efficient means for content designers to master combinations of substreams 411, 412, 413, and 414 and to define substream groups for mastered combinations of substreams 411, 412, 413, and 414.
[0080] Thus, a bitstream 500 is described that represents an object-based audio program and allows for resource-efficient personalization. The bitstream 500 includes a sequence of containers 501 for corresponding sequences of audio program frames of the object-based audio program. The first container 501 in the sequence of containers 501 includes the first audio program frame of the object-based audio program. The first audio program frame includes a first set of object channel frames of a set of object channels and a corresponding first set of object relation metadata frames. The set of object channels may represent the audio content of at least some audio signals from a set of audio signals. Furthermore, the first container 501 includes a plurality of substream entities 520 for a plurality of substreams 411, 412, 413, 414 of the object-based audio program. Each of the plurality of substream entities 520 includes a set of object channel substream entities 520 for the first set of object channel frames. The first container 501 further includes a presentation section 510 having presentation data, where the presentation data may represent at least one presentation 401 of an object-based audio program, the at least one presentation 401 including a set of substream entities 520 from the plurality of substream entities 520 to be presented simultaneously.
[0081] The first audio program frame may further include a first speaker channel frame of speaker channel bed 421, where speaker channel bed 421 represents the audio content of one or more audio signals from the set of audio signals. The plurality of substream entities 520 of the bitstream 500 may then include speaker channel substream entities 520 for the first speaker channel frame.
[0082] The bitstream 500 may be received by decoders 7, 20. Decoders 7, 20 may be configured to perform a method for generating a personalized audio program from the bitstream 500. This method may include extracting presentation data from the presentation section 501. As described above, the presentation data may represent a presentation 401 for the personalized audio program. Furthermore, the method may include, based on the presentation data, extracting one or more object channel frames and corresponding one or more object relation metadata frames from the set of object channel substream entities 520 of the first container 501 in order to generate and / or render the personalized audio program. Depending on the contents of the bitstream, the method may further include, based on the presentation data, extracting first speaker channel frames from the speaker channel substream entities 520 of the first container 501.
[0083] The methods and bitstreams described in this paper are useful in view of generating personalized audio programs for general object-based audio programs. In particular, the methods and bitstreams described allow parts of the bitstream to be stripped or extracted in a resource-efficient manner. For example, if only a portion of the bitstream needs to be transmitted, this can be done without transmitting / processing the full set of metadata and / or the full set of audio data. Only the required portion of the bitstream needs to be processed and transmitted. The decoder may only be required to parse the presentation section of the bitstream (e.g., TOC data) to identify the content contained within the bitstream. Furthermore, the bitstream may provide a “default” presentation (e.g., “standard mix”) that can be used by the decoder to start rendering the program without further parsing. In addition, the decoder only needs to decode the portion of the bitstream required to render a particular personalized audio program. This is achieved by appropriate clustering of the audio data into substreams and substream entities. The audio program may potentially include an unlimited number of substreams and substream entities, thereby giving the bitstream format a high degree of flexibility.
[0084] The methods and systems described in this paper may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on a digital signal processor or microprocessor. Other components may be implemented, for example, as hardware and / or as application-specific integrated circuits. Signals encountered in the methods and systems described may be stored on a medium such as random-access memory or optical storage media, and may be transmitted over a network such as a radio network, satellite network, wireless network, or wired network, such as the Internet. Typical devices utilizing the methods and systems described in this paper are portable electronic devices or other consumer equipment used to store and / or render audio signals.
[0085] Embodiments of the present invention may relate to one or more of the following numbered examples (EE). [EEE1] A method (600) for generating a bitstream (500) representing an object-based audio program, wherein the bitstream (500) comprises a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program; the first container (501) of the sequence of containers (501) comprises a plurality of substream entities (520) for a plurality of substreams (411, 412, 413, 414) of the object-based audio program; the first container (501) further comprises a presentation section (510); the method (600) comprises, Step (601) of determining a set of object channels (424) representing the audio content of at least some audio signals from a set of audio signals, wherein the set of object channels (424) includes a sequence of object channel frames; Step (602) of providing a set of object relation metadata (434, 444) for the set of object channels (424), wherein the set of object relation metadata (434, 444) comprises a sequence of set of object relation metadata frames; and a first audio program frame of the object-based audio program comprises a first set of object channel frames and a corresponding first set of object relation metadata frames; Step (603) of inserting the first set of object channel frames and the first set of object relation metadata frames into each set of object channel substream entities (520) of the plurality of substream entities (520) of the first container (501); The step (604) includes inserting presentation data into the presentation section (510), wherein the presentation data represents at least one presentation (401); and the presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously. method. [EEE2] The method (600) according to EEE1, wherein the presentation (401) includes one or more object channel substream entities (520) from the set of object channel substream entities. [EEE3] The method (600) according to EEE1 or 2, wherein the presentation data represents a plurality of presentations (401, 402, 403) which include different sets of substream entities (520), and the different sets of substream entities (520) include different combinations of object channel substream entities (520) of the set. [EEE4] A method (600) according to any one of EEE1 to 3, wherein the aforementioned presentation data is segmented into different presentation data entities (512) for different presentations (401, 402, 403). [EEE5] The process further includes inserting table of contents data, referred to as TOC data, into the presentation section (510), wherein the TOC data is · The location of the different presentation data entities (512) within the presentation section (510); and / or - Identifiers for the different presentation data entities (512) contained within the presentation section (510), Method described in EEE4 (600). [EEE6] The method (600) described in any one of EEE1 to 5, wherein the substream entity (520) includes a content subentity (521) for audio content (424) and a metadata subentity (522) for related metadata. [EEE7] • The metadata frame for the corresponding channel frame contains multiple different variations of the metadata (434, 444); • Presentation (401) indicates which transformation (434) of the metadata should be used to render the corresponding channel frame. The method described in any one of EEE1 to 6 (600). [EEE8] The step of determining a speaker channel bed (421) representing the audio content of one or more audio signals from the set of audio signals, wherein the speaker channel bed (421) includes a sequence of speaker channel frames; and the first audio program frame of the object-based audio program includes the first speaker channel frame of the speaker channel bed (421); The process further includes the step of inserting the first speaker channel frame into the speaker channel substream entities (520) of the plurality of substream entities (520) of the first container (501), The method described in any one of EEE1 to EEE7 (600). [EEE9] The method (600) according to EEE8, wherein the presentation (401) also includes the speaker channel substream entity (520). [EEE10] The method (600) according to EEE8 or 9, wherein the bed (421) of the speaker channel includes one or more speaker channels which are to be presented by one or more speakers in the presentation environment, respectively. [EEE11] The method (600) further includes providing speaker-related metadata (431, 441) for the speaker channel bed (421); The aforementioned speaker-related metadata (431, 441) includes a sequence of speaker-related metadata frames; - A first speaker relation metadata frame from the sequence of speaker relation metadata frames is inserted into the speaker channel substream entity (520). The method described in any one of EEE8 to 10 (600). [EEE12] The method (600) according to any one of EEE8 to 11, wherein the speaker channel bed (421) includes one or more of 2.0 channels, 5.1 channels and / or 7.1 channels. [EEE13] The method (600) according to any one of EEE1 to 12, wherein the set of object channels (424) includes multiple object channels (424). [EEE14] A method (600) according to any one of EEE1 to 13, wherein the object channel (424) is presented by a combination of speakers (36) of the presentation environment. [EEE15] The method (600) described in EEE14, wherein the object relation metadata (434, 444) of the object channel (424) indicates the location from which the object channel (424) should be rendered within the presentation environment. [EEE16] The position of the aforementioned object channel (424) changes over time; - The combination of speakers (36) for rendering the object channel (424) changes along the sequence of object channel frames of the object channel (424); and / or - The pan of the speaker (36) in the combination of the speaker (36) changes according to the sequence of the object channel frames of the object channel (424). The method described in EEE14 or 15 (600). [EEE17] The method (600) described in any one of EEE1 to 16, wherein the bitstream (500) is an AC-4 bitstream. [EEE18] The method (600) according to any one of EEE1 to 17, wherein the set of audio signals indicates the captured audio content. [EEE19] • Presentation (401) includes target device configuration data for the target device configuration; The target device configuration data indicates a set of substream entities (520) from the plurality of substream entities (520) and / or metadata transformations (434) which should be used to render the presentation (401) in the target device configuration. The method described in any one of EEE1 to EEE18 (600). [EEE20] One, two, or three or more of the aforementioned substreams form a substream group; Presentation (401) indicates the aforementioned substream group. The method described in any one of EEE1 to EEE19 (600). [EEE21] The method (600) described in EEE20 further includes the step of inserting data for identifying one, two, or three or more of the plurality of substreams into the substream group entity of the presentation section (510), wherein the substream group entity includes data for defining the substream group. [EEE22] A bitstream (500) representing an object-based audio program, The bitstream (500) includes a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program; A first container (501) in the sequence of containers (501) includes a first audio program frame of the object-based audio program; The first audio program frame comprises a first set of object channel frames and a corresponding first set of object relation metadata frames; The first set of object channel frames represents the audio content of at least some of the audio signals from the set of audio signals; The first container (501) includes multiple substream entities (520) for multiple substreams (411, 412, 413, 414) of the object-based audio program; Each of the plurality of substream entities (520) includes a set of object channel substream entities (520) for the first set of object channel frames; The first container (501) further includes a presentation section (510) having presentation data; The presented data represents at least one presentation (401) of the object-based audio program; The presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously. Bitstream. [EEE23] The first audio program frame includes the first speaker channel frame of the speaker channel bed (421); The speaker channel bed (421) represents the audio content of one or more audio signals from the set of audio signals; The plurality of substream entities (520) include a speaker channel substream entity (520) for the first speaker channel frame. Bitstream as specified in EEE22. [EEE24] A method for generating a personalized audio program from a bitstream (500) containing an object-based audio program, The bitstream (500) includes a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program; A first container (501) in the sequence of containers (501) includes a first audio program frame of the object-based audio program; The first audio program frame includes a first set of object channel frames for a set of object channels (424) and a corresponding first set of object relation metadata frames; The set of object channels (424) represents the audio content of at least some of the audio signals from the set of audio signals; The first container (501) includes multiple substream entities (520) for multiple substreams (411, 412, 413, 414) of the object-based audio program; Each of the plurality of substream entities (520) includes a set of object channel substream entities (520) for the first set of object channel frames; The first container (501) further includes a presentation section (510); This method is The step of extracting presentation data from the presentation section (510), wherein the presentation data indicates a presentation (401) for the personalized audio program, and the presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously; The step of extracting one or more object channel frames and corresponding one or more object relation metadata frames from the set of object channel substream entities (520) of the first container (501) based on the presented data, method. [EEE25] The first audio program frame includes the first speaker channel frame of the speaker channel bed (421); The speaker channel bed (421) represents the audio content of one or more audio signals from the set of audio signals; The plurality of substream entities (520) include a speaker channel substream entity (520) for the first speaker channel frame, The method further includes the step of extracting the first speaker channel frame from the speaker channel substream entity (520) of the first container (501) based on the presented data, Method as described in EEE24. [EEE26] A system (3) that generates a bitstream (500) representing an object-based audio program, wherein the bitstream (500) includes a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program; a first container (501) of the sequence of containers (501) includes a plurality of substream entities (520) for a plurality of substreams (411, 412, 413, 414) of the object-based audio program; the first container (501) further includes a presentation section (510); the system (3) The step of determining a set of object channels (424) representing the audio content of at least some audio signals from a set of audio signals, wherein the set of object channels (424) includes a sequence of object channel frames; The step of determining a set of object relation metadata (434, 444) for the set of object channels (424), wherein the set of object relation metadata (434, 444) includes a sequence of sets of object relation metadata frames; and a first audio program frame of the object-based audio program includes a first set of object channel frames and a corresponding first set of object relation metadata frames; - The step of inserting the first set of object channel frames and the first set of object relation metadata frames into each set of object channel substream entities (520) of the plurality of substream entities (520) of the first container (501); The steps are configured to include inserting presentation data into the presentation section (510), wherein the presentation data represents at least one presentation (401); and the at least one presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously. system. [EEE27] A system (7) for generating a personalized audio program from a bitstream (500) containing an object-based audio program, The bitstream (500) includes a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program; A first container (501) in the sequence of containers (501) includes a first audio program frame of the object-based audio program; The first audio program frame includes a first set of object channel frames for a set of object channels (424) and a corresponding first set of object relation metadata frames; The set of object channels (424) represents the audio content of at least some of the audio signals from the set of audio signals; The first container (501) includes multiple substream entities (520) for multiple substreams (411, 412, 413, 414) of the object-based audio program; Each of the plurality of substream entities (520) includes a set of object channel substream entities (520) for the first set of object channel frames; The first container (501) further includes a presentation section (510); The system (7) is The step of extracting presentation data from the presentation section (510), wherein the presentation data indicates a presentation (401) for the personalized audio program, and the presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously; The system is configured to perform the steps of extracting one or more object channel frames and corresponding one or more object relation metadata frames from the set of object channel substream entities (520) of the first container (501) based on the presented data. system.
[0086] Several aspects are described below. [Aspect 1] A method (600) for generating a bitstream (500) representing an object-based audio program, wherein the object-based audio program comprises a plurality of substreams; the bitstream (500) comprises a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program; a first container (501) of the sequence of containers (501) comprises a plurality of substream entities (520) for each of the plurality of substreams (411, 412, 413, 414); the substream entities comprise data relating to the frames of the corresponding substreams; the first container (501) further comprises a presentation section (510); the method (600) comprises, Step (601) of determining a set of object channels (424) representing the audio content of at least some audio signals from a set of audio signals, wherein the set of object channels (424) includes a sequence of object channel frames; Step (602) of providing a set of object relation metadata (434, 444) for the set of object channels (424), wherein the set of object relation metadata (434, 444) comprises a sequence of set of object relation metadata frames; a first audio program frame of the object-based audio program comprises a first set of object channel frames and a corresponding first set of object relation metadata frames, wherein the object channels are presented by a combination of speakers in a presentation environment, and the object relation metadata of the object channels indicates the location in the presentation environment from which the object channels should be rendered; Step (603) of inserting the first set of object channel frames and the first set of object relation metadata frames into each set of object channel substream entities (520) of the plurality of substream entities (520) of the first container (501); The step (604) includes inserting presentation data into the presentation section (510), wherein the presentation data represents at least one presentation (401); and the presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously. method. [Aspect 2] The aforementioned presentation data is segmented into different presentation data entities (512) for different presentations (401, 402, 403), The process further includes inserting table of contents data, referred to as TOC data, into the presentation section (510), wherein the TOC data is · The location of the different presentation data entities (512) within the presentation section (510); and / or - Identifiers for the different presentation data entities (512) contained within the presentation section (510), The method described in Embodiment 1 (600). [Aspect 3] • The metadata frame for the corresponding channel frame contains multiple different variations of the metadata (434, 444); • Presentation (401) indicates which transformation (434) of the metadata should be used to render the corresponding channel frame. The method according to embodiment 1 or 2 (600). [Aspect 4] The step of determining a speaker channel bed (421) representing the audio content of one or more audio signals from the set of audio signals, wherein the speaker channel bed (421) includes a sequence of speaker channel frames; and the first audio program frame of the object-based audio program includes the first speaker channel frame of the speaker channel bed (421); The process further includes the step of inserting the first speaker channel frame into the speaker channel substream entities (520) of the plurality of substream entities (520) of the first container (501), The method described in any one of the descriptions in 1 to 3 (600). [Aspect 5] The method (600) according to embodiment 4, wherein the speaker channel bed (421) includes one or more speaker channels to be presented by one or more speakers in the presentation environment, respectively. [Aspect 6] The method (600) further includes providing speaker-related metadata (431, 441) for the speaker channel bed (421); The aforementioned speaker-related metadata (431, 441) includes a sequence of speaker-related metadata frames; - A first speaker relation metadata frame from the sequence of speaker relation metadata frames is inserted into the speaker channel substream entity (520). The method (600) described in aspect 4 or 5. [Aspect 7] The position of the aforementioned object channel (424) changes over time; - The combination of speakers (36) for rendering the object channel (424) changes along the sequence of object channel frames of the object channel (424); and / or - The pan of the speaker (36) in the combination of the speaker (36) changes according to the sequence of the object channel frames of the object channel (424). The method described in any one of the descriptions in 1 to 6 (600). [Aspect 8] • Presentation (401) includes target device configuration data for the target device configuration; The target device configuration data indicates a set of substream entities (520) from the plurality of substream entities (520) and / or metadata transformations (434) which should be used to render the presentation (401) in the target device configuration. The method described in any one of the descriptions in paragraphs 1 to 7 (600). [Aspect 9] One, two, or three or more of the aforementioned substreams form a substream group; Presentation (401) indicates the aforementioned substream group, The method further includes the step of inserting data for identifying one, two, or three or more of the plurality of substreams into the substream group entity of the presentation section (510), wherein the substream group entity includes data for defining the substream group. The method described in any one of the descriptions in paragraphs 1 to 8 (600). [Aspect 10] A bitstream (500) representing an object-based audio program, The bitstream (500) includes a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program, and the object-based audio program includes a plurality of substreams; A first container (501) in the sequence of containers (501) includes a first audio program frame of the object-based audio program; The first audio program frame comprises a first set of object channel frames and a corresponding first set of object relation metadata frames; An object channel frame is presented by a combination of speakers in the presentation environment, and the object relationship metadata frame of the object channel frame indicates the position within the presentation environment from which that object channel frame should be rendered; The first set of object channel frames represents the audio content of at least some of the audio signals from the set of audio signals; The first container (501) includes a plurality of substream entities (520) for each of the plurality of substreams (411, 412, 413, 414); each substream entity includes data relating to the frame of the corresponding substream; Each of the plurality of substream entities (520) includes a set of object channel substream entities (520) for the first set of object channel frames; The first container (501) further includes a presentation section (510) having presentation data; The presented data represents at least one presentation (401) of the object-based audio program; The presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously. Bitstream. [Aspect 11] The first audio program frame includes the first speaker channel frame of the speaker channel bed (421); The speaker channel bed (421) represents the audio content of one or more audio signals from the set of audio signals; The plurality of substream entities (520) include a speaker channel substream entity (520) for the first speaker channel frame. The bitstream described in aspect 10. [Aspect 12] A method for generating a personalized audio program from a bitstream (500) containing an object-based audio program, The bitstream (500) includes a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program, and the object-based audio program includes a plurality of substreams; A first container (501) in the sequence of containers (501) includes a first audio program frame of the object-based audio program; The first audio program frame includes a first set of object channel frames for a set of object channels (424) and a corresponding first set of object relation metadata frames; An object channel frame is presented by a combination of speakers in the presentation environment, and the object relationship metadata frame of the object channel frame indicates the position within the presentation environment from which that object channel frame should be rendered; The set of object channels (424) represents the audio content of at least some of the audio signals from the set of audio signals; The first container (501) includes a plurality of substream entities (520) for each of the plurality of substreams (411, 412, 413, 414); each substream entity includes data relating to the frame of the corresponding substream; Each of the plurality of substream entities (520) includes a set of object channel substream entities (520) for the first set of object channel frames; The first container (501) further includes a presentation section (510); This method is The step of extracting presentation data from the presentation section (510), wherein the presentation data indicates a presentation (401) for the personalized audio program, and the presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously; The step of extracting one or more object channel frames and corresponding one or more object relation metadata frames from the set of object channel substream entities (520) of the first container (501) based on the presented data, method. [Aspect 13] The first audio program frame includes the first speaker channel frame of the speaker channel bed (421); The speaker channel bed (421) represents the audio content of one or more audio signals from the set of audio signals; The plurality of substream entities (520) include a speaker channel substream entity (520) for the first speaker channel frame, The method further includes the step of extracting the first speaker channel frame from the speaker channel substream entity (520) of the first container (501) based on the presented data, The method described in aspect 12. [Aspect 14] A system (3) that generates a bitstream (500) representing an object-based audio program, wherein the bitstream (500) includes a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program; the object-based audio program includes a plurality of substreams; a first container (501) having a sequence of containers (501) includes a plurality of substream entities (520) for each of the plurality of substreams (411, 412, 413, 414); the substream entities include data relating to the frames of the corresponding substreams; the first container (501) further includes a presentation section (510); and the system (3) The step of determining a set of object channels (424) representing the audio content of at least some audio signals from a set of audio signals, wherein the set of object channels (424) includes a sequence of object channel frames; The step of determining a set of object relation metadata (434, 444) for a set of object channels (424), wherein the set of object relation metadata (434, 444) comprises a sequence of object relation metadata frames; a first audio program frame of the object-based audio program comprises a first set of object channel frames and a corresponding first set of object relation metadata frames, where the object channels are presented by a combination of speakers in a presentation environment, and the object relation metadata of the object channels indicates the location in the presentation environment from which the object channels should be rendered; - The step of inserting the first set of object channel frames and the first set of object relation metadata frames into each set of object channel substream entities (520) of the plurality of substream entities (520) of the first container (501); The steps are configured to include inserting presentation data into the presentation section (510), wherein the presentation data represents at least one presentation (401); and the at least one presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously. system. [Aspect 15] A system (7) for generating a personalized audio program from a bitstream (500) including an object-based audio program, wherein the object-based audio program includes a plurality of substreams; The bitstream (500) includes a sequence of containers (501) for corresponding sequences of audio program frames of the object-based audio program; A first container (501) in the sequence of containers (501) includes a first audio program frame of the object-based audio program; The first audio program frame includes a first set of object channel frames for a set of object channels (424) and a corresponding first set of object relation metadata frames; An object channel frame is presented by a combination of speakers in the presentation environment, and the object relationship metadata frame of the object channel frame indicates the position within the presentation environment from which that object channel frame should be rendered; The set of object channels (424) represents the audio content of at least some of the audio signals from the set of audio signals; The first container (501) includes a plurality of substream entities (520) for each of the plurality of substreams (411, 412, 413, 414); each substream entity includes data relating to the frame of the corresponding substream; Each of the plurality of substream entities (520) includes a set of object channel substream entities (520) for the first set of object channel frames; The first container (501) further includes a presentation section (510); The system (7) is The step of extracting presentation data from the presentation section (510), wherein the presentation data indicates a presentation (401) for the personalized audio program, and the presentation (401) includes a set of substream entities (520) from the plurality of substream entities (520) to be presented simultaneously; The system is configured to perform the steps of extracting one or more object channel frames and corresponding one or more object relation metadata frames from the set of object channel substream entities (520) of the first container (501) based on the presented data. system.
Claims
[Claim 1] A method for decoding an audio program from an encoded bitstream, wherein the encoded bitstream includes a sequence of containers for corresponding sequences of audio program frames, each container including presentation data and a plurality of substream entities for each audio program frame, each substream entity including object channel audio data and metadata, and the method: The step of receiving the encoded bitstream; A step of extracting presentation data from the container for the audio program frames of the encoded bitstream, wherein the presentation data indicates a presentation of the audio program, the presentation data is used to identify one or more substream entities from the plurality of substream entities, and the one or more substream entities are used for the audio program to be rendered; The steps include decoding metadata and object channel audio data corresponding to each of the one or more substream entities identified using the presentation data, The metadata indicates the location within the presentation environment from which the corresponding object channel audio data should be rendered. method.