Seamless and scalable decoding of channels, objects, and HOA audio content
By using adaptive spatial coding and crossfade-in/fade-out techniques, the number and encoding method of audio scene elements are dynamically adjusted, solving the decoding and rendering problems of multi-channel and high-fidelity audio content under limited bandwidth, and achieving high-quality audio playback.
Patent Information
- Application Number
- CN202180065769.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-25
- Filing Date
- 2021-09-10
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2041-09-10
AI Technical Summary
Existing technologies struggle to effectively decode and render multi-channel audio objects and high-fidelity stereo reproductions with limited bandwidth, leading to spatial artifacts and a decline in audio quality.
Adaptive spatial coding technology is adopted to dynamically adjust the number and encoding method of audio scene elements according to the target bit rate and the priority of channels, objects, and HOAs. Crossfade-in and crossfade-out technology is used to process mixed encoded frames of different content types, and overlapping and additive synthesis technology is combined to eliminate spatial artifacts.
It achieves high-quality audio decoding and rendering under limited bandwidth, reduces computational complexity and latency, and improves the continuity and consistency of audio playback.
Smart Images

Figure CN116324980B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 083,794, filed September 25, 2020, the disclosure of which is incorporated by reference herein in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates to the field of audio communication; and more specifically, to a digital signal processing method designed to decode immersive audio content that has been encoded using an adaptive spatial coding technique. Other aspects are also described. BACKGROUND
[0004] Consumer electronics devices are providing increasingly complex and improving performance digital audio encoding and decoding capabilities. Traditionally, audio content has been produced, distributed, and consumed primarily using a two-channel stereo format that provides left and right audio channels. Recent market developments aim to provide a more immersive listener experience using richer audio formats (e.g., Dolby Atmos or MPEG-H) that support multi-channel audio, object-based audio, and / or Ambisonics.
[0005] The communication of immersive audio content is associated with greater bandwidth needs, i.e., greater data rates are needed for streaming and downloading compared to for stereo content. If bandwidth is limited, techniques are needed that can reduce the size of the audio data while maintaining the best possible audio quality. A common approach to reduce bandwidth in perceptual audio coding is to exploit perceptual properties of the human hearing to maintain audio quality. For example, spatial encoders corresponding to different content types such as multi-channel audio, audio objects, higher order Ambisonics (HOA), or stereo format can use spatial parameters to achieve bitrate-efficient coding of the soundfield. To efficiently use limited bandwidth, audio scenes of different complexity can be spatially encoded using different content types for transmission. However, the decoding and rendering of audio scenes encoded using different content types can introduce spatial artifacts, such as when transitioning between rendered audio scenes encoded using content types of different spatial resolution. To communicate richer and more immersive audio content using limited bandwidth, stronger audio encoding and decoding (codec) techniques are needed. SUMMARY
[0006] Aspects of a scalable decoder that decodes and renders immersive audio content represented using an adaptive number of elements of various content types are disclosed. An audio scene of the immersive audio content can be represented by an adaptive number of scene elements in one or more content types encoded by adaptive spatial coding and baseline coding techniques; and an adaptive channel configuration that supports a target bitrate of a transport channel or a user. For example, the audio scene can be represented by an adaptive number of scene elements for channels, objects, and / or higher order ambisonics (HOA), etc. HOA describes a soundfield based on spherical harmonics. When recreated at the decoder, the different content types have different bandwidth requirements and correspondingly different audio quality. Adaptive channel and object spatial coding techniques can generate an adaptive number of channels and objects, and adaptive HOA spatial coding or HOA compression techniques can generate an adaptive order of HOA. This adaptivity can be made according to a target bitrate associated with a desired quality and an analysis that determines a priority of channels, objects, and HOA. The target bitrate can change dynamically based on channel conditions or bitrate requirements of one or more users. The priority decisions can be made based on spatial saliency of soundfield components represented by channels, objects, and HOA.
[0007] In one aspect, a scalable decoder can decode an audio stream that represents an audio scene by an adaptive number of scene elements for channels, objects, HOA, and / or stereo based immersive coding (STIC). The scalable decoder can also render the decoded stream using a fixed speaker configuration. Cross-fading of rendered channels, objects, HOA, or stereo based signals between consecutive frames can be performed for the same speaker layout. For example, frame-by-frame audio bitstreams encoded for channels / objects, HOA, and STIC can be decoded using a channel / object spatial decoder, a spatial HOA decoder, and a STIC decoder, respectively. The decoded bitstreams are presented to the speaker configuration of the playback device. If a newly rendered frame contains a different mix of channels, objects, HOA, and STIC signals than a previously rendered frame, for the same speaker layout, the new frame can be faded in and the old frame can be faded out. In an overlap period for cross-fading, the same soundfield can be represented by two different mixes of channels, objects, HOA, and STIC signals.
[0008] In one aspect, at an audio decoder, a bitstream is decoded that represents an audio scene using an adaptive number of scene elements for channel, object, HOA, and / or STIC encoding. The audio decoder can perform crossfading between channel, object, HOA, and stereo format signals in channel, object, HOA, and stereo format. A mixer in the same playback device as the audio decoder or in another playback device can render the crossfaded channel, object, HOA, and stereo format signals based on its respective loudspeaker layout. In one aspect, the crossfaded output of the audio decoder and time-synchronized channel, object, HOA, and STIC metadata can be transmitted to the other playback device where the PCM and metadata are provided to the mixer. In one aspect, the crossfaded output of the audio decoder and time-synchronized metadata can be compressed into a bitstream and transmitted to the other playback device where the bitstream is decompressed and provided to the mixer. In one aspect, the output of the audio decoder can be stored as a file for future rendering.
[0009] In one aspect, at an audio decoder, a bitstream is decoded that represents an audio scene using an adaptive number of scene elements for channel, object, HOA, and / or STIC encoding. A mixer in the same playback device can perform crossfading between channel, object, HOA, and stereo format signals in channel, object, HOA, and stereo format. The mixer can then render the crossfaded channel, object, HOA, and stereo format signals based on its loudspeaker layout. In one aspect, the output of the audio decoder can be PCM channels and its time-synchronized channel, object, HOA, and STIC metadata. The output of the audio decoder can be compressed and transmitted to other playback devices for crossfading and rendering.
[0010] In one aspect, at an audio decoder, a bitstream is decoded that represents an audio scene using an adaptive number of scene elements for channel, object, HOA, and / or STIC encoding. Crossfading between previous and current frames can be performed between transport channels at the output of a baseline decoder prior to spatial decoding. A mixer in one or more devices can render the crossfaded channel, object, HOA, and stereo format signals based on its respective loudspeaker layout. In one aspect, the output of the audio decoder can be PCM channels and its time-synchronized channel, object, HOA, and STIC metadata.
[0011] In one aspect of the technology for cross-fading between signals in channel, object, HOA, and stereo formats, if a current frame contains a bitstream encoded using a different mix of content types than the mix of content types of a previous frame, the transition frame can start with a stream referred to as an immediate fade-in and fade-out frame (IFFF). The IFFF can contain not only a bitstream of a current frame encoded using a mix of channel, object, HOA, and stereo format signals for fade-in, but also a bitstream of a previous frame encoded using a different mix of channel, object, HOA, and stereo format signals for fade-out. In one aspect, the cross-fade using the stream of IFFFs can be performed between transport channels as output of a baseline decoder, between spatial decompressed signals as output of a spatial decoder, or between loudspeaker signals as output of a renderer.
[0012] In one aspect, the cross-fade of two streams can be performed using an overlap-add synthesis technique such as that used by a modified discrete cosine transform (MDCT). Instead of using an IFFF for the transition frame, the time-domain aliasing cancellation (TDAC) of the MDCT can be used as an implicit fade-in and fade-out frame for the spatial mix of streams. In one aspect, the implicit spatial mix of streams with TDAC of the MDCT can be performed before spatial decoding, between transport channels as output of a baseline decoder.
[0013] In one aspect, a method for decoding audio content represented by an adaptive number of scene elements for different content types to perform a cross-fade of content types is disclosed. The method includes receiving frames of the audio content. The audio content is represented by one or more content types such as channel, object, HOA, stereo-based signals, etc. The frames contain an audio stream encoded using an adaptive number of scene elements in one or more content types. The method also includes processing two consecutive frames to generate decoded audio streams for the two consecutive frames, the two consecutive frames containing an audio stream encoding different mixes of an adaptive number of scene elements in one or more content types. The method also includes performing a cross-fade of the decoded audio streams in the two consecutive frames based on a loudspeaker configuration for driving a plurality of loudspeakers. In one aspect, the output of the cross-fade can be provided to headphones or used for applications such as binaural rendering.
[0014] The above summary does not include an exhaustive list of all aspects of the present application. It is contemplated that the application includes all systems and methods that can be practiced from all suitable combinations of the various aspects summarized above, as well as those disclosed in the following detailed description and particularly pointed out in the claims filed with the application. Such combinations have particular advantages not specifically recited in the above summary. BRIEF DESCRIPTION OF DRAWINGS
[0015] Aspects of the disclosure are illustrated by way of example and not by way of limitation in the figures of the accompanying drawings in which like references indicate similar elements. It should be noted that "a" or "one" aspect as referred to herein does not necessarily pertain to the same aspect, and that "a" or "one" aspect can refer to at least one. Additionally, to be succinct, and to reduce the overall number of drawings, a given drawing can illustrate features of more than one aspect of the disclosure, and not every element in such a drawing can be necessary for a given aspect.
[0016] Figure 1 is a functional block diagram of a hierarchical spatial resolution codec that adaptively adjusts encoding of immersive audio content as a target bitrate changes, in accordance with one aspect of the disclosure.
[0017] Figure 2 An audio decoding architecture is shown that decodes and renders a bitstream based on a fixed loudspeaker configuration, in accordance with one aspect of the disclosure, such that cross-fade of the bitstream between successive frames can be performed in the same loudspeaker layout, the bitstream representing an audio scene using different mixtures of encoded content types.
[0018] Figure 3 A functional block diagram of two audio decoders implementing the audio decoding architecture of Figure 2 to perform spatial mixing with redundant frames is shown, in accordance with one aspect of the disclosure.
[0019] Figure 4 An audio decoding architecture is shown that decodes a bitstream such that cross-fade of the bitstream between successive frames can be performed in channel, object, HOA, and stereo format signals in one device, and the output of the cross-fade can be transmitted to multiple devices for rendering, the bitstream representing an audio scene using different mixtures of encoded content types, in accordance with one aspect of the disclosure.
[0020] Figure 5 An audio decoding architecture is shown that decodes a bitstream in one device, and the decoded output can be transmitted to multiple devices for cross-fade of the bitstream between successive frames in channel, object, HOA, and stereo format signals and for rendering, the bitstream representing an audio scene using different mixtures of encoded content types, in accordance with one aspect of the disclosure.
[0021] Figure 6An audio decoding architecture is shown that decodes a bitstream in one device and can transmit the decoded output to multiple devices for rendering and then cross-fade between successive frames of the bitstream in the respective speaker layouts of the multiple devices using different mixtures of encoded content types to represent the audio scene, in accordance with one aspect of the disclosure.
[0022] Figure 7A Cross-fading of two streams using an immediate fade-in and fade-out frame (IFFF) that not only contains a current frame of the bitstream encoded using a mixture of signals for the fade-in, but also a previous frame of the bitstream encoded using a different mixture of signals for the fade-out, in accordance with one aspect of the disclosure, is shown, where the IFFF can be an independent frame.
[0023] Figure 7B Cross-fading of two streams using an IFFF, where the IFFF can be a predictive coded frame, in accordance with one aspect of the disclosure, is shown.
[0024] Figure 8 A functional block diagram of an audio decoder implementing the audio decoding architecture of Figure 6 to perform spatial mixing with the IFFF, in accordance with one aspect of the disclosure, is shown.
[0025] Figure 9A Cross-fading of two streams using an IFFF based on an overlap-add synthesis technique, such as time-domain aliasing cancellation (TDAC) of modified discrete cosine transform (MDCT), in accordance with one aspect of the disclosure, is shown.
[0026] Figure 9B Cross-fading of two streams using an IFFF that spans N frames of the two streams, in accordance with one aspect of the disclosure, is shown.
[0027] Figure 10 A functional block diagram of an audio decoder implementing the audio decoding architecture of Figure 6 to perform implicit spatial mixing with the TDAC of MDCT, in accordance with one aspect of the disclosure, is shown.
[0028] Figure 11An audio decoding architecture is shown that, according to one aspect of the disclosure, performs cross-fade between successive frames of output of a baseline decoder prior to spatial decoding, enabling a mixer in one or more devices to render cross-faded signals in channel, object, HOA, and stereo formats based on their respective speaker layouts, the bitstream representing an audio scene using different mixes of encoded content types.
[0029] Figure 12 A functional block diagram of two audio decoders is shown that, according to one aspect of the disclosure, implement the audio decoding architecture of Figure 11 to perform spatial mixing of redundant frames between transport channels of output of a baseline decoder.
[0030] Figure 13 A functional block diagram of an audio decoder is shown that, according to one aspect of the disclosure, implement the audio decoding architecture of Figure 11 to perform spatial mixing of IFFF between transport channels of output of a baseline decoder.
[0031] Figure 14 A functional block diagram of an audio decoder is shown that, according to one aspect of the disclosure, implement the audio decoding architecture of Figure 11 to perform implicit spatial mixing of TDAC of MDCT between transport channels of output of a baseline decoder.
[0032] Figure 15 is a flowchart of a method of decoding an audio stream to perform cross-fade of content types in the audio stream, the audio stream representing an audio scene by an adaptive number of scene elements for different content types, according to one aspect of the disclosure. DETAILED DESCRIPTION
[0033] It is desirable to provide immersive audio content from an audio source to a playback system over a transport channel while maintaining the best audio quality possible. When the bandwidth of the transport channel changes due to changing channel conditions or changing target bitrates of the playback system, the encoding of the immersive audio content can adapt to improve the tradeoff between audio playback quality and bandwidth. The immersive audio content can include multi-channel audio, audio objects, or spatial audio reconstruction (known as Ambisonics), which describes a soundfield based on spherical harmonics that can be used to recreate the soundfield for playback. Ambisonics can include first order or higher order spherical harmonics, also known as higher order Ambisonics (HOA). The immersive audio content can be adaptively encoded into audio content of different bitrates and spatial resolutions according to the target bitrates and priority ordering of channels, objects, and HOA. The adaptively encoded audio content and its metadata can be transmitted over the transport channel to allow one or more decoders with varying target bitrates to recreate the immersive audio experience.
[0034] Systems and methods for audio decoding techniques that decode immersive audio content encoded with an adaptive number of scene elements for channels, audio objects, HOA, and / or other soundfield representations such as STIC encoding are disclosed. The decoding techniques can present the decoded audio to a speaker configuration of a playback device. For bitstreams that represent an audio scene using different mixtures of channels, objects, HOA, or stereo-based signals received in consecutive frames, a fade-in of a new frame and a fade-out of an old frame can be performed. Cross-fade between consecutive frames encoded using different mixtures of content types can occur between transport channels that are output of a baseline decoder, between spatial decompressed signals that are output of a spatial decoder, or between speaker signals that are output of a renderer.
[0035] In one aspect, techniques for cross-fading between consecutive frames encoded using different mixtures of channels, objects, HOA, or stereo-based signals can use an immediate fade-in and fade-out frame (IFFF) for transition frames. The IFFF can contain a bitstream of a current frame for fade-in and a bitstream of a previous frame for fade-out to eliminate redundant frames needed for cross-fade. In one aspect, cross-fade can use an overlap-add synthesis technique such as time-domain aliasing cancellation (TDAC) of MDCT without using explicit IFFFs. Advantageously, spatial mixing of audio streams using the disclosed cross-fade techniques can eliminate spatial artifacts associated with cross-fade and can reduce computational complexity, latency, and the number of decoders for decoding immersive audio content encoded with an adaptive number of scene elements for channels, audio objects, and / or HOA.
[0036] The following description sets forth numerous specific details. It should be understood, however, that aspects of the disclosure might be practiced without these specific details. In other instances, well-known circuits, structures, and techniques have not been shown in detail in order not to obscure the understanding of this description.
[0037] The terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting of the disclosure. Spatially relative terms, such as "beneath", "below", "lower", "above", "upper", and the like, can be used herein for ease of description to describe one element's or feature's relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if a device including an element is turned over, elements described as "below" or "beneath" other elements or features would then be oriented "above" the other elements or features. Thus, the exemplary term "below" can encompass both an orientation of above and below. The device can be otherwise oriented (e.g., rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly.
[0038] As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and "comprising", when used in this specification, specify the presence of stated features, steps, operations, elements, or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, or groups thereof.
[0039] The terms "or" and "and / or" as used herein are to be interpreted as inclusive or meaning any one or any combination. Therefore, "A, B or C" or "A, B and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B and C." An exception to this definition will occur only when two elements are directly exclusive with each other according to their definitions. That is, the definition "A or B" does not allow the inclusion of both A and B, unless specifically stated otherwise.
[0040] Figure 1is a functional block diagram of a hierarchical spatial resolution codec that adaptively adjusts encoding of immersive audio content as a target bitrate changes according to an aspect of the present disclosure. Immersive audio content 111 can include various immersive audio input formats, also referred to as soundfield representations, such as multi-channel audio, audio objects, HOA, dialogue. In the case of multi-channel input, there can be M channels of a known input channel layout, such as a 7.1.4 layout (7 speakers located in a mid-plane, 4 speakers located in an upper plane, 1 low frequency effects (LFE) speaker). It should be understood that HOA can also include first order ambisonics (FOA). In the following description of the adaptive encoding technique, audio objects can be similarly treated as channels, and for simplicity, channels and objects can be grouped together in the operation of the hierarchical spatial resolution codec.
[0041] An audio scene of immersive audio content 111 can be represented by a plurality of channels / objects 150, HOA 154, and dialogue 158 accompanied by channel / object metadata 151, HOA metadata 155, and dialogue metadata 159, respectively. The metadata can be used to describe properties of the associated soundfield, such as a layout configuration or direction parameters of the associated channels, or a position, size, direction, or spatial image parameters of the associated objects or HOA, to help a renderer achieve a desired source image or recreate a perceived location of a dominant sound. To allow the hierarchical spatial resolution codec to improve the tradeoff between spatial resolution and target bitrate, the channels / objects and HOA can be ordered such that higher ordered channels / objects and HOA are spatially encoded to maintain a higher quality soundfield representation, while lower ordered channels / objects and HOA can be converted and spatially encoded to a lower quality soundfield representation as the target bitrate decreases.
[0042] The channel / object prioritization decision module 121 can receive channels / objects 150 and channel / object metadata 151 of an audio scene to provide a prioritization 162 of the channels / objects 150. In one aspect, the prioritization 162 can be rendered based on spatial salience of the channels and objects, such as location, direction, movement, density, etc. of the channels / objects 150. For example, channels / objects with greater movement near the perceived location of a dominant sound can be more spatially salient, and thus can be ordered higher than channels / objects with less movement away from the perceived location of the dominant sound. To minimize degradation in overall audio quality of the channels / objects as the target bit rate is reduced, the audio quality of the channels / objects represented as higher ordered can be maintained, while the audio quality of the lower ordered channels / objects can be reduced. In one aspect, the channel / object metadata 151 can provide information to guide the channel / object prioritization decision module 121 to render the prioritization 162. For example, the channel / object metadata 151 can include priority metadata for ordering certain channels / objects 150 provided through human input. In one aspect, the channels / objects 150 and channel / object metadata 151 can pass through the channel / object prioritization decision module 121 as channels / objects 160 and channel / object metadata 161, respectively.
[0043] The channel / object spatial encoder 131 can spatially encode the channels / objects 160 and channel / object metadata 161 based on the channel / object prioritization 162 and the target bit rate 190 to generate a channel / object audio stream 180 and associated metadata 181. For example, for a highest target bit rate, all channels / objects 160 and metadata 161 can be spatially encoded into the channel / object audio stream 180 and channel / object metadata 181 to provide the highest audio quality of the resulting transport stream. The target bit rate can be determined by the channel conditions of the transmission channel or the target bit rate of the decoding device. In one aspect, the channel / object spatial encoder 131 can transform the channels / objects 160 into the frequency domain to perform the spatial encoding. The number of frequency subbands and quantization of the encoding parameters can be adjusted according to the target bit rate 190. In one aspect, the channel / object spatial encoder 131 can cluster the channels / objects 160 and metadata 161 to accommodate a reduced target bit rate 190.
[0044] In one aspect, when the target bit rate 190 is reduced, the channels / objects 160 and metadata 161 with lower priority ordering can be converted to another content type and spatially encoded using another encoder to generate a lower quality transport stream. The channel / object spatial encoder 131 can not encode these lower ordered channels / objects 160 as low priority channels / objects 170 and associated metadata 171 that are output. The HOA conversion module 123 can convert the low priority channels / objects 170 and associated metadata 171 to HOA 152 and associated metadata 153. As the target bit rate 190 is gradually reduced, progressively more channels / objects 160 and metadata 161 starting from the lowest priority ordering 162 can be output as low priority channels / objects 170 and associated metadata 171 to be converted to HOA 152 and associated metadata 153. The HOA 152 and associated metadata 153 can be spatially encoded to generate a lower quality transport stream compared to a transport stream that fully encodes all channels / objects 160 but has the advantage of requiring a lower bit rate and lower transmission bandwidth.
[0045] There can be multiple tiers for converting and encoding the channels / objects 160 to another content type to accommodate a lower target bit rate. In one aspect, some low priority channels / objects 170 and associated metadata 171 can be encoded using parametric encoding such as a stereo based immersive coding (STIC) encoder 137. The STIC encoder 137 can render a two-channel stereo audio stream 186 from the immersive audio signals, such as by channel downmixing or rendering objects or HOA to a stereo signal. The STIC encoder 137 can also generate metadata 187 based on a perceptual model that derives parameters describing the perceptual direction of the dominant sound. By converting and encoding some channels / objects to a stereo audio stream 186 instead of HOA, further reduction in bit rate can be accommodated, even at a lower quality transport stream. While the STIC encoder 137 is described as rendering channels, objects, or HOA to a two-channel stereo audio stream 186, the STIC encoder 137 is not so limited and can render channels, objects, or HOA to an audio stream of more than two channels.
[0046] In one aspect, at a medium target bit rate, some low priority channels / objects 170 and their associated metadata 171 with the lowest priority ranking can be encoded into a stereo audio stream 186 and associated metadata 187. The remaining low priority channels / objects 170 and their associated metadata with a higher priority ranking can be converted into HOA 152 and associated metadata 153, which can be prioritized and encoded into a HOA audio stream 184 and associated metadata 185 along with other HOA 154 and associated metadata 155 from immersive audio content 111. The remaining channels / objects 160 and their metadata with the highest priority ranking are encoded into a channels / objects audio stream 180 and associated metadata 181. In one aspect, at a minimum target bit rate, all channels / objects 160 can be encoded into a stereo audio stream 186 and associated metadata, leaving no encoded channels, objects, or HOA in the transport stream.
[0047] Similar to channels / objects, HOAs can also be ranked such that higher ranked HOAs are spatially encoded to maintain a higher quality soundfield representation of the HOAs, while lower ranked HOAs are rendered into a lower quality soundfield representation, such as a stereo signal. HOA priority decision module 125 can receive HOAs 154 and associated metadata 155 of a soundfield representation of an audio scene from immersive audio content 111, and converted HOAs 152 and associated metadata 153 that have been converted from low priority channels / objects 170 to provide a priority ranking 166 between the HOAs. In one aspect, the priority ranking can be presented based on spatial saliency of the HOAs, such as position, direction, movement, density, etc. of the HOAs. To minimize degradation of overall audio quality of the HOAs as the target bit rate is reduced, the audio quality of higher ranked HOAs can be maintained, while the audio quality of lower ranked HOAs can be reduced. In one aspect, HOA metadata 155 can provide information to guide HOA priority decision module 125 to present HOA priority ranking 166. HOA priority decision module 125 can combine HOAs 154 from immersive audio content 111 and converted HOAs 152 that have been converted from low priority channels / objects 170 to generate HOAs 164, and combine associated metadata of the combined HOAs to generate HOA metadata 165.
[0048] The hierarchical HOA spatial encoder 135 can spatially encode the HOA 164 and the HOA metadata 165 based on the HOA prioritization 166 and the target bitrate 190 to generate the HOA audio stream 184 and the associated metadata 185. For example, for a high target bitrate, all of the HOA 164 and the HOA metadata 165 can be spatially encoded into the HOA audio stream 184 and the HOA metadata 184 to provide a high quality transport stream. In one aspect, the hierarchical HOA spatial encoder 135 can convert the HOA 164 into a frequency domain to perform the spatial encoding. The number of frequency subbands and the quantization of the encoding parameters can be adjusted according to the target bitrate 190. In one aspect, the hierarchical HOA spatial encoder 135 can cluster the HOA 164 and the HOA metadata 165 to accommodate a reduced target bitrate 190. In one aspect, the hierarchical HOA spatial encoder 135 can perform compression techniques to generate adaptive order HOA 164.
[0049] In one aspect, as the target bitrate 190 decreases, the HOA 164 and metadata 165 with lower prioritization can be encoded as stereo signals. The hierarchical HOA spatial encoder 135 can not encode these lower prioritized HOA that are output as low priority HOA 174 and associated metadata 175. As the target bitrate 190 gradually decreases, progressively more HOA 164 and HOA metadata 165 starting from the lowest prioritization 166 can be output as low priority HOA 174 and associated metadata 175 to be encoded into stereo audio streams 186 and associated metadata 187. The stereo audio streams 186 and associated metadata 187 require lower bitrates and lower transmission bandwidths compared to the transport stream that fully encodes all of the HOA 164, even at lower audio quality. Thus, as the target bitrate 190 decreases, the transport stream for an audio scene can have a greater mixture of hierarchical structures of content types with lower audio quality. In one aspect, the hierarchical mixture of content types can adaptively change from scene to scene, frame to frame, or packet to packet. Advantageously, the hierarchical spatial resolution codec adaptively adjusts the hierarchical encoding of immersive audio content to generate a varying mixture of channels, objects, HOA, and stereo signals based on the target bitrate and the prioritization of components of the soundfield representation, thereby improving the tradeoff between audio quality and target bitrate.
[0050] In one aspect, the audio scene of immersive audio content 111 can contain dialogue 158 and associated metadata 159. Dialogue spatial encoder 139 can encode dialogue 158 and associated metadata 159 based on target bitrate 190 to generate speech stream 188 and speech metadata 189. In one aspect, when target bitrate 190 is high, dialogue spatial encoder 139 can encode dialogue 158 into a two-channel speech stream 188. When target bitrate 190 is reduced, dialogue 158 can be encoded into a one-channel speech stream 188.
[0051] Baseline encoder 141 can encode channel / object audio stream 180, HOA audio stream 184, and stereo audio stream 186 into audio stream 191 based on target bitrate 190. Baseline encoder 141 can use any known encoding technique. In one aspect, baseline encoder 141 can adapt the rate and quantization of the encoding to target bitrate 190. Speech encoder 143 can separately encode speech stream 188 for audio stream 191. Channel / metadata 181, HOA metadata 185, stereo metadata 187, and speech metadata 189 can be combined into a single transport channel of audio stream 191. Audio stream 191 can be transmitted through the transport channel to allow one or more decoders to reconstruct immersive audio content 111.
[0052] Figure 2 An audio decoding architecture is shown that decodes and renders a bitstream based on a fixed speaker configuration according to one aspect of the disclosure, such that cross-fade between successive frames of the bitstream can be performed in the same speaker layout, which represents an audio scene using different mixtures of encoded content types. Three packets are received by a packet receiver. Packets 1, 2, and 3 can contain bitstreams encoded at 1000 kbps (16 objects), 512 kbps (4 objects + 8 HOA), and 64 kbps (2 STIC), respectively. The channel / object, HOA, and stereo-based parametric encoded frame-by-frame audio bitstreams can be decoded using a channel / object spatial decoder / renderer, a spatial HOA decoder / renderer, and a stereo decoder / renderer, respectively. The decoded bitstreams can be presented to a speaker configuration (e.g., 7.1.4 of a user device).
[0053] If the new packet contains a different mix of channels, objects, HOA, and stereo-based signals than the previous packet, the new packet can be faded in and the old packet can be faded out. In an overlap period for cross-fade, the same soundfield can be represented by two different mixes of channels, objects, HOA, and stereo-based signals. For example, at frame #9, the same audio scene is represented by 4 objects + 8 HOA or 2 STICs. In a 7.1.4 speaker domain, the 4 objects + 8 HOA of the old packet can be faded out and the 2 STICs of the new packet can be faded in.
[0054] Figure 3 A functional block diagram of two audio decoders implementing the audio decoding architecture of Figure 2 is shown according to one aspect of the disclosure to perform spatial mixing with redundant frames. Packet 1 (301) contains frames 1-4. Each frame in packet 1 includes a plurality of objects and HOA. Packet 2 (302) contains frames 3-6. Each frame in packet 2 includes a plurality of objects and STIC signals. The two packets contain a bitstream that can represent one or more audio scenes using an adaptive number of scene elements encoding for channels, objects, HOA, and STIC encoding. The two packets contain overlapping and redundant frames 3-4 that represent an overlap period for cross-fade. A baseline decoder 309 of a first audio decoder performs baseline decoding of the bitstream in frames 1-4 of packet 1 (301). A baseline decoder 359 of a second audio decoder performs baseline decoding of the bitstream in frames 3-6 of packet 2 (302).
[0055] An object spatial decoder 303 of the first audio decoder decodes the encoded objects in frames 1-4 of packet 1 (301) into a N1 number of decoded objects 313. An object renderer 323 in the first audio decoder renders the N1 decoded objects 313 into a speaker configuration (e.g., 7.1.4) of the first audio decoder. The rendered objects can be represented by an O1 number of speaker outputs 333.
[0056] An HOA spatial decoder 305 in the first audio decoder decodes the encoded HOA in frames 1-4 of packet 1 (301) into a N2 number of decoded HOA 315. An HOA renderer 325 in the first audio decoder renders the N2 decoded HOA 315 into a speaker configuration. The rendered HOA can be represented by an O1 number of speaker outputs 335. The rendered objects in the O1 number of speaker outputs 333 and the rendered HOA in the O1 number of speaker outputs 335 can be faded out using a fade-out window 309 at frame 4 to generate a speaker output containing O1 objects 343 and O1 HOA 345.
[0057] Accordingly, the object spatial decoder 353 of the second audio decoder decodes the encoded objects in frames 3-6 of packet 2 (302) into N3 number of decoded objects 363. The object renderer 373 in the second audio decoder renders the N3 decoded objects 363 into the same loudspeaker configuration as the same audio decoder. The rendered objects can be represented by Ol number of loudspeaker outputs 383.
[0058] The STIC decoder 357 in the second audio decoder decodes the encoded STIC signals in frames 3-6 of packet 2 (302) into decoded STIC signals 367. The STIC renderer 377 in the second audio decoder renders the decoded STIC signals 367 into the loudspeaker configuration. The rendered STIC signals can be represented by Ol number of loudspeaker outputs 387. The rendered objects in Ol number of loudspeaker outputs 383 and the rendered STIC signals in Ol number of loudspeaker outputs 387 can be faded in at frame 4 using a fade-in window 359 to generate the loudspeaker outputs containing Ol objects 393 and Ol STIC signals 397. The mixer can mix the loudspeaker outputs containing Ol objects 343 and Ol HOA 345 of frames 1-4 with the loudspeaker outputs containing Ol objects 393 and Ol STIC signals 397 of frames 4-6 to generate Ol loudspeaker outputs 350 with cross-fade occurring at frame 4. Thus, the cross-fade of objects, HOA, and STIC signals is performed in the same loudspeaker layout.
[0059] Figure 4 An audio decoding architecture is shown that decodes a bitstream such that a cross-fade of the bitstream between consecutive frames can be performed in the signals of channels, objects, HOA, and stereo format in one device, and the output of the cross-fade can be transmitted to multiple devices for rendering, the bitstream representing an audio scene using different mix of encoded content types. Packets 1, 2, and 3 can contain the same encoded bitstream as in Figure 2
[0060] The frame-by-frame audio bitstreams of channel / object, HOA, and STIC signals can be decoded using a channel / object spatial decoder, a spatial HOA decoder, and a STIC decoder, respectively. For example, the spatial HOA decoder can decode the spatially compressed representation of the HOA signal into HOA coefficients. The HOA coefficients can then be rendered. Prior to rendering, the decoded bitstreams can cross-fade at frame #9 of the spatially decoded channel / object, HOA, and STIC signals. A mixer in the same playback device as the audio decoder or in another playback device can render the cross-faded channel / object, HOA, and STIC signals based on their respective speaker layouts. In one aspect, the cross-faded output of the audio decoder can be compressed into a bitstream and transmitted to the other playback device where the bitstream is decompressed and provided to the mixer for rendering based on its respective speaker layout. In one aspect, the output of the audio decoder can be stored as a file for future rendering.
[0061] Figure 5 An audio decoding architecture is shown that decodes a bitstream in one device and can transmit the decoded output to multiple devices for cross-fade between consecutive frames of the bitstream in channel, object, HOA, and stereo format signals and rendering using different mixtures of encoded content types to represent an audio scene, according to one aspect of the disclosure. Packets 1, 2, and 3 can contain the same encoded bitstream as in Figure 2 and Figure 4 .
[0062] The frame-by-frame audio bitstreams of channel / object, HOA, and STIC signals can be decoded using a channel / object spatial decoder, a spatial HOA decoder, and a STIC decoder, respectively. Prior to rendering, a mixer in the same playback device as the decoders can perform a cross-fade between the spatially decompressed signals that are the output of the spatial decoders. The mixer can then render the cross-faded channel, object, HOA, and stereo format signals based on the speaker layout. In one aspect, the output of the audio decoders can be compressed into a bitstream and transmitted to the other playback device where the bitstream is decompressed and provided to the mixer for cross-fade and rendering based on its respective speaker layout. In one aspect, the output of the audio decoders can be stored as a file for future rendering.
[0063] Figure 6An audio decoding architecture is shown that decodes bitstreams in one device and can transmit the decoded output to multiple devices for rendering and then cross-fade between successive frames of the bitstream in the respective speaker layouts of the multiple devices that represent an audio scene using different mixtures of encoded content types. Packets 1, 2, and 3 can contain the same encoded bitstream as in Figure 2 , Figure 4 and Figure 5 .
[0064] Frame-by-frame audio bitstreams of channel / object, HOA, and STIC signals can be decoded using a channel / object spatial decoder, a spatial HOA decoder, and a STIC decoder, respectively. A mixer in the same playback device as the decoders can render the decoded bitstreams based on the speaker configuration. The mixer can perform cross-fading between channels, objects, HOA, and STIC signals among the speaker signals. In one aspect, the output of the audio decoders can be compressed into bitstreams and transmitted to other playback devices where the bitstreams are decompressed and provided to mixers for rendering based on their respective speaker layouts and cross-fading. In one aspect, the output of the audio decoders can be stored as files for future rendering.
[0065] Figure 7A Cross-fading of two streams using an immediate fade-in and fade-out frame (IFFF) is shown according to one aspect of the disclosure, which not only contains a bitstream of a current frame encoded using a mixture of channel, object, HOA, and stereo format signals for fade-in, but also a bitstream of a previous frame encoded using a different mixture of channel, object, HOA, and stereo format signals for fade-out, where the IFFF can be an independent frame.
[0066] For immediate fade-in and fade-out of two different streams, a transition frame can start from the IFFF. The IFFF can contain a bitstream of a current frame for fade-in and a bitstream of a previous frame for fade-out to eliminate redundant frames for cross-fade, such as overlap and redundant frames used in Figure 3 . If the IFFF is encoded as an independent frame (I-frame), it can be decoded immediately. However, if it is encoded using predictive encoding (P-frame), the bitstream of the previous frame needs to be decoded. In this case, the IFFF can contain these redundant previous frames starting from an I-frame.
[0067] Figure 7B Cross-fading of two streams using an IFFF is shown according to one aspect of the disclosure, where the IFFF can be a predictive encoded frame. The IFFF can contain redundant previous frames 2-3 in frame 2 starting from an I-frame, since frame 3 is also a predictive encoded frame.
[0068] Figure 8 A functional block diagram of an audio decoder implementing the audio decoding architecture of Figure 6 FIG. 1 illustrates an audio decoder according to one aspect of the disclosure. The audio decoder implements the audio decoding architecture of FIG. 1 to perform spatial mixing with IFFF. Group 1 (801) contains frames 1-4. Each frame in group 1 (801) includes multiple objects and HOA. Group 2 (802) contains frames 5-8. Each frame in group 2 (802) includes multiple objects and STIC signals. The two groups contain a bitstream that can represent one or more audio scenes using an adaptive number of scene elements encoding for channels, objects, HOA, and STIC encoding. The first frame or frame 5 of group 2 (802) is an IFFF that represents a transition frame for cross-fade. A baseline decoder 809 of the audio decoder performs baseline decoding of the bitstream in the two groups.
[0069] An object spatial decoder 803 of the audio decoder decodes the encoded objects in frames 1-4 of group 1 (801) and frames 5-8 of group 2 (802) into N1 number of decoded objects 813. An object renderer 823 renders the N1 decoded objects 813 into a speaker configuration (e.g., 7.1.4) of the audio decoder. The rendered objects can be represented by O1 number of speaker outputs 833.
[0070] An HOA spatial decoder 805 in the audio decoder decodes the encoded HOA in frames 1-4 of group 1 (801) and the IFFF of group 2 (802) into N2 number of decoded HOA 815. An HOA renderer 825 renders the N2 decoded HOA 815 into a speaker configuration. The rendered HOA can be represented by O1 number of speaker outputs 835.
[0071] An STIC decoder 807 in the audio decoder decodes the encoded STIC signals in the IFFF (frame 5) of group 2 (802) and the remaining frames 6-8 of group 2 (802) into decoded STIC signals 817. An STIC renderer 827 renders the decoded STIC signals 817 into a speaker configuration. The rendered STIC signals can be represented by O1 number of speaker outputs 837. A cross-fade window 809 performs a cross-fade of the speaker outputs containing O1 objects 833, O1 HOA 835, and O1 STIC signals 837 to generate O1 speaker outputs 850 where the cross-fade occurs at frame 5. Thus, the cross-fade of the objects, HOA, and STIC signals is performed in the same speaker layout.
[0072] Because the IFFF contains the bitstream for the current frame for the fade-in and the bitstream for the previous frame for the fade-out, it eliminates the dependency on redundant frames for cross-fade, such as Figure 3 the overlapping and redundant frames used in Figure 3 Compared to the two audio decoders of Figure 4 and Figure 5 another advantage of using IFFF for cross-fade includes reduced latency and the ability to use only one audio decoder. In one aspect, cross-fade between objects, HOA, and STIC signals of consecutive frames using IFFF can be performed in signals in channel, object, HOA, and stereo formats, such as the audio decoding architecture shown in and
[0073] Figure 9A Cross-fade of two streams using IFFF based on overlapping-add synthesis techniques, such as time-domain aliasing cancellation (TDAC) of modified discrete cosine transform (MDCT), is shown according to one aspect of the disclosure. For a new packet fade-in, if TDAC of MDCT is needed, one more redundant frame is added in the IFFF. For example, to obtain decoded audio output for frame 4, MDCT coefficients for frame 3 are needed. However, because frame 3 is a P-frame, not only frame 3 but also frame 2 as an I-frame is added in the IFFF.
[0074] Figure 9B Cross-fade of two streams using IFFF spanning N frames of the two streams is shown according to one aspect of the disclosure. If N frames are used for cross-fade, the bitstream representing N frames of the previous and current packets is contained in the IFFF.
[0075] Figure 10 A functional block diagram of an audio decoder implementing the audio decoding architecture of Figure 6 to perform implicit spatial mixing with TDAC of MDCT is shown according to one aspect of the disclosure. Packet 1 (1001) contains frames 1-4. Each frame in packet 1 (1001) includes multiple objects and HOA. Packet 2 (1002) contains frames 5-8. Each frame in packet 2 (1002) includes multiple objects and STIC signals. The two packets contain bitstreams that can represent one or more audio scenes using an adaptive number of scene elements coding for channel, object, HOA, and STIC coding. The first frame or frame 5 of packet 2 (1002) is an implicit IFFF that represents a transition frame for cross-fade based on TDAC of MDCT. A baseline decoder 1009 of the audio decoder performs baseline decoding of the bitstreams in the two packets.
[0076] The object spatial decoder 1003 of the audio decoder decodes the encoded objects in frames 1-4 of packet 1 (1001) and frames 5-8 of packet 2 (1002) into N1 number of decoded objects 1013. The object renderer 1023 renders the N1 decoded objects 1013 into the speaker configuration of the audio decoder (e.g., 7.1.4). The rendered objects can be represented by O1 number of speaker outputs 1033.
[0077] The HOA spatial decoder 1005 in the audio decoder decodes the encoded HOAs in frames 1-4 of packet 1 (1001) and the implicit IFFF in frames 5-8 of packet 2 (1002) into N2 number of decoded HOAs 1015. The HOA renderer 1025 renders the N2 decoded HOAs 1015 into the speaker configuration. The rendered HOAs can be represented by O1 number of speaker outputs 1035.
[0078] The STIC decoder 1007 in the audio decoder decodes the encoded STIC signals in frames 5-8 of packet 2 (802) into decoded STIC signals 1017. The STIC signals 1017 include MDCT TDAC windows starting at frame 5. The STIC renderer 1027 renders the decoded STIC signals 1017 into the speaker configuration. The rendered STIC signals can be represented by O1 number of speaker outputs 1037. The implicit fade-in and fade-out at frame 5 introduced by the MDCT TDAC performs a crossfade of the O1 speaker outputs 1033 of the objects, 1035 of the HOAs, and 1037 of the STIC signals to generate O1 speaker outputs 1050 where the crossfade occurs at frame 5. Thus, the crossfade of the objects, HOAs, and STIC signals is performed in the same speaker layout. Advantages of using the TDAC of the MDCT as the implicit IFFF for the crossfade include eliminating the reliance on redundant frames for the crossfade and the reduced delay of the audio decoding because the windowing function is already introduced by the TDAC. Figure 3 The ability to use only one audio decoder compared to two audio decoders of. Because the TDAC already introduces the windowing function, the crossfade speaker outputs of the current and future frames can be performed by simple addition without the need for explicit fade-in and fade-out windows, thus reducing the delay of the audio decoding.
[0079] Figure 11 An audio decoding architecture is shown that performs a crossfade of a bitstream between consecutive frames that are output of a baseline decoder prior to spatial decoding, such that a mixer in one or more devices can render the crossfaded channels, objects, HOAs, and stereo format signals based on their respective speaker layouts, the bitstream representing an audio scene using different mixes of encoded content types, in accordance with one aspect of the disclosure. Packets 1, 2, and 3 can contain different mixes of encoded content types.Figure 2 , Figure 4 and Figure 5 the same encoded bitstream.
[0080] At an audio decoder, a bitstream is decoded that represents an audio scene using an adaptive number of scene elements for channel, object, HOA, and / or STIC encoding. Crossfades between previous and current frames can be performed between transport channels as output of a baseline decoder and before spatial decoding and rendering to reduce computational complexity. A channel / object spatial decoder, a spatial HOA decoder, and a STIC decoder can spatially decode the crossfaded channel / object, HOA, and STIC signals, respectively. A mixer can render the decoded and crossfaded bitstream based on a loudspeaker configuration. In one aspect, the output of the audio decoder can be compressed into a bitstream and transmitted to other playback devices where it is decompressed and provided to a mixer for rendering based on their respective loudspeaker layouts. In one aspect, the output of the audio decoder can be stored as a file for future rendering. If the number of transport channels is low compared to the number of channel / object, HOA, and STIC signals after spatial decoding, it can be advantageous to perform crossfading of the bitstream in consecutive frames between transport channels as output of a baseline decoder.
[0081] Figure 12 A functional block diagram of two audio decoders implementing the audio decoding architecture of Figure 11 to perform spatial mixing of redundant frames between transport channels as output of a baseline decoder is shown in accordance with one aspect of the disclosure. Group 1 (1201) contains frames 1-4. Each frame in group 1 includes multiple objects and HOA. Group 2 (1202) contains frames 3-6. Each frame in group 2 includes multiple objects and STIC signals. The two groups contain a bitstream that can represent one or more audio scenes using an adaptive number of scene elements encoding for channel, object, HOA, and STIC encoding. The two groups contain overlapping and redundant frames 3-4 that represent an overlap period for crossfading.
[0082] The baseline decoder 1203 of the first audio decoder decodes packet 1 (1201) into a baseline decoded packet 1 (1205), which can be faded out at frame 4 using a fade-out window 1207 to generate a faded-out packet 1 (1209) between transport channels as output of the baseline decoder before spatial decoding and rendering. The object spatial decoder and renderer 1213 of the first audio decoder spatially decodes the encoded objects in the faded-out packet 1 (1209) and renders the decoded objects into the loudspeaker configuration of the first audio decoder (e.g., 7.1.4). The rendered objects can be represented by the Oi number of loudspeaker outputs 1243. The HOA spatial decoder and renderer 1215 of the first audio decoder spatially decodes the encoded HOA in the faded-out packet 1 (1209) and renders the decoded HOA into the loudspeaker configuration of the first audio decoder. The rendered HOA can be represented by the Oi number of loudspeaker outputs 1245.
[0083] Accordingly, the baseline decoder 1253 of the second audio decoder decodes packet 2 (1202) into a baseline decoded packet 2 (1255), which can be faded out at frame 3 and frame 4 using a fade-out window 1257 to generate a faded-out packet 2 (1259) between transport channels as output of the baseline decoder before spatial decoding and rendering. The object spatial decoder and renderer 1263 of the second audio decoder spatially decodes the encoded objects in the faded-out packet 2 (1259) and renders the decoded objects into the same loudspeaker configuration as the first audio decoder. The rendered objects can be represented by the Oi number of loudspeaker outputs 1293. The STIC decoder and renderer 1267 of the second audio decoder spatially decodes the encoded STIC signal in the faded-out packet 1 (1209) and renders the decoded STIC signal into the loudspeaker configuration. The rendered STIC signal can be represented by the Oi number of loudspeaker outputs 1297. A mixer can mix the loudspeaker outputs containing the Oi objects 1243 and Oi HOA 1245 of frames 1-4 with the loudspeaker outputs containing the Oi objects 1293 and Oi STIC signals 1297 of frames 4-6 to generate the Oi loudspeaker outputs 1250 that cross-fade occurs at frame 4.
[0084] Figure 13 A functional block diagram of an audio decoder implementing the techniques of the present disclosure is shown in accordance with one aspect of the present disclosure. Figure 11audio decoding architecture to perform a spatial mix of IFFF between the transport channels that are the output of the baseline decoder. Packet 1 (1301) contains frames 1-4. Each frame in packet 1 (1301) includes a number of objects and HOA. Packet 2 (1302) contains frames 5-8. Each frame in packet 2 (1302) includes a number of objects and STIC signals. These two packets contain a bitstream that can represent one or more audio scenes using an adaptive number of scene elements coding for channels, objects, HOA, and STIC encoding. The first frame or frame 5 of packet 2 (1302) is an IFFF that represents a transition frame for cross-fade.
[0085] The baseline decoder 1303 of the audio decoder decodes packet 1 (1301) and packet 2 (1302) into a baseline decoded packet 1305. A cross-fade window performs a cross-fade of the baseline decoded packet 1305 to generate a cross-faded packet 1309 between the transport channels that are the output of the baseline decoder prior to spatial decoding and rendering, where the cross-fade occurs at frame 5. If the STIC signals in the IFFF are encoded with a prediction frame, the STIC encoded signals in the IFFF can contain the STIC encoded signals from frame 3 and frame 4 of packet 1 (1301).
[0086] The object spatial decoder and renderer 1313 of the audio decoder spatially decodes the encoded objects in the cross-faded packet 1309 and renders the decoded objects into the speaker configuration of the audio decoder (e.g., 7.1.4). The rendered objects can be represented by the Ol number of speaker outputs 1323. The HOA spatial decoder and renderer 1315 of the audio decoder spatially decodes the encoded HOA in the cross-faded packet 1309 and renders the decoded HOA into the speaker configuration of the audio decoder. The rendered HOA can be represented by the Ol number of speaker outputs 1325. The STIC decoder and renderer 1317 of the audio decoder spatially decodes the encoded STIC signals in the cross-faded packet 1309 and renders the decoded STIC signals into the speaker configuration. The rendered STIC signals can be represented by the Ol number of speaker outputs 1327.
[0087] The mixer can mix the speaker outputs containing Ol objects 1323, Ol HOA 1325, and Ol STIC signals 1327 to generate Ol speaker output signals where the cross-fade occurs at frame 5. Because the IFFF contains a bitstream for the current frame for the fade-in and a bitstream for the previous frame for the fade-out, it eliminates the dependency on redundant frames for cross-fade, such as the overlap and redundant frames used in the Figure 12 Figure 12 Another advantage of using IFFF for cross-fade, as compared to using two audio decoders, includes reduced latency and the ability to use only one audio decoder.
[0088] Figure 14 A functional block diagram of an audio decoder implementing an audio decoding architecture to perform implicit spatial mixing of TDAC of MDCT between transport channels that are output of a baseline decoder is shown in accordance with one aspect of the disclosure. Figure 11 Group 1 (1401) contains frames 1-4. Each frame in group 1 (1401) includes a plurality of objects and HOA. Group 2 (1402) contains frames 5-8. Each frame in group 2 (1402) includes a plurality of objects and STIC signals. The two groups contain a bitstream that can represent one or more audio scenes using an adaptive number of scene elements encoding for channel, object, HOA, and STIC encoding. The first frame or frame 5 of group 2 (1402) is an implicit IFFF that represents a transition frame for cross-fade of TDAC based on MDCT.
[0089] The baseline decoder 1303 of the audio decoder decodes group 1 (1401) and group 2 (1402) into a baseline decoded group 1405. The implicit IFFF in frame 5 of the baseline decoded group 1405 introduced by TDAC of MDCT causes the audio decoder to perform a cross-fade between transport channels that are output of the baseline decoder prior to spatial decoding and rendering of the baseline decoded group 1405, where the cross-fade occurs at frame 5.
[0090] The object spatial decoder and renderer 1313 of the audio decoder spatially decodes encoded objects in the cross-faded group 1405 and renders the decoded objects into a speaker configuration (e.g., 7.1.4) of the audio decoder. The rendered objects can be represented by Ol number of speaker outputs 1423. The HOA spatial decoder and renderer 1315 of the audio decoder spatially decodes encoded HOA in the cross-faded group 1405 and renders the decoded HOA into a speaker configuration of the audio decoder. The rendered HOA can be represented by Ol number of speaker outputs 1425. The STIC decoder and renderer 1317 of the audio decoder spatially decodes encoded STIC signals in the cross-faded group 1405 and renders the decoded STIC signals into a speaker configuration. The rendered STIC signals can be represented by Ol number of speaker outputs 1427.
[0091] The mixer can mix the loudspeaker output comprising Oi objects 1423, Oi HOA 1425, and Oi STIC signals 1427 to generate Oi loudspeaker output signals that cross-fade at frame 5. Advantages of using TDAC of MDCT as implicit IFFF for cross-fade include elimination of the dependency on redundant frames for cross-fade and reduced delay of audio decoding because TDAC already introduces windowing functions. Because the cross-fade of the current and future frames' loudspeaker output can be performed by simple addition without the need for explicit fade-in and fade-out windows, the delay of audio decoding is reduced. Figure 12 The ability to use only one audio decoder compared to two audio decoders of
[0092] Figure 15 is a flowchart of a method 1500 of decoding an audio stream to perform cross-fade of content types in the audio stream, the audio stream representing an audio scene with an adaptive number of scene elements for different content types, according to one aspect of the disclosure. The method 1500 can be practiced by a decoder of Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 8 , Figure 10 , Figure 11 , Figure 12 , Figure 13 or Figure 14 .
[0093] In operation 1501, the method 1500 receives frames of audio content. The audio content is represented by one or more content types such as channels, objects, HOA, stereo-based signals, etc. The frames contain an audio stream that encodes the audio content using an adaptive number of scene elements in the one or more content types. For example, the frames can contain an audio stream that encodes an adaptive number of scene elements for channel / object, HOA, and / or STIC encoding.
[0094] In operation 1503, the method 1500 processes two consecutive frames to generate a decoded audio stream for the two consecutive frames, the two consecutive frames containing an audio stream that encodes the audio content using a different mix of an adaptive number of scene elements in the one or more content types.
[0095] In operation 1505, the method 1500 generates a cross-fade of the decoded audio streams in two consecutive frames based on a speaker configuration for driving a plurality of speakers. For example, the decoded audio stream of an old frame of two consecutive frames can be faded in and the decoded audio stream of a new frame of two consecutive frames can be faded in such that the content type of the cross-fade can be mixed to generate a speaker output signal based on the same speaker configuration. In one aspect, the output of the cross-fade can be provided to headphones or used for applications such as binaural rendering.
[0096] Embodiments of the scalable decoder described herein can be implemented in a data processing system, such as a network computer, a network server, a tablet computer, a smartphone, a laptop computer, a desktop computer, other consumer electronics devices, or other data processing systems. In particular, the described operations for decoding and cross-fading a bitstream that represents an audio scene with an adaptive number of scene elements for channel, object, HOA, and / or STIC encoding are digital signal processing operations performed by a processor executing instructions stored in one or more memories. The processor can read the stored instructions from the memories and execute the instructions to perform the described operations. These memories represent examples of machine-readable non-transitory storage media that can store or contain computer program instructions that, when executed, cause the data processing system to perform one or more of the methods described herein. The processor can be a processor in a local device such as a smartphone, a processor in a remote server, or a distributed processing system of multiple processors in a local device and a remote server, with their respective memories containing portions of the instructions needed to perform the described operations.
[0097] The processes and blocks described herein are not limited to the particular examples described, and are not limited to the particular order described in the examples used herein. Rather, any of the process blocks can be reordered, combined or removed, performed in parallel or in series, as desired, to achieve the results described above. The process blocks associated with implementing an audio processing system can be performed by one or more programmable processors executing one or more computer programs stored on non-transitory computer-readable storage media to perform the functions of the system. All or part of the audio processing system can be implemented as special purpose logic circuitry (e.g., an FPGA (field programmable gate array) and / or an ASIC (application-specific integrated circuit)). All or part of the audio system can be implemented with electronic hardware circuitry that includes at least one of electronic devices such as, for example, a processor, a memory, a programmable logic device, or a logic gate. Additionally, the processes can be implemented in any combination of hardware devices and software components.
[0098] While certain example embodiments have been described and shown in the accompanying drawings, it is to be understood that these embodiments are merely exemplary of the application and are not to be considered as limiting the application, and that the application is not limited to the specific constructions and arrangements shown and described, since various other modifications can occur to those ordinarily skilled in the art. Therefore, the application is not to be restricted or limited by the described embodiments, but is to be accorded the full scope of the claims.
[0099] To assist the Patent Office and any readers of this application in interpreting the claims appended hereto, applicants wish to state that they do not intend any of the appended claims to invoke 35 U.S.C. 112(f) unless the words "means for" are explicitly used in the particular claim.
Claims
1. A method of decoding audio content, the method comprising: receiving, by a decoding device, frames of the audio content, the audio content represented by a plurality of content types, the frames containing an audio stream that encodes the audio content using an adaptive number of scene elements from the plurality of content types; generating a decoded audio stream by processing two consecutive frames, the two consecutive frames containing the audio stream that encodes the audio content using different mixtures of the adaptive number of scene elements from the plurality of content types; and generating a cross-fade of the decoded audio stream in the two consecutive frames based on a speaker configuration for driving a plurality of speakers.
2. The method of claim 1, wherein generating the decoded audio stream comprises: generating, for each frame of the two consecutive frames, a spatially decoded audio stream of the plurality of content types with at least one scene element; and rendering the spatially decoded audio stream of the plurality of content types to generate, for each frame of the two consecutive frames, speaker output signals of the plurality of content types based on the speaker configuration of the decoding device; and wherein generating the cross-fade of the decoded audio stream comprises: generating a cross-fade of the speaker output signals of the plurality of content types from an earlier frame to a later frame of the two consecutive frames; and mixing the cross-fade of the speaker output signals of the plurality of content types to drive the plurality of speakers.
3. The method of claim 2, further comprising: transmitting the spatially decoded audio stream of the plurality of content types and time-synchronized metadata to a second device for rendering based on a speaker configuration of the second device.
4. The method of claim 1, wherein generating the decoded audio stream comprises: generating, for each frame of the two consecutive frames, a spatially decoded audio stream of the plurality of content types with at least one scene element, and wherein generating the cross-fade of the decoded audio stream comprises: generating a cross-fade of the spatially decoded audio stream of the plurality of content types from an earlier frame to a later frame of the two consecutive frames; rendering the cross-fade of the spatially decoded audio stream of the plurality of content types to generate speaker output signals of the plurality of content types based on the speaker configuration of the decoding device; and mixing the speaker output signals of the plurality of content types to drive the plurality of speakers.
5. The method of claim 4, further comprising: transmitting the cross-fade of the spatially decoded audio stream of the plurality of content types and time-synchronized metadata to a second device for rendering based on a speaker configuration of the second device.
6. The method of claim 4, further comprising: transmitting the spatially decoded audio stream of the plurality of content types and time-synchronized metadata to a second device for cross-fading and rendering based on a speaker configuration of the second device.
7. The method of claim 1 or 2 or 4, wherein a later frame of the two consecutive frames includes an immediate fade-in and fade-out frame (IFFF) for generating the crossfaded decoded audio stream, wherein the IFFF contains a bitstream that encodes the audio content of the later frame for immediate fade-in and the audio content of an earlier frame of the two consecutive frames for immediate fade-out.
8. The method of claim 7, wherein generating the decoded audio stream comprises: generating a decoded audio stream having the plurality of content types of at least one scene element for each of the two consecutive frames, wherein the decoded audio streams of the two consecutive frames have a different mix of the adaptive number of scene elements of the plurality of content types, and wherein generating the crossfade of the decoded audio streams of the two consecutive frames comprises: generating a transition frame based on the IFFF, wherein the transition frame includes an immediate fade-in of the decoded audio stream of the plurality of content types for the later frame and an immediate fade-out of the decoded audio stream of the plurality of content types for the earlier frame.
9. The method of claim 7, wherein the IFFF includes a first frame of a current packet and the earlier frame includes a last frame of a previous packet.
10. The method of claim 9, wherein the IFFF further includes an independent frame that is decoded into the decoded audio stream for the first frame of the current packet.
11. The method of claim 9, wherein the IFFF further includes a predictively encoded frame and one or more previous frames that enable the IFFF to be decoded into the decoded audio stream for the first frame of the current packet, wherein the one or more previous frames begin with an independent frame.
12. The method of claim 9, wherein for time domain aliasing cancellation (TDAC) of modified discrete cosine transform (MDCT), the IFFF further includes one or more previous frames that enable the IFFF to be decoded into the decoded audio stream for the first frame of the current packet, wherein the one or more previous frames begin with an independent frame.
13. The method of claim 9, wherein the IFFF further includes a plurality of frames of the current packet and a plurality of frames of the earlier packet to enable multiple transition frames when generating the crossfade of the decoded audio stream.
14. The method of claim 1, wherein generating the crossfade of the decoded audio streams of the two consecutive frames comprises: performing a fade-in of the decoded audio stream for a later frame of the two consecutive frames and a fade-out of the decoded audio stream for an earlier frame of the two consecutive frames based on a window function associated with time domain aliasing cancellation (TDAC) of modified discrete cosine transform (MDCT).
15. The method of claim 1, wherein generating the decoded audio stream comprises: generating, for each frame of the two consecutive frames, a baseline-decoded audio stream of the plurality of content types having at least one scene element, and wherein generating the cross-fade of the decoded audio stream comprises: generating a cross-fade of the baseline-decoded audio stream of the plurality of content types from an earlier frame to a later frame of the two consecutive frames between transport channels; generating a spatially-decoded audio stream of the cross-fade of the baseline-decoded audio stream of the plurality of content types; rendering the spatially-decoded audio stream of the plurality of content types to generate speaker output signals of the plurality of content types based on the speaker configuration of the decoding device; and mixing the speaker output signals of the plurality of content types to drive the plurality of speakers.
16. The method of claim 15, further comprising: transmitting the spatially-decoded audio stream of the cross-fade of the baseline-decoded audio stream of the plurality of content types and its temporal synchronization metadata to a second device for rendering based on a speaker configuration of the second device.
17. The method of claim 15, wherein generating the cross-fade of the baseline-decoded audio stream of the plurality of content types from the earlier frame to the later frame between transport channels comprises: generating a transition frame based on an immediate fade-in and fade-out frame (IFFF), wherein the IFFF contains a bitstream that encodes the audio content of the later frame and the audio content of the earlier frame to enable an immediate fade-in for the baseline-decoded audio stream of the plurality of content types of the later frame and an immediate fade-out for the baseline-decoded audio stream of the plurality of content types of the earlier frame between the transport channels.
18. The method of claim 15, wherein generating the cross-fade of the baseline-decoded audio stream of the plurality of content types from the earlier frame to the later frame between transport channels comprises: performing a fade-in for the baseline-decoded audio stream of the plurality of content types of the later frame and a fade-out for the baseline-decoded audio stream of the plurality of content types of the earlier frame based on a window function associated with a time-domain aliasing cancellation (TDAC) of a modified discrete cosine transform (MDCT).
19. The method of claim 1, wherein the plurality of content types comprises audio channels, channel objects, or higher order ambisonics (HOA), and wherein the adaptive number of scene elements in the plurality of content types comprises an adaptive number of channels, an adaptive number of channel objects, or an adaptive order of HOA.
20. A system configured to decode audio content, the system comprising: a memory configured to store instructions; a processor coupled to the memory and configured to execute the instructions stored in the memory to: receive frames of the audio content, the audio content represented by a plurality of content types, the frames containing an audio stream encoding the audio content using an adaptive number of scene elements from the plurality of content types; process two consecutive frames to generate a decoded audio stream, the two consecutive frames containing the audio stream, the audio stream encoded using a different mix of the adaptive number of scene elements from the plurality of content types; and generate a cross-fade of the decoded audio stream in the two consecutive frames based on a loudspeaker configuration for driving a plurality of loudspeakers.
Citation Information
Patent Citations
Device and method for manipulating an audio signal having a transient event
CA2821036A1
Apparatus and method for converting an audio signal into a parameterized representation, apparatus and method for modifying a parameterized representation, apparatus and method for synthensizing a parameterized representation of an audio signal
CN102150203A