Hierarchical spatial resolution codec
By adaptively adjusting the encoding of channels, objects, and HOAs using a hierarchical spatial resolution codec, the problem of maintaining immersive audio quality under limited bandwidth is solved, achieving stable audio quality and bandwidth optimization when the target bit rate changes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2021-08-31
- Publication Date
- 2026-08-04
AI Technical Summary
Existing technologies struggle to maintain high audio quality with limited bandwidth when transmitting immersive audio content, especially when the target bit rate changes, leading to audio quality degradation.
Employing a hierarchical spatial resolution codec, it dynamically adjusts the encoding quality and bandwidth requirements of audio scenes by adaptively adjusting the encoding methods of channels, objects, and high-order high-fidelity stereo reproduction, based on the target bit rate and priority order.
Maintain the quality of an immersive audio experience at different target bitrates, reduce audio quality degradation, adapt to bandwidth changes, and provide flexible audio content delivery.
Smart Images

Figure CN116324978B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 083,788, filed on September 25, 2020, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to the field of audio communications; and more specifically, to a digital signal processing method designed to deliver immersive audio content using adaptive spatial coding techniques. Other aspects are also described. Background Technology
[0004] Consumer electronics devices are offering increasingly sophisticated and performance-enhancing digital audio encoding and decoding capabilities. Traditionally, audio content has been produced, distributed, and consumed primarily using two-channel stereo formats that provide left and right audio channels. Recent market developments aim to provide a more immersive listening experience using richer audio formats (such as Dolby Atmos or MPEG-H) that support multi-channel audio, object-based audio, and / or high-fidelity stereo reproduction (Ambisonics).
[0005] The delivery of immersive audio content is associated with greater bandwidth requirements, necessitating higher data rates for streaming and downloading compared to stereo content. If bandwidth is limited, techniques are needed to reduce audio data size while maintaining the best possible audio quality. A common approach to bandwidth reduction in perceptual audio coding leverages auditory perception to maintain audio quality. For example, spatial encoders corresponding to different content types (such as multichannel audio, audio objects, or higher-order high-fidelity stereo reproductions (HOAs)) can use spatial parameters to achieve bit-rate efficient encoding of certain sonic features, allowing these features to be approximately recreated in the decoder. Spatial encoders can be selectively designed to represent different points along a tradeoff curve between spatial resolution and bandwidth requirements to fit the target bandwidth. In some techniques, audio scenes can be predefined as represented by higher-bandwidth multichannel audio / audio objects or lower-bandwidth stereo signals. Further audio encoding and decoding (codec) techniques are required to deliver richer and more immersive audio content using limited bandwidth. Summary of the Invention
[0006] A hierarchical spatial resolution codec is disclosed, which adaptively adjusts the representation of immersive audio content as the bandwidth of the channels used to deliver the immersive audio content changes. The audio scene of the immersive audio content can be represented by an adaptive number of content types encoded through adaptive spatial coding and baseline coding techniques, and an adaptive channel configuration supporting a target bit rate for the transmission channels or users. For example, the audio scene can be represented by an adaptive number of channels, an adaptive number of objects, an adaptive order higher-order high-fidelity stereo reproduction (HOA), or an adaptive number of other sound field representations. HOAs describe a sound field based on spherical harmonics. Different content types have different bandwidth requirements and corresponding audio qualities when recreated at the decoder. Adaptive spatial coding techniques can include adaptive channel and object spatial coding techniques for generating an adaptive number of channels and objects, and adaptive order adaptive HOA spatial coding or HOA compression techniques for generating HOAs. This adaptation can be based on a target bit rate associated with the desired quality and an analysis determining the priority of channels, objects, and HOAs. The target bitrate can be dynamically changed based on channel conditions or the bitrate requirements of one or more users. Prioritization decisions can be made based on the spatial salience of scene elements in the sound field represented by channels, objects, and HOAs.
[0007] In one aspect, the channel and object priority decision module operates on the channels and audio objects of multichannel audio to provide a priority ordering of channels and objects to the spatial encoder. Based on the priority ordering and the target bitrate, the channel and object spatial encoder can encode only high-priority channels and objects to generate a high-quality bitstream with high spatial resolution. The remaining low-priority channels and objects can be converted to lower-quality content types such as HOA and spatially encoded by the HOA spatial encoder to generate a lower-quality bitstream with low spatial resolution that requires lower bandwidth. To adapt to even lower target bitrates, some or all of the low-priority channels and objects can be rendered to even lower-quality content types, such as two-channel stereo signals that require even lower bandwidth. The adaptive encoding capability of the hierarchical spatial resolution codec allows the same audio scene to be represented by different content types according to the target bitrate, for example, by converting some objects in the object to HOA and encoding the converted objects in the HOA domain according to the target bitrate.
[0008] In one aspect, the HOA priority decision module operates on the HOA content to provide a priority ordering of the HOAs to the HOA spatial encoder. Based on the priority ordering and the target bitrate, the HOA spatial encoder can encode only high-priority HOAs to generate a high-quality bitstream with high spatial resolution. The remaining low-priority HOAs can be rendered as lower-quality content types, such as two-channel stereo signals requiring lower bandwidth. The hierarchical structure of the spatial encoder can thus adaptively generate a mixture of bitstreams of audio content types with different quality and bandwidth requirements as the target bitrate changes.
[0009] In one aspect, one or a set of spatial encoders and baseline encoders transform selective scene elements of channels, objects, HOAs, and other sound field representations (such as stereo signals and speech in audio scenes) to generate a set of bitstreams with varying audio quality at a set of bitrates. This set of bitstreams can be generated in real-time or offline. Based on the end user's target bitrate, different scene elements of the channel and object bitstreams, HOA bitstreams, stereo signal bitstreams, and speech bitstreams are selected and adaptively transmitted to the end user.
[0010] In one aspect, for peer-to-peer audio signal transmission, the hierarchical structure of the spatial encoder can adaptively generate different mixes of transport streams with channels, objects, HOAs, and other scene elements as the user's target bit rate changes. Mixes of different audio content types can be generated in real-time or offline.
[0011] In one aspect, a method for encoding audio content is disclosed. The method includes receiving audio content. The audio content is represented by multiple content types, including a first content type and a second content type. The first content type may include multiple scene elements. The method further includes determining the priority of the scene elements of the first content type. Based on the determined priority of the scene elements and a target bit rate for the transmission of the audio content, the method encodes an adaptive number of scene elements of the first content type into a first content stream. The method also encodes the remaining scene elements of the first content type into a second content stream based on the target bit rate; these remaining scene elements are scene elements that have not yet been encoded into the first content stream. The second content stream represents a spatial encoding of the second content type. The method also generates a transport stream including the first content stream and the second content stream for transmission based on the target bit rate.
[0012] The above overview does not constitute an exhaustive list of all aspects of the invention. The invention is envisioned to encompass all systems and methods that can be practiced from all suitable combinations of the aspects outlined above, as well as those disclosed in the detailed descriptions below and specifically pointed out in the claims filed with this patent application. Such combinations have specific advantages not specifically described in the above overview. Attached Figure Description
[0013] The aspects of this disclosure are illustrated by way of example and are not limited to the illustrations in the accompanying drawings, in which similar reference numerals indicate similar elements. It should be noted that references to “a” or “an” aspect in this disclosure do not necessarily refer to the same aspect, and each refers to at least one. Furthermore, for the sake of brevity and to reduce the total number of drawings, a given drawing may be used to illustrate more than one aspect of this disclosure, and for a given aspect, not all elements in that drawing may be necessary.
[0014] Figure 1 This is a functional block diagram of a hierarchical spatial resolution codec that adaptively adjusts the encoding of immersive audio content as the target bit rate changes, according to one aspect of this disclosure.
[0015] Figure 2 A hierarchical spatial resolution codec is described, according to one aspect of this disclosure, for real-time encoding of an audio scene to generate a set of candidate audio bitstreams for a set of bitrates, such that the candidate audio bitstreams can be selected to adapt to a changing target bitrate for one or more users.
[0016] Figure 3 A hierarchical spatial resolution codec is described in a file that stores an audio scene offline to generate a set of candidate audio bitstreams for a set of bitrates, according to one aspect of this disclosure. This file can be read to adapt the transport stream to a target bitrate changed by one or more users.
[0017] Figure 4 A hierarchical spatial resolution codec is described, according to one aspect of this disclosure, to adaptively encode audio scenes in real time to generate transport streams in peer-to-peer transmission that adapt to varying target bit rates for users.
[0018] Figure 5 This is a flowchart of a method for adaptively adjusting the encoding of audio content as the target bit rate changes, according to one aspect of this disclosure, to generate a hierarchical structure of content types. Detailed Implementation
[0019] The goal is to deliver immersive audio content from an audio source to a playback system via transmission channels while maintaining the best possible audio quality. When the bandwidth of the transmission channels changes due to altered channel conditions or a change in the playback system's target bitrate, the encoding of the immersive audio content can be adapted to improve the trade-off between audio playback quality and bandwidth. Immersive audio content can include multichannel audio, audio objects, or spatial audio reconstructions (referred to as high-fidelity stereo reproductions), which describe a sound field based on spherical harmonics and can be used to recreate the sound field for playback. High-fidelity stereo reproductions can include first-order or higher-order spherical harmonics, also known as higher-order high-fidelity stereo reproductions (HOAs). Immersive audio content can be adaptively encoded into audio content with different bitrates and spatial resolutions, based on the target bitrate and the priority order of channels, objects, and HOAs. The adaptively encoded audio content and its metadata can be transmitted over the transmission channels to allow one or more decoders with varying target bitrates to reconstruct the immersive audio experience through spatial decoding and rendering of the adaptively encoded audio content, aided by the metadata.
[0020] Systems and methods for immersive audio coding techniques are disclosed, which adaptively adjust the number of channels, the number of audio objects, the order of HOAs, or other sound field representations of the audio scene of immersive audio content to adapt to varying target bitrates or transmission channel bandwidths of the decoder. The sound field representation of the audio scene can be adaptively encoded using a hierarchical spatial resolution codec that adaptively adjusts the spatial coding resolution or compression of channels, objects, HOAs, etc., as well as the quantization of metadata. This adaptation can be based on the target bitrate and an analysis determining the priorities of channels, objects, HOAs, etc. Prioritization decisions can be made based on the spatial saliency of scene elements in the sound field representation, such that higher-priority scene elements are encoded to maintain a higher-quality sound field representation, while remaining low-quality scene elements can be converted and encoded into a lower-quality sound field representation. Advantageously, hierarchical spatial resolution coding techniques can maintain an immersive audio experience by reducing audio quality degradation of the transport stream when the decoder's target bitrate fluctuates.
[0021] The following description illustrates many specific details. However, it should be understood that aspects of this disclosure can be practiced without requiring these specific details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0022] The terminology used herein is for the purpose of describing particular aspects only and is not intended to limit the invention. Spatially related terms, such as “below,” “under,” “down,” “above,” “above,” etc., may be used herein for the convenience of describing the relationship of one element or feature to one or more other elements or features, as illustrated in the accompanying drawings. It should be understood that spatially related terms are intended to cover different orientations of elements or features during use or operation other than those shown in the drawings. For example, if a device comprising multiple elements in the figures is flipped, an element described as “below” or “under” other elements or features may then be oriented “above” other elements or features. Thus, the exemplary term “below” can cover both the orientations of above and below. The device may be oriented in other ways (e.g., rotated 90 degrees or in other orientations), and the spatially related descriptors used herein are interpreted accordingly.
[0023] As used herein, the singular forms “a” (“a”, “an”) and “the” are intended to include the plural forms as well, unless the context otherwise indicates. It should be further understood that the terms “comprising” and “including” define the presence of the said feature, step, operation, element, or component, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, or groups thereof.
[0024] The terms “or” and “and / or” as used herein should be interpreted as including or referring to any one or any combination thereof. Therefore, “A, B, or C” or “A, B, and / or C” means “any one of the following: A; B; C; A and B; A and C; B and C; A, B, and C.” Exceptions to this definition will only occur if the combination of elements, functions, steps, or actions is inherently mutually exclusive in some way.
[0025] Figure 1 This is a functional block diagram of a hierarchical spatial resolution codec that adaptively adjusts the encoding of immersive audio content as the target bit rate changes, according to one aspect of this disclosure. Immersive audio content 111 may include various immersive audio input formats, also known as sound field representations, such as multichannel audio, audio objects, HOAs, dialogues, etc. In the case of multichannel input, there may be M channels with a known input channel layout, such as a 7.1.4 layout (7 speakers in the center plane, 4 speakers in the top plane, and 1 low-frequency effects (LFE) speaker). It should be understood that the HOA may also include first-order high-fidelity stereo reproduction (FOA). In the following description of the adaptive coding technique, audio objects can be similarly considered as channels, and for simplicity, channels and objects can be grouped together in the operation of the hierarchical spatial resolution codec.
[0026] The audio scene of immersive audio content 111 can be represented by multiple channels / objects 150, HOAs 154, and dialogues 158, each accompanied by channel / object metadata 151, HOA metadata 155, and dialogue metadata 159, respectively. The metadata can be used to describe attributes of the associated sound field, such as the layout configuration or orientation parameters of the associated channels, or the position, size, orientation, or spatial image parameters of the associated objects or HOAs, to help the renderer achieve the desired source image or recreate the perceived location of the dominant sound. To allow a hierarchical spatial resolution codec to improve the trade-off between spatial resolution and the target bitrate, channels / objects and HOAs can be ordered such that when the target bitrate decreases, higher-ordered channels / objects and HOAs are spatially encoded to maintain a higher-quality sound field representation, while lower-ordered channels / objects and HOAs are transformed and spatially encoded into a lower-quality sound field representation.
[0027] The channel / object priority decision module 121 can receive channels / objects 150 and channel / object metadata 151 of an audio scene to provide a priority ranking 162 for the channels / objects 150. In one aspect, the priority ranking 162 can be determined based on the spatial salience of the channels and objects (such as the position, orientation, movement, density, etc. of the channels / objects 150). For example, a channel / object with greater movement near the perceived location of the dominant sound may be spatially more salient than a channel / object with less movement away from the perceived location of the dominant sound, and therefore ranked higher. To reduce the degradation of the overall audio quality of the channels / objects when reducing the target bit rate, the audio quality of the spatial resolution of higher-ranked channels / objects can be maintained, while the spatial resolution of lower-ranked channels / objects can be reduced. In one aspect, the channel / object metadata 151 can provide information to guide the channel / object priority decision module 121 in determining the priority ranking 162. For example, the channel / object metadata 151 may contain priority metadata for ranking certain channels / objects 150 provided by human input. In one aspect, channel / object 150 and channel / object metadata 151 can be passed through channel / object priority decision module 121 as channel / object 160 and channel / object metadata 161, respectively.
[0028] The channel / object spatial encoder 131 can spatially encode channels / objects 160 and channel / object metadata 161 based on channel / object priority sorting 162 and a target bit rate 190 to generate a channel / object audio stream 180 and associated metadata 181. For example, for the highest target bit rate, all channels / objects 160 and metadata 161 can be spatially encoded into the channel / object audio stream 180 and channel / object metadata 181 to provide the highest audio quality of the resulting transport stream. The target bit rate can be determined by the channel conditions of the transmitted channels or the target bit rate of the decoding device. In one aspect, the channel / object spatial encoder 131 can transform channels / objects 160 into the frequency domain to perform spatial coding. The number of frequency subbands and the quantization of coding parameters can be adjusted according to the target bit rate 190. In one aspect, the channel / object spatial encoder 131 can cluster channels / objects 160 and metadata 161 to adapt to a reduced target bit rate 190.
[0029] In one aspect, as the target bitrate 190 decreases, channels / objects 160 and metadata 161 with lower priority ordering can be converted to another content type and spatially encoded using another encoder to generate a lower quality transport stream. The channel / object spatial encoder 131 may not encode these low-ordered channels / objects output as low-priority channels / objects 170 and associated metadata 171. The HOA conversion module 123 can convert the low-priority channels / objects 170 and associated metadata 171 into HOA 152 and associated metadata 153. As the target bitrate 190 gradually decreases, progressively more channels / objects 160 and metadata 161, starting from the lowest priority ordering of 162, can be output as low-priority channels / objects 170 and associated metadata 171 to be converted into HOA 152 and associated metadata 153. HOA 152 and associated metadata 153 can be spatially encoded to generate a lower quality transport stream compared to a transport stream that fully encodes all channels / objects 160, but with the advantages of requiring a lower bit rate and lower transmission bandwidth.
[0030] Multiple levels of a hierarchical structure may exist for converting and encoding channels / objects 160 into another content type to accommodate a lower target bitrate. In one aspect, parametric encoding, such as a stereo-based immersive coding (STIC) encoder 137, can be used to encode some low-priority channels / objects 170 and associated metadata 171. The STIC encoder 137 can render a two-channel stereo audio stream 186 from an immersive audio signal, such as by downmixing channels or rendering objects or HOAs as stereo signals. The STIC encoder 137 can also generate metadata 187 based on a perceptual model that derives parameters describing the perceptual direction of the dominant sound. Further reductions in bitrate, albeit at lower quality transport streams, can be accommodated by converting and encoding some channels / objects into the stereo audio stream 186 instead of HOAs. Although the STIC encoder 137 is described as rendering a channel, object, or HOA into a two-channel stereo audio stream 186, the STIC encoder 137 is not limited to this and can render a channel, object, or HOA into an audio stream with more than two channels.
[0031] In one aspect, at a medium target bitrate, some of the low-priority channels / objects 170 with the lowest priority order and their associated metadata 171 can be encoded into the stereo audio stream 186 and associated metadata 187. The remaining low-priority channels / objects 170 with higher priority order and their associated metadata can be converted into HOAs 152 and associated metadata 153, which can be prioritized together with other HOAs 154 and associated metadata 155 from the immersive audio content 111 and encoded into the HOA audio stream 184 and associated metadata 185. The remaining channels / objects 160 with the highest priority order and their metadata are encoded into the channel / object audio stream 180 and associated metadata 181. In another aspect, at the lowest target bitrate, all channels / objects 160 can be encoded into the stereo audio stream 186 and associated metadata without leaving encoded channels, objects, or HOAs in the transport stream.
[0032] Similar to channels / objects, HOAs can also be sorted such that higher-ranked HOAs are spatially encoded to maintain a higher-quality sound field representation, while lower-ranked HOAs are rendered as lower-quality sound field representations such as stereo signals. The HOA priority decision module 125 can receive HOAs 154 representing the sound field of the audio scene and associated metadata 155 from the immersive audio content 111, as well as converted HOAs 152 and associated metadata 153 that have been converted from low-priority channels / objects 170, to provide a priority ordering 166 among the HOAs. In one aspect, the priority ordering can be determined based on the spatial salience of the HOAs (such as the location, orientation, movement, density, etc. of the HOA). To reduce the overall audio quality degradation of the HOAs when reducing the target bit rate, the audio quality of higher-ranked HOAs can be maintained while the audio quality of lower-ranked HOAs can be reduced. In one aspect, the HOA metadata 155 can provide information to guide the HOA priority decision module 125 in determining the HOA priority ordering 166. HOA priority decision module 125 can combine HOA 154 from immersive audio content 111 with converted HOA 152 that has been converted from low-priority channels / objects 170 to generate HOA 164, and combine the associated metadata of the combined HOA to generate HOA metadata 165.
[0033] The hierarchical HOA spatial encoder 135 can spatially encode HOA 164 and HOA metadata 165 based on HOA priority ordering 166 and a target bit rate 190 to generate a HOA audio stream 184 and associated metadata 185. For example, for a high target bit rate, all HOA 164 and HOA metadata 165 can be spatially encoded into the HOA audio stream 184 and HOA metadata 184 to provide a high-quality transport stream. In one aspect, the hierarchical HOA spatial encoder 135 can transform HOA 164 into the frequency domain to perform spatial coding. The number of frequency subbands and the quantization of coding parameters can be adjusted according to the target bit rate 190. In one aspect, the hierarchical HOA spatial encoder 135 can cluster HOA 164 and HOA metadata 165 to adapt to a reduced target bit rate 190. In one aspect, the hierarchical HOA spatial encoder 135 can perform compression techniques to generate an adaptive order of HOA 164.
[0034] In one aspect, as the target bitrate 190 is reduced, lower-priority HOAs 164 and metadata 165 can be encoded as stereo signals. The hierarchical HOA spatial encoder 135 may not encode these low-ranked HOAs, which output as low-priority HOAs 174 and associated metadata 175. As the target bitrate 190 gradually decreases, progressively more HOAs 164 and HOA metadata 165, starting from the lowest priority of priority 166, can be output as low-priority HOAs 174 and associated metadata 175 to be encoded into the stereo audio stream 186 and associated metadata 187. Compared to a transport stream that fully encodes all HOAs 164, the stereo audio stream 186 and associated metadata 187 require a lower bitrate and lower transmission bandwidth, despite lower audio quality. Therefore, as the target bitrate 190 decreases, the transport stream of the audio scene can have a larger mixture of hierarchical structures for content types with lower audio quality. In one aspect, the hierarchical mixing of content types can be adaptively changed scene-by-scene, frame-by-frame, or group-by-group. Advantageously, the hierarchical spatial resolution codec, based on the priority ordering of scene elements according to the target bit rate and sound field representation, adaptively adjusts the hierarchical coding of immersive audio content to generate a mixture of variations in channel, object, HOA, and stereo signals, in order to improve the trade-off between audio quality and target bit rate.
[0035] In one aspect, the audio scene of the immersive audio content 111 may include dialogue 158 and associated metadata 159. Dialogue space encoder 139 may encode dialogue 158 and associated metadata 159 based on a target bitrate 190 to generate speech stream 188 and speech metadata 189. In one aspect, when the target bitrate 190 is high, dialogue space encoder 139 may encode dialogue 158 into a two-channel speech stream 188. When the target bitrate 190 is lowered, dialogue 158 may be encoded into a one-channel speech stream 188.
[0036] Baseline encoder 141 can encode channel / object audio stream 180, HOA audio stream 184, and stereo audio stream 186 into audio stream 191 based on a target bit rate 190. Baseline encoder 141 can use any known encoding technique. In one aspect, baseline encoder 141 can adapt the encoding rate and quantization to the target bit rate 190. Speech encoder 143 can encode speech stream 188 separately for audio stream 191. Channel / metadata 181, HOA metadata 185, stereo metadata 187, and speech metadata 189 can be combined into a single transport channel of audio stream 191. Audio stream 191 can be transmitted over the transport channel to allow one or more decoders to reconstruct immersive audio content 111. Audio stream 191 is also referred to as a transport stream.
[0037] Figure 2 A hierarchical spatial resolution codec is described, according to one aspect of this disclosure, to encode an audio scene in real time to generate a set of candidate audio bitstreams 203 for a set of target bitrates, such that the candidate audio bitstreams 203 can be selected to adapt to changing target bitrates for one or more users. A set of encoders 201 can provide this set of candidate audio bitstreams 203. Each candidate audio bitstream may include a channel / object audio stream 180, a HOA stream 184, a stereo audio stream 186, a speech stream 188, and metadata, such as... Figure 1 This is described for a possible target bit rate.
[0038] The possible target bitrate ranges are labeled in descending order as highest, high, high-medium, medium, medium-low, low, and lowest. In one aspect, the target bitrate range may include discrete values of 1 Mbps (megabits per second), 768 Kbps (kilobits per second), 512 Kbps, 384 Kbps, 256 Kbps, 128 Kbps, and 64 Kbps. For each of the possible target bitrates, the group of encoders 201 may include a separate audio encoder, which may include… Figure 1 The hierarchical spatial resolution codec. However, the group of encoders 201 is not limited to this. In one aspect, a single high-rate hierarchical spatial resolution codec can be time-division multiplexed to generate the group of candidate audio bitstreams 203 for all possible target bit rates.
[0039] As shown in the figure, for the highest target bitrate, the audio encoder can generate a candidate bitstream that includes a channel / object audio stream 180 encoding the L1 channels / objects of the immersive audio content 111, but does not include audio streams for HOA, stereo signals, or speech. In another example, the candidate bitstream for the highest target bitrate may include a HOA audio stream 184 encoding a certain order of HOA, a stereo audio stream 186, and / or a speech stream 188. Moving to the next step in the target bitrate range towards a higher target bitrate, some L1 channels / objects with lower priority can be converted and encoded into the order M1 HOA audio stream 184, leaving the channel / object audio stream 180 to encode the higher priority L2 channels / objects. Moving to the next step towards a high-to-medium target bitrate, the number of channels / objects in the channel / object audio stream 180 is merged into L3, where L3 is less than L2. The next step is to move to the target bit rate, where the order of the HOA in the HOA audio stream 184 is merged into M2, where M2 is less than M1.
[0040] Stepping further down to the medium-low target bitrate, some of the L3 channels / objects with lower priority ordering are converted and encoded into the HOA, leaving channel / object audio stream 180 to encode the higher priority ordering L4 channels / objects. The additional converted HOAs are prioritized together with the existing HOAs of order M2, resulting in some HOAs with lower priority ordering being encoded into the stereo audio stream 186. The HOA audio stream 184 remains at order M2 to encode the higher priority ordering HOAs. The stereo audio stream 186 is shown as having N1 channels to show that it is not limited to two channels. The audio stream for the medium-low target bitrate also includes a speech stream 188.
[0041] Further stepping to a lower target bit rate, some of the L4 channels / objects with lower priority ordering are converted and encoded into a HOA, leaving channel / object audio stream 180 to encode the higher priority L5 channels / objects. The additional converted HOA is prioritized together with the existing HOA of order M2, and the orders of the HOAs are merged to keep the HOA audio stream 184 at order M2.
[0042] For the lowest target bit rate, all channels / objects are converted and encoded into HOAs. Additional converted HOAs, along with the existing HOA of order M2, are encoded into a two-channel stereo audio stream 186. There is no channel / object audio stream 180 or HOA stream 184. Note that the candidate bitstreams for all target bit rates have a single meta-data transfer stream. In one aspect, the set of encoders can further encode the set of candidate audio bitstreams 203 using a baseline encoder 141 based on the range of target bit rates.
[0043] The statistical multiplexing module 205 selects a candidate bitstream, which may include a channel / object audio stream 180, a HOA stream 184, a stereo audio stream 186, a speech stream 188, and a meta-data transmission stream, based on each user's target bitrate 190, to adaptively generate the transport stream. The user's target bitrate 190 can be adaptively changed scene-by-scene, frame-by-frame, or group-by-group. For example, for group-adaptive processing, when the user's target bitrate 190 is at its highest, the grouping of the user's transport stream may include a channel / object audio stream 180 encoded with L1 channel / object and a meta-data transmission stream. When the user's target bitrate changes to a medium level, the grouping of the user's transport stream may change to a channel / object audio stream 180 encoded with L3 channel / object, a HOA audio stream 184 of order M2, and a meta-data transmission stream. When a user's target bitrate changes to low, the packetization of the transport stream for that user can be changed to a channel / object audio stream 180 encoded with L5 channels / objects, a HOA audio stream 184 of order M2, a stereo audio stream 186 of N1 channels, a speech stream 188, and a metadata data stream. Transport streams for multiple users (such as transport stream 210 for user A, transport stream 212 for user B, and transport stream 214 for user C) can be individually customized according to each user's target bitrate 190 to provide real-time streaming of immersive audio content 111.
[0044] Figure 3 A hierarchical spatial resolution codec is described, according to one aspect of this disclosure, for offline encoding of an audio scene to generate a set of candidate audio bitstreams 203 for a set of bitrates, which are then stored in a file that can be read to adapt the transport stream to a target bitrate changed by one or more users. Figure 2 As shown, a set of encoders 201 can provide the set of candidate audio bitstreams 203. For a possible target bitrate, each candidate audio bitstream may include a channel / object audio stream 180, a HOA stream 184, a stereo audio stream 186, a speech stream 188, and metadata encoded from immersive audio content 111.
[0045] However, instead of streaming immersive audio content 111 in real time, this set of candidate audio bitstreams 203 can be generated offline and stored in a bitstream manifest file 207. When a user is ready to stream immersive audio content 111, the statistical multiplexing module 205 can read the bitstream manifest file 207 to select a candidate bitstream based on a target bitrate 190, which may include channel / object audio stream 180, HOA stream 184, stereo audio stream 186, speech stream 188, and meta data transmission stream, for the user to adaptively generate the transport stream. Transport streams for multiple users (such as transport stream 210 for user A, transport stream 212 for user B, and transport stream 214 for user C) can be individually customized according to each user's target bitrate 190.
[0046] Figure 4 A hierarchical spatial resolution codec is described, according to one aspect of this disclosure, to adaptively encode audio scenes in real time to generate transport streams in peer-to-peer transmission that adapt to varying target bit rates for the user. Instead of... Figure 2 and Figure 3 In generating a set of candidate bitstreams for the target bitrate range, spatial and baseline encoders 301 (such as...) Figure 1 A hierarchical spatial resolution codec encodes immersive audio content 111 into a transport stream, which may include a channel / object audio stream 180, a HOA stream 184, a stereo audio stream 186, a speech stream 188, and a meta data transmission stream, to adapt in real time to the user's target bitrate 190. In one aspect, the encoded audio stream may be generated offline, stored in a file, and retrieved at a later time to adapt to the user's target bitrate.
[0047] The spatial and baseline encoder 301 can adapt the encoded audio stream to the user's target bit rate 190 based on packets, frames, or audio scenarios. For example, when each packet comprises four frames, at packet 1, when the target bit rate 190 is highest, the packetization for the user's transport stream can include a channel / object audio stream 180 encoding the L1 channel / object and a meta-data transmission stream for four frames. At packet 2, when the target bit rate is high-medium, the packetization for the user's transport stream can be changed to a channel / object audio stream 180 encoding the L3 channel / object, a HOA audio stream 184 of order M1, and a meta-data transmission stream for four frames. At packet 3, when the target bit rate is lowest, the packetization for the user's transport stream can be changed to a two-channel stereo audio stream 186, a one-channel speech stream 188, and a meta-data transmission stream for four frames.
[0048] Figure 5This is a flowchart of a method 500 for adaptively adjusting the encoding of audio content as the target bit rate changes to generate a hierarchical structure of content types, according to one aspect of this disclosure. Method 500 can be derived from... Figure 1 , Figure 2 , Figure 3 or Figure 4 We will implement a hierarchical spatial resolution codec.
[0049] In operation 501, method 500 receives audio content. The audio content is represented by multiple content types, including a first content type and a second content type. The first content type may include multiple scene elements. In one aspect, the first content type may include channels / objects, and the second content type may include HOAs. The number of scene elements may represent the number of channels or objects.
[0050] In operation 503, method 500 determines the priority of scene elements of the first content type. In one aspect, the priority of scene elements of the first content type can be ranked based on the spatial salience of the scene elements.
[0051] In operation 505, method 500 encodes an adaptive number of scene elements of the first content type into the first content stream based on the priority of the scene elements and the target bit rate for the transmission of the audio content. The number of scene elements of the first content type encoded into the first content stream can vary with the target bit rate.
[0052] In operation 507, method 500 encodes remaining scene elements of the first content type into a second content stream based on a target bit rate. These remaining scene elements are scene elements that have not yet been encoded into the first content stream. The second content stream represents the spatial encoding of the second content type. The number of scene elements of the second content type encoded into the second content stream can change with the target bit rate.
[0053] In operation 509, method 500 generates a transport stream comprising a first content stream and a second content stream based on a target bit rate for transmission.
[0054] The implementation of the hierarchical spatial resolution codec described herein can be implemented in a data processing system, for example, via a network computer, network server, tablet computer, smartphone, laptop computer, desktop computer, other consumer electronic device, or other data processing system. Specifically, the operation of adaptively encoding an audio scene according to a changing target bit rate, as described for the hierarchical spatial resolution codec, is a digital signal processing operation performed by a processor executing instructions stored in one or more memories. The processor can read the stored instructions from the memory and execute the instructions to perform the operations. These memories represent examples of machine-readable, non-transitory storage media that can store or contain computer program instructions that, when executed, cause the data processing system to perform one or more of the methods described herein. The processor can be a processor in a local device such as a smartphone, a processor in a remote server, or a distributed processing system of multiple processors in a local device and a remote server, wherein their respective memories contain portions of the instructions required to perform the operations.
[0055] The processes and blocks described herein are not limited to the specific examples described, nor are they limited to the specific order in which they are used as examples herein. Rather, any processing blocks can be reordered, combined, or removed, executed in parallel or serially, as needed to achieve the results described above. Processing blocks associated with implementing an audio processing system can be executed by one or more programmable processors executing one or more computer programs stored on a non-transitory computer-readable storage medium to perform the functions of the system. All or part of the audio processing system can be implemented as special-purpose logic circuitry (e.g., FPGA (Field-Programmable Gate Array) and / or ASIC (Application-Specific Integrated Circuit)). All or part of the audio system can be implemented using electronic hardware circuitry including at least one of electronic devices such as processors, memories, programmable logic devices, or logic gates. Additionally, processes can be implemented in any combination of hardware devices and software components.
[0056] While certain exemplary examples are described and illustrated in the accompanying drawings, it should be understood that these examples are merely illustrative and not limiting to the invention in a broader sense, and the invention is not limited to the specific constructions and arrangements shown and described, as various other modifications can be made by those skilled in the art. Therefore, the description should be regarded as exemplary and not limiting.
[0057] In order to assist the Patent Office and any reader of any patent published in this application in interpreting the appended claims, the applicant wishes to note that they do not intend any of the appended claims or claim elements to invoke 35U.SC112(f) unless the words “means for…” or “steps for…” are expressly used in a particular claim.
Claims
1. A method for encoding audio content, the method comprising: The audio content is received by an encoding device, and the audio content is represented by one or more content types, the first content type including multiple scene elements; Determine the priority of the plurality of scene elements of the first content type; Based on the priority of the plurality of scene elements and the target bit rate for transmitting the audio content, an adaptive number of the plurality of scene elements of the first content type are encoded into the first content stream; The remaining scene elements of the first content type that were not selected for encoding into the first content stream are encoded into the second content stream, where the second content stream represents the encoding of the second content type; As the target bit rate changes, the adaptive number of scene elements of the first content type is selected based on selecting scene elements with a higher priority than the remaining scene elements. as well as A transport stream comprising the first content stream and the second content stream is generated based on the target bit rate for transmission.
2. The method of claim 1, wherein the first content type has a sound field representation of the audio content of a higher quality than the second content type.
3. The method of claim 1, wherein the bit rate for supporting the transmission of the first content type is higher than the bit rate for supporting the transmission of the second content type.
4. The method according to claim 1 or 3, wherein determining the priority of the plurality of scene elements of the first content type includes: A priority ranking of the multiple scene elements of the first content type is generated based on the spatial saliency of the multiple scene elements, wherein scene elements with higher spatial saliency have higher quality sound field representation than other scene elements with lower spatial saliency.
5. The method of claim 1, wherein encoding the remaining scene elements of the first content type that were not selected for encoding into the first content stream into the second content stream, based on the target bit rate and the priority of scene elements of the second content type, comprises: Convert the remaining scene elements of the first content type into scene elements of the second content type; as well as Based on the target bit rate, the converted scene elements, combined with scene elements of the second content type received from the audio content, are encoded to generate the second content stream.
6. The method of claim 5, wherein encoding the converted scene element combined with scene elements of the second content type received from the audio content comprises: Determine the priority of a plurality of scene elements of the second content type, the plurality of scene elements of the second content type including the converted scene elements and the scene elements of the second content type received from the audio content; Based on the priority and target bit rate of the plurality of scene elements of the second content type, an adaptive number of the plurality of scene elements of the second content type are encoded into the second content stream; Based on the target bit rate, the remaining scene elements of the second content type that were not selected for encoding into the second content stream are encoded into a third content stream, the third content stream representing the encoding of the third content type; as well as The transport stream is generated to include the third content stream.
7. The method of claim 6, wherein the first content type has a sound field representation of the audio content of higher quality than the second content type, and the second content type has a sound field representation of the audio content of higher quality than the third content type.
8. The method of claim 6, wherein the bit rate for supporting the transmission of the first content type is higher than the bit rate for supporting the transmission of the second content type, and the bit rate for supporting the transmission of the second content type is higher than the bit rate for supporting the transmission of the third content type.
9. The method of claim 5 or 6, wherein determining the priority of the plurality of scene elements of the second content type comprises: A priority ranking of the multiple scene elements of the second content type is generated based on the spatial saliency of the multiple scene elements, wherein scene elements with higher spatial saliency have higher quality sound field representations than other scene elements with lower spatial saliency.
10. The method of claim 5 or 6, wherein encoding the adaptive number of the plurality of scene elements of the second content type into the second content stream comprises: As the target bit rate changes, an adaptive number of scene elements of the second content type are selected based on the fact that the selected scene elements have a higher priority than the remaining scene elements of the second content type that were not selected for encoding into the second content stream.
11. The method of claim 1 or 6, wherein encoding the remaining scene elements of the first content type that were not selected for encoding into the first content stream into the second content stream based on the target bit rate comprises: Convert the first subset of the remaining scene elements of the first content type into scene elements of the second content type; The converted scene elements are encoded into the second content stream based on the target bit rate; Based on the target bit rate, a second subset of the remaining scene elements of the first content type that have not been converted to the second content type is encoded into a third content stream, wherein the third content stream represents the encoding of the third content type; as well as The transport stream is generated to include the third content stream.
12. The method of claim 1 or 6, wherein generating the transport stream comprises: Baseline coding and spatial coding of the first content stream and the second content stream are performed based on the target bit rate.
13. The method of claim 1 or 6, wherein the audio content includes voice dialogue as one of the content types, and wherein the method further comprises: The voice dialogue is encoded into a voice stream based on the target bit rate; as well as The transport stream is generated to include the voice stream.
14. The method of claim 1 or 6, wherein the first content type is associated with metadata describing the attributes of the plurality of scene elements of the first content type. Encoding the adaptive number of the plurality of scene elements of the first content type into the first content stream includes: Based on the target bitrate, the metadata associated with the plurality of scene elements of the adaptive number is encoded into the metadata of the first content stream. Encoding the remaining scene elements of the first content type into the second content stream based on the target bit rate includes: Based on the target bitrate, the metadata associated with the remaining scene elements is encoded into the metadata of the second content stream. And generating the transport stream includes: Based on the target bit rate, the metadata of the first content stream and the metadata of the second content stream are combined into a metadata data transmission stream.
15. The method of claim 14, wherein the metadata associated with the first content type includes metadata for assisting the encoding device in determining the priority of the plurality of scene elements of the first content type and for assisting the decoding device in spatially decoding and rendering the plurality of scene elements of the first content type.
16. The method of claim 1 or 6, wherein encoding the adaptive number of the plurality of scene elements of the first content type into the first content stream comprises: Multiple candidate first content streams are generated based on the priorities of the multiple scene elements and multiple target bit rates. The multiple candidate first content streams encode an adaptive number of the scene elements of the first content type. Encoding the remaining scene elements of the first content type that were not selected for encoding into the first content stream into the second content stream based on the target bit rate includes: Multiple candidate second content streams are generated based on the multiple target bitrates. These multiple candidate second content streams encode an adaptive number of scene elements of the second content type. The adaptive number of scene elements of the second content type includes the remaining scene elements of the first content type that have been converted into scene elements of the second content type and combined with scene elements of the second content type received from the audio content. And generating the transport stream includes: Based on the user's target bit rate, one of the plurality of candidate first content streams and one of the plurality of candidate second content streams are selected for use in the transport stream.
17. The method of claim 16, further comprising: The file stores the plurality of candidate first content streams and the plurality of candidate second content streams. And generating the transport stream includes: Based on the user's target bit rate, one of the plurality of candidate first content streams and one of the plurality of candidate second content streams are selected from the file for use in the transport stream.
18. The method of claim 1 or 6, wherein encoding the adaptive number of the plurality of scene elements of the first content type into the first content stream comprises: Based on the priority of the plurality of scene elements and as the user's target bitrate changes, the first content stream is generated to encode an adaptive number of scene elements of the first content type; Furthermore, encoding the remaining scene elements of the first content type that were not selected for encoding into the first content stream into the second content stream based on the target bit rate includes: As the user's target bitrate changes, a second content stream is generated to encode an adaptive number of scene elements of the second content type, the adaptive number of scene elements of the second content type including the remaining scene elements of the first content type that have been converted into scene elements of the second content type and combined with scene elements of the second content type received from the audio content.
19. The method of claim 1 or 6, wherein the first content type includes audio channels or audio objects, wherein the plurality of scene elements of the first content type includes a plurality of audio channels or a plurality of audio objects, and wherein the second content type includes higher-order high-fidelity stereo reproduction (HOA).
20. A system configured to encode audio content, the system comprising: A memory configured to store instructions; A processor, coupled to the memory, and configured to execute the instructions stored in the memory to: The audio content is received, and the audio content is represented by multiple content types, the first content type including multiple scene elements; Determine the priority of the plurality of scene elements of the first content type; Based on the priority of the plurality of scene elements and the target bit rate for transmitting the audio content, an adaptive number of the plurality of scene elements of the first content type are encoded into the first content stream; The remaining scene elements of the first content type that were not selected for encoding into the first content stream are encoded into the second content stream, where the second content stream represents the encoding of the second content type; As the target bit rate changes, the adaptive number of scene elements of the first content type is selected based on selecting scene elements with a higher priority than the remaining scene elements. as well as A transport stream comprising the first content stream and the second content stream is generated based on the target bit rate for transmission.