Audio decoder, audio encoder, method of providing decoded audio signal, method of providing encoded audio signal, audio stream, audio stream provider, and computer program using stream identifier
The audio decoder addresses the challenge of seamless transitions between audio streams by using a stream identifier to differentiate between streams, even when encoding configurations are the same, thereby enhancing the quality of adaptive audio streaming.
Patent Information
- Application Number
- JP2025011286
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-01-11
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2038-01-10
AI Technical Summary
Existing audio decoding technologies face challenges in seamlessly transitioning between different audio streams, particularly when the bitrates are different but the encoding configurations are the same, leading to potential audible artifacts.
An audio decoder that adjusts decoding parameters based on configuration information, including a stream identifier, to distinguish between different audio streams and perform seamless transitions, even when the decoding configurations appear identical.
Enables seamless switching between different audio streams without audible artifacts, by using the stream identifier to differentiate between streams, thus improving the quality of adaptive audio streaming.
Smart Images

Figure 2025081335000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments according to the present invention relate to an audio decoder for supplying a decoded audio signal representation based on an encoded audio signal representation.
[0002] Also, embodiments according to the present invention relate to an audio encoder for supplying an encoded audio signal representation.
[0003] Also, embodiments according to the present invention relate to a method for supplying a decoded audio signal representation.
[0004] Also, embodiments according to the present invention relate to a method for supplying an encoded audio signal representation.
[0005] Also, embodiments according to the present invention relate to an audio stream.
[0006] Also, embodiments according to the present invention relate to an audio stream provider.
[0007] Also, embodiments according to the present invention relate to a computer program for executing one of the methods.
Background Art
[0008] Hereinafter, the problems underlying the aspects of the present invention and possible usage scenarios of the embodiments according to the present invention will be described.
[0009] There are situations where there are transitions between various audio streams or between various sequences of encoded audio frames. For example, various sequences of audio frames can include various audio contents, and transitions should be made between them.
[0010] For example, when MPEG-D USAC (ISO / IEC 23003-3 + Amd.1 + Amd.2 + Amd.3) is used in a case where adaptive streaming is employed, a situation can occur where two streams within a so-called adaptation set (e.g., two or more streams that a user can switch between can be grouped together) have exactly the same configuration structure even if their bitrates are different. This can occur, for example, when the encoder simply chooses to operate the encoder using exactly the same set of encoding tools for both bitrates.
[0011] For example, an audio encoder can use the same basic encoding settings (which are also notified to the audio decoder), yet still supply various representations of the audio values. For example, the audio encoder can use a coarser quantization of the spectral values, which results in a smaller bit requirement when it is desired to achieve a lower bitrate while the basic encoder or decoder settings remain unchanged.
[0012] However, this (e.g., the occurrence of a situation where two streams within an adaptation set have exactly the same configuration structure even if their bitrates are different) is not a problem in itself.
[0013] However, it has been found that when using adaptive streaming, the decoder should know whether the subsequently received access unit (or "frame") is from the same stream or whether a stream change has occurred.
[0014] When a stream change is detected, it has been found that the audio decoder, in some cases, performs a specific series of operational steps to ensure the following. · One decoder instance is properly shut down and the partially decoded signal that was temporarily stored internally is sent to the decoder output. - A process called "flushing" · The decoder re-instantiates and reconfigures itself using the configuration information associated with the changed stream. · The decoder "pre-rolls" the embedded access units piggybacked in the instant playback frame (IPF). This pre-rolling of the access units brings the decoder to a fully initialized state, such that the output of decoding the first frame results in a fully compliant decoded audio signal. · Optionally, depending on, for example, the corresponding bitstream signaling element, the audio output from the decoder flushing process and the output from decoding the first access unit of the reconfigured decoder are cross-faded over a very short period.
[0015] For example, all of the above steps can be performed to achieve the sole purpose of obtaining a "seamless" transition from the decoded audio of one stream to the decoded audio of another stream. "Seamless" means that there are no audible artifacts or glitches from the stream transition itself. In practice, the stream transition may be perceptible. This can be due to, for example, variations in overall encoding quality or audio bandwidth and sound quality. However, the actual point in time (within the time) of the transition does not, by itself, cause an auditory impression. In other words, there are no disturbing sounds such as "clicks" or "noise bursts" at the transition point.
[0016] It has been found that information on whether a stream change has occurred is obtained by analyzing the configuration structure embedded in the instant playback frame and comparing it with the configuration of the currently decoded stream. For example, an audio decoder can assume a stream change only if the received configuration is different from the current configuration.
[0017] For example, when a decoder receives an instant playback frame (IPF) of a stream at various bitrates, it will detect the presence of an audio preroll extension payload, extract the configuration structure, and compare this new configuration with the current configuration. For further details, also refer to the section "Bitrate Adaptation" in ISO / IEC 23003-3:2012 / Amd.3.
[0018] However, if both the current and new configuration structures are the same, the decoder cannot recognize that it is receiving access units from a different stream than before, and thus does not reconfigure the decoder and does not decode the audio preroll in the IPF's extension payload.
[0019] Instead, the decoder will attempt to continue decoding as if it were receiving consecutive access units from the previous active stream. This can lead to a likely situation where (in a conventional case where, for example, the streamID is not used or evaluated) the window boundaries and coding modes of the last decoded frame and the new frames of the new stream do not match, resulting in the generation of audible artifacts such as clicks or noise bursts. This would defeat the main purpose of the IPF and the idea of adaptive audio streaming based on the concept of seamless transitions between streams.
[0020] In the following, some conventional techniques will be described.
[0021] It should be noted that in the case of Unified Speech and Audio Coding (USAC), there is no known solution.
[0022] In MPEG-H 3D Audio (ISO / IEC 23008-3 + all revisions), the problem can be solved when audio data is transmitted in the MPEG-H Audio Stream ("MHAS") packetized stream format. Since the MHAS package contains packet labels that can potentially differ between streams, it can serve the purpose of differentiating configurations. However, the MHAS format is not defined in MPEG-D USAC.
[0023] In MPEG-4 HE-AAC (ISO / IEC 14496-3 + all revisions), there are workarounds that require the encoder to ensure that all streams have the same window shape and window sequence, as well as additional constraints regarding the signal processing tools used, at potential transition points (so-called stream access points (SAPs)). This can potentially have an adverse effect on audio quality. The above IPF is designed to completely free the new codec from all these constraints.
[0024] In conclusion, there is a need for a concept that enables switching between different audio streams and provides an improved trade-off between the amount of overhead and ease of implementation.
Summary of the Invention
Problems to be Solved by the Invention
[0025] Embodiments according to the present invention create an audio decoder for supplying a decoded audio signal representation based on an encoded audio signal representation. The audio decoder is configured to adjust decoding parameters depending on configuration information. The audio decoder is configured to decode one or more audio frames using the current configuration (e.g., using currently active configuration information). Further, if the configuration information within the configuration structure associated with the one or more frames to be decoded or a relevant part of the configuration information within the configuration structure associated with the one or more frames to be decoded (e.g., up to and including the stream identifier) is different from the current configuration information, the audio decoder compares the configuration information within the configuration structure associated with the one or more frames to be decoded with the current configuration information and transitions to perform the decoding using the configuration information within the configuration structure associated with the one or more frames to be decoded as the new configuration information. The audio decoder is configured to consider the stream identifier information included in the configuration structure when comparing the configuration information such that a difference between the stream identifier represented by the stream identifier previously obtained by the audio decoder and the stream identifier information within the configuration structure associated with the one or more frames to be decoded causes a transition.
[0026] This embodiment according to the present invention enables the audio decoder to distinguish different streams based on the presence and evaluation of stream identifier information included in the configuration structure. As a result, even when the actual decoding configuration (e.g., describable by the remaining configuration information within the configuration structure) is the same for both streams, the execution of transitions becomes possible, based on the idea that the stream identifier can be used as a criterion for distinguishing different streams capable of making transitions. Since the stream identifier information is included in the configuration structure (e.g., together with other configuration information for adjusting the decoding parameters of the audio decoder), there is no need to evaluate information from different protocol layers when determining whether to perform a transition. For example, the stream identifier information is included in a sub-data structure of the data structure that defines the decoding parameters ("configuration structure") so that there is no need to transfer information from the packet level to the actual audio decoder. By including the stream identifier information in the configuration structure, the audio decoder can recognize a transition from the first stream to the second stream without affecting the decoding parameters when decoding consecutive portions of a single stream, and can recognize the switching between different streams on the audio decoder side without accessing information from different protocol layers even in a situation where the same decoding parameters are used for different streams. Also, at positions where switching between different streams is allowed, it is not necessary to use the same decoding parameters for different streams.
[0027] In conclusion, the concept defined by independent claim 1 enables the recognition of switching between different streams with a reasonable implementation complexity (e.g., without extracting dedicated signaling information from different protocol layers and transferring it to the audio decoder) while avoiding the need to enforce specific encoding / decoding settings (e.g., window selection, etc.) during transitions. Therefore, excessive overhead and degradation of audio quality can also be avoided.
[0028] In a preferred embodiment, the audio decoder is configured to check whether the configuration structure includes stream identifier information, and if the stream identifier information is included in the configuration structure, selectively consider the stream identifier information in the comparison. Therefore, it is not necessary to include stream identifier information in each configuration structure. Rather, it is possible to omit the stream identifier in the configuration structure of an audio frame where the possibility of switching between different streams is not required. Therefore, some bits can be saved, and the evaluation of the stream identification information can be avoided in that switching between different streams is not allowed.
[0029] In a preferred embodiment, the audio decoder is configured to check whether the configuration structure includes a configuration extension structure and to check whether the configuration extension structure includes a stream identifier. The audio decoder can be configured to selectively consider the stream identifier information in the comparison if the stream identifier information is included in the configuration extension structure.
[0030] Therefore, the stream identifier can be placed within a configuration extension structure whose existence is optional, and the existence of the stream identifier information can even be regarded as optional even if the configuration extension structure exists. Therefore, the audio decoder can flexibly recognize whether the stream identifier information exists and can avoid including unnecessary information in the audio encoder. When the stream identifier is placed in a data structure that can be activated and deactivated (for example, by a flag in the fixed (always present) part of the configuration structure), when the stream identifier information is not required, the stream identifier information can be accurately placed where it is required while saving bits. Since switching between streams is usually only possible at a specified time, this is advantageous because each frame with a configuration structure does not need to include stream identifier information.
[0031] In a preferred embodiment, the audio decoder is configured to accept a variable ordering of configuration information items within a configuration extension structure. For example, when comparing the configuration information within a configuration structure associated with one or more frames to be decoded with the current configuration information, the audio decoder is configured to consider configuration information items (e.g., configuration extensions) placed within the configuration extension structure before the stream identifier information (e.g., before an item named "streamID") (e.g., similar to the stream identifier information). Further, the audio decoder may be configured to leave configuration information items (e.g., configuration extensions) placed within the configuration extension structure (e.g., "UsacConfigExtension()") after the stream identifier information is not considered when comparing the configuration information within a configuration structure associated with one or more frames to be decoded with the current configuration information.
[0032] By using such a concept, the detection of transitions between different streams can be performed in a very flexible way. For example, all such configuration information items indicating "important" changes in an audio stream can be placed in a configuration extension structure in front of the stream identifier information such that a change in these parameters causes a transition from one stream to another. On the other hand, when comparing the information in the configuration structure related to one or more frames to be decoded with the current configuration information, some configuration information items can be left out of consideration, thereby changing the "dependent" configuration parameters of the audio decoder without triggering a "transition", i.e., a switch from one stream to another that may be linked to re-initialization. In other words, in the comparison, by evaluating only the configuration information items placed in the configuration extension structure in front of the stream identification information and the stream identification information itself, it is possible to avoid a change in the "dependent" decoding parameters causing a "transition". Rather, it is possible for the audio encoder to place such "dependent" configuration information items (related to the dependent decoding parameters) after the stream identifier information in the configuration extension structure. Then, the audio encoder can change such "dependent" configuration information items within the stream without triggering a "transition" (or re-initialization) with each change. On the other hand, these configuration information items that remain unchanged in the stream and changes to such "highly relevant" configuration information items (e.g., which may indicate a "significant" change in the audio stream) in front of the stream identifier information in the configuration extension structure will result in a "transition" (and typically re-initialization of the audio decoder). Since the audio decoder can also accept a variable ordering of the configuration information items in the configuration extension structure, the audio encoder can determine, depending on the signal characteristics or other criteria, which changes in the configuration information items cause a "transition" or re-initialization of the audio decoder and which changes in the configuration information items are possible within the stream without causing a "transition" or re-initialization of the audio decoder.
[0033] In a preferred embodiment, the audio decoder is configured to identify one or more configuration information items within a configuration extension structure based on one or more configuration extension type identifiers preceding each configuration information item. By using such configuration extension type identifiers, it becomes possible to implement a variable ordering of the configuration information items.
[0034] In a preferred embodiment, the configuration extension structure is a sub-data structure of the configuration structure, and the presence of the configuration extension structure is indicated by bits of the configuration structure that are evaluated by the audio decoder. The stream identifier information is a sub-data item of the configuration extension structure, and the presence of the stream identifier information is indicated by a configuration extension type identifier associated with the stream identifier information that is evaluated by the audio decoder. Thus, it is possible to flexibly determine when to add stream identifier information to an audio stream, and the audio decoder can easily determine when such stream identifier information is available. Thus, it is sufficient to include the stream identifier information (which requires a large number of bits) of the audio stream at points where there can be a switch between different streams. The immediate playback frames (IPFs) within a continuous audio stream do not need to convey the stream identifier information at positions where there is no possibility of switching between different streams, thus saving bitrate.
[0035] In a preferred embodiment, the audio decoder is configured to obtain and process an audio frame representation (e.g., an instant playback frame, IPF) that includes random access information (e.g., an "audio preroll extension payload" also denoted as "AudioPreRoll()"). The random access information includes a configuration structure (e.g., denoted as "Config()") and information (e.g., denoted as "AccessUnit()") for bringing the state of the audio decoder's processing chain to a desired state. When the audio decoder detects that the configuration structure of the random access information and the configuration information within (e.g., "Config()"), or a relevant portion of the configuration information of the configuration structure of the random access information, is different from the current configuration information, the audio decoder uses the configuration structure of the random access information to initialize the audio decoder and then uses the information for bringing the state of the processing chain to a desired state to adjust the state of the audio decoder. After that, a crossfade is performed between the audio information represented by the processed (decoded) audio frames before reaching the audio frame representation that includes the random access information (e.g., direct playback frame, IPF) and the audio information derived based on the audio frame representation that includes the random access information. For example, if the value of "numPreRollFrames" is zero, the decoding of the preroll frames can be omitted.
[0036] In other words, by evaluating the configuration information within the configuration structure, or its related parts (e.g., up to and including the stream identifier information), the audio decoder can recognize whether there is a transition between different streams, and if there is a transition between different streams, the audio decoder can utilize the random access information. The random access information helps to put the processing chain of the audio decoder in an appropriate state (usually affected by one or more previous frames in the absence of a transition), thereby avoiding artifacts during the transition. In conclusion, this concept enables artifact-free switching between different streams, and the audio decoder does not require any information from different protocol levels except for a series of frame representations.
[0037] In a preferred embodiment, when the audio decoder decodes the audio frame immediately preceding the audio frame represented by the audio frame representation including the random access information (e.g., the instant playback frame), and when the audio decoder detects that the related part of the configuration information in the configuration structure of the random access information is equal to the current configuration information, the audio decoder is configured to continue decoding without performing initialization of the audio decoder and without using information (e.g., pre-roll extended payload) for putting the state of the processing chain of the audio decoder in a desired state. Thus, when the audio decoder recognizes through comparing the related part of the configuration information within the configuration structure with the current configuration information that there is continuous playback of the same stream rather than a transition between different streams, the overhead (e.g., processing overhead or computational overhead) that would be caused by performing the initialization of the audio decoder is avoided. Thus, a high level of efficiency is achieved, and the initialization of the audio decoder is only performed when it is required.
[0038] In a preferred embodiment, the audio decoder uses the configuration structure of the random access information to perform initialization of the audio decoder, and when the audio decoder has not decoded the audio frame immediately preceding the audio frame represented by the audio frame representation including the random access information, the audio decoder is configured to adjust the state of the audio decoder using information for setting the state of the processing chain to a desired state. In other words, if there is an actual "random access" (where the audio decoder knows that the preceding audio frame has not been decoded), initialization is also performed. Thus, the random access information is used in the case of an actual "random access" (i.e., when jumping to a specific frame) and when switching between different streams (when an "actual" random access can be signaled to the audio decoder and when the switching between different streams can be recognized only by the audio decoder based on the evaluation of the stream identifier information).
[0039] It should be noted that the audio decoder described herein can be optionally added by any individual or combination of the features, functions, and details described herein.
[0040] Embodiments according to the present invention create an audio encoder for supplying an encoded audio signal representation. The audio encoder is configured to encode overlapping or non-overlapping frames of an audio signal using encoding parameters and obtain an encoded audio signal representation. The audio encoder is configured to supply a configuration structure that describes the encoding parameters (or, similarly, the decoding parameters to be used by the audio decoder). The configuration structure also includes a stream identifier.
[0041] Accordingly, the audio encoder supplies an audio signal representation that can be fully used by the above-described audio decoder. For example, the audio encoder can include different stream identifiers in different stream configurations. Accordingly, the stream identifier may not describe the decoder configuration (or decoding parameters) to be used by the audio decoder, but rather can be information that identifies the stream. Accordingly, the encoded audio signal representation includes a stream identifier, and different streams can be identified based on the encoded audio signal information itself without requiring information from different protocol levels. For example, since the stream identifier information is an essential part of the audio signal representation or the configuration included within the audio signal representation, the use of information supplied at the packet level is not necessary. As a result, as discussed herein, the audio decoder can recognize a switch between different streams even while the actual configuration parameters of the decoder remain unchanged.
[0042] In a preferred embodiment, the audio encoder is configured to include a stream identifier in an extended configuration of the configuration, and the extended configuration structure including the stream identifier can be enabled and disabled by the audio encoder. Accordingly, on the audio encoder side, it is possible to flexibly determine whether or not to include the stream identifier information. For example, the inclusion of the stream identifier information can be selectively omitted for audio frames where the audio encoder knows that there is no stream switch.
[0043] In a preferred embodiment, the audio encoder is configured to include an extended configuration type identifier that designates a stream identifier in the extended configuration structure in order to notify the presence of the stream identifier in the extended configuration structure. Accordingly, when there is other extended configuration information in the extended configuration structure, it is even possible to omit the stream identifier information. In other words, not all extended configuration structures necessarily need to include a stream identifier, which helps to save bits.
[0044] In a preferred embodiment, the audio encoder is configured to supply at least one configuration structure including a stream identifier and at least one configuration structure not including a stream identifier. Thus, if the audio encoder recognizes that this is necessary, the stream identifier is only included in the configuration structure. For example, the audio encoder may only include the stream identifier in the configuration structure of a frame where switching between streams is possible. By doing so, the bit rate can be kept quite small.
[0045] In a preferred embodiment, the audio encoder is configured to switch between supplying first encoded audio information represented by a sequence of first audio frames and supplying second encoded audio information represented by a sequence of second frames, where proper rendering of the first audio frame of the second sequence of audio frames after rendering of the last frame of the first sequence of audio frames requires re-initialization of the audio decoder. In this case, the audio encoder is configured to include in the audio frame representation representing the first audio frame of the second sequence of audio frames a configuration structure including a stream identifier associated with the second sequence of audio frames. The stream identifier associated with the second sequence of audio frames is selected to be different from the stream identifier associated with the first sequence of frames. Thus, the audio encoder can supply, within the configuration structure, signaling that enables the audio decoder to distinguish different streams and recognize when re-initialization (also referred to as "transition") should be performed.
[0046] In a preferred embodiment, the audio encoder does not supply any other signaling information indicating a switch from a first sequence of audio frames to a second sequence of audio frames, except for the stream identifier. Therefore, the bitrate can be kept quite small. In particular, it is possible to avoid having signaling included at different protocol levels other than the encoded audio information. Further, the audio encoder does not know in advance when the switch from the first sequence of audio frames to the second sequence of audio frames will actually occur. For example, the audio decoder first requests audio frames from the first sequence of audio frames, and when the audio decoder recognizes some necessity (for example, when there is an increase or decrease in the available bitrate), the audio decoder (or another control device that controls the supply of audio frames) can determine that audio frames from the second stream should be processed by the audio decoder. However, in some cases, it may happen that the audio decoder itself does not know when (or exactly when) there is a switch between the supply of audio frames from the first sequence and the supply of audio frames from the second sequence, and by evaluating the stream identifier included in the configuration structure, it will only be possible to recognize from which sequence of audio frames the currently received audio frames originated.
[0047] In a preferred embodiment, the audio encoder is configured to supply a first sequence of audio frames (e.g., a first stream) and a second sequence of audio frames (e.g., a second stream) using different bitrates (where the first stream and the second stream can represent the same audio content). Further, the audio encoder can be configured to indicate to the audio decoder the same decoder configuration information for decoding the first sequence of audio frames and for decoding the second sequence of audio frames, except for different bitstream identifiers. In other words, the audio encoder can indicate to the audio decoder to use the same decoder parameters, but the first stream and the second stream can still include different bitrates. This can be caused, for example, by using different quantization resolutions or different psychoacoustic models when supplying the first audio stream and the second audio stream. However, these different quantization resolutions or different psychoacoustic models do not affect the decoding parameters to be used by the audio decoder and only affect the actual bitrate. Thus, different bitstream identifiers can be the only possibility for the audio decoder to distinguish whether the audio frames to be decoded are from the first stream or from the second stream, and the evaluation of the bitstream identifier also makes it possible for the audio decoder to recognize when to perform a transition (or re-initialization).
[0048] Accordingly, the audio encoder can function in an environment where changes in the available bitrate can occur and the signaling overhead can be kept reasonably small.
[0049] Furthermore, it should be noted that the audio encoder described herein can optionally add any of the features, functions, and details described herein.
[0050] Another embodiment according to the present invention relates to a method of supplying a decoded audio signal representation based on an encoded audio signal representation. The method includes adjusting decoding parameters depending on configuration information, and the method includes decoding one or more audio frames using current configuration information (e.g., currently active configuration information). The method also includes comparing the configuration information in the configuration structure associated with the one or more frames to be decoded with the current configuration information, and the method includes the configuration information in the configuration structure associated with the one or more frames to be decoded or the relevant part of the configuration information in the configuration structure associated with the one or more frames to be decoded (e.g., up to and including the stream identifier) is different from the current configuration information, performing a transition (e.g., including re-initialization of decoding) to perform decoding using the configuration information in the configuration structure associated with the one or more frames to be decoded as the new configuration. This method also includes taking into account the stream identifier information included in the configuration structure when comparing the configuration information, and as a result, a difference between the stream identifier previously obtained in audio decoding and the stream identifier represented by the stream identifier information in the configuration structure associated with the one or more frames to be decoded causes a transition. This method is based on the same considerations as the audio decoder described above.
[0051] This method can add any of the features, functions, and details described herein, either individually or in combination.
[0052] Another embodiment according to the present invention creates a method of supplying an encoded audio signal representation. The method includes obtaining an encoded audio signal representation by encoding an overlapping or non-overlapping frame of an audio signal using encoding parameters. The method includes supplying a configuration structure that describes the encoding parameters (or equivalently, the decoding parameters to be used by an audio decoder), and the configuration structure includes a stream identifier. This method is based on the same considerations as the audio encoder as described above.
[0053] Furthermore, it should be noted that the methods described herein can have any of the features and functions described above added with respect to corresponding audio decoders and audio encoders. Additionally, the method can have any of the features, functions, and details described herein added, either individually or in combination.
[0054] Embodiments according to the present invention create an audio stream. The audio stream includes an encoded representation of either overlapping or non - overlapping frames of audio signals. The audio stream also includes a configuration structure that describes encoding parameters (or, equivalently, decoding parameters to be used by an audio decoder). The configuration structure includes stream identifier information that represents a stream identifier (e.g., in the form of an integer value).
[0055] The audio stream is based on the above considerations. In particular, the stream identifier included in the configuration structure of the audio stream that describes encoding parameters (or, similarly, decoding parameters) enables an audio decoder to distinguish different streams when the same encoding parameters (or decoding parameters) are used.
[0056] In a preferred embodiment, the stream identifier information is included in a configuration extension structure. In this case, the configuration extension structure is preferably a sub - data structure of the configuration structure, and the presence of the configuration extension structure is indicated by bits of the configuration structure. Further, the stream identifier information is a sub - data item of the configuration extension structure, and the presence of the stream identifier information is indicated by a configuration extension type identifier associated with the stream identifier information. The use of such an audio stream allows for flexible inclusion of stream identifier information whenever it is needed, while omission of the inclusion of stream identifier information is possible when it is not needed (e.g., in the case of frames where switching between multiple streams is not permitted). Thus, bitrate can be saved.
[0057] In a preferred embodiment, the stream identifier is embedded in a sub-data structure of the representation of the audio frame (and can be extracted from such a sub-data structure by an audio decoder). By embedding the stream identifier in the sub-data structure of the representation of the audio frame, the audio decoder can be avoided from having to use information from a higher protocol level. Rather, to decode the audio frame, the audio decoder only needs the representation of the audio frame and can determine whether there has been a switch between different streams.
[0058] In a preferred embodiment, the stream identifier is only embedded in a sub-data structure of the representation of the audio frame including a configuration structure (and can be extracted from the sub-data structure of the representation of the audio frame including the configuration structure by an audio decoder). This idea is based on the finding that switching between streams (without significant artifacts) can only be performed on frames including a configuration structure. Thus, it is sufficient to embed the stream identifier in the sub-data structure of the representation of the audio frame including the configuration structure, while it has been found that there is no stream identifier included in the representation of the audio frame not including the configuration structure.
[0059] The audio streams described herein can have any of the features, functions, and details described herein added to them, either individually or in combination. In particular, such functions described with respect to the audio encoder, audio decoder, and stream provider can also be applied to the audio streams.
[0060] Embodiments according to the present invention create an audio stream provider for supplying an encoded audio signal representation. The audio stream provider is configured to supply, as part of the encoded audio signal representation, an encoded version of frames of the audio signal that are temporally overlapping or non-overlapping, encoded using encoding parameters. The audio stream provider is configured to supply a configuration structure that describes the encoding parameters (or, equivalently, the decoding parameters to be used by the audio decoder) as part of the encoded audio signal representation, and the configuration structure includes a stream identifier. This audio stream provider is based on the same considerations as the above-described audio encoder and the above-described audio decoder.
[0061] In a preferred embodiment, the audio stream provider is configured to supply the encoded audio signal representation such that the stream identifier is included in a configuration extension structure of the configuration structure, and the configuration extension structure including the stream identifier can be enabled and disabled by one or more bits within the configuration structure. This embodiment is based on the same idea as described above for both the audio encoder and the audio decoder. In other words, the audio stream provider supplies an audio stream corresponding to the audio stream supplied by the audio encoder (even if the audio stream provider is configured to switch the supply of different streams, such as supplied by a plurality of audio encoders operating in parallel or supplied from a storage medium).
[0062] In a preferred embodiment, the audio stream provider is configured to supply the encoded audio signal representation such that the configuration extension structure includes a configuration extension type identifier that designates the stream identifier to indicate the presence of the stream identifier within the configuration extension structure. This embodiment is based on the same considerations as described above for the audio encoder and the audio stream.
[0063] In a preferred embodiment, the audio stream provider is configured to supply an encoded audio signal representation such that the encoded audio signal representation includes at least one structural component that includes a stream identifier and at least one structural component that does not include a stream identifier. As described above, it is not necessary for the stream identifier to be included in each structural component. Rather, there can be a flexible adjustment as to which structural components to include the stream identifier in. Typically, the stream identifier will be included in the structural components of the audio frames where there is (or is expected or permitted to be) a switch between streams. In other words, switching between different streams that include the same structural components, except for different stream identifiers, will only be performed by the stream provider at frames where the stream identifier is present. Thus, even if the decoding parameters (indicated by the structural components) are substantially identical or even completely identical, an audio decoder (which receives the encoded audio representation from the audio stream provider) may recognize a switch between different streams.
[0064] In a preferred embodiment, the audio stream provider is configured to switch between providing a first portion of encoded audio information represented by a first sequence of audio frames and providing a second portion of encoded audio information represented by a second sequence of audio frames, and properly rendering the first audio frame of the second sequence of audio frames after rendering the last frame of the first sequence of audio frames requires re-initialization of the audio decoder. The audio stream provider is configured to provide an encoded audio signal representation such that the audio frame representation representing the first frame of the second sequence of audio frames includes a structural configuration including a stream identifier associated with the second sequence of audio frames, where the stream identifier associated with the second sequence of audio frames is different from the stream identifier associated with the first sequence of audio frames. In other words, the audio stream provider switches between two audio streams (sequences of audio frames) having different associated stream identifiers. Thus, the audio decoder typically knows (e.g., by evaluating the structural configuration associated with the first sequence of audio frames) the stream identifier associated with the first sequence of audio frames, and when the audio decoder receives the first frame of the second sequence of audio frames, the audio decoder can evaluate a structural configuration including the stream identifier associated with the second sequence of audio frames and can recognize the switch from the first stream to the second stream by comparing the stream identifiers (which are different for each stream). Thus, the audio stream provider provides audio frames from the first stream and then switches to providing audio frames from the second stream, and provides appropriate signaling information, i.e., a stream identifier, within the structural configuration of the first frame of the second audio stream provided after the switch. Thus, no additional signaling is required to signal the switch between different audio streams.
[0065] In a preferred embodiment, the audio stream provider is configured to supply an encoded audio signal representation such that it does not supply other signaling information indicating a switch from a first sequence of audio frames to a second sequence of audio frames, excluding the stream identifier, of the encoded audio signal representation. Thus, a significant savings in bitrate can be achieved. Also, since it contains information at different protocol levels and there is no need to extract such information from different protocol levels on the audio decoder side, the protocol complexity is also kept low.
[0066] In a preferred embodiment, the audio stream provider is configured to supply an encoded audio signal representation such that a first sequence of audio frames (e.g., a first stream) and a second sequence of audio frames (e.g., a second stream) are encoded using different bitrates. Further, the audio stream provider is configured to supply the encoded audio signal representation such that the encoded audio signal representation, excluding different bitstream identifiers, indicates the same audio decoder for decoding the first sequence of audio frames and for decoding the second sequence of audio frames (or decoder parameters, or decoding parameters). Thus, the audio stream provider supplies very similar configuration information for different streams (the first stream and the second stream), which may differ only, for example, by the bitstream identifier. In this scenario, using the bitstream identifier is particularly useful because it can reliably distinguish different bitstreams while minimizing signaling overhead.
[0067] In a preferred embodiment, the audio stream provider is configured to switch between supplying a first sequence of audio frames (e.g., a first stream) to an audio decoder and a second sequence of audio frames (e.g., a second stream), where the first sequence of audio frames and the second sequence of audio frames are encoded using different bitrates. The audio stream provider is configured to selectively switch between supplying the first sequence of audio frames and the second sequence of audio frames in an audio frame representation (e.g., an instant playback frame, IPF) while avoiding switching between sequences in audio frames that do not contain random access information. The audio stream provider is configured to supply an encoded audio signal representation such that it is included in the composition structure of the audio frames supplied when the stream identifier switches from the first sequence of audio frames to the second sequence of audio frames. For example, when the first frame of the second sequence of audio frames includes a composition structure that also has a stream identifier and random access information, such a configuration of the audio stream provider ensures that there is only a switch between supplying frames from the first sequence of audio frames and supplying frames from the second sequence of audio frames. As a result, the audio decoder can detect a switch between different audio streams and thus can recognize that random access information should be evaluated (whereas random access information is not normally evaluated when there is no switch between different audio streams and when the audio decoder is assumed to be rendering a continuous sequence of audio frames of a single stream).
[0068] Thus, good audio quality without artifacts can be achieved by such a concept when switching between different audio streams.
[0069] In a further embodiment, the audio stream provider is configured to obtain a plurality of parallel sequences of audio frames encoded using different bitrates, the audio stream provider is configured to switch the supply of frames to an audio decoder from different parallel sequences, and the audio stream provider is configured to use a stream identifier included in the configuration structure of the first audio frame representation supplied after the switch to indicate to the audio decoder which sequence one or more frames are associated with. Thus, the audio decoder can recognize transitions between different streams with little overhead and without using information from other protocol layers.
[0070] It should be noted that the audio stream provider described herein can add any of the features, functions, and details described herein, either individually or in combination.
[0071] Another embodiment according to the present invention creates a method for supplying an encoded audio signal representation. The method includes supplying, as part of the encoded audio signal representation, an encoded version of an overlapping or non-overlapping frame of the audio signal encoded using encoding parameters. The method includes supplying a configuration structure that describes encoding parameters (or, equivalently, decoding parameters to be used by an audio decoder) as part of the encoded audio signal representation, the configuration structure including a stream identifier.
[0072] This method is based on the same considerations as the stream provider described above. This method can add any of the other features, functions, and details described herein, for example, with respect to an audio encoder, an audio decoder, or an audio stream, rather than with respect to a stream provider.
[0073] Another embodiment according to the present invention creates a computer program for executing the method described herein.
Brief Description of the Drawings
[0074] Embodiments according to the present invention will be described hereinafter with reference to the accompanying drawings.
[0075]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10a
Figure 10b
Figure 10c
Figure 10d
Figure 11a
Figure 11b
Figure 11c
Mode for Carrying Out the Invention
[0076] 1. Audio Decoder According to FIG. 1 FIG. 1 shows a block schematic diagram of an audio decoder according to a (simple) embodiment of the present invention.
[0077] The audio decoder 100 receives the encoded audio signal representation 110 and, based thereon, supplies the decoded audio signal representation 112. For example, the encoded audio signal representation 110 can be an audio stream including a series of Unified Speech and Audio Coding (USAC) frames. However, the encoded audio signal representation can take different forms and can be, for example, an audio representation defined by the bitstream syntax of any of the known audio coding standards. The encoded audio signal representation can, for example, be included in a configuration structure and can include configuration information 110 that can include, for example, a stream identifier. The stream identifier can be included, for example, in the configuration information or the configuration structure. The configuration information or the configuration structure can be associated, for example, with one or more frames to be decoded and can describe, for example, decoding parameters used by the audio decoder.
[0078] Here, the decoder 100 can include a decoder core 130 that is configured to decode one or more audio frames, for example, using current configuration information (the current configuration information can define, for example, decoding parameters). The audio decoder is also configured to adjust the decoding parameters depending on the configuration information 110a.
[0079] For example, an audio decoder is configured to compare configuration information within a configuration structure associated with one or more frames to be decoded with current configuration information (e.g., configuration information used for decoding one or more previously decoded frames). Further, if the configuration information within the configuration structure associated with one or more frames to be decoded, or a relevant portion of the configuration information within the configuration structure associated with one or more frames to be decoded, is different from the current configuration information, the audio decoder may be configured to transition to performing decoding using the configuration information within the configuration structure associated with one or more frames to be decoded as new configuration information. When making the "transition", the audio decoder can, for example, re-initialize the decoder core 130 using random access information, which is intended to describe the state of the decoder core to be used for properly decoding the audio frame (or the first audio frame) after the "transition".
[0080] In particular, the audio decoder is configured to consider the stream identifier included in the configuration structure (i.e., within the configuration information) when comparing the configuration information such that a difference between a stream identifier previously obtained by the audio decoder and a stream identifier represented by stream identifier information within the configuration structure associated with one or more frames to be decoded causes a transition (i.e., when comparing the configuration information within the configuration structure associated with one or more frames to be decoded with the current configuration information).
[0081] In other words, the audio decoder may include memory for the current configuration (or current configuration information) that can be specified, for example, by 140. The audio decoder 100 may also include a comparator (or any other means for performing the comparison) 150 that can compare at least the relevant part of the current configuration information including the stream identifier with the corresponding part of the configuration information associated with the next (audio) frame to be decoded that includes the stream identifier. The relevant part is, for example, up to and including the stream identifier, and in some embodiments, the configuration information after the stream identifier in the bitstream representing the configuration information may be ignored.
[0082] If this comparison that can be performed by the comparator 150 indicates a difference between the current configuration information (or its relevant part) and the configuration information associated with the next (audio) frame to be decoded (or its relevant part), it may be recognized that a "transition" should be made.
[0083] Performing a transition may include, for example, re-initializing the decoder core even if the decoding parameters described by the configuration information associated with the next (audio) frame to be decoded are the same as the decoder configuration (decoding parameters) described by the current configuration information (where the configuration information associated with the next audio frame to be decoded differs from the current configuration information only in that the stream identifier is different). On the other hand, for example, by defining different decoding parameters, if the configuration information associated with the next audio frame to be decoded is further different from the current configuration information, the audio decoder 100 will of course also perform a "transition" in the ordinary sense of re-initializing the decoder core 130 and changing the decoding parameters.
[0084] In conclusion, the audio decoder 100 according to FIG. 1 can recognize transitions between frames of different audio streams even if the decoding parameters to be used by the decoder core 130 remain unchanged by evaluating the stream identifier included in the structural configuration of the audio frame, which eliminates the need for dedicated signaling of transitions between audio streams and / or conditions for re-initializing the decoder core. Therefore, the audio decoder can appropriately handle such transitions, for example, by re-initializing the audio decoder and (if necessary) reconfiguring the audio decoder with new setting parameters, so that the decoder 100 can appropriately decode audio frames even if there is a transition from one stream to another.
[0085] It should be noted that the audio decoder 100 according to FIG. 1 can be optionally added by any of the features, functions, and details described herein, individually or in combination.
[0086] 2. Audio Decoder According to FIG. 2 FIG. 2 shows a block schematic diagram of an audio decoder 200 according to an embodiment of the present invention.
[0087] The audio decoder 200 is configured to receive an encoded audio signal representation 210 and, based thereon, supply a decoded audio signal representation 212. The encoded audio signal representation 210 can be, for example, an audio stream including a series of Unified Speech and Audio Coding (USAC) frames. However, a sequence of audio frames encoded using different audio coding concepts may also be input to the audio decoder 200. For example, the audio decoder can receive an audio frame 220 of a first stream and subsequently (as the next audio frame) an audio frame 222 of a second stream. The audio frames 220, 222 can be supplied, for example, by an audio stream provider. The audio frame 220 can include, for example, an encoded representation 220a of the audio signal in the form of encoded spectral values and encoded scale factors and / or in the form of encoded spectral values and encoded transform coding (TXC) coefficients and / or in the form of encoded excitation and encoded transform coding coefficients. The audio frame 222 can also include, for example, an encoded representation 222a of the audio signal that can be in the same format as the encoded representation 220a of the audio signal included in the frame 220. However, furthermore, the frame 222 can also include random access information 222b, which can include information 222d for bringing about a configuration structure 222c and the state of a processing chain (e.g., a decoder core) to a desired state. This information 222d can be indicated, for example, as "AudioPreRoll".
[0088] The audio decoder 200 can, for example, extract the configuration structure 222c from the encoded audio signal representation 210, which can also be regarded as configuration information. The configuration structure 222c can include, for example, information or a flag (or bit) indicating whether a configuration extension structure 226 exists as part of the configuration structure. This information or flag or bit is indicated by 224a.
[0089] The configuration extension structure 226 may include, for example, information or a flag or a bit or an identifier indicating whether a stream identifier exists. The latter information, flag, bit or identifier is indicated by 228. If the information or flag or bit or identifier 228 indicates the existence of a stream identifier, the stream identifier 230 also exists, which may typically be part of the configuration extension structure 226.
[0090] Furthermore, the configuration extension structure can include information as to whether there is other information, such as appropriate bits, flags, or identifiers, and can also include other information (if applicable).
[0091] The audio decoder 100 can include, for example, a memory 240 that can store current configuration information (e.g., configuration information used for decoding the previous frame and extracted from the configuration structure of the previous frame or a preceding frame). The audio decoder 200 can also include a comparator or comparer 250 configured to compare the configuration information related to the audio frame to be decoded with the current configuration information stored in the memory 240. For example, the comparator or comparer 250 can be configured to compare the configuration information of the configuration structure 222c of the audio frame to be decoded with the current configuration information stored in the memory up to and including the stream identifier. In other words, any information item of the configuration structure 222c up to and including the stream identifier is compared with the current configuration information from the memory 240 to determine whether the configuration information (up to and including the stream identifier) within the frame 222 is the same as the current configuration information extracted from one of the previous audio frames. In this comparison, it is naturally checked whether the configuration structure 222c actually includes the configuration extension structure 226 and the stream identifier 230. If the configuration extension structure 226 does not exist, it cannot naturally be considered in the comparison. Also, if the stream identifier 230 does not exist (e.g., to indicate that the flag 228 is not included in the frame 222), it is naturally not evaluated in the comparison. Also, the configuration information after the stream identifier 230 within the configuration structure 222c has a low importance, and a change in such configuration information after the stream identifier 230 within the configuration structure 222c is assumed to occur not only between different streams but also within a single stream and does not indicate a switch between different streams, so it is usually ignored in the comparison.
[0092] As a result, comparison 250 typically compares the configuration information including up to and including the stream identifier of the audio frame to be decoded (where preferably the configuration placed after the stream identifier in the configuration extension structure is omitted) with the current configuration information (obtained from the previously decoded audio frame). Thus, comparison 250 detects a new stream (or sub-stream) if there is a difference in the configuration information found in the comparison. Thus, the comparison is used to control the transition from the first stream (or sub-stream) to the second stream (or sub-stream).
[0093] For example, such a transition may occur by causing the decoding of the last frame of the first stream, the reconstruction, the initialization of the processing chain to a desired state, and the execution of a cross-fading, for example, between the time-domain representation of the last frame of the first stream and the first frame of the second stream.
[0094] The audio decoder 200 also includes a decoder core 216 that may be configured to decode frames of the first stream (or the first sequence of frames) using the first configuration (which may be described by the current configuration information). Further, the decoder core 216 may be configured to decode the second stream or the second sequence of frames using the second configuration (for example, using the new configuration described by the configuration information 222c of the audio frame to be decoded). For example, re-initialization of the decoder core may be triggered by comparison 250 when a significant difference between the configuration information 222c of the audio frame 222 to be decoded and the current configuration information in the memory 240 is detected.
[0095] For example, re-initialization of the decoder may be used between decoding of the last frame of the first stream and decoding of the first frame of the second stream. Alternatively, for example, if the decoder is (at least partially) implemented in software, a "new instance" of the decoder may be used. Further, when switching from decoding of the first stream to decoding of the second stream ("transition"), the state of the processing chain of the decoder core can be brought to a desired state using some side information. For example, the context state of arithmetic decoding can be set to a desired state, or the content of a temporal discrete filter can be set to a desired state. This can be done using dedicated information also referred to as "audio preroll" APR. Since the first frame of the second stream processed (decoded) by the audio decoder may not be the actual first frame of the second audio stream, it is important to bring the state of the processing chain to a desired state. Rather, the first frame of the second audio stream processed by the audio decoder may be some frames between the second audio streams when the audio stream provider switches from supplying frames from the first audio stream to supplying frames from the second audio stream. Thus, the "first frame of the second audio stream" processed by the audio decoder may depend on a specific setting of the state of the decoding chain normally caused by decoding of frames preceding the second audio stream (which is the first audio frame of the second audio stream to be handled by the audio decoder after the transition). Thus, when switching from decoding of the audio frames of the first audio stream to decoding of the audio frames of the second audio stream, the lack of setting of the state of the audio decoder normally brought about by decoding of the frames preceding the second audio stream is created using "audio preroll" information that defines an appropriate setting of the state of audio decoding.
[0096] As can be seen in reference numeral 270, the decoding of the last frame of the first audio stream supplies a decoded portion 272 (also shown as the "useful portion"). Optionally, the decoding of the last frame of the first audio stream can supply a longer decoded portion, which is partially discarded. Further, when decoding the first frame of the second audio stream, a "preroll portion" 274 is provided while the decoder state is initialized for proper decoding of the first frame of the second audio stream. Further, the decoder core 260 also supplies a useful portion 276 of the first frame of the second audio stream handled by the decoder 200, and the useful portion 276 of the first frame of the second audio stream temporally overlaps with the useful portion 272 of the last frame of the first stream. Thus, an optional crossfade can be performed between the end of the useful portion 272 of the last frame of the first stream and the start of the useful portion of the first frame of the second stream. Thus, a decoded output signal 212 can be derived, with an artifact-free transition between the last frame of the first stream (handled by the audio decoder 200) and the first frame of the second stream (handled by the audio decoder 200).
[0097] In summary, the audio decoder 200 is an audio encoder or an audio s The team supplier can recognize when switching from supplying audio frames of the first stream to supplying audio frames of the second stream. For this purpose, the audio decoder evaluates the configuration information 222c (also called the configuration structure) and performs a comparison with the current configuration information stored in the memory 240. When it is recognized that the audio frames to be decoded belong to a different audio stream compared to the previously decoded audio frames, re-initialization of the decoder core is performed, which usually includes evaluating the "audio preroll" information to bring the state of the processing chain of the decoder core to a desired state. Thus, the audio decoder can appropriately handle the situation where the audio encoder or the audio stream supplier supplies audio frames from a new stream (the second audio stream) without further notice (except for the supply of the configuration structure 222c including the stream identifier 230).
[0098] It should be noted that the audio decoder 200 described herein can add any of the features, functions, and details described herein individually or in combination.
[0099] 3. Audio Encoder According to FIG. 3 FIG. 3 shows a block schematic diagram of an audio encoder according to an embodiment of the present invention.
[0100] The audio encoder 300 receives an input audio signal 310 (e.g., in the form of a time-domain representation) and supplies an encoded audio signal representation 312 based thereon. The audio encoder 300 includes an encoder core 320 configured to encode overlapping or non-overlapping frames of the input audio signal 310 using encoding parameters and obtain the encoded audio signal representation. The audio encoder 320 may include, for example, a conversion from the time domain to the spectral domain and an encoding of the spectral domain representation. The processing may be performed, for example, on a per-frame basis.
[0101] Furthermore, the audio encoder may include a configuration structure supply 330 configured to supply, for example, a configuration structure 332 that describes encoding parameters (or equivalently, decoding parameters to be used by the audio decoder). The configuration structure 332 may correspond to, for example, the configuration structure 222c. In particular, the configuration structure 332 may include encoding parameters (such as an encoding format) or equivalently, decoding parameters (such as an encoding format) that describe settings to be used by a decoder (or decoder core) when decoding the encoded audio signal representation 312. Examples of the configuration structure 332 will be described below. Furthermore, the configuration structure 332 may include a stream identifier corresponding to the stream identifier 230. For example, the stream identifier can specify an audio stream (for example, a continuous portion of audio content continuously encoded using specific encoder settings). For example, the stream identifier supplied by the configuration structure supply 330 can be selected such that all audio streams that can be switched without artifacts and without explicitly notifying the audio decoder about the switch convey different stream identifiers. However, in some cases, it may be sufficient if different stream identifiers are included in streams having the same relevant encoding parameters (or equivalently, decoding parameters to be used in the audio decoder). In other words, different stream identifiers may only be required for streams where other encoding or decoding parameters are the same.
[0102] Accordingly, the encoder control 340 can control, for example, both the encoder core 320 and the configuration structure supply 330. The encoder control 340 can determine, for example, the encoding parameters (which may at least partially correspond to the decoding parameters to be used by the audio decoder) to be used by the encoder core 320, and can also notify the configuration structure provision 330 regarding the encoding / decoding parameters to be included in the configuration structure 332. Accordingly, the encoded audio representation 312 includes the encoded audio content and the configuration structure 332 as well. Accordingly, an audio decoder (e.g., audio decoder 100 or audio decoder 200) can immediately recognize when different audio streams encoded using different encoding parameters are supplied (even if not all encoding parameters are reflected in the decoding parameters included in the configuration structure).
[0103] Regarding this issue, it should be noted that it is not usually necessary to show all encoding parameters to the audio decoder. For example, it is only necessary to show the encoding parameters to the audio decoder that affect the decoding algorithm. The encoding parameters sent to the audio decoder to determine the settings of the audio decoder are also shown as decoding parameters. On the other hand, some important encoding parameters are usually not notified to the audio decoder, but rather are implicitly reflected in the encoded audio signal representation. For example, the desired bit rate can be an important encoding parameter, and it may be possible to determine how coarsely the audio encoder quantizes the spectral values and / or how many spectral values the audio quantizes even to small or zero values. However, for the audio decoder, it is sufficient to simply check the result of the encoding, and it is not necessary to know the specific strategy of the encoder that keeps the bit rate moderately small. Also, depending on the type of audio content and the actually required bit rate, there can be various approaches on the encoder side to achieve a sufficiently small bit rate. These parameters can be regarded as "encoding parameters", but they are not reflected in the set of "decoding parameters" (nor are they included in the encoded representation of the audio frame). Decoding parameters (and the encoding parameters incorporated into these encoded audio representations) usually only describe the settings used by the decoder, i.e., how to process the encoded information supplied by the encoder.
[0104] Therefore, in practice, even if the encoder core is using different encoding parameters, the decoding parameters that may be included in the setting structure 332 may be the same (e.g., regarding the target bit rate or regarding parameters that affect the target bit rate such as quantization resolution or psychoacoustic model).
[0105] In other words, the audio encoder may use different encoding parameters to encode a particular audio content, even though the decoding parameters to be used by the decoder, for example, may be the same (for processing and decoding the encoded representation of the audio content).
[0106] In such a case, the audio encoder may supply different stream identifiers within the configuration structure 332 so that the audio decoder can still distinguish such different encoded representations of the audio content.
[0107] Furthermore, it should be noted that the audio encoder 300 according to FIG. 3 can be optionally added by any of the features, functions, and details described herein.
[0108] 4. Audio Stream Provider According to FIG. 4 FIG. 4 shows a block schematic diagram of an audio stream provider according to an embodiment of the present invention.
[0109] The audio stream provider 400 is configured to supply an encoded audio signal representation 412. The audio stream provider is configured to supply, as part of the encoded audio signal representation 412, an encoded version 422 of the (temporarily) superimposed or non-superimposed frames of the audio signal encoded using the encoding parameters.
[0110] Furthermore, the audio stream provider is configured to supply a configuration structure 424 that describes the encoding parameters (or, equivalently, the decoding parameters to be used by the audio decoder) as part of the encoded audio signal representation, and the configuration structure 424 includes a stream identifier.
[0111] For example, an audio stream provider may include a supply (or provider) of an encoded version of an audio signal's superimposed or non-superimposed frames. Further, the audio stream provider may comprise a configuration structure supply or a configuration structure provider 423 for supplying the configuration structure 424.
[0112] Accordingly, the audio stream provider may supply, as part of the encoded audio signal representation 412, a part of various audio streams that the audio stream provider may store in a memory, for example, or receive from an audio encoder. When supplying a part of the first audio stream and then switching to supplying a part of the second audio stream, the configuration structure 424 may be associated with the first audio frame of the second audio stream supplied after switching from the first audio stream to the second audio stream. The configuration structure 424 may be, for example, a part of each audio stream received by the audio stream provider from an audio encoder or stored in the memory of the audio stream provider. Accordingly, the audio stream provider may, for example, store a continuous sequence of audio frames of the first audio stream and a continuous sequence of audio frames of the second audio stream. At least some of the frames of the first audio stream and some of the frames of the second audio stream may have respective associated configuration structures that describe the decoding parameters to be used by the audio decoder. The configuration structure may also include a respective stream identifier, for example, an integer that identifies the audio stream. For example, the audio stream provider may be configured to supply frames 1 to n - 1 (1 to n - 1 may be time indices) for the first audio frame and supply frames n to n + x (n to n + x may be time indices) of the second audio stream as part of the encoded audio signal representation 412, where the frames 1 to n - 1 of the second audio stream may not be supplied as part of the encoded audio signal representation 412 directed to a particular audio decoder or a particular group of audio decoders. The first audio stream and the second audio stream may represent, for example, the same content encoded at different bitrates.Accordingly, frames 1 to n-1 of the audio content are represented by an encoded audio signal representation 412 directed to a specific device or group of devices encoded at a first bitrate by a first audio stream, and frames n to n+x of the audio content are represented by frames n to n+x of a second audio stream encoded at a second bitrate different from the first bitrate.
[0113] For example, an audio stream provider 400, or some external control, may ensure that the first frame n of the second audio stream included in the encoded audio signal representation 412 contains a configuration structure. In other words, for example, the switching between the supply of audio frames from the first audio stream and the supply of audio frames from the second audio stream may be guaranteed to occur only at "appropriate" frames, which includes a configuration structure and preferably also includes some information (such as an audio preroll, etc.) for initializing the audio decoder.
[0114] Accordingly, the audio stream provider can supply, for example, a part of the audio content encoded at a first bitrate (e.g., by supplying frames 1 to n - 1 of the first audio stream) and another part of the audio stream encoded using a second bitrate (e.g., by supplying audio frames n to n + x of the second audio stream). Probably, the structural composition of the first audio stream and the second audio stream will be the same except for the fact that the stream identifiers are different. This is actually only the stream identifier, which is also included in the structural composition, so that the audio decoder can determine whether to perform a "transition" (e.g., by re-initializing the decoder core). Thus, the decoding parameters reflected in the structural composition 424 do not necessarily have to reflect the different encoding parameters (or all encoding parameters) used for the encoding of the first audio stream and the encoding of the second audio stream, due to the fact that the stream identifier is the only thing that actually matters and is included in the structural composition.
[0115] In some embodiments, the decision of whether to supply an audio frame from the first audio stream or the second audio stream may be made by the audio stream provider (e.g., based on knowledge of network conditions, such as network load or the available network bitrate of the network between the audio stream provider and the audio decoder). However, alternatively, the audio decoder, or an intermediate device (e.g., a network management device), may determine the audio stream to use.
[0116] However, it should be noted that the audio decoder or at least the audio decoder core may not be explicitly notified by the audio stream provider and / or the intermediate network that a change in the stream has occurred. In other words, except for the configuration structure 424, the audio decoder does not receive additional information indicating that frames n through n+x are from a second audio stream and frames 1 through n-1 are from a first audio stream.
[0117] In conclusion, the audio stream provider can flexibly supply the encoded representation of the audio content to the audio decoder in the form of the encoded audio signal representation. The audio stream provider can, for example, flexibly switch the supply of encoded frames from a first audio stream and encoded frames from a second audio stream, where the switch between audio streams is indicated by a change in the stream identifier included in the configuration structure 424 that is part of the encoded audio signal representation 412.
[0118] Here, it should be noted that the audio stream provider 400 can be optionally added by any of the features, functions, and details described herein.
[0119] Hereinafter, an example of the function of the audio stream provider 400 will be described with reference to FIG. 5, which shows a block schematic diagram of the audio stream provider according to an embodiment of the present invention.
[0120] The audio stream provider shown in FIG. 5 is denoted by 500 and may correspond to the audio stream provider 400 according to FIG. 4. The audio stream provider 500 is configured to supply an encoded audio signal representation 512 that may correspond to the encoded audio signal representation 412.
[0121] In particular, the audio stream provider may be configured to switch between supplying frames from a first audio stream and supplying frames from a second audio stream. For example, the audio stream provider 500 may be configured to switch between supplying frames from a first audio stream and supplying frames from a second audio stream only at so-called "independent playout frames" (also referred to as "IPF").
[0122] The audio stream provider 500 may be stored in a memory or may receive a first audio stream 520 and a second audio stream 530 from an audio encoder. The first audio stream may be encoded, for example, at a first bitrate and may include a first stream identifier in its configuration structure (e.g., of immediate playout frames). The second audio stream 530 may be encoded at a second bitrate and may include a second stream identifier in its configuration structure (e.g., of immediate playout frames). However, the first audio stream and the second audio stream may represent, for example, the same audio content. However, the first audio stream and the second audio stream can also represent different audio contents.
[0123] For example, the first audio stream 520 may include independent playout frames in the frames indicated by n 1 、n 2 、n 3 、and n 4 . For example, one or more "normal" audio frames that are not independent playout frames can be placed between two adjacent independent playout frames. However, depending on the situation, independent playout frames may also be adjacent.
[0124] Similarly, the second audio stream 530 has frame positions n 1 、n 2 、n 3 and n4 also includes a playback frame independent thereof.
[0125] It should be noted that the positions of the independent playback frames in the two streams 520, 530 may optionally be the same or different. For simplicity, it is assumed here that the frame positions of the independent playback frames are the same in both streams.
[0126] However, in principle, only that the first frame after the switch is an independent playback frame is important. For example, when switching from the supply of audio frames of the first audio stream to the supply of audio frames from the second audio stream, the audio stream supplier 500 needs to ensure that the first frame of a part of the frames supplied from the second audio stream is an independent playback frame.
[0127] The embodiment will be described with reference to the encoded audio signal representation denoted by reference numeral 550. As can be seen by reference, the encoded audio signal representation 512 includes, at its start, a portion 552 including one or more frames of the first audio stream. However, the index n of the first audio stream 1 After supplying the audio frame having - 1, the audio stream supplier 500 may decide to switch to the second audio stream (based on an internal decision or based on some control information received from outside). Accordingly, a part 554 of the audio frames of the second audio stream is supplied into the encoded audio signal representation 512. For example, the frames having frame indices from n to n 1 to n 2 - 1 of the second audio stream are supplied to the portion 554 within the encoded audio signal representation 512. The first frame of the portion 554 is an independent playback frame, and it should be noted that it is at the frame index n within the second audio stream 530. However, the frame index n 1 to2 - When a frame with - 1 is supplied into the encoded audio signal representation 512, the audio stream provider may decide to return to supplying audio frames from the first audio stream 520 again. Thus, the frame index n (based on the second audio stream 530) 2 - After (or immediately after) an audio frame with - 1, the frame index n obtained from the first audio stream 520 2 A frame with may be supplied into the encoded audio signal representation. The index n 2 It should be noted that the frame with index n is also an independent playback frame. Thus, the part from the first audio stream 2 starts from a frame with index n and ends with frame index n 4 - 1.
[0128] In conclusion, the encoded audio signal representation 512 is a concatenation of parts of one or more frames, where some parts of the frames are obtained from the first audio stream 520 and some parts of the frames are obtained from the second audio stream 530. The first frame of each part is preferably an independent playback frame, which is preferably guaranteed by the operation of the audio stream provider.
[0129] Such an independent playback frame preferably includes a structural configuration with a stream identifier, and the stream identifier may be included in, for example, a configuration extension structure. For example, the configuration information of the first stream and the second stream may be the same except for the stream identifier (and perhaps, except for the configuration information included in the configuration extension structure after the stream identifier).
[0130] For example, the independent playback frame may correspond to frame 220 as described above with respect to the audio decoder 200.
[0131] As a further conclusion, the audio stream provider 500 can access a plurality of audio streams (e.g., a first audio stream 520 and a second audio stream 530, and optionally further audio streams), and select portions of frames to be transferred to an audio decoder (e.g., via a communication network) for inclusion in an encoded audio signal representation 512 from two or more of these audio streams. When selecting portions of frames to be included in the encoded audio signal representation 512, the audio stream provider can ensure that the first frame of each portion is an independent playback frame containing sufficient information for (artifact-free) rendering without decoding the previous frame of that audio stream. Further, the audio stream provider supplies an encoded audio signal representation such that a switch between portions of audio frames from different streams is recognizable by the audio decoder receiving the encoded audio signal representation 512 from differences within the relevant portions of the configuration structure. In some transitions, the configuration structure may differ with respect to decoder configuration parameters, but in the case of one or more other transitions, the configuration structure may differ only in the stream identifier and other decoding configuration parameters may be the same.
[0132] As a result, the audio decoder can recognize a switch between different audio streams and, when appropriate, perform re-initialization ("transition") at any time.
[0133] 5. Audio Frame According to FIG. 6 FIG. 6 shows a representation of an audio frame that enables random access and includes a component with a stream identifier in an extended configuration portion.
[0134] For example, FIG. 6 shows an example of an audio frame that can take over the role of the audio frame 222 described with reference to FIG. 2. For example, the audio frame can be a "USAC frame". The audio frame in FIG. 6 can be regarded as a "stream access point" or an "intermediate playback frame".
[0135] The frame can, for example, comply with the syntax rules of an audio coding standard including available modifications, but can also be adapted to the bitstream syntax of other or newer audio standards.
[0136] For example, the USAC frame 600 may include a USAC independent flag 610. Further, the USAC frame may include an extension element designated as a "USAC ExtElement". The extension element 620 may be an extension element with configuration information and preroll data.
[0137] Optionally, there may be a flag "USAC ExtElementPresent" indicating the presence of additional data. For example, in the case of an IPF (e.g., a stream access point), this flag is preferably 1. However, this flag can be regarded as optional.
[0138] Further, optionally, there may be a flag "USAC ExtElementUseDefaultLength" that can be used to encode whether to use the default length of the extension element or to encode the length of the extension element. For example, in the case of an IPF, the value of this flag is preferably zero (but not mandatory).
[0139] Furthermore, there is extended element segment data, also denoted as "USACExtElementSegmentData". These extended element segment data include audio preroll information, also denoted as "AudioPreRoll()" in the USAC standard revision. The audio preroll optionally includes configuration length information "configLen" and configuration information "Config()", and the configuration information may be the same as the "USAC configuration information", also denoted as "UsacConfig()". When the configuration information exists, "configLen" needs to take a value greater than zero, but preferably it is not necessarily so. For example, a zero value of "config Len" may indicate the absence of configuration information. The configuration information can include some basic configuration information such as information regarding the sampling frequency, information regarding the SBR frame length, information regarding the channel configuration, and the number of other (optional) decoder configuration items. The other decoder configuration items can include, for example, one or more or all of the configuration items described in the definition of the "UsacDecoderConfig()" syntax element of the USAC standard.
[0140] Furthermore, the configuration information includes a configuration extension structure as a sub-data structure. The configuration extension structure can, for example, follow the syntax of the syntax element "UsacConfigExtension()". For example, the configuration extension structure may include information regarding a number of configuration extensions "numConfigExtensions". In the typical case of a configuration extension of typeID_Config_Ext_Stream_ID according to an embodiment of the present invention, the stream identifier can be represented by, for example, the bit stream syntax element "streamId()", which can be represented by a 16-bit value.
[0141] In conclusion, the configuration structure included in the USAC frame of the extended element includes some configuration information for setting decoder parameters, and further includes a stream identifier that can be represented, for example, as a 16-bit integer value as a configuration extension.
[0142] The audio preroll information optionally includes a flag "applyCrossfade" indicating whether to apply crossfade (for example, a zero value may indicate not to apply crossfade), information indicating the number of preroll frames, and further information such as information regarding the preroll frames that can be specified as "auLen" and "AccessUnit()".
[0143] The USAC frame optionally further includes additional extension elements and typically comprises one or more of a single channel element, a channel pair element, or a low-frequency effect element.
[0144] In conclusion, a USAC frame (for example, one of USAC frame 222 or the instant playback frame IPF) can include, for example, extension syntax elements, and the extension syntax elements can include configuration structure (for example, 222c) and information regarding one or more preroll frames. The configuration structure and the information regarding one or more preroll frames are used, for example, to bring the state of the processing chain to a desired state and can correspond to, for example, information 222d. Further, the USAC frame also comprises encoded audio information such as a single channel element, a channel pair element, or a low-frequency effect element. Therefore, the audio decoder can recognize a change in the audio stream based on the stream identifier "streamId()". Also, the decoding parameters can be set based on the configuration information included in the configuration structure, and since the appropriate state of the audio decoding can be set based on the preroll frame information, it is possible for the audio decoder to perform artifact-free decoding of the USAC frame 600. Therefore, the described USAC frame enables switching the decoding of frames from different audio streams and also enables detection of the switching by the audio decoder without additional control information.
[0145] The USAC frame 600 described in this specification can correspond to the audio frame 222, to the first frame of the second audio stream included in the encoded audio signal representation 312, to the first frame of the second audio stream included in the encoded signal representation 412, or to an instant playback frame IPF as shown in FIG. 5.
[0146] 6. Example of an audio stream according to FIG. 7 FIG. 7 shows an exemplary representation of an audio stream that can be supplied by one of the audio encoders described in this specification and decoded by one of the audio decoders described in this specification. The audio stream of FIG. 7 can also be supplied by an audio stream provider, as described in this specification.
[0147] The audio stream 700 includes, for example, decoder configuration information as a first information block. The decoder configuration information may include, for example, the bitstream element "UsacConfig()" as defined in the USAC standard. The decoder configuration information may indicate, for example, a stream identifier of 1 and may be regarded as a stream access point at the beginning of the stream.
[0148] The audio stream may also include, for example, an audio frame data information unit 720 that may not include preroll data and may not include stream identifier information. For example, the information unit 720 may be a USAC frame and may correspond to, for example, the bitstream syntax element "UsacFrame()" defined in the USAC standard.
[0149] The information units 710 and 720 may both belong to the first audio stream, for example.
[0150] The audio stream 700 can also include an information unit 730 that can represent, for example, the first frame of a second stream included in the audio stream 700. The information unit 730 may include, for example, audio frame data, preroll data, and stream identifier information. The stream identifier information may indicate, for example, two stream identifiers different from the stream identifier included in the information unit 710.
[0151] The information unit 730 can be regarded as, for example, a stream access point.
[0152] For example, the information unit 730 can follow the syntax of the bitstream element "UsacFrame()" as defined in the USAC standard. However, the information unit 730 may include an extension element of type "id_ext_ele_audiopreroll". This extension element can include, for example, a configuration structure by the bitstream syntax "UsacConfig" with a configuration extension structure by the bitstream syntax "UsacConfigExtension". The configuration extension structure may include, for example, an extension element of type "ID_CONFIG_EXT_STREAM_ID" that encodes a stream identifier. Therefore, the information item or information unit 730 may include, for example, the information of the USAC frame 600 as described above.
[0153] Therefore, the information unit 730 represents an audio frame of the second stream and can supply complete configuration information for configuring an audio decoder to properly decode the audio frame. In particular, the configuration information also includes audio preroll information for setting the state of the audio decoder, and the configuration information includes a stream identifier that enables the audio decoder to recognize whether the information unit 730 is associated with a different audio stream when compared with the information units 700 and 710.
[0154] The audio stream 700 also includes an information unit 740 that follows the information unit 700. The information unit 740 may be a "normal" audio frame that contains only audio frame data, such as preroll data, configuration data, and no stream identifier. For example, the information unit 740 may conform to the bitstream syntax "UsacFrame()" without using extended elements.
[0155] The audio stream 700 can include, for example, audio frame data and preroll data, and can also include an information unit 750 that may not contain a stream identifier. Thus, the information unit 750 can be used as a stream access point, but may not be able to detect a switch between different streams.
[0156] For example, the information unit 750 can conform to the bitstream syntax "UsacFrame()" with the extended element ID_ext_ele_audiopreroll. However, in the information unit 750, the configuration information that is part of the audio preroll extended element does not contain a stream identifier. Thus, the information unit 750 cannot be reliably used as the first information unit after switching between different audio streams. On the other hand, the information unit 730, with the stream identifier it contains, enables the detection of a switch between different streams, and since the information unit also contains complete information for decoding including configuration information and preroll information, it can be reliably used as the first information unit after switching between different audio streams.
[0157] In conclusion, the audio stream 700 may comprise "information units" having different information content or encoded audio frames. There may be "very simple" audio frames that contain only encoded audio data without configuration data and without preroll data. There may also be audio frames that contain configuration information including not only the encoded audio information but also a stream identifier and preroll information. Such frames enable the identification of switching between different audio streams and completely independent decoding.
[0158] Furthermore, as an option, there may be frames that have only partial information, for example, no stream identifier information, and thus do not enable reliable identification of switching between different streams.
[0159] The audio decoders according to FIGS. 1 and 2 can typically utilize the audio stream 700, and it should be noted that the audio encoder and the audio stream provider according to FIGS. 3 and 4 can typically provide the audio stream 700 as shown in FIG. 7 (for example, as the encoded audio signal representations 312, 314).
[0160] 7. The audio stream according to FIG. 8 FIG. 8 shows a representation of an exemplary audio stream according to another embodiment of the present invention.
[0161] The audio stream in FIG. 8 is denoted as 800 in its entirety.
[0162] It should be noted that information units 810a to 810e belong to the first audio stream. For example, information unit 810a may have a decoder configuration, and may follow, for example, the bitstream syntax "UsacConfig()" defined by the USAC standard. The decoder configuration may have a configuration structure similar to, for example, configuration structure 222c. For example, information unit 810 can include a stream identifier extension, and the stream identifier can be included, for example, in the configuration extension structure of the configuration structure.
[0163] Information unit 810b can include, for example, preroll data and audio frame data without a stream identifier (such as encoded spectral values and encoded scale factor information). Information unit 810d may have a structure similar to or identical to that of information unit 810b, and may also represent audio frame data without preroll data and a stream identifier.
[0164] Furthermore, the audio stream can include a portion 820 following portion 810, and portion 820 is associated with a second audio stream different from the first audio stream. Portion 820 includes information unit 820a, and information unit 820a includes audio frame data with preroll data, and the preroll data includes a stream identifier extension (for example, within the configuration structure). Therefore, information unit 820a represents an audio frame. When an audio decoder detects based on the extension of the stream identifier that a previously decoded audio frame is from a different audio stream, the preroll data is used by the audio decoder to set the audio decoder to an appropriate state before decoding the audio frame data within information unit 820a. Therefore, information unit 820a is suitable to be the first information unit after switching between different audio streams.
[0165] Block 820 also includes one, two, or more information units 820b, 820d, which contain audio frame data but do not contain preroll data and do not contain a stream identifier.
[0166] Data stream 800 also includes a portion 830 related to a third audio stream. Portion 830 includes an information unit 830a, and information unit 830a includes audio frame data with preroll data and includes a stream identifier extension. Portion 830 further includes an information unit 830b that contains audio frame data without preroll data and without a stream identifier. The third portion 830 also includes an information unit 830d that contains audio frame data with preroll data but without a stream identifier.
[0167] Accordingly, audio stream 800 includes subsequent portions resulting from different audio streams, and at each transition from one stream to another, there is an information unit (e.g., an encoded audio frame) that contains audio frame data with preroll data and a stream identifier. Accordingly, since there is stream identifier information available at each switch from one audio stream to another within the encoded audio frame, the audio decoder can easily recognize the transition by evaluating the stream identifier (e.g., with respect to a comparison with a previously acquired and stored stream identifier).
[0168] It should be noted that the audio stream can be supplied by the audio encoder or bitstream provider described herein, and audio stream 800 can be evaluated by the audio decoder described herein.
[0169] 8. Decoder Function According to FIG. 9 FIG. 9 shows a schematic diagram of possible decoder functions of the audio decoder described herein.
[0170] For example, the functions described with reference to FIG. 9 can be implemented in the audio encoder 100 according to FIG. 1 or the audio decoder 200 according to FIG. 2. For example, using the functions described in FIG. 5, a method for continuing decoding can be determined.
[0171] However, it should be noted that the functions described with reference to FIG. 9 are merely examples, and for example, as long as the overall functions are the same, the order of determination can be changed. Also, as long as the overall function is not changed, the determinations can be combined.
[0172] The functions described in FIG. 9 are assumed to evaluate a new audio frame that has knowledge of information about previously decoded frames and conforms to the syntax described herein.
[0173] For example, in the first check 110, the audio decoder can check whether there is a "random access", that is, a jump operation to a stream access point. If it is recognized that there is a jump to a stream access point where the "normal" order of the frames is intentionally changed, the decoder function proceeds to step 920 of evaluating the configuration data of the stream access point to re-initialize the decoder. Optionally, a crossfade can be performed to avoid a sudden switch. It should be noted that random access means a "jump" from the first frame to the second frame, and the second frame has a frame index that is not immediately after the frame index of the previously decoded frame. In other words, random access is a jump from a frame with frame index n to a frame with frame index o, where o is different from n + 1.
[0174] In step 920, the jump is executed, and the jump target is an immediate playback frame, a frame that contains sufficient information to re-initialize the decoder.
[0175] However, if it is found in check 910 that there is "sequential playback" instead of "random access", further check 930 can be executed. In other words, if the decoding proceeds from a frame having frame index n to a frame having frame index n + 1, check 930 is executed.
[0176] In check 930, it is checked whether the (relevant) configuration defined by the constituent structure of the stream access point (or intermediate playback frame) (e.g., up to the stream identifier without including the stream identifier) is different from the current configuration without considering the stream identifier. If the (relevant) configuration described in the constituent structure of the stream access point is different from the current configuration (path "yes"), the decoding may proceed in step 940. However, it should be noted that step 930 can only be executed naturally if the next frame is a stream access point including the constituent structure. If the next frame does not include the constituent structure, step 930 cannot be executed naturally and the difference from the current configuration cannot be found.
[0177] However, if at step 930 it is detected that the configuration of the next frame's structure is the same as the current configuration (without considering the stream identifier), the next check shown in block 950 is performed. At step 950, it is determined whether the stream access point (e.g., within the structure configuration) contains a stream identifier. For example, the stream identifier does not necessarily have to be included, but is included in the configuration structure only if there is a configuration extension structure and this configuration extension structure actually contains a data structure element that is the stream identifier. In comparison 950, if it is found that the stream access point contains a stream identifier (branch "yes"), the stream identifier included in the stream access point of the next frame (the frame to be decoded) is compared with the current (stored) stream identifier. If it is found that the stream identifier included in the next frame (the frame to be decoded) is different from the current stream identifier (branch "yes" in decision 960), a jump is made to block 940. On the other hand, if it is detected that the stream identifier of the next frame is the same as the stored stream identifier, the additional configuration information (e.g., configuration extensions, etc.) following the configuration extension structure after the stream identifier remains unconsidered in order to determine whether to perform "transition" or the first initialization (branch "no" in step 960).
[0178] However, if at check 950 it is found that the stream access point (the next frame to be decoded) does not contain a stream identifier, or if it is found that the stream identifier of the next frame to be decoded is equal to the stored stream identifier, the procedure continues at step 970.
[0179] Furthermore, it should be noted that step 940 includes fading between an audio frame using the old configuration and an audio frame using the new configuration. To decode an audio frame using the new configuration, there is a re-initialization of the audio decoder (which may include initialization of a new decoder instance). Also, the old decoder instance is "flushed" and a crossfade is performed.
[0180] On the other hand, step 970 includes decoding the next frame without re-initializing the decoder, and any preroll information that may be included in the next frame is discarded (left out of consideration).
[0181] In conclusion, there are various possibilities that can be executed each time the audio decoder reaches an "intermediate playback frame" that can also be regarded as a "stream access point". Also, such an audio frame has no available configuration structure or preroll information, and such an audio frame does not permit re-initialization of the audio decoder. Therefore, it should be noted that in a frame that is not an "intermediate playback frame" or a "stream access point", specific processing is usually not performed.
[0182] When the decoder recognizes a "jump", i.e., a deviation from the normal frame order, usually, the preroll information and re-initialization of the audio decoder using the new configuration structure (jumping within the same stream) are naturally performed.
[0183] If such a jump exists, there are different cases.
[0184] If the audio decoder detects that the configuration information of the next stream to be decoded up to and including the configuration identifier is different from the stored information, the audio decoder is also re-initialized. On the other hand, if the audio decoder detects that the configuration information of the next frame to be decoded up to and including the stream identifier (if present) is the same as the stored information obtained from the previously decoded frame, re-initialization is not performed. In any case, when determining whether to perform re-initialization, the configuration information placed after the stream identifier in the configuration structure is ignored by the audio decoder. Also, if the audio decoder detects that there is no stream identifier in the configuration structure, the audio decoder naturally does not consider the stream identifier in the comparison with the stored information.
[0185] However, in order to perform the evaluation in a computationally efficient way, the decoder can first check the configuration information before the stream identifier against the stored configuration information, then check whether there is a stream identifier included in the configuration structure, and proceed to compare the stream identifier (if present in the configuration structure) with the stored stream identifier. As soon as the audio decoder detects a difference, it may decide to re-initialize. On the other hand, if the audio decoder cannot detect a difference between the configuration information until it includes the stream identifier, the audio decoder can decide to omit re-initialization.
[0186] Therefore, after the stream identifier in the configuration extension structure by the audio encoder, a minor configuration change that does not result in re-initialization can be notified, in which case the audio decoder can proceed with decoding by simply changing the configuration slightly (without requiring re-initialization).
[0187] In conclusion, the decoder function described with reference to FIG. 9 can be used with any of the audio decoders described in this specification, but should be considered optional.
[0188] 9. Bitstream Syntax According to FIGS. 10a, 10b, 10c, and 10d The syntax of the bitstream will be described below. In particular, the syntax of the constituent structure will be described. As an example, the syntax of the constituent structure "UsacConfig()" will be described, which can be used in place of the constituent structure 222c or the constituent structure 332 or the constituent structure 424 or the constituent structure "Config()" shown in FIG. 6 or the constituent structure "UsacConfig()" shown in FIG. 7 or the constituent structure "Config" shown in FIG. 8.
[0189] FIG. 10 shows the representation of the constituent structure "UsacConfig()". As can be seen from the figure, the constituent structure may include, for example, sampling frequency index information 1020a and optionally sampling frequency information 1020b. The sampling frequency index information 1020a (possibly in combination with the sampling frequency information 1020b) describes, for example, the sampling frequency used by the encoder and thus also describes the sampling frequency to be used by the audio decoder.
[0190] Furthermore, the constituent structure can also include frame length index information for spectral band replication (SBR). For example, the index may determine some parameters of the spectral bandwidth replication, as defined in the USAC standard, for example.
[0191] Furthermore, the constituent structure can also include a channel configuration index 1024 that can determine, for example, the channel configuration. The channel configuration index information may define, for example, a speaker mapping associated with a number of channels. For example, the channel configuration index information may have a meaning as defined in the USAC standard. For example, when the channel configuration index information is equal to zero, details regarding the channel configuration may be included in the "UsacChannelConfig()" data structure 1024b.
[0192] Furthermore, the configuration structure may include decoder configuration information 1026a that can describe (or enumerate) information elements existing in, for example, an audio frame data structure. For example, the decoder configuration information can include one or more of the elements described in the USAC standard.
[0193] Furthermore, the configuration structure 1010 also includes a flag (e.g., named "UsacConfigExtensionPresent") indicating the presence of a configuration extension structure (e.g., configuration extension structure 226). The configuration structure 1010 also includes a configuration extension structure indicated by, for example, "UsacConfigExtension()" 1028a. The configuration extension structure is preferably part of the configuration structure 1010 and can be represented, for example, by a bit sequence immediately following the bits representing other configuration items of the configuration structure 1010. The configuration extension structure can convey, for example, stream identifier information, as will be described below.
[0194] Hereinafter, the possible syntax of the configuration extension structure will be described with reference to FIG. 10b. The configuration extension structure is entirely indicated by 1030 and corresponds to the configuration extension structure 1028a.
[0195] A configuration extension structure (also denoted as "UsacConfigExtension()") may encode, for example, some configuration extensions within the syntax element 1040a. It should be noted that for each configuration extension item, there is configuration extension type information 1042a and configuration extension length information 1044a, so the order of different configuration extension information items can be arbitrarily selected. Therefore, the configuration extension structure 1030 can convey a plurality of configuration extension items (or configuration extension information items) in a variable order, and the audio encoder can determine which configuration extension item is encoded first and which is encoded later. For example, for each configuration information item, there may first be a configuration extension type identifier 1042a, followed by configuration extension length information 1044, and then the "payload" of each configuration extension information item. The encoding of the payload of each configuration extension information item may vary depending on, for example, the type of the configuration extension information item indicated by the configuration extension type information, and the length of the payload of each configuration extension information item can be determined by the value of the respective configuration extension length information 1044a. For example, if the configuration extension information item is padding information, there may be one or more padding bytes. On the other hand, if the configuration extension information item is configuration extension loudness information, there may be a data structure containing information regarding loudness (e.g., denoted as "loudnessInfoSet()").
[0196] Furthermore, if the configuration extension information item is a stream identifier, there may be a numerical representation of the stream identifier specified as "streamId()". Syntax examples of different types of configuration extension information items are shown by reference numerals 1046a, 1048a, and 1050a.
[0197] In conclusion, the syntax of the configuration extension structure is such that the order of different configuration information items can be changed. For example, the stream identifier configuration extension information item can be placed before or after other configuration extension information items by the audio encoder. Therefore, depending on the placement of the stream identifier configuration extension information item within the configuration extension structure, where other information items of the configuration extension structure should be considered in the comparison between the configuration indicated by the current configuration structure and the configuration information previously obtained by the audio decoder, the audio encoder is controllable. Usually, all configuration extension information items up to and including the configuration information item preceding the configuration extension structure and the stream identifier information are considered in such a comparison, but all configuration extension information items encoded in the bitstream after the stream identifier configuration extension information item are ignored in the comparison.
[0198] As described above, the configuration structures explained with respect to FIGS. 10a and 10b are very suitable for the concept according to the present invention.
[0199] FIG. 10 shows the syntax of the stream identifier (configuration extension) information item, which is also denoted as "StreamId()" (or "streamId()"). As shown, the stream identifier can be represented by a 16-bit binary representation. Therefore, different values exceeding 65000 can be encoded as stream identifiers, which is usually sufficient to recognize transitions between different audio streams.
[0200] FIG. 10d shows an example of the assignment of type identifiers to different configuration extension information items. For example, a configuration extension information item of type "stream identifier" can be represented by the value 7 of the configuration extension type information 1042a. Other types of configuration extension information items can be represented, for example, by other values of the configuration extension type identifier 1042a.
[0201] In conclusion, FIGS. 10a - 10d depict possible syntax (or syntax extensions) of a structural configuration that can be used by an audio encoder to encode stream identifier information that can be used by an audio decoder to extract stream identifier information.
[0202] However, it should be noted that the structural configurations described herein are to be considered merely as examples and can be varied widely. For example, the sampling frequency index information and / or the sampling frequency information and / or the spectral bandwidth replication frame length index information and / or the channel configuration index information can be encoded in different ways. Also, optionally, one or more of the above information items can be dropped. Furthermore, the UsacDecoderConfig information item can also be omitted.
[0203] Furthermore, the encoding of the configuration extension number of the configuration extension type and the configuration extension length can be modified. Also, different configuration extension information items should be considered as optional and can probably be encoded in different ways.
[0204] Furthermore, the stream identifier can be encoded with more or fewer bits, where different types of number representations can be used. Additionally, the assignment of the identifier number to different configuration extension types should be considered as a preferred example rather than an essential feature.
[0205] 9. Conclusion
[0206] In the following, several aspects of the present invention will be described, which can be used individually or in combination with the embodiments described herein.
[0207] In particular, the solution according to the present invention is described herein.
[0208] It should be noted that the aspects of the embodiments according to the present invention are defined by the appended claims.
[0209] However, the embodiments defined by the claims may optionally be further augmented by any of the features described herein, either individually or in combination. It should also be noted that definitions within parentheses "()" or brackets "[]" should be regarded as optional, especially when used in the claims.
[0210] Notwithstanding, it should be noted that the features of the present invention described below may also be used separately from the features of the claims.
[0211] Furthermore, the features and functions recited in the claims and described below can optionally be combined with the features and functions described in the section that explains the problems, embodiments, and possible usage scenarios for conventional approaches underlying the aspects of the present invention. In particular, the features and functions described herein can be used in a USAC audio decoder compliant with ISO / IEC 23003-3:2012, including Revision 3, subsection "Bitrate Adaptation" (e.g., as standardized as of the filing date of the priority application of the present application, or as standardized as of the filing date of the present invention, but possibly including further future revisions).
[0212] According to one aspect of the present invention, it is proposed to introduce a new configuration extension of USAC with usacConfigExtType == ID_CONFIG_EXT_STREAM_ID together with a related bitstream structure including a simple universal 16-bit identifier bitfield (e.g., into the USAC bitstream syntax). This identifier is different between any two configuration structures for all streams within a set of streams intended for seamless switching between them (e.g., can be selected differently by an audio encoder or an audio stream provider). An example of such a set of streams is the so-called "adaptation set" in the use case of MPEG-DASH delivery.
[0213] The proposed unique stream ID configuration extension is ensured, for example, when comparing the current (or current configuration) with a new configuration structure (e.g., on the audio encoder side or the audio decoder side), such that the new configuration (and thus the new stream) is correctly identified and the decoder operates as expected. For example, the decoder will perform an appropriate decoder flush, preroll the access unit, and perform a crossfade (if applicable).
[0214] The following is a proposal for a (modified) specification text (which is standardized as of the filing date of this application or the filing date of the priority application and optionally includes future amendments (e.g., of MPEG-D USAC (ISO / IEC 23003-3+AMD.1+AMD-2+AMD.3))).
[0215] The sections referred to in the aspects of the present invention described below can be used individually or in combination with a USAC audio decoder or within another frame-based audio decoder.
[0216] As shown in Table 15 below, the configuration extension can be used by an audio encoder to provide an audio bitstream and by an audio decoder to extract information from the audio bitstream.
[0217] When using audio encoding and decoding according to the above USAC standard, Table 15 in Section 5.2 needs to be replaced with the following updated version of Table 15.
[0218] Table 15 - Syntax of UsacConfigExtension() JPEG2025081335000002.jpg172162
[0219] Also, when considering audio encoding or audio decoding according to the USAC standard, a new table AMD.01 as follows needs to be added at the end of Section 5.2 of the USAC standard (encoding details, number of bits are optional).
[0220] Table AMD.01 - Syntax of StreamId() JPEG2025081335000003.jpg32162
[0221] However, in the above table, encoding details and for example a large number of bits should be considered optional.
[0222] Also, when considering encoding or decoding according to the USAC standard, the following subordinate section 6.1.15 needs to be added after "6.1.14 UsacConfigExtension()".
[0223] "6.1.15 Unique Stream Identifier (streamID) 6.1.15.1 Terms, Definitions, and Meanings
[0224] Stream Identifier A 2 - byte unsigned integer stream identifier (streamID) that uniquely identifies the configuration of a stream within a related series of streams for the purpose of seamless switching between these streams. The streamIdentifier can take values from 0 to 65535 (encoding details are optional).
[0225] If it is part of an MPEG-DASH compliant set defined in, for example, ISO / IEC 23009, all stream IDs of the streams within that DASH compliant set are different in pairs.
[0226] 6.1.15.2 Explanation of Stream Identifiers The configuration extension of type ID_CONFIG_EXT_STREAM_ID provides a container for indicating a stream identifier (abbreviation: "stream ID"). The stream ID configuration extension enables adding a unique integer to the configuration structure so that the audio bitstream configurations of two streams can be distinguished even if the rest of the configuration structure is (bitwise) the same.
[0227] The usacConfigExtLength of the configuration extension of type ID_CONFIG_EXT_STREAM_ID shall have a value of 2 (2) (optionally, it may be different).
[0228] Any default audio bitstream cannot (optionally) have multiple configuration extensions of type ID_CONFIG_EXT_STREAM_ID.
[0229] For example, when a normally operating decoder instance receives a new configuration structure by Config() in the ID_EXT_ELE_AUDIOPREROLL extension payload, it must compare this new configuration structure with the currently active configuration (see, for example, 7.18.3.3). Such a comparison can be performed, for example, by a bitwise comparison of the corresponding configuration structures.
[0230] If the configuration structure includes configuration extensions, for example, all configuration extensions up to and including the configuration extension of type ID_CONFIG_EXT_STREAM_ID must be included in the comparison. All configuration extensions following the configuration extension of type ID_CONFIG_EXT_STREAM_ID may be considered not to be considered during the comparison (optionally).
[0231] Note that the above rules enable the encoder to control whether changes in a particular configuration extension cause decoder reconfiguration.
[0232] It should be noted that the definitions and details from this text to be added to the standard can be used, optionally, individually or in combination, in embodiments according to the present invention.
[0233] When considering USAC encoding or decoding, Table 74 of Clause 6 needs to be replaced with the table shown in Figure 10d.
[0234] It has been explained to conclude some possible changes that may be introduced into the USAC standard. However, concepts as described here can also be used in relation to other audio coding standards. In other words, as described herein, it would also be possible to introduce stream identifier information into some configuration structures of any other audio coding standard.
[0235] The features described herein with respect to stream identifier information can also be applied when adopted in combination with other coding standards. In this case, the terms should be adapted to the terms of each audio coding standard.
[0236] The following describes some of the effects, advantages, or features of options according to the present invention.
[0237] The proposed configuration extension provides an easily implementable solution for distinguishing between configuration structures that are otherwise bit-identical. The distinguishability obtained between configurations enables, for example, the accurate and originally intended functionality of dynamic adaptive streaming with seamless transitions between streams.
[0238] Some alternative solutions are described below.
[0239] For example, if the encoder guarantees that all streams within a set of streams have different configurations, i.e., they use different coding tools or different parameterizations, the above problem can be avoided. If the differences in the bitrates of individual streams are large enough, this will usually result in different settings for each pair. This is common, but in cases where a fine grid of bitrates is required, the (conventional) solution often does not work well.
[0240] In contrast, by using stream identifiers included in the configuration parts (also called configuration structures) to distinguish different streams, it is possible to distinguish streams even when the remaining configuration structures are the same (and the bitrates may be similar).
[0241] Alternatively (e.g., instead of using stream identifiers), it is possible to create appropriate unspecified configuration extensions that are different for each stream but are structured differently for some reason. The effect is the same. Since it cannot be guaranteed that all decoder implementations will evaluate this unspecified configuration extension when configurations are compared in the above scenario, the correct functionality cannot be guaranteed.
[0242] In contrast, embodiments according to the present invention give rise to the concept that stream identifiers are clearly specified within the configuration structure, enabling clear distinction between different streams.
[0243] It should be noted that the implementation of the concept of the present invention can be recognized by analyzing the constituent structure of the USAC stream. Furthermore, the implementation of the concept of the present invention can be recognized by testing for the presence of configuration extensions as described above.
[0244] Hereinafter, some possible application fields of aspects according to the present invention will be described.
[0245] Embodiments according to the present invention provide the identifiability of the same data structure in other respects.
[0246] Further embodiments according to the present invention provide the identifiability of the same audio codec configuration structure in other respects.
[0247] Embodiments according to the present invention enable seamless and dynamic adaptive streaming of audio over any transmission network.
[0248] Hereinafter, some further aspects will be described, which should be regarded as optional.
[0249] For example, the behavior of an audio encoder / audio stream provider will be described below. Hereinafter, some options details regarding an audio encoder (which can also take the form of an audio stream provider) will be described.
[0250] An audio encoder typically does not generate a single stream that suddenly changes its configuration, but an encoder framework including an encoder or multiple encoder instances generates multiple streams in parallel, each including an IPF (plural) ("instant playback frame") (plural) at the synchronization position (time point) within the stream.
[0251] Next, the decoder framework selects, according to specific and / or predetermined criteria, for example the quality of the Internet connection, one of the streams generated in parallel, and "requests" (or demands) from the encoder-side server to accurately transmit the stream. Then it transfers that stream to the decoder. All further encoded streams are simply ignored. Changes between streams are only permitted in the IPF.
[0252] The audio decoder does not initially recognize such changes and / or is not notified about such changes, for example, by the decoder framework. Rather, the audio decoder needs to detect stream changes by comparing the embedded configuration structures ("Config-structures"). From the decoder's perspective, the encoder appears to have simply generated a stream with a changed configuration ("Config"). In fact, this is usually not the case. Rather, multiple variants (including different bitrates) are always (continuously) generated in parallel by the encoder, and only the decoder framework and the encoder-side server (or stream provider) split the stream and rearrange (reconnect) parts of the stream (or the stream).
[0253] Further option details are illustrated.
[0254] Furthermore, it should be noted that the devices shown in the drawings can be augmented by any of the features and functions described herein, either individually or in combination.
[0255] As a conclusion, an audio encoder or an audio stream provider can switch the supply of different streams to a specific audio decoder (or audio decoding device), and this switching can be performed, for example, in response to a request from the audio decoder, or in response to a request from the audio decoding device or other network management device, or even by a decision of the audio encoder or audio stream provider. The switching between the supply of frames from different audio streams can be used to adapt the actual bitrate to the available bitrate. The decoder configuration indicated from the audio encoder (or audio stream provider) to the audio decoder can be the same between different streams, but the stream identifiers should be different between different streams. Therefore, the audio decoder can use the stream identifier and the additional information (e.g., setting information and preroll information) included in the immediate playback frame to recognize when the re-initialization of the audio decoder should be performed.
[0256] As a further conclusion, using a stream identifier ("streamID") as described in the present specification can overcome the problems described in the section explaining the problems underlying the aspects of the present invention and the possible usage scenarios of the embodiments.
[0257] 10. Method
[0258] Figures 11a to 11c show the flowcharts of the method according to an embodiment of the present invention.
[0259] The method shown in Figures 11a to 11c can be supplemented by any of the features and functions described herein.
[0260] 11. Alternative Implementations
[0261] Although several aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent corresponding method descriptions, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent descriptions of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.
[0262] The encoded audio signal of the present invention can be stored in a digital storage medium or transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0263] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation can be carried out using a digital storage medium storing electronically readable control signals, such as a floppy disk (floppy is a registered trademark), DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, which cooperate (or can cooperate) with a programmable computer system so that the respective methods are executed. Thus, the digital storage medium can be computer-readable.
[0264] Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system so that one of the methods described herein is executed.
[0265] In general, embodiments of the present invention can be implemented as a computer program product having program code, the program code being operative to execute one of the methods when the computer program product is run on a computer. The program code may be stored, for example, on a machine-readable carrier.
[0266] Other embodiments include a computer program stored on a machine-readable carrier for executing one of the methods described herein.
[0267] In other words, one embodiment of the method of the present application is thus a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.
[0268] Accordingly, a further embodiment of the method of the present application is a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for executing one of the methods described herein. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory.
[0269] Accordingly, a further embodiment of the method of the present application is a data stream or a series of signals representing a computer program for executing one of the methods described herein. The data stream or series of signals may be configured to be transferred via a data communication connection such as, for example, the Internet.
[0270] Further embodiments include processing means, such as a computer or a programmable logic device, configured or adapted to execute one of the methods described herein.
[0271] Further embodiments include a computer having installed thereon a computer program for executing one of the methods described herein.
[0272] Further embodiments according to the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for executing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transferring the computer program to the receiver.
[0273] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to execute some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to execute one of the methods described herein. Generally, the method is preferably executed by any hardware device.
[0274] The apparatuses described herein can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0275] The apparatuses described herein, or any component of the apparatuses described herein, can be implemented at least partially in hardware and / or software.
[0276] The methods described herein can be executed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0277] The methods described herein, or any component of the apparatuses described herein, can be executed at least partially by hardware and / or software.
[0278] The above embodiments are merely examples for explaining the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to other persons skilled in the art. Accordingly, it is intended to be limited only by the scope of the impending claims and not by the specific details presented for the description or illustration of the embodiments herein.
Claims
1. 1. An audio decoder (100; 200) for providing a decoded audio signal representation (112; 212) on the basis of an encoded audio signal representation (110; 210; 312; 412; 550; 600; 700; 800), comprising: said audio decoder being configured to adjust decoding parameters in dependence on configuration information (110a; 222c; 332; 424; 1010, 1030), the audio decoder is configured to decode one or more audio frames using current configuration information (140; 240); and the audio decoder is configured to compare configuration information (110a; 222c; 332; 424; 1010, 1030) in a configuration structure associated with one or more frames to be decoded (222) with current configuration information (140; 240) and, if the configuration information in the configuration structure associated with the one or more frames to be decoded or a relevant part (1020a, 1020b, 1022a, 1024a, 1024b, 1026a, 1050a) of the configuration information in the configuration structure associated with the one or more frames to be decoded differs from the current configuration information, to perform a transition and to perform decoding using the configuration information in the configuration structure associated with the one or more frames to be decoded as new configuration information, the audio decoder is configured to take into account stream identifier information (230; streamID, 1050a, streamIdentifier) included in the configuration structure when comparing the configuration information, such that the transition is performed depending on a difference between a stream identifier previously acquired by the audio decoder and a stream identifier represented by stream identifier information in the configuration structure associated with the one or more frames to be decoded, The stream identifier is represented by a bitstream syntax element represented by a 16-bit value. Audio decoder.
2. 1. A method for providing a decoded audio signal representation based on an encoded audio signal representation, comprising: The method comprises the step of adjusting decoding parameters in dependence on configuration information (110a; 222c; 332; 424; 1010, 1030), The method includes the step of decoding one or more audio frames using current configuration information (140; 240), the method comprises a step of comparing configuration information (110a; 222c; 332; 424; 1010, 1030) in a configuration structure associated with one or more frames to be decoded (222) with the current configuration information, and if the configuration information in the configuration structure associated with the one or more frames to be decoded or a relevant part of the configuration information (1020a, 1020b, 1022a, 1024a, 1024b, 1026a, 1050a) in the configuration structure associated with the one or more frames to be decoded differs from the current configuration information, performing a transition and performing decoding using the configuration information in the configuration structure associated with the one or more frames to be decoded as new configuration information, The method includes the step of taking into account stream identifier information (230; streamID, 1050a, streamIdentifier) included in the configuration structure when comparing the configuration information, such that the transition is performed depending on a difference between a stream identifier previously obtained in the audio decoding and a stream identifier represented by the stream identifier information in the configuration structure associated with the one or more frames to be decoded, The stream identifier is represented by a bitstream syntax element represented by a 16-bit value. method.
3. A computer program for carrying out the method according to claim 2 when the computer program runs on a computer.
Citation Information
Patent Citations
Audio decoder, apparatus for generating encoded audio output data, and method enabling decoder initialization
JP2016539357A