Audio decoder, audio encoder, method and computer readable storage medium
By using a layered audio decoder and encoder, and based on multi-channel decoding and encoding technology, the bandwidth expansion of multi-channel audio signals is processed, solving the encoding and decoding problems of three-dimensional audio scenes, and achieving high-quality audio reproduction and auditory impression.
Patent Information
- Application Number
- CN201911131913.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2013-10-18
- Filing Date
- 2014-07-14
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2034-07-14
AI Technical Summary
Existing audio coding technologies struggle to effectively handle the encoding and decoding of three-dimensional audio scenes, particularly lacking high-quality solutions for bandwidth expansion of multi-channel audio signals.
By employing a layered audio decoder and encoder, and through multi-channel decoding and encoding, based on the joint encoding representation of different submixed signals, signals associated with perceptually important locations in the audio scene are separated, multi-channel bandwidth is extended separately, and prediction and residual signal-assisted techniques are used to improve audio quality.
It achieves excellent reproduction of audio scenes, especially in distinguishing between horizontal and vertical positions, improving auditory impression and audio quality, while maintaining efficient bitrate usage.
Smart Images

Figure CN111128205B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese invention patent application number 201480041693.7, entitled "Audio Decoder, Audio Encoder, Method and Computer-Readable Storage Medium", which entered the Chinese national phase of international application PCT / EP2014 / 065021, with an international filing date of July 14, 2014 and a priority date of July 22, 2013. Technical Field
[0002] An audio decoder is created according to an embodiment of the invention for providing at least four bandwidth extended channel signals based on an encoded representation.
[0003] According to another embodiment of the invention, an audio encoder is created for providing an encoded representation based on at least four audio channel signals.
[0004] According to another embodiment of the invention, a method is created for providing at least four audio channel signals based on an encoded representation.
[0005] According to another embodiment of the invention, a method is created for providing an encoded representation based on at least four audio channel signals.
[0006] According to another embodiment of the present invention, a computer program is created for performing one of the methods.
[0007] Generally, embodiments of the present invention involve the joint encoding of n audio channels. Background Technology
[0008] In recent years, the demand for the storage and transmission of audio content has been steadily increasing. Furthermore, the quality requirements for the storage and transmission of audio content have also been steadily increasing. Therefore, the concepts for encoding and decoding audio content have been enhanced. For example, the so-called "Advanced Audio Coding" (AAC) has been developed, described in, for example, the international standard ISO / IEC 13818-7:2003. In addition, spatial extensions have been created, such as the so-called "MPEG Surround Sound," described in, for example, the international standard ISO / IEC 23003-1:2007. Furthermore, additional improvements for encoding and decoding spatial information in audio signals are described in the international standard ISO / IEC 23003-2:2010, which relates to the so-called Spatial Audio Object Coding (SAOC).
[0009] Furthermore, the concept of flexible audio coding / decoding is defined in the international standard ISO / IEC 23003-3:2012. The flexible audio coding / decoding concept provides the possibility of encoding both general audio signals and speech signals with good coding efficiency and processing multi-channel audio signals. This international standard describes the so-called "Unified Speech and Audio Coding" (USAC) concept.
[0010] In MPEG USAC [1], joint stereo coding of two channels is performed using complex prediction with band-limited residual signals or full-band residual signals, MPS 2-1-1 or unified stereo.
[0011] MPEG surround sound[2] combines OTT frames and TTT frames in layers to perform joint encoding of multichannel audio with or without the transmission of residual signals.
[0012] However, there is a desire for even more advanced concepts to provide efficient encoding and decoding for 3D audio scenarios. Summary of the Invention
[0013] According to embodiments of the present invention, an audio decoder is created for providing at least four bandwidth-extended channel signals based on an encoded representation. The audio encoder is configured to use (first) multichannel decoding to provide a first and second submixed signal based on a joint encoded representation of a first and a second submixed signal. The audio decoder is configured to use (second) multichannel decoding to provide at least a first and a second audio channel signal based on the first submixed signal, and to use (third) multichannel decoding to provide at least a third and a fourth audio channel signal based on the second submixed signal. The audio decoder is configured to perform multichannel bandwidth extension based on the first and third audio channel signals to obtain channel signals with first and third bandwidth extensions. Furthermore, the audio decoder is configured to perform multichannel bandwidth extension based on the second and fourth audio channel signals to obtain channel signals with second and fourth bandwidth extensions.
[0014] This embodiment of the invention is based on the discovery that particularly good bandwidth expansion results can be obtained in a hierarchical audio decoder if audio channel signals obtained based on different downmixed signals in the second stage of the audio decoder are used in the multichannel bandwidth expansion, wherein different downmixed signals are derived from the joint coded representation in the first stage of the audio decoder. It has been found that particularly good audio quality can be obtained if downmixed signals associated with locations that are particularly important to the perception of the audio scene are separated in the first stage of the hierarchical audio decoder, while spatial locations that are not so important to the auditory impression are separated in the second stage of the hierarchical audio decoder. Furthermore, it has been found that audio channel signals associated with different locations that are perceptually important to the audio scene (e.g., the location of the audio scene, where the relationship between signals from these locations is perceptually important) should be jointly processed in the multichannel bandwidth expansion, because the multichannel bandwidth expansion can therefore take into account the dependencies and differences between signals from these auditorily important locations. This is achieved by performing multichannel bandwidth expansion based on a first audio channel signal (derived from a first downmixed signal in the second stage of the layered audio decoder) and a third audio channel signal (derived from a second downmixed signal in the second stage of the layered audio decoder) to obtain two bandwidth-expanded channel signals (i.e., a first bandwidth-expanded channel signal and a third bandwidth-expanded channel signal). Therefore, the (joint) multichannel bandwidth expansion is performed based on audio channel signals derived from different downmixed signals in the second stage of the layered multichannel decoder, such that the relationship between the first and third audio channel signals is analogous to (or determined by) the relationship between the first and second downmixed signals. Thus, multichannel bandwidth expansion can utilize this relationship (e.g., the relationship between the first and third audio channel signals), which is substantially determined by deriving the first and second downmixed signals from a joint encoded representation of the first and second downmixed signals using multichannel decoding, performed in the first stage of the audio decoder. Therefore, multichannel bandwidth extension can utilize this relationship to reproduce it with good accuracy in the first stage of the layered audio decoder, resulting in a particularly good auditory impression.
[0015] In a preferred embodiment, the first and second submixed signals are associated with different horizontal (or azimuth) positions of the audio scene. It has been found that distinguishing between different horizontal audio positions (or azimuth positions) is particularly relevant because the human auditory system is especially sensitive to different horizontal positions. Therefore, it is advantageous to separate the submixed signals associated with different horizontal positions of the audio scene in the first stage of the layered audio decoder, since processing in the first stage of the layered audio decoder is generally more precise than processing in subsequent stages. Furthermore, the first and third audio channel signals thus used jointly in the (first) multichannel bandwidth extension are associated with different horizontal positions of the audio scene (because in the second stage of the layered audio decoder, the first audio channel signal is derived from the first submixed signal, and the third audio channel signal is derived from the second mix signal), thereby allowing the (first) multichannel bandwidth extension to be highly adapted to the human ability to distinguish different horizontal positions. Similarly, the (second) multichannel bandwidth extension, performed based on the second and fourth audio channel signals, operates on the audio channel signals associated with different horizontal positions in the audio scene, making the (second) multichannel bandwidth extension highly suitable for psychoacoustic relationships between audio channel signals associated with different horizontal positions in the audio scene. Therefore, a particularly good auditory impression can be achieved.
[0016] In a preferred embodiment, the first submixed signal is associated with the left side of the audio scene, and the second submixed signal is associated with the right side of the audio scene. Therefore, the first audio channel signal is typically also associated with the left side of the audio scene, and the third audio channel signal is associated with the right side of the audio scene, such that the (first) multichannel bandwidth extension operates (preferably in conjunction) on audio channel signals from different sides of the audio scene, and is thus highly suitable for human left / right perception. This also applies to the (second) multichannel bandwidth extension, which operates based on the second and fourth audio channel signals.
[0017] In a preferred embodiment, the first and second audio channel signals are associated with vertically adjacent positions in the audio scene. Similarly, the third and fourth audio channel signals are associated with vertically adjacent positions in the audio scene. It has been found advantageous to separate the audio channel signals associated with vertically adjacent positions in the second stage of the layered audio decoder. Furthermore, it has been found that audio channel signals are generally not severely degraded by separating them with vertically adjacent positions, making the input signal for multichannel bandwidth expansion still highly suitable for multichannel bandwidth expansion (e.g., stereo bandwidth expansion).
[0018] In a preferred embodiment, the first and third audio channel signals are associated with a first common horizontal plane (or first common height) of the audio scene, but with different horizontal positions (or azimuth positions) of the audio scene, and the second and fourth audio channel signals are associated with a second common horizontal plane (or second common height) of the audio scene, but with different horizontal positions (or azimuth positions) of the audio scene. In this case, the first common horizontal plane (or height) is different from the second common horizontal plane (or height). It has been found that multichannel bandwidth expansion can be performed with particularly good quality results based on two audio channel signals associated with the same horizontal plane (or height).
[0019] In a preferred embodiment, the first and second audio channel signals are associated with a first common vertical plane (or common azimuth position) of the audio scene, but with different vertical positions (or heights) of the audio scene. Similarly, the third and fourth audio channel signals are associated with a second common vertical plane (or common azimuth position) of the audio scene, but with different vertical positions (or heights) of the audio scene. In this case, the first common vertical plane (or azimuth position) is preferably different from the second common vertical plane (or azimuth position). It has been found that the second stage of the layered audio decoder can be used to perform the segmentation (or separation) of audio channel signals associated with a common vertical plane (or azimuth position) with good results, while the first stage of the layered audio decoder can be used to perform the separation (or segmentation) of audio channel signals associated with different vertical planes (or azimuth positions) with good quality results.
[0020] In a preferred embodiment, the first and second audio channel signals are associated with the left side of the audio scene, and the third and fourth audio channel signals are associated with the right side of the audio scene. This configuration takes into account a particularly good multichannel bandwidth extension, which utilizes the relationship between the audio channel signals associated with the left side and the audio channel signals associated with the right side, and is therefore extremely well-suited to the human ability to distinguish between sounds from the left and sounds from the right.
[0021] In a preferred embodiment, the first and third audio channel signals are associated with the lower part of the audio scene, and the second and fourth audio channel signals are associated with the upper part of the audio scene. This spatial arrangement of the audio channel signals has been found to produce particularly good auditory results.
[0022] In a preferred embodiment, the audio decoder is configured to perform horizontal partitioning when providing the first and second submixed signals using a joint coded representation based on the first and second submixed signals via multichannel decoding. It has been found that performing horizontal partitioning in the first stage of the hierarchical audio decoder results in a particularly good auditory impression because the processing performed in the first stage of the hierarchical audio decoder can generally be performed with greater efficiency compared to the processing performed in the second stage of the hierarchical audio decoder. Furthermore, performing horizontal partitioning in the first stage of the audio decoder results in a good auditory impression because the human auditory system is more sensitive to the horizontal position of an audio object than its vertical position.
[0023] In a preferred embodiment, the audio decoder is configured to perform vertical partitioning when using multichannel decoding to provide at least a first audio channel signal and a second audio channel signal based on a first submixed signal. Similarly, the audio decoder is preferably configured to perform vertical partitioning when using multichannel decoding to provide at least a third audio channel signal and a fourth audio channel signal based on a second submixed signal. It has been found that performing vertical partitioning in the second stage of the layered decoder provides a good auditory impression because the human auditory system is not very sensitive to the vertical position of the audio source (or audio object).
[0024] In a preferred embodiment, the audio decoder is configured to perform stereo bandwidth expansion based on a first audio channel signal and a third audio channel signal to obtain channel signals with first bandwidth expansion and third bandwidth expansion, wherein the first audio channel signal and the third audio channel signal represent a first left / right channel pair. Similarly, the audio decoder is configured to perform stereo bandwidth expansion based on a second audio channel signal and a fourth audio channel signal to obtain channel signals with second bandwidth expansion and fourth bandwidth expansion, wherein the second audio channel signal and the fourth audio channel signal represent a second left / right channel pair. It has been found that stereo bandwidth expansion results in a particularly good auditory impression because stereo bandwidth expansion takes into account the relationship between the left stereo channel and the right stereo channel and performs bandwidth expansion based on this relationship.
[0025] In a preferred embodiment, the audio decoder is configured to provide the first and second submixed signals using prediction-based multichannel decoding, based on a joint coded representation of the first and second submixed signals. It has been found that using prediction-based multichannel decoding in the first stage of the layered audio decoder results in a good trade-off between bitrate and quality. It has also been found that the use of prediction leads to a good reconstruction of the differences between the first and second submixed signals, which is important for distinguishing the left / right sides of audio objects.
[0026] For example, an audio decoder can be configured to estimate prediction parameters that describe the contribution of signal components derived using signal components from previous frames to the submixed signal providing the current frame. Therefore, the contribution strength of the signal components derived using signal components from previous frames can be adjusted based on parameters included in the encoded representation.
[0027] For example, prediction-based multichannel decoding can operate in the MDCT domain, making it highly suitable for and easy to interface with the audio decoding stage, which provides the input signal to the multichannel decoder that derives the first and second submixed signals. Preferably, but not necessarily, prediction-based multichannel decoding can be USAC complex stereo prediction, which facilitates the implementation of the audio decoder.
[0028] In a preferred embodiment, the audio decoder is configured to use residual signal-assisted multichannel decoding to provide the first and second submixed signals based on a joint encoded representation of the first and second submixed signals. The use of residual signal-assisted multichannel decoding takes into account a particularly accurate reconstruction of the first and second submixed signals, which is further based on the audio channel signals and therefore on the bandwidth-extended channel signals to improve left-right position perception.
[0029] In a preferred embodiment, the audio decoder is configured to use parameter-based multichannel decoding to provide at least a first audio channel signal and a second audio channel signal based on a first submixed signal. Furthermore, the audio decoder is configured to use parameter-based multichannel decoding to provide at least a third audio channel signal and a fourth audio channel signal based on a second submixed signal. The use of parameter-based multichannel decoding has been found to be extremely suitable for the second stage of a layered audio decoder. Parameter-based multichannel decoding has been found to provide a good trade-off between audio quality and bit rate. Although the reconstruction quality of parameter-based multichannel decoding is generally not as good as that of prediction-based (and possibly residual signal-assisted) multichannel decoding, it has been found that the use of parameter-based multichannel decoding is often sufficient because the human auditory system is not particularly sensitive to the vertical position (or height) of an audio object, which is preferably determined by the distribution (or separation) between the first and second audio channel signals or between the third and fourth audio channel signals.
[0030] In a preferred embodiment, parameter-based multichannel decoding is configured to estimate one or more parameters describing the desired correlation (or covariance) between two channels and / or the step difference between the two channels to provide two or more audio channel signals based on the corresponding mixed signal. It has been found that the use of these parameters describing, for example, the desired correlation between two channels and / or the step difference between two channels is extremely suitable for partitioning (or separating) signals from a first audio channel and a second audio channel (which are typically associated with different vertical positions in an audio scene), and extremely suitable for partitioning (or separating) signals from a third audio channel and a fourth audio channel (which are also typically associated with different vertical positions).
[0031] For example, parametric multichannel decoding can operate in the QMF domain. Therefore, parametric multichannel decoding is highly suitable for multichannel bandwidth extension and easy to interface with multichannel bandwidth extension, which is preferred but not required to operate in the QMF domain.
[0032] For example, parameter-based multichannel decoding could be MPEG surround sound 2-1-2 decoding or unified stereo decoding. The use of this encoding concept can facilitate implementation because these decoding concepts may already exist in traditional audio decoders.
[0033] In a preferred embodiment, the audio decoder is configured to use residual signal-assisted multichannel decoding to provide at least a first audio channel signal and a second audio channel signal based on a first downmixed signal. Furthermore, the audio decoder can be configured to use residual signal-assisted multichannel decoding to provide at least a third audio channel signal and a fourth audio channel signal based on a second downmixed signal. By using residual signal-assisted multichannel decoding, audio quality can be further improved because separation between the first and second audio channel signals and / or separation between the third and fourth audio channel signals can be performed with particularly high quality.
[0034] In a preferred embodiment, the audio decoder may be configured to use multi-channel decoding to provide the first and second residual signals based on a joint encoded representation of the first and second residual signals. The first residual signal is used to provide at least a first audio channel signal and a second audio channel signal, and the second residual signal is used to provide at least a third audio channel signal and a fourth audio channel signal. Therefore, the concept for layered decoding can be extended to providing two residual signals, one of which is used to provide the first and second audio channel signals (but the residual signal is generally not used to provide the third and fourth audio channel signals), and the other of the two residual signals is used to provide the third and fourth audio channel signals (but preferably not used to provide the first and second audio channel signals).
[0035] In a preferred embodiment, the first and second residual signals can be associated with different horizontal (or azimuth) positions of the audio scene. Therefore, the provision of horizontal segmentation (or separation) of the first and second residual signals, which can be performed in the first stage of the layered audio decoder, has been found to be particularly effective in this process (when compared to the processing performed in the second stage of the layered audio decoder). Thus, performing horizontal separation, which is particularly important for human listeners, in the first stage of the layered audio decoding provides particularly good reproduction, resulting in a good auditory impression.
[0036] In a preferred embodiment, the first residual signal is associated with the left side of the audio scene, and the second residual signal is associated with the right side of the audio scene, which conforms to human position sensitivity.
[0037] According to embodiments of the present invention, an audio encoder is created for providing an encoded representation based on at least four audio channel signals. The audio encoder is configured to obtain a first set of common bandwidth extension parameters based on a first audio channel signal and a third audio channel signal. The audio encoder is also configured to obtain a second set of common bandwidth extension parameters based on a second audio channel signal and a fourth audio channel signal. The audio encoder is configured to use multichannel coding to jointly encode at least the first and second audio channel signals to obtain a first submixed signal, and to use multichannel coding to jointly encode at least the third and fourth audio channel signals to obtain a second submixed signal. Furthermore, the audio encoder is configured to use multichannel coding to jointly encode the first and second submixed signals to obtain an encoded representation of the submixed signal.
[0038] This embodiment is based on the idea that the first set of common bandwidth extension parameters should be obtained from audio channel signals represented by different submixed signals jointly encoded only in the second stage of the layered audio encoder. In parallel with the audio decoder described above, the relationship between the audio channel signals combined only in the second stage of the layered audio decoder can be reproduced with particularly high accuracy on the audio decoder side. Therefore, it has been found that two audio signals effectively combined only in the second stage of the layered encoder are extremely suitable for obtaining the set of common bandwidth extension parameters, because multichannel bandwidth extension is optimally applied to the audio channel signals, and the relationship between these audio channel signals can be well reconstructed on the audio decoder side. Therefore, it has been found that, in terms of achievable audio quality, deriving the set of common bandwidth extension parameters from the audio channel signals combined only in the second stage of the layered audio encoder is better than obtaining the set of common bandwidth extension parameters from such audio channel signals combined in the first stage of the layered audio encoder. However, it has also been found that optimal audio quality can be obtained by deriving the set of common bandwidth extension parameters from the audio channel signals before jointly encoding the audio channel signals in the first stage of the layered audio encoder.
[0039] In a preferred embodiment, the first and second submixed signals are associated with different horizontal (or azimuth) positions of the audio scene. This concept is based on the idea that optimal auditory impression can be achieved if the signals associated with different horizontal positions are jointly encoded only in the second stage of the hierarchical audio encoder.
[0040] In a preferred embodiment, the first submixed signal is associated with the left side of the audio scene, and the second submixed signal is associated with the right side of the audio scene. Thus, these multi-channel signals associated with different sides of the audio scene provide a set of common bandwidth extension parameters. Therefore, this set of common bandwidth extension parameters is highly suitable for human abilities to distinguish audio sources at different sides.
[0041] In a preferred embodiment, the first and second audio channel signals are associated with vertically adjacent positions in the audio scene. Furthermore, the third and fourth audio channel signals are also associated with vertically adjacent positions in the audio scene. It has been found that a good auditory impression can be obtained by jointly encoding the audio channel signals associated with vertically adjacent positions in the audio scene in the first stage of the layered encoder, while preferably deriving a set of common bandwidth extension parameters from audio channel signals not associated with vertically adjacent positions (but associated with different horizontal positions or different azimuth positions).
[0042] In a preferred embodiment, the first and third audio channel signals are associated with a first common horizontal plane (or first common height) of the audio scene, but with different horizontal positions (or azimuth positions) of the audio scene, and the second and fourth audio channel signals are associated with a second common horizontal plane (or second common height) of the audio scene, but with different horizontal positions (or azimuth positions) of the audio scene, wherein the first horizontal plane is different from the second horizontal plane. It has been found that this spatial association of the audio channel signals can be used to achieve particularly good audio coding results (and therefore, audio decoding results).
[0043] In a preferred embodiment, the first and second audio channel signals are associated with a first vertical plane (or a first azimuth position) of the audio scene, but with different vertical positions (or different heights) of the audio scene. Furthermore, the third and fourth audio channel signals are preferably associated with a second vertical plane (or a second azimuth position) of the audio scene, but with different vertical positions (or different heights) of the audio scene, wherein the first common vertical plane is different from the second common vertical plane. This spatial association of the audio channel signals has been found to result in better audio coding quality.
[0044] In a preferred embodiment, the first and second audio channel signals are associated with the left side of the audio scene, and the third and fourth audio channel signals are associated with the right side of the audio scene. Therefore, a good auditory impression can be achieved while decoding remains bit-rate efficient.
[0045] In a preferred embodiment, the first and third audio channel signals are associated with the lower part of the audio scene, and the second and fourth audio channel signals are associated with the upper part of the audio scene. This arrangement also helps to obtain effective audio coding with a good auditory impression.
[0046] In a preferred embodiment, the audio encoder is configured to perform horizontal combination when providing an encoded representation of the submixed signal based on a first submixed signal and a second submixed signal using multichannel encoding. In parallel with the above description of the audio decoder, it has been found that a particularly good auditory impression can be obtained if horizontal combination is performed in the second stage of the audio encoder (when compared with the first stage of the audio encoder), because the horizontal position of the audio object is particularly relevant to the listener, and because the second stage of the layered audio encoder generally corresponds to the first stage of the layered audio decoder described above.
[0047] In a preferred embodiment, the audio encoder is configured to perform vertical combination when providing a first downmixed signal based on the first and second audio channel signals using multichannel decoding. Furthermore, the audio decoder is preferably configured to perform vertical combination when providing a second downmixed signal based on the third and fourth audio channel signals. Thus, vertical combination is performed in the first stage of the audio encoder. This is advantageous because the vertical position of an audio object is generally less important to a human listener than its horizontal position, allowing the degradation in reproduction caused by layered encoding (and therefore layered decoding) to remain reasonably small.
[0048] In a preferred embodiment, the audio encoder is configured to use prediction-based multichannel coding to provide a joint coded representation of the first and second submixed signals based on the first and second submixed signals. This prediction-based multichannel coding has been found to be particularly suitable for joint coding performed in the second stage of the layered encoder. Referring to the above description of the audio decoder, this description can also be applied in parallel.
[0049] In a preferred embodiment, prediction-based multichannel coding is used to provide prediction parameters that describe the contribution of signal components derived from signal components of previous frames to the provision of the downmixed signal for the current frame. Therefore, good signal reconstruction can be achieved on the audio encoder side, where the audio encoder can apply these prediction parameters, which describe the contribution of signal components derived from signal components of previous frames to the provision of the downmixed signal for the current frame.
[0050] In a preferred embodiment, prediction-based multichannel coding can be operated in the MDCT domain. Therefore, prediction-based multichannel coding is particularly suitable for the final coding of the output signal (e.g., a common submixed signal) of prediction-based multichannel coding, which is typically performed in the MDCT domain to keep blocking artifacts reasonably small.
[0051] In a preferred embodiment, the prediction-based multichannel coding is USAC Complex Stereo Predictive Coding. The use of USAC Complex Stereo Predictive Coding facilitates implementation because existing hardware and / or program code can be easily reused to implement a layered audio encoder.
[0052] In a preferred embodiment, the audio encoder is configured to use residual signal-assisted multichannel encoding to provide a joint encoded representation of the first and second submixed signals based on the first and second submixed signals. Therefore, particularly good reproduction quality can be achieved at the audio decoder side.
[0053] In a preferred embodiment, the audio encoder is configured to use parametric multichannel coding to provide a first submixed signal based on a first audio channel signal and a second audio channel signal. Furthermore, the audio encoder is configured to use parametric multichannel coding to derive a second submixed signal based on a third audio channel signal and a fourth audio channel signal. It has been found that the use of parametric multichannel coding provides a good trade-off between reproduction quality and bit rate when applied to the first stage of a layered audio encoder.
[0054] In a preferred embodiment, parameter-based multichannel coding is configured to provide one or more parameters describing the desired correlation between two channels and / or the step difference between the two channels. Therefore, efficient coding with a suitable bit rate is possible without significantly degrading audio quality.
[0055] In a preferred embodiment, parameter-based multichannel coding operates in the QMF domain, which is highly suitable for preprocessing that can be performed on audio channel signals.
[0056] In a preferred embodiment, the parameter-based multichannel coding is MPEG Surround 2-1-2 coding or Unified Stereo coding. The use of this coding concept can significantly reduce implementation effort.
[0057] In a preferred embodiment, the audio encoder is configured to use residual signal-assisted multichannel encoding to provide a first downmixed signal based on the first and second audio channel signals. Furthermore, the audio encoder can be configured to use residual signal-assisted multichannel encoding to provide a second downmixed signal based on the third and fourth audio channel signals. Therefore, even better audio quality may be obtained.
[0058] In a preferred embodiment, the audio encoder is configured to provide a joint coded representation of a first residual signal and a second residual signal using multichannel coding. The first residual signal is obtained by jointly coding at least a first audio channel signal and a second audio channel signal, and the second residual signal is obtained by jointly coding at least a third audio channel signal and a fourth audio channel signal. It has been found that the layered coding concept can even be applied to the residual signal provided in the first stage of layered audio coding. By using the joint coding of the residual signals, the dependency (or correlation) between the audio channel signals can be utilized, as this dependency (or correlation) is typically also reflected in the residual signal.
[0059] In a preferred embodiment, the first and second residual signals are associated with different horizontal (or azimuth) positions of the audio scene. Therefore, the coherence between the residual signals can be encoded with good accuracy in the second stage of hierarchical coding. This takes into account the reproduction of the coherence (or correlation) between different horizontal (or azimuth) positions on the audio decoder side, while maintaining a good auditory impression.
[0060] In a preferred embodiment, the first residual signal is associated with the left side of the audio scene, and the second residual signal is associated with the right side of the audio scene. Therefore, joint encoding of the first and second residual signals associated with different horizontal (or azimuth) positions is performed in the second stage of the audio encoder, taking into account high-quality reproduction on the audio decoder side.
[0061] According to a preferred embodiment of the present invention, a method for providing at least four audio channel signals based on an encoded representation is provided. The method includes: providing a first submixed signal and a second submixed signal based on a joint encoded representation of a first submixed signal and a second submixed signal using (first) multichannel decoding. The method further includes: providing at least a first audio channel signal and a second audio channel signal based on the first submixed signal using (second) multichannel decoding; and providing at least a third audio channel signal and a fourth audio channel signal based on the second submixed signal using (third) multichannel decoding. The method further includes: performing (first) multichannel bandwidth expansion based on the first and third audio channel signals to obtain channel signals with first and third bandwidth expansions. The method further includes: performing (second) multichannel bandwidth expansion based on the second and fourth audio channel signals to obtain channel signals with second and fourth bandwidth expansions. This method is based on the same considerations as the audio decoder described above.
[0062] According to a preferred embodiment of the present invention, a method is created for providing an encoded representation based on at least four audio channel signals. The method includes: obtaining a first set of common bandwidth extension parameters based on a first audio channel signal and a third audio channel signal. The method further includes: obtaining a second set of common bandwidth extension parameters based on a second audio channel signal and a fourth audio channel signal. The method further includes: jointly encoding at least the first and second audio channel signals using multichannel coding to obtain a first submixed signal; and jointly encoding at least the third and fourth audio channel signals using multichannel coding to obtain a second submixed signal. The method further includes: jointly encoding the first and second submixed signals using multichannel coding to obtain an encoded representation of the submixed signal. This method is based on the same considerations as the audio encoder described above.
[0063] According to other embodiments of the present invention, a computer program is created for performing the methods mentioned herein. Attached Figure Description
[0064] Embodiments of the invention will then be described with reference to the accompanying drawings, in which:
[0065] Figure 1 A schematic block diagram of an audio encoder according to an embodiment of the present invention is shown;
[0066] Figure 2 A schematic block diagram of an audio decoder according to an embodiment of the present invention is shown;
[0067] Figure 3 A schematic block diagram of an audio decoder according to another embodiment of the present invention is shown;
[0068] Figure 4 A schematic block diagram of an audio encoder according to an embodiment of the present invention is shown;
[0069] Figure 5 A schematic block diagram of an audio decoder according to an embodiment of the present invention is shown;
[0070] Figure 6 shows a schematic block diagram of an audio decoder according to another embodiment of the present invention;
[0071] Figure 7 A flowchart is shown below illustrating a method for providing an encoded representation based on at least four audio channel signals according to an embodiment of the present invention.
[0072] Figure 8 A flowchart is shown below illustrating a method for providing at least four audio channel signals based on an encoded representation according to an embodiment of the present invention;
[0073] Figure 9 A flowchart illustrating a method for providing an encoded representation based on at least four audio channel signals according to an embodiment of the present invention is shown; and
[0074] Figure 10 A flowchart is shown below illustrating a method for providing at least four audio channel signals based on an encoded representation according to an embodiment of the present invention;
[0075] Figure 11 A schematic block diagram of an audio encoder according to an embodiment of the present invention is shown;
[0076] Figure 12 A schematic block diagram of an audio encoder according to another embodiment of the present invention is shown;
[0077] Figure 13A schematic block diagram illustrating an audio decoder according to an embodiment of the present invention;
[0078] Figure 14a illustrates the syntax representation of the bitstream, which can be compared with... Figure 13 Used together with an audio encoder;
[0079] Figure 14b shows a tabular representation of the different values of the parameter qceIndex;
[0080] Figure 15 A schematic block diagram of a 3D audio encoder that can be used according to the concept of the present invention is shown;
[0081] Figure 16 A schematic block diagram illustrating a 3D audio decoder that can be used according to the concept of the present invention is shown; and
[0082] Figure 17 A schematic block diagram of a format converter is shown.
[0083] Figure 18 A schematic representation of the topology of a four-channel unit (QCE) according to an embodiment of the present invention is shown;
[0084] Figure 19 A schematic block diagram of an audio decoder according to an embodiment of the present invention is shown;
[0085] Figure 20 A detailed schematic block diagram of a QCE decoder according to an embodiment of the present invention is shown; and
[0086] Figure 21 A detailed schematic block diagram of a four-channel encoder according to an embodiment of the present invention is shown. Detailed Implementation
[0087] 1. According to Figure 1 audio encoder
[0088] Figure 1A schematic block diagram of an audio encoder, all designated as 100, is shown. The audio encoder 100 is configured to provide an encoded representation based on at least four audio channel signals. The audio encoder 100 is configured to receive a first audio channel signal 110, a second audio channel signal 112, a third audio channel signal 114, and a fourth audio channel signal 116. Furthermore, the audio encoder 100 is configured to provide an encoded representation of a first submixed signal 120 and an encoded representation of a second submixed signal 122, as well as a joint encoded representation 130 of the residual signal. The audio encoder 100 includes a residual signal-assisted multichannel encoder 140, which is configured to jointly encode the first audio channel signal 110 and the second audio channel signal 112 using residual signal-assisted multichannel encoding to obtain the first submixed signal 120 and the first residual signal 142. The audio signal encoder 100 further includes a residual signal-assisted multichannel encoder 150, which is configured to jointly encode at least a third audio channel signal 114 and a fourth audio channel signal 116 using residual signal-assisted multichannel encoding to obtain a second downmixed signal 122 and a second residual signal 152. The audio decoder 100 also includes a multichannel encoder 160, which is configured to jointly encode a first residual signal 142 and a second residual signal 152 using multichannel encoding to obtain a joint encoded representation 130 of the residual signals 142 and 152.
[0089] Regarding the function of the audio encoder 100, it should be noted that the audio encoder 100 performs layered encoding, wherein the first audio channel signal 110 and the second audio channel signal 112 are jointly encoded using a residual signal-assisted multichannel encoder 140, wherein both a first submix signal 120 and a first residual signal 142 are provided. The first residual signal 142 may, for example, describe the difference between the first audio channel signal 110 and the second audio channel signal 112, and / or may describe some or any signal characteristics that cannot be represented by the first submix signal 120 and optional parameters provided by the residual signal-assisted multichannel encoder 140. In other words, the first residual signal 142 may be a refined residual signal taking into account the decoding results obtainable based on the first submix signal 120 and any possible parameters provided by the residual signal-assisted multichannel encoder 140. For example, compared to a pure reconstruction of higher-order signal characteristics (such as, for example, correlation characteristics, covariance characteristics, order difference characteristics, etc.), the first residual signal 142 can at least take into account a portion of the waveform reconstruction of the first audio channel signal 110 and the second audio channel signal 112 on the audio decoder side. Similarly, the residual signal-assisted multichannel encoder 150 provides both the second downmixed signal 122 and the second residual signal 152 based on the third audio channel signal 114 and the fourth audio channel signal 116, such that the second residual signal takes into account the refinement of the signal reconstruction of the third audio channel signal 114 and the fourth audio channel signal 116 on the audio decoder side. The second residual signal 152 can therefore function as the same as the first residual signal 142. However, if the audio channel signals 110, 112, 114, 116 include some correlation, the first residual signal 142 and the second residual signal 152 are usually also correlated to some extent. Therefore, the joint encoding of the first residual signal 142 and the second residual signal 152 using the multichannel encoder 160 typically involves high efficiency because multichannel encoding of related signals usually reduces the bit rate by utilizing compliance. Thus, the first residual signal 142 and the second residual signal 152 can be encoded with good accuracy while keeping the bit rate of the joint encoded representation 130 of the residual signals reasonably small.
[0090] In short, according to Figure 1 The embodiments provide layered multichannel encoding, wherein good reproduction quality can be achieved by using a multichannel encoder 140, 150 assisted by residual signals, and wherein a moderate bit rate requirement can be maintained by jointly encoding a first residual signal 142 and a second residual signal 152.
[0091] Another possible improvement to the audio encoder 100 is possible. (See reference...) Figure 4 , Figure 11 and Figure 12Some of these improvements are described below. However, it should be noted that the audio encoder 100 can also be adapted to run in parallel with the audio decoder described herein, wherein the function of the audio encoder is generally opposite to that of the audio decoder.
[0092] 2. According to Figure 2 audio decoder
[0093] Figure 2 A schematic block diagram of an audio decoder is shown, all of which are specified as 200.
[0094] The audio decoder 200 is configured to receive an encoded representation including a joint encoded representation 210 of a first residual signal and a second residual signal. The audio decoder 200 also receives representations of a first downmixed signal 212 and a second downmixed signal 214. The audio decoder 200 is configured to provide a first audio channel signal 220, a second audio channel signal 222, a third audio channel signal 224, and a fourth audio channel signal 226.
[0095] The audio decoder 200 includes a multichannel decoder 230 configured to provide the first residual signal 232 and the second residual signal 234 based on a joint encoded representation 210 of the first residual signal 232 and the second residual signal 234. The audio decoder 200 also includes a (first) residual signal-assisted multichannel decoder 240 configured to use multichannel decoding to provide a first audio channel signal 220 and a second audio channel signal 222 based on a first downmixed signal 212 and the first residual signal 232. The audio decoder 200 further includes a (second) residual signal-assisted multichannel decoder 250 configured to provide a third audio channel signal 224 and a fourth audio channel signal 226 based on a second downmixed signal 214 and the second residual signal 234.
[0096] Regarding the functionality of the audio decoder 200, it should be noted that the audio signal decoder 200 provides the first audio channel signal 220 and the second audio channel signal 222 based on a (first) common residual signal-assisted multichannel decoding 240, wherein the decoding quality of the multichannel decoding is improved by the first residual signal 232 (compared to non-residual signal-assisted decoding). In other words, the first downmixed signal 212 provides “coarse” information about the first audio channel signal 220 and the second audio channel signal 222, wherein, for example, the difference between the first audio channel signal 220 and the second audio channel signal 222 can be described by (optional) parameters and by the first residual signal 232, which can be received by the residual signal-assisted multichannel decoder 240. Therefore, the first residual signal 232 can, for example, take into account partial waveform reconstruction of the first audio channel signal 220 and the second audio channel signal 222.
[0097] Similarly, the (second) residual signal-assisted multichannel decoder 250 provides a third audio channel signal 224 and a fourth audio channel signal 226 based on a second downmixed signal 214, wherein the second downmixed signal 214 may, for example, "coarsely" describe the third audio channel signal 224 and the fourth audio channel signal 226. Furthermore, the difference between the third audio channel signal 224 and the fourth audio channel signal 226 may, for example, be described by an optional parameter and by a second residual signal 234, which may be received by the (second) residual signal-assisted multichannel decoder 250. Therefore, the estimation of the second residual signal 234 may, for example, take into account the partial waveform reconstruction of the third audio channel signal 224 and the fourth audio channel signal 226. Thus, the second residual signal 234 may take into account the enhancement of the reconstruction quality of the third audio channel signal 224 and the fourth audio channel signal 226.
[0098] However, the first residual signal 232 and the second residual signal 234 are derived from the joint coded representation 210 of the first and second residual signals. This multi-channel decoding performed by the multi-channel decoder 230 takes into account high decoding efficiency because the first audio channel signal 220, the second audio channel signal 222, the third audio channel signal 224, and the fourth audio channel signal 226 are generally similar or "correlated". Therefore, the first residual signal 232 and the second residual signal 234 are also generally similar or "correlated", and this can be taken advantage of by deriving the first residual signal 232 and the second residual signal 234 from the joint coded representation 210 using multi-channel decoding.
[0099] Therefore, it is possible to decode the residual signals by means of the joint encoding representation 210 based on the residual signals 232 and 234, and to obtain high decoding quality with a moderate bit rate by using each of the residual signals for decoding two or more audio channel signals.
[0100] In summary, the audio decoder 200 takes high coding efficiency into account by providing high-quality audio channel signals 220, 222, 224, and 226.
[0101] It should be noted that references will follow. Figure 3 , Figure 5 Figure 6 and Figure 13 This describes the additional features and functions that can be optionally implemented in the audio decoder 200. However, it should be noted that the audio encoder 200 can include the advantages mentioned above without any additional modifications.
[0102] 3. According to Figure 3 audio decoder
[0103] Figure 3 A schematic block diagram of an audio decoder according to another embodiment of the present invention is shown. Figure 3 All audio decoders are specified as 300. Audio decoder 300 is similar to that specified according to... Figure 2 The audio decoder 200 makes the above explanation applicable. However, the audio decoder 300 adds additional features and functions compared to the audio decoder 200, as will be explained below.
[0104] Audio decoder 300 is configured to receive a joint encoded representation 310 of a first residual signal and a second residual signal. Furthermore, audio decoder 300 is configured to receive a joint encoded representation 360 of a first downmixed signal and a second downmixed signal. Additionally, audio decoder 300 is configured to provide a first audio channel signal 320, a second audio channel signal 322, a third audio channel signal 324, and a fourth audio channel signal 326. Audio decoder 300 includes a multichannel decoder 330 configured to receive the joint encoded representation 310 of the first residual signal and the second residual signal, and to provide a first residual signal 332 and a second residual signal 334 based on the joint encoded representation. Audio decoder 300 also includes a (first) residual signal-assisted multichannel decoder 340, which receives the first residual signal 332 and the first downmixed signal 312, and provides the first audio channel signal 320 and the second audio channel signal 322. The audio decoder 300 also includes a (second) residual signal-assisted multichannel decoder 350, which is configured to receive a second residual signal 334 and a second downmixed signal 314, and to provide a third audio channel signal 324 and a fourth audio channel signal 326.
[0105] The audio decoder 300 also includes another multichannel decoder 370, which is configured to receive a joint coded representation 360 of the first submixed signal and the second submixed signal, and to provide the first submixed signal 312 and the second submixed signal 314 based on the joint coded representation.
[0106] Some other specific details of the audio decoder 300 will be described below. However, it should be noted that an actual audio decoder does not need to implement all of these additional features and functions. Instead, the features and functions described below can be added individually to the audio decoder 200 (or any other audio decoder) to progressively improve the audio decoder 200 (or any other audio decoder).
[0107] In a preferred embodiment, the audio decoder 300 receives a joint coded representation 310 of the first residual signal and the second residual signal, wherein the joint coded representation 310 may include a downmixed signal of the first residual signal 332 and the second residual signal 334, and a common residual signal of the first residual signal 332 and the second residual signal 334. Additionally, the joint coded representation 310 may, for example, include one or more prediction parameters. Therefore, the multichannel decoder 330 may be a multichannel decoder assisted by predicted residual signals. For example, the multichannel decoder 330 may be USAC complex stereo prediction as described, for example, in the “Complex Stereo Prediction” section of the international standard ISO / IEC 23003-3:2012. For example, the multichannel decoder 330 may be configured to estimate prediction parameters that describe the contribution of signal components derived using signal components from previous frames to providing the first residual signal 332 and the second residual signal 334 for the current frame. Furthermore, the multichannel decoder 330 can be configured to apply a common residual signal (included in the joint coding representation 310) with a first symbol to obtain a first residual signal 332, and apply the common residual signal (included in the joint coding representation 310) with a second symbol opposite to the first symbol to obtain a second residual signal 334. Thus, the common residual signal can at least partially describe the difference between the first residual signal 332 and the second residual signal 334. However, the multichannel decoder 330 can estimate the mixed signal, the common residual signal, and one or more prediction parameters (all included in the joint coding representation 310) to obtain the first residual signal 332 and the second residual signal 334, as described in the above-referenced international standard ISO / IEC 23003-3:2012. Furthermore, it should be noted that the first residual signal 332 can be associated with a first horizontal position (or azimuth position) (e.g., a left horizontal position), and the second residual signal 334 can be associated with a second horizontal position (or azimuth position) of the audio scene (e.g., a right horizontal position).
[0108] The joint encoded representation 360 of the first and second submixed signals preferably includes a submixed signal of the first and second submixed signals, a common residual signal of the first and second submixed signals, and one or more prediction parameters. In other words, there exists a "common" submixed signal formed by submixing the first submixed signal 312 and the second submixed signal 314, and there exists a "common" residual signal that at least partially describes the difference between the first submixed signal 312 and the second submixed signal 314. The multichannel decoder 370 is preferably a multichannel decoder assisted by the predicted residual signal, for example, a USAC complex stereo predictive decoder. In other words, the multichannel decoder 370 providing the first submixed signal 312 and the second submixed signal 314 can be substantially the same as the multichannel decoder 330 providing the first residual signal 332 and the second residual signal 334, making the above explanation and references applicable as well. Furthermore, it should be noted that the first downmix signal 312 is preferably associated with a first horizontal or azimuth position of the audio scene (e.g., a left horizontal or azimuth position), and the second downmix signal 314 is preferably associated with a second horizontal or azimuth position of the audio scene (e.g., a right horizontal or azimuth position). Therefore, the first downmix signal 312 and the first residual signal 332 can be associated with the same first horizontal or azimuth position (e.g., a left horizontal position), and the second downmix signal 314 and the second residual signal 334 can be associated with the same second horizontal or azimuth position (e.g., a right horizontal position). Thus, both the multichannel decoder 370 and the multichannel decoder 330 can perform horizontal partitioning (or horizontal separation or horizontal distribution).
[0109] The residual signal-assisted multichannel decoder 340 is preferably parameter-based and can therefore receive one or more parameters 342 describing the desired correlation between two channels (e.g., between the first audio channel signal 320 and the second audio channel signal 322) and / or the step difference between the two channels. For example, the residual signal-assisted multichannel decoder 340 may be based on MPEG surround sound coding with residual signal extension (as described, for example, in ISO / IEC 23003-1:2007), or a "unified stereo decoder" (as described, for example, in ISO / IEC 23003-3, Chapter 7.11 (decoder) and Annex B.21 (description of encoder and definition of the term "unified stereo")). Thus, the residual signal-assisted multichannel decoder 340 can provide a first audio channel signal 320 and a second audio channel signal 322, wherein the first audio channel signal 320 and the second audio channel signal 322 are associated with vertically adjacent positions in the audio scene. For example, the first audio channel signal can be associated with the lower left position of the audio scene, and the second audio channel signal can be associated with the upper left position of the audio scene (such that the first audio channel signal 320 and the second audio channel signal 322 are associated, for example, with the same horizontal or azimuth position of the audio scene, or with azimuth positions no more than 30 degrees apart). In other words, the residual signal-assisted multichannel decoder 340 can perform vertical partitioning (or distribution, or separation).
[0110] The residual signal-assisted multichannel decoder 350 functions similarly to the residual signal-assisted multichannel decoder 340, wherein the third audio channel signal may be associated, for example, with the lower right position of the audio scene, and the fourth audio channel signal may be associated, for example, with the upper right position of the audio scene. In other words, the third and fourth audio channel signals may be associated with vertically adjacent positions in the audio scene, and may also be associated with the same horizontal or azimuth position in the audio scene, wherein the residual signal-assisted multichannel decoder 350 performs vertical division (or separation, or distribution).
[0111] In conclusion, according to Figure 3The audio decoder 300 performs layered audio decoding, wherein left-right partitioning is performed in the first stage (multi-channel decoder 330, multi-channel decoder 370), and top-bottom partitioning is performed in the second stage (residual signal-assisted multi-channel decoders 340, 350). Furthermore, residual signals 332, 334 are encoded using joint coding representation 310, and downmixed signals 312, 314 are encoded (using joint coding representation 360). Thus, the correlation between different channels is used in both the encoding (and decoding) of downmixed signals 312, 314 and the encoding (and decoding) of residual signals 332, 334. Therefore, high coding efficiency is achieved, and the correlation between signals is also utilized.
[0112] 4. According to Figure 4 audio encoder
[0113] Figure 4 A schematic block diagram of an audio encoder according to another embodiment of the present invention is shown. Figure 4 The audio encoder 400 is designated entirely by 400. The audio encoder 400 is configured to receive four audio channel signals: a first audio channel signal 410, a second audio channel signal 412, a third audio channel signal 414, and a fourth audio channel signal 416. Furthermore, the audio encoder 400 is configured to provide encoded representations based on the audio channel signals 410, 412, 414, and 416, wherein the encoded representations include a joint encoded representation 420 of two downmixed signals, and encoded representations of a first set 422 and a second set 424 of common bandwidth extension parameters. The audio encoder 400 includes a first bandwidth extension parameter extractor 430, which is configured to obtain a first set 422 of common bandwidth extraction parameters based on the first audio channel signal 410 and the third audio channel signal 414. The audio encoder 400 also includes a second bandwidth extension parameter extractor 440, which is configured to obtain a second set 424 of common bandwidth extension parameters based on the second audio channel signal 412 and the fourth audio channel signal 416.
[0114] Furthermore, the audio encoder 400 includes a (first) multichannel encoder 450 configured to jointly encode at least a first audio channel signal 410 and a second audio channel signal 412 using multichannel encoding to obtain a first submixed signal 452. Additionally, the audio encoder 400 includes a (second) multichannel encoder 460 configured to jointly encode at least a third audio channel signal 414 and a fourth audio channel signal 416 using multichannel encoding to obtain a second submixed signal 462. Furthermore, the audio encoder 400 includes a (third) multichannel encoder 470 configured to jointly encode the first submixed signal 452 and the second submixed signal 462 using multichannel encoding to obtain a jointly encoded representation 420 of the submixed signal.
[0115] Regarding the function of the audio encoder 400, it should be noted that the audio encoder 400 performs layered multichannel encoding, wherein the first audio channel signal 410 and the second audio channel signal 412 are combined in a first stage, and the third audio channel signal 414 and the fourth audio channel signal 416 are also combined in the first stage to obtain a first submixed signal 452 and a second submixed signal 462. The first submixed signal 452 and the second submixed signal 462 are then jointly encoded in a second stage. However, it should be noted that the first bandwidth extension parameter extractor 430 provides a first set 422 of common bandwidth extraction parameters based on the audio channel signals 410 and 414 processed by different multichannel encoders 450 and 460 in the first stage of layered multichannel encoding. Similarly, the second bandwidth extension parameter extractor 440 provides a second set 424 of common bandwidth extraction parameters based on the different audio channel signals 412 and 416 processed by different multichannel encoders 450 and 460 in the first processing stage. This particular processing order offers the advantage that the set of bandwidth expansion parameters 422, 424 is based on channels combined only in the second stage of layered coding (i.e., in the multichannel encoder 470). This is advantageous because combining such audio channels in the first stage of layered coding is desirable, as the relationship between these audio channels is not highly relevant to the perception of the sound source location. Instead, it is desirable that the relationship between the first and second submix signals primarily determines the perception of the sound source location, because the relationship between the first submix signal 452 and the second submix signal 462 is better maintained compared to the relationship between the corresponding audio channel signals 410, 412, 414, 416. In other words, it has been found that the first set 422 of common bandwidth extension parameters is based on two audio channels (audio channel signals) that contribute to the difference between the lower mixed signals 452 and 462, and the second set 424 of common bandwidth extension parameters is based on audio channel signals 412 and 416 that also contribute to the difference between the lower mixed signals 452 and 462, which is achieved by the processing of audio channel signals in the aforementioned layered multichannel coding. Therefore, when compared with the channel relationship between the first lower mixed signal 452 and the second lower mixed signal 462, the first set 422 of common bandwidth extension parameters is based on a similar channel relationship, where the channel relationship between the first lower mixed signal and the second lower mixed signal generally dominates the spatial impression generated on the audio decoder side. Therefore, the provision of the first set 422 of bandwidth extension parameters and the provision of the second set 424 of bandwidth extension parameters are extremely suitable for the spatial auditory impression generated on the audio decoder side.
[0116] 5. According to Figure 5 audio decoder
[0117] Figure 5A schematic block diagram of an audio decoder according to another embodiment of the present invention is shown. Figure 5 All audio decoders are specified with 500.
[0118] The audio decoder 500 is configured to receive a joint encoded representation 510 of a first submixed signal and a second submixed signal. Furthermore, the audio decoder 500 is configured to provide a first bandwidth-extended channel signal 520, a second bandwidth-extended channel signal 522, a third bandwidth-extended channel signal 524, and a fourth bandwidth-extended channel signal 526.
[0119] The audio decoder 500 includes a (first) multichannel decoder 530 configured to use multichannel decoding to provide a first submixed signal 532 and a second submixed signal 534 based on a joint encoded representation 510 of a first submixed signal and a second submixed signal. The audio decoder 500 also includes a (second) multichannel decoder 540 configured to use multichannel decoding to provide at least a first audio channel signal 542 and a second audio channel signal 544 based on the first submixed signal 532. The audio decoder 500 further includes a (third) multichannel decoder 550 configured to use multichannel decoding to provide at least a third audio channel signal 556 and a fourth audio channel signal 558 based on the second submixed signal 544. Furthermore, the audio decoder 500 includes a (first) multichannel bandwidth extension 560, which is configured to perform multichannel bandwidth extension based on a first audio channel signal 542 and a third audio channel signal 556 to obtain a first bandwidth extended channel signal 520 and a third bandwidth extended channel signal 524. Additionally, the audio decoder includes a (second) multichannel bandwidth extension 570, which is configured to perform multichannel bandwidth extension based on a second audio channel signal 544 and a fourth audio channel signal 558 to obtain a second bandwidth extended channel signal 522 and a fourth bandwidth extended channel signal 526.
[0120] Regarding the functionality of the audio decoder 500, it should be noted that the audio decoder 500 performs layered multi-channel decoding. The division between the first submix signal 532 and the second submix signal 534 is performed in the first stage of layered decoding. In the second stage of layered decoding, the first audio channel signal 542 and the second audio channel signal 544 are derived from the first submix signal 532, and the third audio channel signal 556 and the fourth audio channel signal 558 are derived from the second submix signal 550. However, both the first multi-channel bandwidth extension 560 and the second multi-channel bandwidth extension 570 each receive one audio channel signal derived from the first submix signal 532 and one audio channel signal derived from the second submix signal 534. Because good channel separation is typically achieved by the (first) multichannel decoder 530 (performed as the first stage of layered multichannel decoding), when compared with the second stage of layered decoding, it can be seen that each multichannel bandwidth extension 560, 570 receives a well-separated input signal (because the input signal originates from the well-separated first submix signal 532 and second submix signal 534). Therefore, the multichannel bandwidth extensions 560, 570 can take into account stereo characteristics, which are important for auditory impression, and these stereo characteristics are well represented by the relationship between the first submix signal 532 and the second submix signal 534, thus providing a good auditory impression.
[0121] In other words, the “cross” structure of the audio decoder takes into account good multichannel bandwidth extension, which takes into account the stereo relationship between channels, wherein each of the multichannel bandwidth extension stages 560 and 570 receives input signals from both (second stage) multichannel decoders 540 and 550.
[0122] However, it should be noted that the audio decoder 500 can be derived from this document based on... Figure 2 , Figure 3 According to 6 and Figure 13 The audio decoder may be supplemented by any of the features and functions described herein, wherein it is possible to introduce the corresponding features into the audio decoder 500 to gradually improve the performance of the audio decoder.
[0123] 6. Based on the audio decoder in Figure 6
[0124] Figure 6 shows a schematic block diagram of an audio decoder according to another embodiment of the present invention. All audio decoders according to Figure 6 are designated as 600. The audio decoder 600 according to Figure 6 is similar to that according to... Figure 5 The audio decoder 500 makes the above explanation applicable. However, the audio decoder 600 has been supplemented with some features and functions that can be introduced into the audio decoder 500, either alone or in combination, for improvement.
[0125] Audio decoder 600 is configured to receive a joint encoded representation 610 of a first submixed signal and a second submixed signal, and to provide a first bandwidth-extended signal 620, a second bandwidth-extended signal 622, a third bandwidth-extended signal 624, and a fourth bandwidth-extended signal 626. Audio decoder 600 includes a multichannel decoder 630 configured to receive the joint encoded representation 610 of the first submixed signal and the second submixed signal, and to provide a first submixed signal 632 and a second submixed signal 634 based on the joint encoded representation. Audio decoder 600 further includes a multichannel decoder 640 configured to receive the first submixed signal 632, and to provide a first audio channel signal 542 and a second audio channel signal 544 based on the first submixed signal. Audio decoder 600 also includes a multichannel decoder 650 configured to receive the second submixed signal 634, and to provide a third audio channel signal 656 and a fourth audio channel signal 658. The audio decoder 600 also includes a (first) multichannel bandwidth extension 660, which is configured to receive a first audio channel signal 642 and a third audio channel signal 656, and to provide a first bandwidth extended channel signal 620 and a third bandwidth extended channel signal 624 based on the first and third audio channel signals. Furthermore, a (second) multichannel bandwidth extension 670 receives a second audio channel signal 644 and a fourth audio channel signal 658, and to provide a second bandwidth extended channel signal 622 and a fourth bandwidth extended channel signal 626 based on the second and fourth audio channel signals.
[0126] The audio decoder 600 also includes another multichannel decoder 680, which is configured to receive a joint coded representation 682 of a first residual signal and a second residual signal, and the other multichannel decoder provides a first residual signal 684 for use by the multichannel decoder 640 and a second residual signal 686 for use by the multichannel decoder 650 based on the joint coded representation.
[0127] The multichannel decoder 630 is preferably a multichannel decoder assisted by a predicted residual signal. For example, the multichannel decoder 630 may be substantially the same as the multichannel decoder 370 described above. For example, the multichannel decoder 630 may be a USAC complex stereo predictive decoder as described above and as referenced in the USAC standard above. Therefore, the joint encoded representation 610 of the first and second submixed signals may, for example, include a (common) submixed signal of the first and second submixed signals, a (common) residual signal of the first and second submixed signals, and one or more prediction parameters estimated by the multichannel decoder 630.
[0128] Furthermore, it should be noted that the first downmix signal 632 may be associated, for example, with a first horizontal or azimuth position (e.g., left horizontal position) of the audio scene, and the second downmix signal 634 may be associated, for example, with a second horizontal or azimuth position (e.g., right horizontal position) of the audio scene.
[0129] Furthermore, the multichannel decoder 680 may be, for example, a multichannel decoder associated with predicted residual signals. The multichannel decoder 680 may be substantially the same as the multichannel decoder 330 described above. For example, the multichannel decoder 680 may be a USAC complex stereo predictive decoder, as mentioned above. Therefore, the joint encoded representation 682 of the first and second residual signals may include a (common) submixed signal of the first and second residual signals, a (common) residual signal of the first and second residual signals, and one or more prediction parameters estimated by the multichannel decoder 680. Furthermore, it should be noted that the first residual signal 684 may be associated with a first horizontal or azimuth position (e.g., left horizontal position) of the audio scene, and the second residual signal 686 may be associated with a second horizontal or azimuth position (e.g., right horizontal position) of the audio scene.
[0130] Multichannel decoder 640 may be, for example, a parameter-based multichannel decoder, similar to, for example, MPEG surround sound multichannel decoder as described above and in the referenced standards. However, in the presence of (optional) multichannel decoder 680 and (optional) first residual signal 684, multichannel decoder 640 may be a parameter-based, residual signal-assisted multichannel decoder, similar to, for example, a unified stereo decoder. Thus, multichannel decoder 640 may be substantially the same as multichannel decoder 340 described above, and multichannel decoder 640 may, for example, receive the parameters 342 described above.
[0131] Similarly, multichannel decoder 650 may be substantially the same as multichannel decoder 640. Thus, multichannel decoder 650 may be, for example, parametric and optionally residual signal assisted (in the presence of an optional multichannel decoder 680).
[0132] Furthermore, it should be noted that the first audio channel signal 642 and the second audio channel signal 644 are preferably associated with vertically adjacent spatial positions in the audio scene. For example, the first audio channel signal 642 is associated with the lower left position of the audio scene, and the second audio channel signal 644 is associated with the upper left position of the audio scene. Therefore, the multichannel decoder 640 performs a vertical division (or separation, or distribution) of the audio content described by the first downmix signal 632 (and, optionally, by the first residual signal 684). Similarly, the third audio channel signal 656 and the fourth audio channel signal 658 are associated with vertically adjacent positions in the audio scene, and are preferably associated with the same horizontal or azimuthal position in the audio scene. For example, the third audio channel signal 656 is preferably associated with the lower right position of the audio scene, and the fourth audio channel signal 658 is preferably associated with the upper right position of the audio scene. Therefore, the multichannel decoder 650 performs the vertical division (or separation, or distribution) of the audio content described by the second downmixed signal 634 (and, optionally, by the second residual signal 686).
[0133] However, the first multichannel bandwidth extension 660 receives a first audio channel signal 642 and a third audio channel 656, which are associated with the lower left and lower right positions of the audio scene, respectively. Therefore, the first multichannel bandwidth extension 660 performs multichannel bandwidth extension based on two audio channel signals associated with the same horizontal plane (e.g., lower horizontal plane) or height of the audio scene and different sides (left / right) of the audio scene. Thus, when performing bandwidth extension, the multichannel bandwidth extension can take into account stereo characteristics (e.g., human stereo perception). Similarly, the second multichannel bandwidth extension 670 can also take into account stereo characteristics because the second multichannel bandwidth extension operates on audio channel signals at the same horizontal plane (e.g., upper horizontal plane) or height of the audio scene but at different horizontal positions (different sides) (left / right).
[0134] To further summarize, the layered audio decoder 600 includes the following structure: left / right partitioning (or separation, or distribution) is performed in the first stage (multichannel decoding 630, 680), vertical partitioning (separation or distribution) is performed in the second stage (multichannel decoding 640, 650), and multichannel bandwidth expansion operates on a pair of left / right signals (multichannel bandwidth expansion 660, 670). This "crossing" of the decoding path allows for left / right separation, which is particularly important for auditory impression (e.g., more important than top / bottom partitioning), to be performed in the first processing stage of the layered audio decoder, and also allows for multichannel bandwidth expansion on a pair of left and right audio channel signals, which in turn results in a particularly good auditory impression. Top / bottom partitioning is performed as an intermediate stage between left / right separation and multichannel bandwidth expansion, which allows four audio channel signals (or bandwidth-expanded channel signals) to be derived without significantly degrading the auditory impression.
[0135] 7. According to Figure 7 Method
[0136] Figure 7 A flowchart is shown for a method 700 for providing an encoded representation based on at least four audio channel signals.
[0137] Method 700 includes jointly encoding at least a first audio channel signal and a second audio channel signal using residual signal-assisted multichannel coding 710 to obtain a first submixed signal and a first residual signal. The method also includes jointly encoding at least a third audio channel signal and a fourth audio channel signal using residual signal-assisted multichannel coding 720 to obtain a second submixed signal and a second residual signal. The method further includes jointly encoding the first residual signal and the second residual signal using multichannel coding 730 to obtain an encoded representation of the residual signal. However, it should be noted that method 700 may be supplemented by any of the features and functions described herein with respect to audio encoders and audio decoders.
[0138] 8. According to Figure 8 Method
[0139] Figure 8 A flowchart of a method 800 for providing at least four audio channel signals based on an encoded representation is shown.
[0140] Method 800 includes using multi-channel decoding to provide 810 the first residual signal and the second residual signal based on a joint encoded representation of the first residual signal and the second residual signal. Method 800 also includes using residual signal-assisted multi-channel decoding to provide 820 the first audio channel signal and the second audio channel signal based on the first downmixed signal and the first residual signal. The method further includes using residual signal-assisted multi-channel decoding to provide 830 the third audio channel signal and the fourth audio channel signal based on the second downmixed signal and the second residual signal.
[0141] Furthermore, it should be noted that method 800 may be supplemented by any of the features and functions described herein with respect to audio decoders and audio encoders.
[0142] 9. According to Figure 9 Method
[0143] Figure 9 A flowchart is shown for a method 900 for providing an encoded representation based on at least four audio channel signals.
[0144] Method 900 includes obtaining a first set of common bandwidth extension parameters based on a first audio channel signal and a third audio channel signal. Method 900 also includes obtaining a second set of common bandwidth extension parameters based on a second audio channel signal and a fourth audio channel signal. The method further includes using multichannel coding to jointly encode at least the first audio channel signal and the second audio channel signal to obtain a first downmixed signal, and using multichannel coding to jointly encode at least the third audio channel signal and the fourth audio channel signal to obtain a second downmixed signal. The method further includes using multichannel coding to jointly encode the first downmixed signal and the second downmixed signal to obtain an encoded representation of the downmixed signal.
[0145] It should be noted that some of the steps of method 900, which do not involve specific interdependencies, can be performed in any order or in parallel. Furthermore, it should be noted that method 900 may be supplemented by any of the features and functions described herein with respect to audio encoders and audio decoders.
[0146] 10. According to Figure 10 Method
[0147] Figure 10 A flowchart is shown for a method 1000 for providing at least four audio channel signals based on an encoded representation.
[0148] Method 1000 includes: providing 1010 a first submixed signal and a second submixed signal based on a joint encoded representation of a first submixed signal and a second submixed signal using multichannel decoding; providing 1020 at least a first audio channel signal and a second audio channel signal based on the first submixed signal using multichannel decoding; providing 1030 at least a third audio channel signal and a fourth audio channel signal based on the second submixed signal using multichannel decoding; performing 1040 multichannel bandwidth expansion based on the first audio channel signal and the third audio channel signal to obtain a channel signal with first bandwidth expansion and a channel signal with third bandwidth expansion; and performing 1050 multichannel bandwidth expansion based on the second audio channel signal and the fourth audio channel signal to obtain a channel signal with second bandwidth expansion and a channel signal with fourth bandwidth expansion.
[0149] It should be noted that some of the steps of method 1000 can be performed in any order or in parallel. Furthermore, it should be noted that method 1000 can be supplemented by any of the features and functions described herein with respect to audio encoders and audio decoders.
[0150] 11. According to Figure 11 , Figure 12 and Figure 13 Implementation examples
[0151] In the following sections, some additional embodiments and underlying considerations according to the present invention will be described.
[0152] Figure 11 A schematic block diagram of an audio encoder 1100 according to an embodiment of the present invention is shown. The audio encoder 1100 is configured to receive a lower left channel signal 1110, a upper left channel signal 1112, a lower right channel signal 1114, and an upper right channel signal 1116.
[0153] Audio encoder 1100 includes a first multi-channel audio encoder (or encoder) 1120, which is an MPEG surround sound 2-1-2 audio encoder (or encoder) or a unified stereo audio encoder (or encoder), and receives a lower left channel signal 1110 and a upper left channel signal 1112. The first multi-channel audio encoder 1120 provides a lower left mixed signal 1122 and (optionally) a left residual signal 1124. Furthermore, audio encoder 1100 includes a second multi-channel encoder (or encoder) 1130, which is an MPEG surround sound 2-1-2 encoder (or encoder) or a unified stereo encoder (or encoder), and receives a lower right channel signal 1114 and an upper right channel signal 1116. The second multi-channel audio encoder 1130 provides a lower right mixed signal 1132 and (optionally) a right residual signal 1134. The audio encoder 1100 also includes a stereo encoder (or code) 1140 that receives a left-lower mixed signal 1122 and a right-lower mixed signal 1132. Furthermore, as a first stereo code 1140 for complex predictive stereo coding, it receives psychoacoustic model information 1142 from a psychoacoustic model. For example, the psychoacoustic model information 1142 may describe psychoacoustic correlations such as different frequency bands or sub-bands, psychoacoustic masking effects, etc. The stereo code 1140 provides a channel pair unit (CPE) "low-mix," which is specified as 1144 and describes the left-lower mixed signal 1122 and the right-lower mixed signal 1132 in a joint coding form. Additionally, the audio encoder 1100 optionally includes a second stereo encoder (or code) 1150 configured to receive an optional left residual signal 1124 and an optional right residual signal 1134, as well as the psychoacoustic model information 1142. The second stereo code 1150, as a complex predictive stereo code, is configured to provide a channel pair unit (CPE) "residual" that represents the left residual signal 1124 and the right residual signal 1134 in a joint encoded form.
[0154] Encoder 1100 (and other audio encoders described herein) is based on the idea of utilizing horizontal and vertical signal compliance by layering available USAC stereo tools (i.e., coding concepts available in USAC coding). Vertically adjacent channel pairs are combined using MPEG surround sound 2-1-2 or unified stereo (specified as 1120 and 1130) with either band-limited or full-band residual signals (specified as 1124 and 1134). The output of each vertical channel pair is a downmixed signal 1122, 1132, and for unified stereo, residual signals 1124, 1134. To meet the requirement of unmasked perception in both ears, both downmixed signals 1122 and 1132 are horizontally combined and jointly encoded using complex prediction in the MDCT domain (encoder 1140), including the possibility of left-right and center-side encoding. The same method can be applied to the horizontally combined residual signals 1124, 1134. This concept is... Figure 11 As shown in the image.
[0155] refer to Figure 11 The hierarchical structure of the interpretation can be achieved by enabling two stereo tools (e.g., two USAC stereo tools) and re-sorting the channels between them. Therefore, no additional pre-processing / post-processing steps are required, and the bitstream syntax of the payload used for sending tools remains unchanged (e.g., substantially unchanged compared to the USAC standard). This idea leads to... Figure 12 The encoder structure shown is shown in the figure.
[0156] Figure 12 A schematic block diagram of an audio encoder 1200 according to an embodiment of the present invention is shown. The audio encoder 1200 is configured to receive a first channel signal 1210, a second channel signal 1212, a third channel signal 1214, and a fourth channel signal 1216. The audio encoder 1200 is configured to provide a bitstream 1220 for a first channel pair unit and a bitstream 1222 for a second channel pair unit.
[0157] The audio encoder 1200 includes a first multichannel encoder 1230, which is an MPEG surround sound 2-1-2 encoder or a unified stereo encoder, and receives a first channel signal 1210 and a second channel signal 1212. Furthermore, the first multichannel encoder 1230 provides a first downmixed signal 1232, an MPEG surround sound payload 1236, and (optionally) a first residual signal 1234. The audio encoder 1200 also includes a second multichannel encoder 1240, which is an MPEG surround sound 2-1-2 encoder or a unified stereo encoder, and receives a third channel signal 1214 and a fourth channel signal 1216. The second multichannel encoder 1240 provides a first downmixed signal 1242, an MPEG surround sound payload 1246, and (optionally) a second residual signal 1244.
[0158] The audio encoder 1200 also includes a first stereo code 1250, which is a complex predictive stereo code. The first stereo code 1250 receives a first downmixed signal 1232 and a second downmixed signal 1242. The first stereo code 1250 provides a joint coded representation 1252 of the first downmixed signal 1232 and the second downmixed signal 1242, wherein the joint coded representation 1252 may include a representation of a (common) downmixed signal (of the first downmixed signal 1232 and the second downmixed signal 1242) and a (common) residual signal (of the first downmixed signal 1232 and the second downmixed signal 1242). Furthermore, the (first) complex predictive stereo code 1250 provides a complex predictive payload 1254, which typically includes one or more complex predictive coefficients. Additionally, the audio encoder 1200 also includes a second stereo code 1260, which is a complex predictive stereo code. The second stereo encoder 1260 receives a first residual signal 1234 and a second residual signal 1244 (or a zero input value, if no residual signal is provided by the multichannel encoders 1230, 1240). The second stereo encoder 1260 provides a joint coded representation 1262 of the first residual signal 1234 and the second residual signal 1244, which may, for example, include a (common) submixed signal (of the first residual signal 1234 and the second residual signal 1244) and a (common) residual signal (of the first residual signal 1234 and the second residual signal 1244). Furthermore, the complex prediction stereo encoder 1260 provides a complex prediction payload 1264, which typically includes one or more prediction coefficients.
[0159] Furthermore, the audio encoder 1200 includes a psychoacoustic model 1270, which provides information for controlling the first complex predictive stereo code 1250 and the second complex predictive stereo code 1260. For example, the information provided by the psychoacoustic model 1270 can describe which frequency bands or frequency grids have high psychoacoustic relevance and should be encoded with high precision. However, it should be noted that using the information provided by the psychoacoustic model 1270 is optional.
[0160] Furthermore, the audio encoder 1200 includes a first encoder and multiplexer 1280, which receives a joint encoded representation 1252 from a first complex predictive stereo encoder 1250, a complex predictive payload 1254 from the first complex predictive stereo encoder 1250, and an MPEG surround sound payload 1236 from a first multichannel audio encoder 1230. Additionally, the first encoder and multiplexer 1280 may receive information from a psychoacoustic model 1270 describing, for example, considering psychoacoustic masking effects, which encoding precision should be applied to which frequency bands or sub-bands. Therefore, the first encoder and multiplexer 1280 provides a first channel-to-cell bitstream 1220.
[0161] Furthermore, the audio encoder 1200 includes a second encoding and multiplexing 1290 configured to receive a joint encoded representation 1262 provided by a second complex predictive stereo codec 1260, a complex predictive payload 1264 provided by the second complex predictive stereo codec 1260, and an MPEG surround sound payload 1246 provided by a second multichannel audio encoder 1240. Additionally, the second encoding and multiplexing 1290 can receive information from a psychoacoustic model 1270. Therefore, the second encoding and multiplexing 1290 provides a second channel pair unit bitstream 1222.
[0162] Regarding the functions of the audio encoder 1200, please refer to the above explanation, and also refer to the information regarding... Figure 2 , Figure 3 , Figure 5 And an explanation of the audio encoder in Figure 6.
[0163] Furthermore, it should be noted that this concept can be extended to the joint encoding of multiple MPEG surround sound grids for horizontally correlated, vertically correlated, or other geometrically correlated channels, as well as the combination of downmixed and residual signals into complex predictive stereo pairs, taking into account their geometric and perceptual properties. This leads to a generalized decoder architecture.
[0164] The implementation of the four-channel unit is described below. In a three-dimensional audio coding system, a layered combination of four channels is used to form a four-channel unit (QCE). The QCE consists of two USAC channel pair units (CPEs) (or provides two USAC channel pair units, or receives two USAC channel pair units). Vertical channel pairs are combined using MPS 2-1-2 or unified stereo. The lower mixed channels are jointly coded in the first channel pair unit (CPE). If residual coding is applied, the residual signal is jointly coded in the second channel pair unit (CPE); otherwise, the signal in the second CPE is set to zero. The two channel pair units (CPEs) perform complex predictions for joint stereo coding, including the possibility of left-right coding and center-side coding. To preserve the perceptual stereo properties of the high-frequency components of the signal, a stereo SBR (Spectral Bandwidth Replication) is applied between the top-left / top-right channel pair and the bottom-left / bottom-right channel pair via an additional re-sorting step before applying the SBR.
[0165] Reference Figure 13 Describe possible decoder structures. Figure 13 A schematic block diagram of an audio decoder according to an embodiment of the present invention is shown. The audio decoder 1300 is configured to receive a first bitstream 1310 representing a first channel pair unit and a second bitstream 1312 representing a second channel pair unit. However, the first bitstream 1310 and the second bitstream 1312 may be included in a common total bitstream.
[0166] The audio decoder 1300 is configured to provide a first bandwidth extended channel signal 1320, a second bandwidth extended channel signal 1322, a third bandwidth extended channel signal 1324, and a fourth bandwidth extended channel signal 1326. The first bandwidth extended channel signal 1320 may, for example, represent the lower left position of the audio scene; the second bandwidth extended channel signal 1322 may, for example, represent the upper left position of the audio scene; the third bandwidth extended channel signal 1324 may, for example, be associated with the lower right position of the audio scene; and the fourth bandwidth extended channel signal 1326 may, for example, be associated with the upper right position of the audio scene.
[0167] The audio decoder 1300 includes a first bitstream decoder 1330 configured to receive a bitstream 1310 for a first channel pair unit and, based on the bitstream, provide a joint coded representation of two downmixed signals, a complex prediction payload 1334, an MPEG surround sound payload 1336, and a spectral bandwidth replication payload 1338. The audio decoder 1300 also includes a first complex prediction stereo decoder 1340 configured to receive a joint coded representation 1332 and a complex prediction payload 1334, and, based on the joint coded representation and the complex prediction payload, provide a first downmixed signal 1342 and a second downmixed signal 1344. Similarly, the audio decoder 1300 includes a second bitstream decoder 1350 configured to receive a bitstream 1312 for the second channel unit and, based on this bitstream, provide a joint coded representation 1352 of the two residual signals, a complex prediction payload 1354, an MPEG surround sound payload 1356, and a spectral bandwidth replication bit payload 1358. The audio decoder also includes a second complex prediction stereo decoder 1360, which provides a first residual signal 1362 and a second residual signal 1364 based on the joint coded representation 1352 and the complex prediction payload 1354.
[0168] Furthermore, the audio decoder 1300 includes a first MPEG surround sound multichannel decoder 1370, which is either MPEG surround sound 2-1-2 decoding or unified stereo decoding. The first MPEG surround sound multichannel decoder 1370 receives a first downmixed signal 1342, a first residual signal 1362 (optional), and an MPEG surround sound payload 1336, and provides a first audio channel signal 1372 and a second audio channel signal 1374 based on the first downmixed signal, the first residual signal, and the MPEG surround sound payload. The audio decoder 1300 also includes a second MPEG surround sound multichannel decoder 1380, which is either MPEG surround sound 2-1-2 multichannel decoding or unified stereo multichannel decoding. The second MPEG surround sound multichannel decoder 1380 receives a second downmixed signal 1344 and a second residual signal 1364 (optional), as well as an MPEG surround sound payload 1356, and provides a third audio channel signal 1382 and a fourth audio channel signal 1384 based on the second downmixed signal, the second residual signal, and the MPEG surround sound payload. The audio decoder 1300 also includes a first stereo spectral bandwidth copy 1390, which is configured to receive a first audio channel signal 1372 and a third audio channel signal 1382, as well as a spectral bandwidth copy payload 1338, and provides a first bandwidth-extended channel signal 1320 and a third bandwidth-extended channel signal 1324 based on the first audio channel signal, the third audio channel signal, and the spectral bandwidth copy payload. In addition, the audio decoder includes a second stereo spectrum bandwidth replica 1394, which is configured to receive a second audio channel signal 1374 and a fourth audio channel signal 1384, as well as a spectrum bandwidth replica payload 1358, and to provide a second bandwidth extended channel signal 1322 and a fourth bandwidth extended channel signal 1326 based on the second audio channel signal, the fourth audio channel signal, and the spectrum bandwidth replica payload.
[0169] Regarding the functions of the audio decoder 1300, please refer to the above discussion, and also refer to the information provided. Figure 2 , Figure 3 , Figure 5 And a discussion of the audio decoder in Figure 6.
[0170] In the following text, examples of bitstreams that can be used for the audio encoding / decoding described herein will be described with reference to Figures 14a and 14b. It should be noted that the bitstream may be, for example, an extension of the bitstream used in the Unified Speech and Audio Coding (USAC), which is described in the aforementioned standard (ISO / IEC 23003-3:2012). For example, MPEG surround sound payloads 1236, 1246, 1336, 1356 and complex prediction payloads 1254, 1264, 1334, 1354 can be transmitted as conventional channel pair units (i.e., for channel pair units according to the USAC standard). For the use of a four-channel unit QCE transmitted in a signaled manner, the USAC channel pair configuration can be extended by two bits, as shown in Figure 14a. In other words, the two bits specified by “qceIndex” can be added to the USAC bitstream unit “UsacChannelPairElementConfig()”. The meaning of the parameter represented by the bit “qceindex” can be defined, for example, as shown in the table in Figure 14b.
[0171] For example, the two channel pairs forming the QCE can be transmitted as consecutive units, firstly including the lower mixing channel and the CPE for the MPS payload of the first MPS frame, and secondly including the residual signal (or the zero audio signal for MPS 2-1-2 encoding) and the CPE for the MPS payload of the second MPS frame.
[0172] In other words, there is only a small signaling overhead compared to the regular USAC bitstream used to transmit the four-channel unit QCE.
[0173] However, different bitstream formats can also be used.
[0174] 12. Encoding / Decoding Environment
[0175] The following will describe an audio encoding / decoding environment to which the concepts according to the present invention can be applied.
[0176] A 3D audio codec system based on the concept of the present invention can be used, employing the MPEG-D USAC codec for decoding channel and object signals. To improve the efficiency of encoding large numbers of objects, MPEG SAOC technology has been adapted. Three types of renderers perform the tasks of rendering objects to channels, rendering channels to headphones, or rendering channels to different speaker settings. When object signals are explicitly sent or object signals are encoded using SAOC parameterization, the corresponding object metadata information is compressed and multiplexed into a 3D audio bitstream.
[0177] Figure 15 A schematic block diagram of this audio encoder is shown, and Figure 16A schematic block diagram of this audio decoder is shown. In other words, Figure 15 and Figure 16 Different algorithmic blocks for a 3D audio system are shown.
[0178] refer to Figure 15 Now, some details will be explained. Figure 15 A schematic block diagram of a 3D audio encoder 1500 is shown. The encoder 1500 includes an optional pre-render / mixer 1510 that receives one or more channel signals 1512 and one or more object signals 1514, and provides one or more channel signals 1516 and one or more object signals 1518, 1520 based on the one or more channel signals and the one or more object signals. The audio encoder also includes a USAC encoder 1530 and (optionally) a SAOC encoder 1540. The SAOC encoder 1540 is configured to provide one or more SAOC transport channels 1542 and SAOC sideband information 1544 based on one or more objects 1520 provided to the SAOC encoder. Furthermore, the USAC encoder 1530 is configured to receive channel signals 1516, including channels and pre-rendered objects, from the pre-renderer / mixer, receive one or more object signals 1518 from the pre-renderer / mixer, and receive one or more SAOC transport channels 1542 and SAOC sideband information 1544, and provide an encoded representation 1532 based on the above. Additionally, the audio encoder 1500 includes an object metadata encoder 1550, which is configured to receive object metadata 1552 (which may be estimated by the pre-renderer / mixer 1510) and encode the object metadata to obtain encoded object metadata 1554. The encoded metadata is also received by the USAC encoder 1530 and used to provide the encoded representation 1532.
[0179] The following will describe some details about the various components of the audio encoder 1500.
[0180] Now for reference Figure 16 The audio decoder 1600 will be described. The audio decoder 1600 is configured to receive an encoded representation 1610 and provide a multi-channel speaker signal 1612, an earphone signal 1614, and / or a speaker signal 1616 in an alternative format (e.g., 5.1 format) based on the encoded representation.
[0181] The audio decoder 1600 includes a USAC decoder 1620 and provides one or more channel signals 1622, one or more pre-rendered object signals 1624, one or more object signals 1626, one or more SAOC transport channels 1628, SAOC sideband information 1630, and compressed object metadata information 1632 based on an encoded representation 1610. The audio decoder 1600 also includes an object renderer 1640 configured to provide one or more rendered object signals 1642 based on the object signals 1626 and the object metadata information 1644, wherein the object metadata information 1644 is provided by an object metadata decoder 1650 based on the compressed object metadata information 1632. The audio decoder 1600 also includes (optionally) an SAOC decoder 1660 configured to receive the SAOC transport channels 1628 and the SAOC sideband information 1630, and to provide one or more rendered object signals 1662 based on the SAOC transport channels and the SAOC sideband information. The audio decoder 1600 also includes a mixer 1670 configured to receive channel signals 1622, pre-rendered object signals 1624, rendered object signals 1642 and 1662, and to provide a plurality of mixed channel signals 1672 based on the above, which may, for example, constitute a multi-channel speaker signal 1612. The audio decoder 1600 may also include, for example, a binaural rendering 1680 configured to receive the mixed channel signals 1672 and to provide a headphone signal 1614 based on the mixed channel signals. Furthermore, the audio decoder 1600 may include a format converter 1690 configured to receive the mixed channel signals 1672 and reproduction layout information 1692, and to provide a speaker signal 1616 for an alternative speaker setup based on the mixed channel signals and the reproduction layout information.
[0182] The following text will describe some details about the components of the audio encoder 1500 and the audio decoder 1600.
[0183] Pre-render / mixer
[0184] The pre-render / mixer 1510 is optionally used to convert a channel-plus-object input scene into a channel scene before encoding. Functionally, this pre-render / mixer can be the same as the object renderer / mixer described below. Object pre-rendering can, for example, ensure deterministic signal entropy at the encoder input, which is substantially independent of the number of simultaneously active object signals. No object metadata is sent during object pre-rendering. Discreet object signals are rendered to the channel layout configured for use by the encoder. The weights of the objects for each channel are obtained from the associated object metadata (OAM) 1552.
[0185] USAC Core Codec
[0186] The core codecs 1530 and 1620, used for speaker channel signals, discreet object signals, object submixed signals, and pre-rendered signals, are based on MPEG-D USAC technology. This core codec processes the encoding of a large volume of signals by creating channel and object mapping information based on the geometric and semantic information of the input channels and object assignments. This mapping information describes how the input channels and objects are mapped to USAC channel units (CPE, SCE, LFE) and how the corresponding information is sent to the decoder. All additional payloads (such as SAOC data or object metadata) have been extended and are already considered in the encoder rate control.
[0187] Object encoding can vary depending on renderer speed / distortion requirements and interactivity needs. The following object encoding variations are possible:
[0188] 1. Pre-rendered object: The object signal is pre-rendered and mixed into a 22.2-channel signal before encoding. See the 22.2-channel signal for subsequent encoding chains.
[0189] 2. Cautious Object Waveform: The object is supplied to the encoder as a monophonic waveform. In addition to the channel signal, the encoder uses a monophonic unit (SCE) to transmit the object. The object is rendered and mixed on the receiver side. Compressed object metadata information is sent along the receiver / render.
[0190] 3. Parametric Object Waveform: SAOC parameters describe the properties of objects and their relationships. USAC is used to encode the downmixing of object signals. Parameter information is sent side-by-side. The number of downmixing channels is selected based on the number of objects and the overall data rate. Compressed object metadata information is sent to the SAOC renderer.
[0191] SAOC
[0192] The SAOC encoder 1540 and SAOC decoder 1660 for object signals are based on MPEG SAOC technology. The system can recreate, modify, and render numerous audio objects based on a small number of transmit channels and additional parameter data (object step difference (OLD), inter-object correlation (IOC), and downmixing gain (DMG)). The additional parameter data exhibits a significantly lower data rate than if all objects were transmitted individually, making encoding extremely efficient. The SAOC encoder takes the object / channel signal (e.g., a monotone waveform) as input and outputs parameter information (encapsulated in 3D audio bitstreams 1532 and 1610) and the SAOC transmit channels (encoded and transmitted using mono units).
[0193] The SAOC decoder 1600 reconstructs the object / channel signal based on the decoded SAOC transmission channel 1628 and parameter information 1630, and generates the output audio scene based on the reconstructed layout, decompressed object metadata information, and optionally based on user interaction information.
[0194] Object Metadata Codec
[0195] For each object, associated metadata specifying the object's geometric location and volume in 3D space is efficiently encoded by quantizing the object's properties in time and space. The compressed object metadata cOAM 1554 and 1632 are sent to the receiver as sideband information.
[0196] Object renderer / mixer
[0197] The object renderer uses compressed object metadata to generate object waveforms according to a given reproduction format. Each object is rendered to certain output channels based on its metadata. The output of this box is derived from the sum of partial results. If the channel-based content and carefully selected object / parameter objects are decoded, the channel-based waveform and the rendered object waveform are blended before the generated waveform is output (or before the generated waveform is fed to a post-processor module such as a binaural renderer or speaker renderer module).
[0198] Binocular renderer
[0199] The binaural renderer module 1680 generates binaural submixing of multi-channel audio material, so that each input channel is represented by a virtual sound source. Processing is performed frame-by-frame in the QMF domain. Binauralization is based on measured binaural spatial impulse responses.
[0200] Speaker renderer / format conversion
[0201] The speaker renderer 1690 converts the transmit channel configuration to the desired reproduction format. This speaker renderer is therefore referred to hereinafter as a "format converter". The format converter performs a conversion to a lower number of output channels; that is, the format converter creates a downmix. The system automatically generates an optimal downmixing matrix for a given combination of input and output formats, and applies this matrix during the downmixing process. The format converter takes into account both standard speaker configurations and random configurations with non-standard speaker positions.
[0202] Figure 17A schematic block diagram of a format converter is shown. As can be seen, the format converter 1700 receives a mixer output signal 1710, such as a mixed channel signal 1672, and provides a speaker signal 1712, such as a speaker signal 1616. The format converter includes a downmixing configurator 1730 and a downmixing process 1720 in the QMF domain, wherein the downmixing configurator provides configuration information for the downmixing process 1720 based on mixer output layout information 1732 and reproduction layout information 1734.
[0203] Furthermore, it should be noted that the concepts described above, such as audio encoder 100, audio decoder 200 or 300, audio encoder 400, audio decoder 500 or 600, method 700, 800, 900 or 1000, audio encoder 1100 or 1200, and audio decoder 1300, may be used within audio encoder 1500 and / or audio decoder 1600. For example, the previously mentioned audio encoder / decoder can be used for encoding or decoding channel signals associated with different spatial locations.
[0204] 13. Alternative Embodiments
[0205] Some additional embodiments will be described below.
[0206] For reference Figures 18 to 21 The following will explain additional embodiments according to the present invention.
[0207] It should be noted that the so-called "quad-channel unit" (QCE) can be regarded as a tool for an audio decoder, which can be used, for example, to decode three-dimensional audio content.
[0208] In other words, the Quad-Channel Unit (QCE) is a more efficient four-channel joint coding method for encoding both horizontally and vertically distributed channels. A QCE consists of two consecutive CPEs and is formed by layering and combining joint stereo tools that offer the possibility of complex stereo prediction tools in the horizontal direction and the possibility of MPEG surround sound-based stereo tools in the vertical direction. This is achieved by enabling two stereo tools and swapping the output channels between the applied tools. Stereo SBR is performed in the horizontal direction to preserve high-frequency left-right relationships.
[0209] Figure 18 The topology of the QCE is shown. It should be noted that... Figure 18 QCE is extremely similar to Figure 11 The QCE (Quality, Cost, and Execution) allows for reference to the above explanation. However, it should be noted that... Figure 18In QCE, the use of a psychoacoustic model is not mandatory when performing complex stereo prediction (optional, although such use is certainly possible). Furthermore, it can be seen that a first stereo spectral bandwidth replication (SBR) is performed based on the lower left and lower right channels, and a second stereo spectral bandwidth replication (SBR) is performed based on the upper left and upper right channels.
[0210] In the following text, some terms and definitions will be provided, which may be applied to some embodiments.
[0211] The data unit qceIndex indicates the QCE mode of the CPE. Refer to Figure 14b for the meaning of the bitstream variable qceIndex. Note that qceIndex describes whether two subsequent units of type UsacChannelPairElement() are treated as a quad-channel unit (QCE). Different QCE modes are shown in Figure 14b. qceIndex should be the same for the two subsequent units forming a QCE.
[0212] In the following sections, some helper units will be defined, which can be used in some implementations of the present invention:
[0213] cplx_out_dmx_L[] is the first channel of the first CPE after complex predictive stereo decoding.
[0214] cplx_out_dmx_R[] is the second channel of the first CPE after complex predictive stereo decoding.
[0215] cplx_out_res_L[] is the second CPE after complex predictive stereo decoding (zero if qceIndex=1).
[0216] cplx_out_res_R[] The second channel of the second CPE after complex predictive stereo decoding (zero if qceIndex=1).
[0217] mps_out_L_1[] is the first output channel of the first MPS frame.
[0218] mps_out_L_2[] is the second output channel of the first MPS frame.
[0219] mps_out_R_1[] is the first output channel of the second MPS frame.
[0220] mps_out_R_2[] is the second output channel of the second MPS frame.
[0221] sbr_out_L_1[] is the first output channel of the first stereo SBR frame.
[0222] sbr_out_R_1[] is the second output channel of the first stereo SBR frame.
[0223] sbr_out_L_2[] is the first output channel of the second stereo SBR frame.
[0224] sbr_out_R_2[] is the second output channel of the second stereo SBR frame.
[0225] The decoding process performed in an embodiment according to the present invention will be explained below.
[0226] In `UsacChannelPairElementConfig()`, the syntax unit (or bitstream unit, or data unit) `qceIndex` indicates whether the CPE belongs to the QCE and whether residual encoding is used. If `qceIndex` is not equal to 0, the current CPE forms a QCE together with its subsequent unit, which should be a CPE with the same `qceIndex`. Stereo SBRs are always used for QCEs, therefore the syntax term `stereoConfigIndex` should be 3 and `bsStereoSbr` should be 1.
[0227] When qceIndex == 1, it is used only for the payload of MPEG surround sound and SBR and no related audio signal data is included in the second CPE, and the syntax unit bsResidualCoding is set to 0.
[0228] The presence of a residual signal in the second CPE is indicated by qceIndex == 2. In this case, the syntax unit bsResidualCoding is set to 1.
[0229] However, some different and potentially simplified signal transmission schemes can also be used.
[0230] Decoding of joint stereo with the possibility of complex stereo prediction is performed as described in Section 7.7 of ISO / IEC 23003-3. The output of the first CPE is the mixed signal cplx_out_dmx_L[] and cplx_out_dmx_R[] under MPS. If residual coding is used (i.e., qceIndex == 2), the output of the second CPE is the MPS residual signal cplx_out_res_L[] and cplx_out_res_R[], and if no residual signal has been sent (i.e., qceIndex == 1), a zero signal is inserted.
[0231] Before applying MPEG surround sound decoding, swap the second channel of the first component (cplx_out_dmx_R[]) and the first channel of the second component (cplx_out_res_L[]).
[0232] Decoding of MPEG surround sound is performed as described in Section 7.11 of ISO / IEC 23003-3. However, in some embodiments, the decoding may be modified compared to conventional MPEG surround sound decoding if residual coding is used. This can be modified as defined in Section 7.11.2.7 (Figure 23) of ISO / IEC 23003-3 for decoding of MPEG surround sound without residual coding using SBR, so that the stereo SBR is also used for bsResidualCoding == 1, resulting in... Figure 19 The diagram shows a decoder. Figure 19 A schematic block diagram is shown for an audio encoder with bsResidualCoding == 0 and bsStereoSbr == 1.
[0233] like Figure 19 As can be seen, the USAC core decoder 2010 provides the downmixed signal (DMX) 2012 to the MPS (MPEG Surround Sound) decoder 2020, which provides a first decoded audio signal 2022 and a second decoded audio signal 2024. The stereo SBR decoder 2030 receives the first decoded audio signal 2022 and the second decoded audio signal 2024, and provides a left bandwidth extended audio signal 2032 and a right bandwidth extended audio signal 2034 based on the first decoded audio signal and the second decoded audio signal.
[0234] Before applying stereo SBR, the second channel of the first component (mps_out_L_2[]) and the first channel of the second component (mps_out_R_1[]) are swapped to allow left and right stereo SBR. After applying stereo SBR, the second output channel of the first component (sbr_out_R_1[]) and the first channel of the second component (sbr_out_L_2[]) are swapped again to restore the input channel order.
[0235] exist Figure 20 The QCE decoder structure is illustrated below. Figure 20 A schematic diagram of the QCE decoder is shown.
[0236] It should be noted that Figure 20 The schematic diagram is very similar to Figure 13 The schematic diagram is provided for further reference to the explanation above. Additionally, it should be noted that... Figure 20Some signal identifiers have been added, where reference is made to the definitions in this section. Additionally, the final re-sorting of the channels, performed after the stereo SBR, is shown.
[0237] Figure 21 A schematic block diagram of a four-channel encoder 2200 according to an embodiment of the present invention is shown. In other words, in Figure 21 The example shown is a four-channel encoder (four-channel unit) that can be considered a core encoder tool.
[0238] The four-channel encoder 2200 includes a first stereo SBR 2210, which receives a first left channel input signal 2212 and a second left channel input signal 2214, and provides a first SBR payload 2215, a first left channel SBR output signal 2216, and a first right channel SBR output signal 2218 based on the first left channel input signal and the second left channel input signal. Furthermore, the four-channel encoder 2200 includes a second stereo SBR, which receives a second left channel input signal 2222 and a second right channel input signal 2224, and provides a first SBR payload 2225, a first left channel SBR output signal 2226, and a first right channel SBR output signal 2228 based on the second left channel input signal and the second right channel input signal.
[0239] The four-channel encoder 2200 includes a first MPEG surround sound (MPS 2-1-2 or unified stereo) multichannel encoder 2230, which receives a first left channel SBR output signal 2216 and a second left channel SBR output signal 2226, and provides a first MPS payload 2232, a left channel MPEG surround sound submixed signal 2234, and (optionally) a left channel MPEG surround sound residual signal 2236 based on the first left channel SBR output signal and the second left channel SBR output signal. The four-channel encoder 2200 also includes a second MPEG surround sound (MPS 2-1-2 or unified stereo) multichannel encoder 2240, which receives a first right channel SBR output signal 2218 and a second right channel SBR output signal 2228, and provides a first MPS payload 2242, a right channel MPEG surround sound downmixed signal 2244, and (optionally) a right channel MPEG surround sound residual signal 2246 based on the first right channel SBR output signal and the second right channel SBR output signal.
[0240] The four-channel encoder 2200 includes a first complex predictive stereo codec 2250, which receives a left-channel MPEG surround submix signal 2234 and a right-channel MPEG surround submix signal 2244, and provides a complex predictive payload 2252 and a joint coded representation 2254 of the left-channel MPEG surround submix signal 2234 and the right-channel MPEG surround submix signal 2244 based on the left-channel MPEG surround submix signal and the right-channel MPEG surround submix signal. The four-channel encoder 2200 includes a second complex prediction stereo codec 2260 that receives a left-channel MPEG surround sound residual signal 2236 and a right-channel MPEG surround sound residual signal 2246. The second complex prediction stereo codec provides a complex prediction payload 2262 and a joint encoded representation 2264 of the left-channel MPEG surround sound submixed signal 2236 and the right-channel MPEG surround sound submixed signal 2246 based on the left-channel MPEG surround sound residual signal and the right-channel MPEG surround sound residual signal.
[0241] The four-channel encoder also includes a first bitstream code 2270, which receives a joint encoding representation 2254, a complex prediction payload 2252, an MPS payload 2232, and an SBR payload 2215, and provides a bitstream portion representing the first channel pair units based on the above. The four-channel encoder also includes a second bitstream code 2280, which receives a joint encoding representation 2264, a complex prediction payload 2262, an MPS payload 2242, and an SBR payload 2225, and provides a bitstream portion representing the first channel pair units based on the above.
[0242] 14. Alternative Implementation Plans
[0243] Although some schemes have been described in the context of the device, it is clear that these schemes also represent descriptions of the corresponding methods, where a block or device corresponds to a method step or feature of a method step. Similarly, in the context of method steps, the schemes also represent descriptions of corresponding blocks or items or features of the corresponding device. Some or all of the method steps may be performed by (using) a hardware device, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more steps of the most important method steps may be performed by this device.
[0244] The inventive coded audio signal can be stored on a digital storage medium or transmitted via a transmission medium such as a wireless or wired transmission medium, such as the Internet.
[0245] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software. Implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, storing electronically readable control signals that cooperate with (or are capable of cooperating with) a programmable computer system to enable the execution of corresponding methods. Therefore, the digital storage medium may be computer-readable.
[0246] According to some embodiments of the present invention, a data carrier includes electronically readable control signals that are capable of cooperating with a programmable computer system to enable the execution of one of the methods described herein.
[0247] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when executed on a computer, is operable for performing one of the methods. The program code may, for example, be stored on a machine-readable medium.
[0248] Other embodiments include a computer program for performing one of the methods described herein, the computer program being stored on a machine-readable medium.
[0249] In other words, an embodiment of the inventive method is therefore a computer program having program code that, when executed on a computer, performs one of the methods described herein.
[0250] Another embodiment of the inventive method is therefore a data carrier (or digital storage medium, or computer-readable medium) comprising a computer program recorded on the data carrier for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.
[0251] Another embodiment of the inventive method is therefore a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).
[0252] Another embodiment includes a processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.
[0253] Another embodiment includes a computer having a computer program installed on it for performing one of the methods described herein.
[0254] Another embodiment of the invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may include, for example, a file server for transmitting the computer program to the receiver.
[0255] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0256] The embodiments described above are merely illustrative of the principles of the present invention. It will be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the scope of the forthcoming patent claims and not by the specific details presented through the description and explanation of the embodiments herein.
[0257] 15. Conclusion
[0258] Some conclusions will be provided below.
[0259] Embodiments of the present invention are based on the following considerations: to illustrate the signal compliance between vertically distributed channels and horizontally distributed channels, the four channels can be jointly coded by combining joint stereo coding tools in a hierarchical manner. For example, MPS 2-1-2 and / or unified stereo with band-limited residual coding or full-band residual coding can be used to combine vertical channel pairs. To meet the requirement of unmasked perception of the two ears, the output undermixing can be jointly coded, for example, by using complex prediction in the MDCT domain, which includes the possibility of left-right coding and center-side coding. If residual signals exist, the same method is used to horizontally combine the residual signals.
[0260] Furthermore, it should be noted that embodiments of the invention overcome some or all of the disadvantages of the prior art. Embodiments of the invention are suitable for 3D audio scenarios where speaker channels are distributed in layers of fruit-length height, resulting in horizontal and vertical channel pairs. It has been found that joint encoding of only two channels, as defined in USAC, is insufficient to account for spatial and perceptual relationships between channels. However, embodiments of the invention overcome this problem.
[0261] Furthermore, conventional MPEG surround sound is applied in additional pre-processing / post-processing steps, allowing residual signals to be transmitted separately when joint stereo encoding is not possible, for example, to explore the dependency between the left fundamental tone residual signal and the right fundamental tone residual signal. Conversely, embodiments of the invention take into account efficient encoding / decoding by utilizing this dependency.
[0262] To further summarize, embodiments of the present invention create apparatuses, methods, or computer programs for encoding and decoding as described herein.
[0263] References:
[0264] [1] ISO / IEC 23003-3: 2012 - Information Technology – MPEG AudioTechnologies, Part 3: Unified Speech and Audio Coding;
[0265] [2] ISO / IEC 23003-1: 2007 - Information Technology – MPEG AudioTechnologies, Part 1: MPEG Surround.
Claims
1. An audio decoder for providing at least four bandwidth-extended channel signals based on an encoded representation. in, The audio decoder is configured to provide the first submixed signal and the second submixed signal using multi-channel decoding based on a joint encoded representation of the first submixed signal and the second submixed signal; The audio decoder is configured to use multi-channel decoding to provide at least a first audio channel signal and a second audio channel signal based on the first downmixed signal; The audio decoder is configured to use multi-channel decoding to provide at least a third audio channel signal and a fourth audio channel signal based on the second downmixed signal; The audio decoder is configured to: perform multi-channel bandwidth expansion based on the first audio channel signal and the third audio channel signal to obtain a first bandwidth expanded channel signal and a third bandwidth expanded channel signal; and The audio decoder is configured to perform multi-channel bandwidth expansion based on the second audio channel signal and the fourth audio channel signal to obtain the second bandwidth expanded channel signal and the fourth bandwidth expanded channel signal.
2. The audio decoder according to claim 1, wherein, The first and second submixed signals are associated with different horizontal or azimuth positions of the audio scene.
3. The audio decoder according to claim 1, wherein, The first submixed signal is associated with the left side of the audio scene, and the second submixed signal is associated with the right side of the audio scene.
4. The audio decoder according to claim 1, wherein, The first audio channel signal and the second audio channel signal are associated with the vertically adjacent positions of the audio scene, and The third audio channel signal and the fourth audio channel signal are associated with the vertically adjacent positions of the audio scene.
5. The audio decoder according to claim 1, wherein, The first audio channel signal and the third audio channel signal are associated with a first common horizontal plane or a first common height of the audio scene, but with different horizontal or azimuth positions of the audio scene. The second audio channel signal and the fourth audio channel signal are associated with a second common horizontal plane or a second common height of the audio scene, but with different horizontal or azimuth positions of the audio scene. The first common horizontal plane or the first common height is different from the second common horizontal plane or the second common height.
6. The audio decoder according to claim 5, wherein, The first audio channel signal and the second audio channel signal are associated with the position of a first common vertical plane or a first common azimuth angle of the audio scene, but are associated with different vertical positions or heights of the audio scene. The third and fourth audio channel signals are associated with the second common vertical plane or the second common azimuth angle of the audio scene, but are associated with different vertical positions or heights of the audio scene. The position of the first common vertical plane or the first common azimuth angle is different from the position of the second common vertical plane or the second common azimuth angle.
7. The audio decoder according to claim 1, wherein, The first audio channel signal and the second audio channel signal are associated with the left side of the audio scene, and The third and fourth audio channel signals are associated with the right side of the audio scene.
8. The audio decoder according to claim 1, wherein, The first audio channel signal and the third audio channel signal are associated with the lower part of the audio scene, and The second audio channel signal and the fourth audio channel signal are associated with the upper part of the audio scene.
9. The audio decoder according to claim 1, wherein, The audio decoder is configured to perform level splitting when providing the first and second submixed signals based on a joint coded representation of the first and second submixed signals using the multichannel decoder.
10. The audio decoder according to claim 1, wherein, The audio decoder is configured to perform vertical segmentation when using the multi-channel decoder to provide at least the first audio channel signal and the second audio channel signal based on the first submixed signal; as well as The audio decoder is configured to perform vertical segmentation when using the multi-channel decoder to provide at least the third and fourth audio channel signals based on the second downmixed signal.
11. The audio decoder according to claim 1, wherein, The audio decoder is configured to perform stereo bandwidth expansion based on the first audio channel signal and the third audio channel signal to obtain the first bandwidth-expanded channel signal and the third bandwidth-expanded channel signal. The first audio channel signal and the third audio channel signal represent the first left / right channel pair; and The audio decoder is configured to perform stereo bandwidth expansion based on the second audio channel signal and the fourth audio channel signal to obtain the second bandwidth-expanded channel signal and the fourth bandwidth-expanded channel signal. The second audio channel signal and the fourth audio channel signal represent the second left / right channel pair.
12. The audio decoder according to claim 1, in, The audio decoder is configured to provide the first and second submixed signals using prediction-based multichannel decoding, based on a joint coded representation of the first and second submixed signals.
13. The audio decoder according to claim 1, in, The audio decoder is configured to: use residual signal-assisted multichannel decoding to provide the first and second submixed signals based on a joint encoded representation of the first and second submixed signals.
14. The audio decoder according to claim 1, in, The audio decoder is configured to use parameter-based multichannel decoding to provide at least the first audio channel signal and the second audio channel signal based on the first downmixed signal; The audio decoder is configured to use parameter-based multichannel decoding to provide at least the third audio channel signal and the fourth audio channel signal based on the second downmixed signal.
15. The audio decoder according to claim 14, wherein, The parameter-based multichannel decoding is configured to estimate one or more parameters that describe the desired correlation between two channels and / or the step difference between two channels, in order to provide two or more audio channel signals based on the corresponding submixed signal.
16. The audio decoder according to claim 1, in, The audio decoder is configured to: use residual signal-assisted multichannel decoding to provide at least the first audio channel signal and the second audio channel signal based on the first downmixed signal; as well as The audio decoder is configured to use residual signal-assisted multi-channel decoding to provide at least the third audio channel signal and the fourth audio channel signal based on the second downmixed signal.
17. The audio decoder according to claim 1, in, The audio decoder is configured to provide the first residual signal and the second residual signal using multi-channel decoding, based on a joint encoded representation of the first residual signal and the second residual signal, wherein the first residual signal is used to provide at least the first audio channel signal and the second audio channel signal, and the second residual signal is used to provide at least the third audio channel signal and the fourth audio channel signal.
18. The audio decoder according to claim 17, wherein, The first residual signal and the second residual signal are associated with different horizontal or azimuth positions of the audio scene.
19. The audio decoder according to claim 17, wherein, The first residual signal is associated with the left side of the audio scene, and the second residual signal is associated with the right side of the audio scene.
20. An audio encoder for providing an encoded representation based on at least four audio channel signals. in, The audio encoder is configured to obtain a first set of common bandwidth extension parameters based on the first audio channel signal and the third audio channel signal; The audio encoder is configured to obtain a second set of common bandwidth extension parameters based on the second audio channel signal and the fourth audio channel signal. The audio encoder is configured to jointly encode at least the first audio channel signal and the second audio channel signal to obtain a first submixed signal; The audio encoder is configured to jointly encode at least the third audio channel signal and the fourth audio channel signal to obtain a second submixed signal.
21. The audio encoder according to claim 20, wherein, The first and second submixed signals are associated with different horizontal or azimuth positions of the audio scene.
22. The audio encoder according to claim 20, wherein, The first submixed signal is associated with the left side of the audio scene, and the second submixed signal is associated with the right side of the audio scene.
23. The audio encoder according to claim 20, wherein, The first audio channel signal and the second audio channel signal are associated with the vertically adjacent positions of the audio scene, and The third audio channel signal and the fourth audio channel signal are associated with the vertically adjacent positions of the audio scene.
24. The audio encoder according to claim 20, wherein, The first audio channel signal and the third audio channel signal are associated with a first common horizontal plane or a first height of the audio scene, but with different horizontal or azimuth positions of the audio scene. The second audio channel signal and the fourth audio channel signal are associated with a second common horizontal plane or a second height of the audio scene, but with different horizontal or azimuth positions of the audio scene. The first common horizontal plane or the first height is different from the second common horizontal plane or the second height.
25. The audio encoder according to claim 24, wherein, The first audio channel signal and the second audio channel signal are associated with a first common vertical plane or a first azimuth angle position of the audio scene, but are associated with different vertical positions or heights of the audio scene. The third and fourth audio channel signals are associated with the second common vertical plane or second azimuth angle position of the audio scene, but are associated with different vertical positions or heights of the audio scene. The position of the first common vertical plane or the first azimuth angle is different from the position of the second common vertical plane or the second azimuth angle.
26. The audio encoder according to claim 20, wherein, The first audio channel signal and the second audio channel signal are associated with the left side of the audio scene, and The third and fourth audio channel signals are associated with the right side of the audio scene.
27. The audio encoder according to claim 20, wherein, The first audio channel signal and the third audio channel signal are associated with the lower part of the audio scene, and The second audio channel signal and the fourth audio channel signal are associated with the upper part of the audio scene.
28. The audio encoder according to claim 20, wherein, The audio encoder is configured to perform level combination when using multichannel encoding to provide an encoded representation of the submixed signal based on the first submixed signal and the second submixed signal.
29. The audio encoder according to claim 20, wherein, The audio encoder is configured to perform vertical combining when providing the first downmixed signal based on the first audio channel signal and the second audio channel signal using multichannel encoding; and The audio encoder is configured to perform vertical combination when providing the second downmixed signal based on the third and fourth audio channel signals using multichannel encoding.
30. The audio encoder according to claim 20, in, The audio encoder is configured to provide a joint coded representation of the first and second submixed signals based on the first and second submixed signals using prediction-based multichannel coding.
31. The audio encoder according to claim 20, in, The audio encoder is configured to use residual signal-assisted multichannel encoding to provide a joint encoded representation of the first and second submixed signals based on the first and second submixed signals.
32. The audio encoder according to claim 20, in, The audio encoder is configured to: use parameter-based multichannel encoding to provide the first downmixed signal based on the first audio channel signal and the second audio channel signal; and The audio encoder is configured to use parameter-based multichannel encoding to provide the second downmixed signal based on the third audio channel signal and the fourth audio channel signal.
33. The audio encoder according to claim 32, wherein, The parameter-based multichannel coding is configured to provide one or more parameters that describe the desired correlation between two channels and / or the step difference between two channels.
34. The audio encoder according to claim 20, in, The audio encoder is configured to: use residual signal-assisted multichannel encoding to provide the first downmixed signal based on the first audio channel signal and the second audio channel signal; and The audio encoder is configured to use residual signal-assisted multichannel encoding to provide the second downmixed signal based on the third audio channel signal and the fourth audio channel signal.
35. The audio encoder according to claim 20, in, The audio encoder is configured to provide a joint encoded representation of a first residual signal and a second residual signal using multichannel encoding, wherein the first residual signal is obtained by jointly encoding at least the first audio channel signal and the second audio channel signal, and the second residual signal is obtained by jointly encoding at least the third audio channel signal and the fourth audio channel signal.
36. The audio encoder according to claim 35, wherein, The first residual signal and the second residual signal are associated with different horizontal or azimuth positions of the audio scene.
37. The audio encoder according to claim 35, wherein, The first residual signal is associated with the left side of the audio scene, and the second residual signal is associated with the right side of the audio scene.
38. A method for providing at least four audio channel signals based on an encoded representation, wherein, The method includes: Multi-channel decoding is used to provide the first and second submixed signals based on a joint encoded representation of the first and second submixed signals; Using multi-channel decoding, at least a first audio channel signal and a second audio channel signal are provided based on the first downmixed signal; Using multi-channel decoding, at least a third audio channel signal and a fourth audio channel signal are provided based on the second downmixed signal; Multichannel bandwidth expansion is performed based on the first audio channel signal and the third audio channel signal to obtain channel signals with first bandwidth expansion and channel signals with third bandwidth expansion; and Multichannel bandwidth expansion is performed based on the second audio channel signal and the fourth audio channel signal to obtain the second bandwidth expanded channel signal and the fourth bandwidth expanded channel signal.
39. A method for providing an encoded representation based on at least four audio channel signals, the method comprising: A first set of common bandwidth extension parameters is obtained based on the first audio channel signal and the third audio channel signal; A second set of common bandwidth extension parameters is obtained based on the second and fourth audio channel signals; At least the first audio channel signal and the second audio channel signal are jointly encoded to obtain a first submixed signal; as well as The third and fourth audio channel signals are jointly encoded to obtain a second submixed signal.
40. A computer program product comprising a computer program, wherein when the computer program is executed on a computer, the computer program is configured to perform the method according to claim 38 or 39.
41. An audio decoder for providing at least four bandwidth-extended channel signals based on an encoded representation. in, The audio decoder is configured to provide at least a first audio channel signal and a second audio channel signal based on a first submixed signal; The audio decoder is configured to provide at least a third audio channel signal and a fourth audio channel signal based on a second downmixed signal; The audio decoder is configured to: perform multi-channel bandwidth expansion based on the first audio channel signal and the third audio channel signal to obtain a first bandwidth expanded channel signal and a third bandwidth expanded channel signal; and The audio decoder is configured to perform multi-channel bandwidth expansion based on the second audio channel signal and the fourth audio channel signal to obtain the second bandwidth expanded channel signal and the fourth bandwidth expanded channel signal.
42. A method for providing at least four audio channel signals based on an encoded representation, wherein, The method includes: Provide at least a first audio channel signal and a second audio channel signal based on the first sub-mixed signal; Provide at least a third audio channel signal and a fourth audio channel signal based on the second down-mixed signal; Multichannel bandwidth expansion is performed based on the first audio channel signal and the third audio channel signal to obtain channel signals with first bandwidth expansion and channel signals with third bandwidth expansion; and Multichannel bandwidth expansion is performed based on the second audio channel signal and the fourth audio channel signal to obtain the second bandwidth expanded channel signal and the fourth bandwidth expanded channel signal.
Citation Information
Patent Citations
Audio decoder, audio encoder, method, and computer-readable storage medium
CN105580073B