Method, apparatus and system for enhancing multi-channel audio in a reduced dynamic range domain

The use of a multi-channel generator in a GAN setting with companding techniques addresses the challenge of coding artifacts in multi-channel audio, enhancing audio quality by jointly processing channels and reducing noise.

JP2026012688APending Publication Date: 2026-01-27DOLBY INTERNATIONAL AB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025157661
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-05-20
Filing Date
2025-09-24
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing audio coding techniques at low bitrates introduce coding artifacts and noise, particularly in multi-channel audio, which are challenging to remove due to their correlation with the desired sound, and current deep learning approaches have limitations in effectively enhancing multichannel audio quality.

Method used

A method using a multi-channel generator in a generative adversarial network (GAN) setting to jointly enhance multiple channels of a reduced dynamic range audio signal, combined with companding techniques, to reduce coding artifacts and improve audio quality.

Benefits of technology

The method effectively reduces coding artifacts and enhances multi-channel audio quality by jointly processing multiple channels, resulting in improved audio fidelity and reduced noise, especially when integrated with companding operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012688000001_ABST
    Figure 2026012688000001_ABST
Patent Text Reader

Abstract

To improve the quality of a multichannel audio signal in a reduced dynamic range region.SOLUTION: The method may include receiving an audio bitstream, core decoding the audio bitstream to obtain a dynamic range reduced raw multi-channel audio signal based on the received audio bitstream, wherein the dynamic range reduced raw multi-channel audio signal comprises two or more channels, inputting the dynamic range reduced raw multi-channel audio signal to a multi-channel generator for jointly processing the dynamic range reduced raw multi-channel audio signal, and jointly enhancing the two or more channels of the dynamic range reduced raw multi-channel audio signal by the multi-channel generator in a dynamic range reduction domain.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Patent Application No. 63 / 018,282 (Reference: D20011USP1), filed April 30, 2020, and European Patent Application No. 20175654.1 (Reference: D20011EP), filed May 20, 2020. [Technical field] The present disclosure relates generally to methods for generating an enhanced multi-channel audio signal from an audio bitstream containing the multi-channel audio signal in a reduced dynamic range domain, and more particularly to jointly enhancing two or more channels of a reduced dynamic range raw multi-channel audio signal using a multi-channel generator in a generative adversarial network setting. Although some embodiments are described herein with particular reference to that disclosure, it will be understood that the present disclosure is not limited to such field of use and is applicable in a broader context. [Background technology]

[0002] Discussion of background art throughout this disclosure should in no way be construed as an admission that such techniques are widely known or form part of the common general knowledge in the art. Audio recording systems are used to encode audio signals into encoded signals suitable for transmission or storage, and then receive or acquire the coded signals and decode them to obtain a version of the original audio signal for playback. Low bit-rate audio coding is a perceptual audio compression technique that can reduce bandwidth and storage requirements. Examples of perceptual audio coding systems include AC-3, Advanced Audio Coding (AAC), and the more recently standardized AC-4 audio coding system standardized by ETSI and included in ATSC 3.0.

[0003] However, low bitrate audio coding inevitably results in coding artifacts: audio coded at low bitrates is particularly sensitive to the details of the audio signal, and the quality of the audio signal can be reduced due to noise introduced by quantization and coding.

[0004] To date, several approaches have been developed to enhance the quality of single-channel and multi-channel audio coded at low bit rates. Multi-channel approaches include, for example, beamforming and multi-channel Wiener filtering. Due to the use of spatial information, multi-channel approaches can generally perform better than single-channel approaches.

[0005] In their publication "Methods for Low Bitrate Coding Enhancement Part II: Spatial Enhancement," AES International Conference on Automotive Audio, 2017, C. Uhle et al. review perceptual coding techniques and discuss the nature and origin of common spatial coding artifacts. Furthermore, they propose a set of dedicated algorithms designed to mitigate common types of artifacts. From this set, a low bitrate coding enhancement (LBCE) engine can be built that adapts specifically to the encoder configuration underlying the coded audio material.

[0006] Companding is a coding tool in the AC-4 coding system that improves the perceptual coding of speech and high-density transient events (e.g., applause). The benefits of companding include reducing the short-term dynamics of the input signal, reducing the bitrate requirements at the encoder side while ensuring adequate temporal noise shaping at the decoder side.

[0007] In recent years, deep learning approaches have become increasingly attractive in various application areas, including speech enhancement. In this context, D. Michelsanti and Z.-H. Tan, in their publication "Conditional Generative Adversarial Networks for Speech Enhancement and Noise-Robust Speaker Verification" published at INTERSPEECH 2017, state that conditional generative adversarial networks (GANs) methods outperform classical short-time spectral amplitude minimum mean squared error speech enhancement algorithms and are comparable to deep neural network-based approaches to speech enhancement.

[0008] N. Tawara, T. Kobayash, and T. Ogawa describe a multi-channel time-domain convolutional denoising autoencoder (TCDAE) and evaluate its speech enhancement performance in a multi-channel configuration in their publication "Multi-channel Speech Enhancement Using Time-Domain Convolutional Denoising Autoencoder" published at INTERSPEECH 2019. TCDAE aims to directly map noisy speech signals to clean signals in the time domain and learn spatial information end-to-end.

[0009] In "Audio Codec Enhancement with Generative Adversarial Networks," A. Biswas et al. describe a GAN-based coded audio enhancer for effectively restoring signals contaminated by coding noise. The concept is codec-independent, as the method operates directly on the decoded waveform.

[0010] Generally, most recent research is based on deep convolutional GANs. While GANs are increasingly being used in speech and audio-related applications, their application to multichannel audio remains limited. Furthermore, most deep learning approaches to date have been related to speech denoising. Restoring audio from coding noise is a challenging problem. Intuitively, reducing coding artifacts and denoising can be considered closely related. However, removing coding artifacts / noises that are highly correlated with the desired sound appears more complex than removing other noise types (denoising applications) that are often less correlated. The characteristics of coding artifacts vary depending on the codec, the coding tool employed, and the selected bitrate. Therefore, it is desirable to combine the advantages of generators trained in a GAN setting with those of companding techniques to significantly reduce coding artifacts in multichannel audio signals and provide users with the benefits of high-quality enhanced audio. Summary of the Invention

[0011] According to a first aspect of the present invention, there is provided a method for generating an enhanced multi-channel audio signal from an audio bitstream containing a multi-channel audio signal in a reduced dynamic range domain. The method may include receiving the audio bitstream. The method may further include core-decoding the audio bitstream to obtain a raw multi-channel audio signal with reduced dynamic range based on the received audio bitstream, where the reduced dynamic range raw multi-channel audio signal includes two or more channels. The method may further include inputting the reduced dynamic range raw multi-channel audio signal to a multi-channel generator for jointly processing the reduced dynamic range raw multi-channel audio signal. The method may further include jointly enhancing the two or more channels of the reduced dynamic range raw multi-channel audio signal by the multi-channel generator in the reduced dynamic range domain. The method may further include obtaining an enhanced reduced dynamic range multi-channel audio signal as an output from the multi-channel generator for subsequent dynamic range expansion, where the enhanced reduced dynamic range multi-channel audio signal has two or more channels.

[0012] The method configured as described above can improve the quality of multi-channel audio signals in reduced dynamic range regions using a multi-channel generator trained in a generative adversarial network setting, where joint restoration and spatial enhancement of coded audio can be performed.

[0013] In some embodiments, the method may further include, after the step of core decoding the audio bitstream, performing a dynamic range reduction operation to obtain a dynamic range reduced raw multi-channel audio signal.

[0014] In some embodiments, the audio bitstream may be in AC-4 format.

[0015] In some embodiments, the method may further include expanding the enhanced dynamic range reduced multi-channel audio signal into an extended dynamic range region by performing an expansion operation on two or more channels.

[0016] In some embodiments, the expansion operation may be a companding operation based on the p-norm of the spectral magnitudes to calculate the respective gain values.

[0017] In some embodiments, the received audio bitstream includes metadata, and receiving the audio bitstream may include demultiplexing the received audio bitstream.

[0018] In some embodiments, the step of co-emphasizing two or more channels of the dynamic range reduced raw multi-channel audio signal by the multi-channel generator may be based on metadata.

[0019] In some embodiments, the metadata may include one or more items of companding control data.

[0020] In some embodiments, the companding control data may include information regarding a companding mode, among one or more companding modes, that was used in encoding the multi-channel audio signal.

[0021] In some embodiments, the companding modes may include a companding on companding mode, a companding off companding mode, and an average companding companding mode.

[0022] In some embodiments, the step of co-emphasizing two or more channels of the dynamic range reduced raw multi-channel audio signal by the multi-channel generator may depend on a companding mode indicated by the companding control data.

[0023] In some embodiments, when the companding mode is companding off, no co-enhancement by the multi-channel generator may be performed.

[0024] In some embodiments, the multi-channel generator may be a generator trained in the reduced dynamic range domain in a generative adversarial network setting.

[0025] In some embodiments, the multi-channel generator includes an encoder stage and a decoder stage arranged with mirror symmetry, wherein the encoder stage and the decoder stage each include L layers with N filters in each layer, where L is a natural number greater than 1 and N is a natural number greater than 1, the N filters in each layer of the encoder stage and the decoder stage are of the same size, and each of the N filters in the encoder stage and the decoder stage may operate with a stride greater than 1.

[0026] In some embodiments, nonlinear operations including one or more of ReLU, PReLU, LReLU, eLU, and SeL may be performed in at least one layer of the encoder stage and at least one layer of the decoder stage.

[0027] In some embodiments, the multi-channel generator may further include a non-strided convolutional layer as an input layer preceding the encoder stage.

[0028] In some embodiments, the multi-channel generator may further include a non-stride permuted convolutional layer as an output layer subsequent to the decoder stage.

[0029] In some embodiments, there may be one or more skip connections between each homogeneous layer of the multi-channel generator.

[0030] In some embodiments, the multi-channel generator may include a stage between the encoder stage and the decoder stage for modifying the multi-channel audio in a reduced dynamic range domain based at least on the reduced dynamic range coded multi-channel audio feature space.

[0031] In some embodiments, a random noise vector z may be used within the reduced dynamic range coded multi-channel audio feature space to modify the multi-channel audio in the reduced dynamic range domain.

[0032] In some embodiments, the use of the random noise vector z may be conditional on the bitrate of the audio bitstream and the number of channels of the multi-channel audio signal.

[0033] In some embodiments, the method further comprises the following steps to be performed before the step of receiving the audio bitstream: inputting a dynamic range reduced raw multi-channel audio training signal into a multi-channel generator, the dynamic range reduced raw multi-channel audio training signal comprising two or more channels; jointly generating, by a multi-channel generator, an enhanced reduced dynamic range multi-channel audio training signal based on the reduced dynamic range raw multi-channel audio training signal; inputting, one at a time, each of the two or more channels of the enhanced dynamic range reduced multi-channel audio training signal and a corresponding channel of the original dynamic range reduced multi-channel audio signal from which the dynamic range reduced raw multi-channel audio training signal is derived into one single-channel discriminator of a group of one or more single-channel discriminators; further inputting the enhanced dynamic range reduced multi-channel audio training signal and the corresponding original dynamic range reduced multi-channel audio signal into a multi-channel discriminator one at a time; determining, by a single-channel discriminator and a multi-channel discriminator, whether the input reduced-dynamic-range multi-channel audio signal is an enhanced reduced-dynamic-range multi-channel audio training signal or an original reduced-dynamic-range multi-channel audio signal; tuning the parameters of the multi-channel generator until the single-channel discriminator and the multi-channel discriminator are no longer able to distinguish the enhanced reduced dynamic range multi-channel audio training signal from the original reduced dynamic range multi-channel audio signal.

[0034] In some embodiments, the set of one or more single-channel discriminators is selected based on the type of the original reduced-dynamic-range multi-channel audio signal, which may include a stereo type multi-channel audio signal, a 5.1 type multi-channel audio signal, a 7.1 type multi-channel audio signal, or a 9.1 type multi-channel audio signal.

[0035] In some embodiments, additionally, the random noise vector z is an input to the multi-channel generator, and the step of co-generating the enhanced dynamic range reduced multi-channel audio training signal by the multi-channel generator may additionally be based on the random noise vector z.

[0036] In some embodiments, the additional metadata is an input to the multi-channel generator, and the step of co-generating the enhanced dynamic range reduced multi-channel audio training signal by the multi-channel generator may be additionally based on the metadata.

[0037] In some embodiments, the metadata may include one or more items of companding control data.

[0038] In some embodiments, The companding control data may include information regarding the companding mode, among one or more companding modes, that was used in encoding the original multi-channel audio signal.

[0039] In some embodiments, the companding modes may include a companding on companding mode, a companding off companding mode, and an average companding companding mode.

[0040] In some embodiments, the step of jointly generating the enhanced dynamic range reduced multi-channel audio training signals by the multi-channel generator may depend on a companding mode indicated by the companding control data.

[0041] In some embodiments, when the companding mode is companding off, no co-enhancement is performed by the multi-channel generator.

[0042] According to a second aspect of the present invention, there is provided a method for training a multi-channel generator in a dynamic range reduced domain in a setting of a multi-channel generator, a group of one or more single-channel discriminators, and a generative adversarial network having the multi-channel discriminator. The method may include inputting a dynamic range reduced raw multi-channel audio training signal to the multi-channel generator, where the dynamic range reduced raw multi-channel audio training signal includes two or more channels. The method may include jointly generating, by the multi-channel generator, an enhanced dynamic range reduced multi-channel audio training signal based on the dynamic range reduced raw multi-channel audio training signal. The method may include inputting, one at a time, each channel of the two or more channels of the enhanced dynamic range reduced multi-channel audio training signal and a corresponding channel of an original dynamic range reduced multi-channel audio signal from which the dynamic range reduced raw multi-channel audio training signal is derived to one single-channel discriminator of the group of one or more single-channel discriminators. The method may further include inputting the enhanced dynamic range reduced multi-channel audio training signal and the corresponding original dynamic range reduced multi-channel audio signal one at a time into a multi-channel discriminator. The method may include determining, by the single-channel discriminator and the multi-channel discriminator, whether the input dynamic range reduced multi-channel audio signal is the enhanced dynamic range reduced multi-channel audio training signal or the original dynamic range reduced multi-channel audio signal; and tuning parameters of the multi-channel generator until the single-channel discriminator and the multi-channel discriminator can no longer distinguish the enhanced dynamic range reduced multi-channel audio training signal from the original dynamic range reduced multi-channel audio signal.

[0043] In some embodiments, the set of one or more single-channel discriminators is selected based on the type of the original reduced-dynamic-range multi-channel audio signal, which may include a stereo type multi-channel audio signal, a 5.1 type multi-channel audio signal, a 7.1 type multi-channel audio signal, or a 9.1 type multi-channel audio signal.

[0044] In some embodiments, additionally, the random noise vector z is an input to a multi-channel generator, and the step of jointly generating the enhanced dynamic range reduced multi-channel audio training signal by the multi-channel generator may additionally be based on the random noise vector z.

[0045] In some embodiments, the additional metadata is an input to the multi-channel generator, and the step of co-generating the enhanced dynamic range reduced multi-channel audio training signal by the multi-channel generator may be additionally based on the metadata.

[0046] In some embodiments, the metadata may include one or more items of companding control data.

[0047] In some embodiments, the companding control data may include information regarding the companding mode, among one or more companding modes, that was used in encoding the original multi-channel audio signal.

[0048] In some embodiments, the companding modes may include a companding on companding mode, a companding off companding mode, and an average companding companding mode.

[0049] In some embodiments, the step of jointly generating, by the multi-channel generator, the enhanced dynamic range reduced multi-channel audio training signals may depend on a companding mode indicated by the companding control data.

[0050] In some embodiments, when the companding mode is companding off, no co-enhancement is performed by the multi-channel generator.

[0051] According to a third aspect of the present invention, there is provided an apparatus for generating an enhanced multi-channel audio signal from an audio bitstream containing a multi-channel audio signal in a reduced dynamic range domain. The apparatus may include a receiver for receiving the audio bitstream. The apparatus may further include a core decoder for core-decoding the audio bitstream to obtain a raw multi-channel audio signal with reduced dynamic range based on the received audio bitstream, the reduced dynamic range raw multi-channel audio signal including two or more channels. The apparatus may further include a multi-channel generator for co-emphasizing two or more channels of the reduced dynamic range raw multi-channel audio signal in the reduced dynamic range domain to obtain an enhanced reduced dynamic range multi-channel audio signal, the enhanced reduced dynamic range multi-channel audio signal having two or more channels.

[0052] In some embodiments, the apparatus may further include a demultiplexer that demultiplexes the received audio bitstream, the received audio bitstream including the metadata.

[0053] In some embodiments, the metadata may include one or more items of companding control data.

[0054] In some embodiments, the companding control data may include information regarding a companding mode, among one or more companding modes, that was used in encoding the multi-channel audio signal.

[0055] In some embodiments, the companding modes may include a companding on companding mode, a companding off companding mode, and an average companding companding mode.

[0056] In some embodiments, the multi-channel generator may be configured to co-emphasize the two or more channels of the dynamic range reduced raw multi-channel audio signal in a dynamic range reduced region depending on the companding mode indicated by the companding control data.

[0057] In some embodiments, when the companding mode is companding off, the multi-channel generator may be configured not to perform co-enhancement.

[0058] In some embodiments, the apparatus may further include an expansion unit configured to perform an expansion operation on two or more channels to expand the enhanced dynamic range reduced multi-channel audio signal to an extended dynamic range region.

[0059] In some embodiments, the apparatus may further include a dynamic range reduction unit configured to perform a domain range reduction operation after core decoding the audio bitstream to obtain a dynamic range reduced raw multi-channel audio signal.

[0060] According to a fourth aspect of the present invention, there is provided a computer program product comprising a computer-readable storage medium having instructions adapted, when executed by a device having processing capability, to cause the device to perform a method for generating an enhanced multi-channel audio signal from an audio bitstream comprising the multi-channel audio signal in a reduced dynamic range domain.

[0061] According to a fifth aspect of the present invention, there is provided a computer program product including a computer-readable storage medium having instructions adapted, when executed by a device having processing capability, to cause the device to perform a method for training a multi-channel generator in a reduced dynamic range domain in a setting of a multi-channel generator, a group of one or more single-channel discriminators and a generative adversarial network having the multi-channel discriminator.

[0062] According to a sixth aspect of the present invention, there is provided a system of an apparatus for generating an enhanced multi-channel audio signal from an audio bitstream in a reduced dynamic range domain and a generative adversarial network, the generative adversarial network comprising a multi-channel generator, a group of one or more single-channel discriminators, and a multi-channel discriminator, the system being configured to perform a method for generating an enhanced multi-channel audio signal from an audio bitstream comprising the multi-channel audio signal in a reduced dynamic range domain.

[0063] According to a seventh aspect of the present invention, there is provided a system of apparatus for applying dynamic range reduction to an input multi-channel audio signal and encoding the reduced dynamic range multi-channel audio signal in an audio bitstream, and apparatus for generating an enhanced multi-channel audio signal in a reduced dynamic range domain from an audio bitstream comprising the multi-channel audio signal. [Brief explanation of the drawings]

[0064] Exemplary embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which: [Figure 1] FIG. 1 is a flow diagram illustrating an example of a method for generating an enhanced multi-channel audio signal from an audio bitstream containing the multi-channel audio signal in a reduced dynamic range domain. [Figure 2] FIG. 2 illustrates an example of a generative adversarial network setup including a multi-channel discriminator for training a multi-channel generator in the reduced dynamic range domain. [Figure 3] FIG. 3 illustrates an example of a generative adversarial network setup including a single-channel discriminator for training a multi-channel generator in the reduced dynamic range domain. [Figure 4] FIG. 4 illustrates an example of a generative adversarial network setup for training a multi-channel discriminator in the reduced dynamic range domain. [Figure 5] FIG. 5 illustrates a further example of a generative adversarial network setup for training a multi-channel discriminator in reduced dynamic range domains. [Figure 6] FIG. 6 illustrates an example of a generative adversarial network setup for training a single-channel discriminator in the reduced dynamic range domain. [Figure 7] FIG. 7 illustrates a further example of a generative adversarial network setup for training a single-channel discriminator in the reduced dynamic range domain. [Figure 8] FIG. 8 is a diagram showing an example of a multi-channel generator architecture. DETAILED DESCRIPTION OF THE INVENTION

[0065] Companding The companding technique described in U.S. Patent No. 9,947,335 achieves temporal noise shaping of quantization noise in audio codecs using a companding algorithm implemented in the QMF (quadrature mirror filter) domain and is incorporated herein by reference in its entirety. Generally, companding is a parametric coding tool operating in the QMF domain that can be used to control the temporal distribution of quantization noise (e.g., quantization noise introduced in the MDCT (modified discrete cosine transform) domain). As such, the companding technique includes a QMF analysis step, followed by application of the actual companding operation / algorithm, and a QMF synthesis step.

[0066] Companding can be seen as an example of a technique for reducing the dynamic range of a signal and, equivalently, removing the temporal envelope from the signal. The methods, apparatus, and systems described herein aim to improve the quality of multi-channel audio signals in reduced dynamic range domains. Such improvement can be particularly beneficial for applications using companding techniques. Accordingly, some embodiments relate to companding, and in particular to improving the quality of multi-channel audio signals in the QMF domain as a reduced dynamic range domain.

[0067] overview Referring to FIG. 1, a method for generating an enhanced multi-channel audio signal from an audio bitstream including a multi-channel audio signal in a reduced dynamic range domain is illustrated. In a first step 101, an audio bitstream including a multi-channel audio signal is received. The codec of the audio bitstream may be any codec used in lossy audio compression, such as, but not limited to, AAC (Advanced Audio Coding), AC-3, HE-AAC, USAC, or AC-4. In one embodiment, the audio bitstream is in AC-4 format. In a second step 102, the audio bitstream is core-decoded to obtain a reduced dynamic range raw multi-channel audio signal based on the received audio bitstream, where the reduced dynamic range raw multi-channel audio signal includes multiple channels. For example, the audio bitstream may be core-decoded to obtain a reduced dynamic range raw multi-channel audio signal including two or more channels based on the audio bitstream including the multi-channel audio signal. The term core-decoding, as used herein, generally refers to audio decoded after waveform coding in the MDCT domain. In AC-4, the core codec is known as ASF (Audio Spectral Frontend) or SSF (Speech Spectral Frontend).

[0068] As used herein, the term "raw" in relation to a dynamic range reduced multi-channel audio signal refers to a dynamic range reduced multi-channel audio signal before co-enhancement by a multi-channel generator (hereinafter also referred to simply as a generator) described below, i.e., an unenhanced dynamic range reduced multi-channel audio signal.

[0069] The reduced dynamic range multi-channel audio signal can be encoded into an audio bitstream.

[0070] Alternatively, the dynamic range reduction can be performed before or after core decoding of the audio bitstream. Thus, in one embodiment, step 102 can further include performing a dynamic range reduction operation, such as companding, after core decoding of the audio bitstream.

[0071] In step 103, the dynamic range-reduced raw multi-channel audio signals are input to a multi-channel generator for jointly processing the dynamic range-reduced raw multi-channel audio signals. Here, "jointly" refers to processing / enhancement or other operations performed simultaneously on two or more channels of the multi-channel audio signals. In this case, jointly refers to the simultaneous enhancement of two or more channels of the dynamic range-reduced raw multi-channel audio signals by the multi-channel generator. In other words, two or more channels of the dynamic range-reduced raw multi-channel audio signals are input to the multi-channel generator simultaneously. In step 104, two or more channels of the dynamic range-reduced raw multi-channel audio signals are then jointly enhanced by the multi-channel generator in a dynamic range reduction domain, as will be described in more detail below. The enhancement process performed by the multi-channel generator is intended to improve the quality of the dynamic range-reduced raw multi-channel audio signals by reducing coding artifacts and quantization noise. In step 105, an enhanced dynamic range-reduced multi-channel audio signal for subsequent dynamic range expansion is obtained as an output from the multi-channel generator, where the enhanced dynamic range-reduced multi-channel audio signal has two or more channels.

[0072] In one embodiment, the method may further include extending the enhanced dynamic range reduced multi-channel audio signal to an extended dynamic range region by performing an extension operation on two or more channels, which may be a companding operation based on the p-norm of the spectral magnitudes to calculate respective gain values.

[0073] In general companding (compression / expansion), compression and expansion gain values ​​are calculated and applied to a filter bank. To solve potential problems associated with applying individual gain values, short prototype filters can be applied. Referring to the companding operation described above, the enhanced dynamic range reduced multi-channel audio signal output by the multi-channel generator is analyzed by a filter bank, and wideband gains can be applied directly to two or more channels of the enhanced dynamic range reduced multi-channel audio signal in the frequency domain. Depending on the shape of the applied prototype filter, the corresponding effect in the time domain is to naturally smooth the gain application. The modified frequency signals are then transformed back to the time domain by respective synthesis filter banks. In this regard, it should be noted that many QMF tools, including, but not limited to, one or more of bandwidth expansion and parametric upmixing, can be subsequently performed before converting from QMF back to the time domain. Analyzing a signal with a filter bank provides access to its spectral content, allowing the calculation of gains that preferentially boost high-frequency contributions (or boost contributions with weaker spectral content), resulting in gain values ​​that are not dominated by the strongest components of the signal, thus solving problems associated with mixed audio sources. In this regard, gain values ​​can be calculated using the p-norm of the spectral magnitude, where p is typically less than 2, which has been found to be more effective at shaping quantization noise than energy-based methods, such as when p=2.

[0074] The above method can be implemented in any decoder. When the above method is applied in combination with companding, the above method can be implemented in an AC-4 decoder.

[0075] Alternatively or additionally, the above method may be performed by a system of an apparatus for generating an enhanced multi-channel audio signal from an audio bitstream in a reduced dynamic range domain and a generative adversarial network, the generative adversarial network having a multi-channel generator, a group of one or more single-channel discriminators, and a multi-channel discriminator.

[0076] The device may be a decoder.

[0077] The above method may also be performed by an apparatus for generating an enhanced multi-channel audio signal from an audio bitstream containing a multi-channel audio signal in a reduced dynamic range domain. The apparatus may include a receiver for receiving the audio bitstream. The apparatus may further include a core decoder for core-decoding the audio bitstream to obtain a dynamic range reduced raw multi-channel audio signal based on the received audio bitstream, where the dynamic range reduced raw multi-channel audio signal includes two or more channels. The apparatus may further include a multi-channel generator for co-emphasizing two or more channels of the dynamic range reduced raw multi-channel audio signal in the reduced dynamic range domain to obtain an enhanced dynamic range reduced multi-channel audio signal, where the enhanced dynamic range reduced multi-channel audio signal has two or more channels. In one embodiment, the apparatus may further include a demultiplexer. In one embodiment, the apparatus may further include an expansion unit. In one embodiment, the apparatus may further include a dynamic range reduction unit.

[0078] Alternatively or additionally, the apparatus may be part of a system of apparatuses for applying dynamic range reduction to an input multi-channel audio signal and encoding the dynamic range reduced multi-channel audio signal in an audio bitstream, and for generating an enhanced multi-channel audio signal in a reduced dynamic range domain from an audio bitstream comprising the multi-channel audio signal. Alternatively or additionally, the above methods may be implemented by a respective computer program product comprising a computer-readable storage medium having instructions adapted, when executed by a device having processing capability, to cause the device to perform the method for generating an enhanced multi-channel audio signal in a reduced dynamic range domain from an audio bitstream comprising the multi-channel audio signal.

[0079] Metadata Alternatively or additionally, the above method may include metadata. In one embodiment, the received audio bitstream includes metadata, and step 101 further includes demultiplexing the received audio bitstream. In one embodiment, in step 104, as described above, co-emphasizing two or more channels of the dynamic range-reduced raw multi-channel audio signal by a multi-channel generator may be based on the metadata. As described above, the methods, apparatus, and systems described herein may be beneficial when applied in combination with companding. In one embodiment, the metadata may therefore include one or more items of companding control data. While companding may generally benefit speech and transient signals, individually modifying each QMF time slot with a gain value may cause discontinuities during encoding, degrading the quality of some stationary signals, and in a companding decoder, causing discontinuities in the envelope of shaped noise, which may lead to audible artifacts. The respective companding control data may selectively switch companding on for transient signals, off for stationary signals, or apply average companding as needed. Average companding, as used herein, refers to applying a constant gain to an audio frame similar to the gain of an adjacent active companding frame. The companding control data can be detected during encoding and transmitted to a decoder via the audio bitstream. In one embodiment, the companding control data can therefore include information regarding a companding mode, among one or more companding modes, that was used to encode the multi-channel audio signal. In one embodiment, the companding modes can include a companding-on companding mode, a companding-off companding mode, and an average companding companding mode.In one embodiment, as described above, in step 104, co-emphasizing two or more channels of the dynamic range reduced raw multi-channel audio signal by the multi-channel generator may depend on the companding mode indicated by the companding control data. In one embodiment, if the companding mode is companding off, no co-emphasis by the multi-channel generator is performed. In embodiments, reference is made to metadata including one or more items of companding control data, but this is not intended to be limiting. Alternatively or additionally, co-emphasizing two or more channels of the dynamic range reduced raw multi-channel audio signal by the multi-channel generator may be based on parameters derived from other metadata or a combination of other metadata.

[0080] Generative Adversarial Network setting In step 104, as described above, the multi-channel generator jointly emphasizes two or more channels of the reduced-dynamic-range raw multi-channel audio signal in a reduced-dynamic-range domain, such that this emphasis reduces coding artifacts and improves the quality of the reduced-dynamic-range multi-channel audio signal compared to the original, uncoded, reduced-dynamic-range multi-channel audio signal, the quality of which has already been enhanced before extending the dynamic range of the two or more channels of the reduced-dynamic-range multi-channel audio signal.

[0081] In one embodiment, The multi-channel generator may be a generator trained in a dynamic range reduction domain in a generative adversarial network setting (GAN setting). For example, the dynamic range reduction domain may be an AC-4 companding domain. In some cases (e.g., AC-4 companding), reducing the dynamic range may be equivalent to removing (or suppressing) the temporal envelope of the signal. Therefore, it can be said that the multi-channel generator may be a generator trained in a domain after removing the temporal envelope from the signal. Furthermore, while the following describes a GAN setting, this should not be understood in a limiting sense, and other generative models are also contemplated and within the scope of the present disclosure.

[0082] The GAN setting includes a multi-channel generator G, a set of one or more single-channel discriminators Dk, and a multi-channel discriminator Dj, each trained by an iterative process. During training in the generative adversarial network setting, the multi-channel generator G generates a dynamic range reduced raw multi-channel audio training signal x, which contains two or more channels and is derived from a corresponding original dynamic range reduced multi-channel audio signal x. [Outside 1] TIFF2026012688000002.tif6170 (hereinafter x ~ ) (core encoding and core decoding), to jointly generate an enhanced dynamic range reduced multi-channel audio training signal x* comprising two or more channels. The dynamic range reduction can be performed by applying a companding operation to two or more channels of the multi-channel audio signal. The companding operation may be the companding operation specified in the AC-4 codec and performed in an AC-4 encoder.

[0083] In one embodiment, the random noise vector z can be input to the multi-channel generator in addition to the dynamic range reduced raw multi-channel audio training signal x*, and the joint generation of the enhanced dynamic range reduced multi-channel audio training signal x* by the multi-channel generator can be additionally based on the random noise vector z. In one embodiment, the additional input of the random noise vector (z) can be conditioned on the bit rate of the audio bitstream containing the original multi-channel audio signal from which the dynamic range reduced multi-channel audio training signal was derived and / or on the dynamic range reduced multi-channel audio training signal. For example, for a stereo signal, the random noise vector z can be used at 36 kbit / s or less. For applause, the random noise vector z can be used for all bit rates. However, the random noise vector can also be set to z=0. If the bit rate is not too low, good results can be achieved in reducing coding artifacts if the random noise vector is set to z=0. Alternatively, training can be performed without inputting the random noise vector z. Alternatively or additionally, in one embodiment, metadata may be input to the multi-channel generator, and the enhanced dynamic range reduced multi-channel audio training signal x* may be jointly generated and further based on the metadata. During training, the joint generation of the enhanced dynamic range reduced multi-channel audio training signal x* may thus be conditioned based on the metadata. In one embodiment, the metadata may include one or more items of companding control data. In one embodiment, the companding control data may include information regarding a companding mode, among one or more companding modes, used to encode the audio data. In one embodiment, the companding modes may include a companding-on companding mode, a companding-off companding mode, and an average companding companding mode.In one embodiment, the joint generation of the enhanced dynamic range reduced multi-channel audio training signal x* by the multi-channel generator may depend on a companding mode indicated by companding control data. In this case, during training, the multi-channel generator may condition the companding mode. In one embodiment, if the companding mode is companding off, this may indicate that the input raw multi-channel audio training signal is not dynamic range reduced and joint enhancement by the multi-channel generator may not be performed. As mentioned above, the companding control data may be detected during encoding of the multi-channel audio signal to selectively apply companding: companding on for transient signals, companding off for stationary signals, and applying average companding as needed.

[0084] During training, the multi-channel generator attempts to output an enhanced dynamic range-reduced multi-channel audio training signal x* that is indistinguishable from the corresponding original dynamic range-reduced multi-channel audio signal x. In a first step, a single-channel discriminator Dk of a group of one or more single-channel discriminators is fed, one at a time, each of two or more channels of the generated enhanced dynamic range-reduced multi-channel audio training signal x* and the corresponding channel of the original dynamic range-reduced multi-channel audio signal x from which the dynamic range-reduced raw multi-channel audio training signal is derived, and determines in a true / false manner whether the input data is a channel of the generated enhanced dynamic range-reduced multi-channel audio training signal x* or the corresponding channel of the original dynamic range-reduced multi-channel audio signal x. Here, the single-channel discriminator Dk attempts to distinguish each channel of the original dynamic range-reduced multi-channel audio signal x from the corresponding channel of the enhanced dynamic range-reduced multi-channel audio training signal x*. During the iterative process, the multi-channel generator adjusts its parameters to generate enhanced dynamic range reduced multi-channel audio training signals x* that are increasingly preferable compared to the original dynamic range reduced multi-channel audio signal x, and the single-channel discriminator Dk learns to make better decisions between two or more channels of the enhanced dynamic range reduced multi-channel audio training signal x* and the corresponding channels of the original dynamic range reduced multi-channel audio signal x.

[0085] It should be noted that the step of determining in a true / false manner by the single-channel discriminator Dk whether the input data is a channel of the generated enhanced dynamic range reduced multi-channel audio training signal x* or a corresponding channel of the original dynamic range reduced multi-channel audio signal x can be performed by the same single-channel discriminator Dk for each channel of the generated enhanced dynamic range reduced multi-channel audio training signal x*. Alternatively or additionally, the step of determining in a true / false manner by the single-channel discriminator Dk whether the input data is a channel of the generated enhanced dynamic range reduced multi-channel audio training signal x* or a corresponding channel of the original dynamic range reduced multi-channel audio signal x can be performed by a group of single-channel discriminators Dk, where each channel of the generated enhanced dynamic range reduced multi-channel audio training signal x* and each corresponding channel of the original dynamic range reduced multi-channel audio signal x are input to an individual single-channel discriminator Dk trained on that channel. In one embodiment, the set of one or more single-channel discriminators Dk may be selected based on the type of the original dynamic range reduced multi-channel audio signal, which may include a stereo type multi-channel audio signal, a 5.1 type multi-channel audio signal, a 7.1 type multi-channel audio signal, or a 9.1 type multi-channel audio signal.

[0086] In the second step, the multi-channel discriminator Dj is supplied with the generated enhanced dynamic range reduced multi-channel audio training signal x* and the corresponding original dynamic range reduced multi-channel audio signal x from which the dynamic range reduced raw multi-channel audio training signal is derived, one at a time, and determines in a true / false manner whether the input data is the generated enhanced dynamic range reduced multi-channel audio training signal x* or the corresponding original dynamic range reduced multi-channel audio signal x. Here, the multi-channel discriminator Dj attempts to distinguish the original dynamic range reduced multi-channel audio signal x from the enhanced dynamic range reduced multi-channel audio training signal x*. During the iterative process, the multi-channel generator adjusts its parameters to generate increasingly preferable enhanced dynamic range reduced multi-channel audio training signals x* compared to the original dynamic range reduced multi-channel audio signal x, and the multi-channel discriminator Dj learns to better distinguish between the enhanced dynamic range reduced multi-channel audio training signal x* and the original dynamic range reduced multi-channel audio signal x.

[0087] Note that, in order to train the multi-channel generator in the final step, the single-channel discriminator Dk and the multi-channel discriminator Dj can be trained first. The training and updating of the discriminator can also be performed in the reduced dynamic range domain. The training and updating of the discriminator can include maximizing the probability of assigning a high score to the original reduced dynamic range multi-channel audio signal x and a low score to the enhanced reduced dynamic range multi-channel audio training signal x*. The goal of training the discriminator can be to recognize the original reduced dynamic range multi-channel audio signal x as true while recognizing the enhanced reduced dynamic range multi-channel audio training signal x* (generated data) as false. The parameters of the multi-channel generator can be fixed while the discriminator is trained and updated.

[0088] Training and updating the multi-channel generator can include minimizing the difference between the original dynamic range reduced multi-channel audio signal x and the generated enhanced dynamic range reduced multi-channel audio training signal x*. The goal of training the multi-channel generator is for the single-channel discriminator Dk to recognize each of two or more channels of the generated enhanced dynamic range reduced multi-channel audio training signal x* as true. Furthermore, the multi-channel discriminator Dj recognizes the generated enhanced dynamic range reduced multi-channel audio training signal x* as true.

[0089] We now describe in more detail the training of a multi-channel generator,G,1,in the dynamic range reduced domain in a generative adversarial network (GAN) setting,with reference to the example in Fig. 2. In the example in Fig. 2, the GAN setting includes a multi-channel generator,G,1,and a multi-channel discriminator,Dj,2j,. The training of the multi-channel generator,G,1,may include the following:

[0090] An original multi-channel audio signal containing two or more channels x,12 can be subjected to dynamic range reduction comp,10 to obtain a dynamic range-reduced original multi-channel audio signal containing two or more channels x,9. The dynamic range reduction can be achieved by applying a companding operation, specifically an AC-4 companding operation, to each of the two or more channels, followed by a QMF (Quadrature Mirror Filter) synthesis step. Because the companding operation is performed in the QMF domain, a subsequent QMF synthesis step is required. In addition to core encoding and decoding, the dynamic range-reduced original multi-channel audio signal x,9 can be subjected to a dynamic range-reduced multi-channel audio training signal x,8 before input to the multi-channel generator G,1. The dynamic range-reduced raw multi-channel audio training signal x,8 is then generated. ~,8 and a random noise vector z,11 are input to a multi-channel generator G,1. Then, based on the inputs, the multi-channel generator G,1 jointly generates an enhanced dynamic range reduced multi-channel audio training signal x*,7 in a reduced dynamic range domain. In one embodiment, the input of the random noise vector z may be conditioned on the bit rate of an audio bitstream containing the original multi-channel audio signal from which the dynamic range reduced multi-channel audio training signal was derived and / or the number of channels of the dynamic range reduced multi-channel audio training signal. In one embodiment, the random noise vector z,11 may be set to z=0. Alternatively, training may be performed without inputting the random noise vector z,11. Additionally or alternatively, the multi-channel generator G,1 may be trained using metadata as input of the reduced dynamic range coded multi-channel audio feature space to modify the enhanced dynamic range reduced multi-channel audio training signal x*,7. One at a time, the original reduced-dynamic-range multi-channel audio signal x,9 from which the reduced-dynamic-range raw multi-channel audio training signal x*,8 is derived and the generated enhanced reduced-dynamic-range multi-channel audio training signal x*,7 are input to a multi-channel discriminator Dj,2j. ~,8 can also be input to the multi-channel discriminator Dj,2j each time. The multi-channel discriminator Dj,2j then determines 3j,4j whether the input data is the enhanced dynamic range reduced multi-channel audio training signal x*,7,(false) or the original dynamic range reduced multi-channel audio signal x,9,(true). In the next step, the parameters of the multi-channel generator G,1 are adjusted until the multi-channel discriminator Dj,2j can no longer distinguish the enhanced dynamic range reduced multi-channel audio training signal x*,7 from the original dynamic range reduced multi-channel audio signal x,9. This can be done in an iterative process 5j.

[0091] Now, referring to the example of Figure 3, the training of the multi-channel generator,G,1,in the dynamic range reduction domain in a generative adversarial network (GAN) setting is described in more detail, where the GAN setting includes a multi-channel generator,G,1,and a single-channel discriminator,D,2k,. The training of the multi-channel generator,G,1,may include the following:

[0092] As described above, the dynamic range reduced raw multi-channel audio training signal x*,8 and the enhanced dynamic range reduced multi-channel audio training signal x*,7 can be obtained. One at a time, the dynamic range reduced raw multi-channel audio training signal x ~ Channel k of the original dynamic range reduced multi-channel audio signal x,9 from which x*,8 is derived and the corresponding channel of the generated enhanced dynamic range reduced multi-channel audio training signal x*,7 are input to a single channel discriminator Dk,2k (note that thin lines indicate individual channels and thick lines indicate the multi-channel signal). As additional information, the dynamic range reduced raw multi-channel audio training signal x*,8 is ~The corresponding channel of,x,8,can also be input to the single-channel discriminator,Dk,2k,each time.,The single-channel discriminator,Dk,2k,then determines whether the input data is a channel of,the enhanced dynamic range reduced multi-channel audio training signal,x*,7,(false) or a corresponding channel of the original,dynamic range reduced multi-channel audio signal,x,9,(true)3k,4k.

[0093] In a next step, the parameters of the multi-channel generator G,1 are adjusted until the single-channel discriminator D,2k can no longer distinguish the channels of the enhanced dynamic range reduced multi-channel audio training signal x*,7 from the corresponding channels of the original dynamic range reduced multi-channel audio signal x,9. This can be done in an iterative process 5k. Note that the determining step as described above can be performed for each channel of each enhanced dynamic range reduced multi-channel audio training signal x*,7 and the original dynamic range reduced multi-channel audio signal x,9 by the same single-channel discriminator Dk,2k. Alternatively or additionally, the determining step can be performed for each channel individually by a respective channel-specific single-channel discriminator Dk,2k of a group of one or more single-channel discriminators Dk. The group of one or more single-channel discriminators can be selected based on the type of the original dynamic range reduced multi-channel audio signal, which may include a stereo type multi-channel audio signal, a 5.1 type multi-channel audio signal, a 7.1 type multi-channel audio signal, or a 9.1 type multi-channel audio signal.

[0094] The decisions made by the single-channel discriminator Dk and the multi-channel discriminator Dj can be based on one or more perceptually motivated objective functions according to the following equation (1), where Nc refers to the total number of channels in the multi-channel audio signal:

number

[0095] The subscript LS means the introduction of the least squares method. Furthermore, as can be seen from the first and second terms of equation (1), the core-decoded dynamic range reduced raw multi-channel audio signal x ~ We apply a conditional generative adversarial network setting by inputting as side information to both the single-channel discriminator Dk and the multi-channel discriminator Dj, which allows the discriminator to learn the conditional classification task, i.e., whether the input of the discriminator is the original signal or the enhanced signal based on the given coded signal.

[0096] The introduction of the last term in equation (1) above, which refers to the single-channel discriminator Dk, helps ensure that these frequencies are not disturbed during the iteration process, since they are usually coded with a higher number of bits. The last term is the 1-norm distance scaled by the λ factor. The value of λ can be chosen between 10 and 100, depending on the application and / or the length of the signal input to the multi-channel generator. For example, λ=100 can be chosen.

[0097] Referring now to the examples in Figures 4 and 5, training a multi-channel discriminator Dj,2j in the reduced dynamic range domain in a generative adversarial network setting involves training a reduced dynamic range raw multi-channel audio training signal x ~,8, the enhanced dynamic range reduced multi-channel audio training signal x*,7 and the original dynamic range reduced multi-channel audio signal x,9 can be input to the multi-channel discriminator Dj,2j one at a time 6j, 14j, according to which the training of the multi-channel generator G,1 can follow the same general iterative process 13j as described above for the training of the multi-channel generator G,1, where the parameters of the multi-channel generator G,1 can be fixed but the multi-channel discriminator Dj,2j varies (indicated by the thicker lines around the discriminators in Figures 2 and 3 compared to Figures 4 and 5). The training of the multi-channel discriminator Dj,2j can be described by the following equation (2), where the multi-channel discriminator Dj,2j can determine the enhanced dynamic range reduced multi-channel audio training signal x*,7 as false.

number

[0098] In the above case, the least squares (LS) and conditional generative adversarial network settings are also used to train the core decoded dynamic range reduced raw multi-channel audio training signal x ~ , is applied by inputting as additional information to the multi-channel discriminator Dj. Referring now to the examples in Figures 6 and 7, training a multi-channel discriminator Dk,2k in the reduced dynamic range domain in a generative adversarial network setting involves inputting the reduced dynamic range raw multi-channel audio training signal x ~,8, the channels of the enhanced dynamic range reduced multi-channel audio training signal x*,7 and the channels corresponding to the original dynamic range reduced multi-channel audio signal x,9, along with the corresponding channels of x*,8, can be input one at a time to a single-channel discriminator Dk,2k,6k,14k, which can follow the same general iterative process 13k as described above for training the multi-channel generator G,1, except that the parameters of the multi-channel generator G,1 can be fixed while the single-channel discriminator Dk,2k varies (indicated by the thicker lines around the discriminator in Figures 2 and 3 compared to Figures 6 and 7). The training of the single-channel discriminator Dk,2k can be described by the following equation (3), where the single-channel discriminator Dk,2k can determine the enhanced dynamic range reduced multi-channel audio training signal x*,7 as false.

number

[0099] In the above case, the least squares (LS) and conditional generative adversarial network settings are also used to train the core decoded dynamic range reduced raw multi-channel audio training signal x ~ The corresponding channels of the multi-channel audio signal are applied by inputting them as additional information to the single-channel discriminator Dk. Furthermore, Nc denotes the number of channels of the multi-channel audio signal that the multi-channel generator emphasizes.

[0100] Based on the above training, the single-channel discriminator Dk can be trained to determine only one channel of the enhanced dynamic range reduced multi-channel audio training signal x*,7 as false, or to determine each channel of the enhanced dynamic range reduced multi-channel audio training signal x*,7 as false, where the enhanced dynamic range reduced multi-channel audio training signal may include a stereotype multi-channel audio signal, a 5.1 type multi-channel audio signal, a 7.1 type multi-channel audio signal, or a 9.1 type multi-channel audio signal.

[0101] In general, training using both a single-channel discriminator and a multi-channel discriminator allows for better control over not only individual channels but also the overall spatial impression. Other training methods besides least squares can be used to train the multi-channel generator, similar to the multi-channel discriminator Dj and single-channel discriminator Dk in the dynamic range reduction domain generative adversarial network setting. This disclosure is not limited to a specific training method. Alternatively or additionally, the so-called Wasserstein approach can be used. In this case, Earth Mover Distance (EMD), also known as Wasserstein distance, can be used instead of least squares distance. In general, different training methods result in more stable training of the multi-channel generator and discriminator. However, the type of training method applied does not affect the architecture of the multi-channel generator described below.

[0102] Architecture of the multi-channel generator Although the architecture of the multi-channel generator is generally not limited, in one embodiment, the multi-channel generator may include an encoder stage and a decoder stage. The encoder stage and the decoder stage of the multi-channel generator are fully convolutive. In one embodiment, the decoder stage may mirror the encoder stage, and similar to the decoder stage, each encoder stage may include a number of L layers, with a number of N filters in each layer L. L may be a natural number greater than 1, and N may be a natural number greater than 1. The size of the N filters (also called kernel size) is not limited and can be selected according to the requirements for improving the quality of the dynamic range-reduced raw multi-channel audio signal by the multi-channel generator. However, the filter size may be the same for each of the L layers.

[0103] 8, which shows a schematic example of an architecture of a multi-channel generator, in a first step 15, a dynamic range reduced raw multi-channel audio signal having multiple channels can be input to the multi-channel generator. In one embodiment, this input layer 15 can be a no-stride (e.g., stride=1 means no stride) convolutional layer preceding or prepending the encoder stage.

[0104] The output of a trained non-strided convolutional layer (e.g., input layer 15) can be viewed as a combination of several individual input channels (the exact number depends on the number of filters or kernels in the non-strided convolutional layer). Therefore, the output of such a layer can be viewed as a multi-channel mid-side signal. For example, for a stereo input signal (e.g., a two-channel input signal), where XL and XR are the left and right channels, the mid-side signal M = 0.5 * (XL + XR) and the side-side signal S = 0.5 * (XL - XR). Therefore, when a multi-channel mid-side signal is created, multiple combinations of XL and XR are generated. Training such a system can provide additional hints about the spatial relationship between XL and XR. For example, considering the simple case of side signal S = 0, XL = XR is most likely. Therefore, the pre-strided non-strided convolutional layer can be set up with information about the audio signals (e.g., both the original audio signal and the coded audio signal) and the corresponding spatial relationships (e.g., the spatial relationships between the original audio signal and the coded audio signal). Therefore, if the spatial width is lost due to coding, it can be restored by the proposed system, and coded audio enhancement including spatial enhancement can be performed jointly.

[0105] Exemplary values ​​of N=16 filters and a filter size of 31 in the input layer yield good results, e.g., a minimal amount of coding artifacts. A nonlinear activation can be performed in the input layer, which can be a parametric rectified linear unit (PReLU). The first illustrated encoder layer 16, layer number L=1, can include N=16 filters with a filter size of 31. The second illustrated encoder layer 17, layer number L=2, can include N=32 filters with a filter size of 31. Subsequent layers are omitted for clarity and brevity. The third illustrated encoder layer 18, layer number L=11, can include N=512 filters with a filter size of 31. Thus, the number of filters may increase at each layer. In one embodiment, each filter can operate on two or more channels of the dynamic range reduced multi-channel audio signal input to each encoder layer with a stride > 1. Each filter can operate on two or more channels of the dynamic range reduced multi-channel audio signal input to each encoder layer with a stride of 2, for example. Thus, we can achieve a learnable downsampling of 2x.

[0106] Alternatively, filters can be run at each encoder layer with a stride of 1, followed by a factor of 2 downsampling (as in known signal processing). Alternatively, for example, each filter can operate on two or more channels of a reduced dynamic range multi-channel audio signal input to each encoder layer with a stride of 4. This may halve the overall number of layers in the multi-channel generator.

[0107] In at least one encoder layer and at least one decoder layer of the multi-channel generator, nonlinear operations can be performed in addition to activations. In one embodiment, the nonlinear operations can include one or more of a parametric rectified linear unit (PReLU), a rectified linear unit (ReLU), a leaky rectified linear unit (LReLU), an exponential linear unit (eLU), and a scaled exponential linear unit (SeLU). In the example of Figure 8, the nonlinear operations are based on PReLU.

[0108] As shown schematically in Figure 8, each decoder layer 22, 21, 20 mirrors the encoder layers 16, 17, 18. The number of filters in each layer and the filter width of each layer can be the same in the decoder stage as in the encoder stage, but upsampling of the multi-channel audio signal in the decoder stage can be performed by two alternative approaches. In one embodiment, fractionally stride convolution (also called transposed convolution) operations can be used in layers 20, 21, 22 of the decoder stage. Alternatively, in each layer of the decoder stage, filters can operate on two or more channels of the multi-channel audio signal input to each layer with a stride of 1, after upsampling and interpolation by an upsampling factor of 2, as in conventional signal processing.

[0109] Furthermore, in one embodiment, the multi-channel generator further includes a non-strided (meaning a transposed convolution with stride=1) transposed convolution layer as the output layer 23, followed by a decoder stage. In this example, the output layer 23 may include N=2 filters with a filter size of 31. Note that the number of filters in the output layer may be the same as the number of channels Nc of the multi-channel audio signal that the multi-channel generator enhances. For example, in the case of stereo enhancement, the output layer may be held at Nc=N=2. In the output layer 23, activations may differ from those performed in at least one of the encoder layers and at least one of the decoder layers. The activations may be based on a tanh operation, for example.

[0110] Between the encoder stage and the decoder stage, the reduced-dynamic-range multi-channel audio signal may be modified to generate an enhanced reduced-dynamic-range multi-channel audio signal. In one embodiment, the modification may be based on a reduced-dynamic-range coded multi-channel audio feature space 25 (also referred to as a bottleneck layer). In one embodiment, a random noise vector z may be used in the reduced-dynamic-range coded multi-channel audio feature space 25 to modify two or more channels of the multi-channel audio signal in the reduced-dynamic-range domain. The modification in the reduced-dynamic-range coded multi-channel audio feature space 25 may be performed, for example, by concatenating the random noise vector (z) with vector representations (c) of two or more channels of the multi-channel audio signal output from the last layer of the encoder stage. In one embodiment, the use of the random noise vector z may be conditional on the bitrate of the audio bitstream and / or the number of channels of the multi-channel audio signal. For example, the random noise vector z may be used for stereo signals up to 36 kbit / s and for all bitrates in the case of applause. However, the random noise vector may also be set to z=0. If the bit rate is not too low, the reduction of coding artifacts will yield good results if the random noise vector is set to z=0. Alternatively or additionally, metadata can be entered at this point to modify two or more channels of the multi-channel audio signal. In this case, the generation of the enhanced dynamic range reduced multi-channel audio signal can be conditioned on the given metadata.

[0111] In one embodiment, skip connections 24 may exist between homogeneous layers of the encoder and decoder stages, and between the input layer preceding the encoder stage and the (additional) output layer following the decoder stage. In this case, the aforementioned reduced dynamic range coded multi-channel audio feature space 25 may be bypassed to prevent loss of information. In one embodiment, skip connections 24 may be implemented using one or more concatenations and signal additions. Implementing skip connections 24 may "virtually" double the number of filter outputs.

[0112] Referring to the example of FIG. 8, the architecture of the multi-channel generator can be summarized as follows: 15 / Input layer: Non-stride convolutional layer: Number of filters N = 16, filter size = 31, activation = PreLU 16 / Encoder layer L=1: Number of filters N=16, filter size=31, activation=PreLU 17 / Encoder layer L=2: Number of filters N=32, filter size=31, activation=PreLU . . . 18 / Encoder layer L=11: Number of filters N=512, filter size=31 19 / Encoder layer L=12: Number of filters N=1024, filter size=31 25 / Dynamic range reduction coded multi-channel audio feature space 20 / Decoder layer L=1: Number of filters N=512, filter size=31 . . . 21 / Decoder layer L=10: Number of filters N=32, filter size=31, activated PreLU 22 / Decoder layer L=11: Number of filters N=16, filter size=31, activated PreLU 23 / Output layer: Number of filters N = 2, filter size = 31, activation tanh 24 / Skip Connection

[0113] The above architectures are merely exemplary. Depending on the application, the number of layers in the encoder and decoder stages of the multi-channel generator may be downscaled or upscaled, respectively.

[0114] In general, the above multi-channel generator architecture offers the possibility of one-shot artifact reduction since it does not require performing complex operations such as Wavenet or sampleRNN.

[0115] Furthermore, the multi-channel generator (e.g., composed of non-strided convolutional layers operating jointly on a multi-channel input signal (with corresponding non-strided transposed convolutional layers generating a multi-channel enhanced output signal)) offers reduced complexity due to better utilization of spatial redundancy compared to applying one or more single-channel generators. For example, for a stereo (e.g., two-channel) input signal (having the best audio quality), the stereo generator (e.g., multi-channel generator) may have 0.14% more parameters than a single-channel generator. This increase in parameters translates to 12.1% more complexity compared to a single-channel generator. However, because the stereo input signal is now jointly processed by a stereo (e.g., multi-channel) generator (rather than two separate single-channel generators), a 44% complexity savings is achieved compared to two separate single-channel generators.

[0116] Architecture of the Discriminator The architectures of both single-channel and multi-channel discriminators are not restricted. The architecture of a multi-channel discriminator can follow the same structure as the encoder stage of the multi-channel generator described above. The architecture of a multi-channel discriminator can mirror the encoder stage of the multi-channel generator. Therefore, a multi-channel discriminator can include a number of L layers and multiple N filters. L is a natural number greater than 1, and N is a natural number greater than 1. The size of the N filters is not restricted and can be selected according to the requirements of the discriminator. However, the filter size can be the same for each of the L layers. The nonlinear operation performed in at least one encoder layer of the discriminator can include LReLU. Prepending the encoder stage, the multi-channel discriminator can include an input layer. The input layer can be a non-strided convolutional layer (stride = 1, representing non-strided) as described above. Following the encoder stage, the multi-channel discriminator can include an output layer. The output layer can have N=1 filters with a filter size of 1 (the discriminator makes a single true / false decision). In this case, the filter size of the output layer and the filter size of the encoder layer can be different. Therefore, the output layer can be a 1D convolutional layer that does not downsample the hidden activations. This means that the filters in the output layer may operate with stride 1, while all layers before the encoder stage of the multi-channel discriminator may use stride 2. Alternatively, each filter in the layer before the encoder stage can operate with stride 4. This allows the overall number of layers in the multi-channel discriminator to be halved. The activations in the output layer can be different from the activations in at least one of the encoder layers. The activations are sigmoidal. However, when using a least-squares training approach, sigmoid activations may not be necessary and are therefore optional.

[0117] A multi-channel discriminator can accept two or more channels as input, while a single-channel discriminator can accept only one channel as input. Therefore, the architecture of a single-channel discriminator is slightly different from that of a multi-channel discriminator in that the single-channel discriminator does not include the pre-pended layer mentioned above.

[0118] Generally, a multi-channel discriminator is meant to evaluate the overall presentation quality (e.g., a multi-channel signal) by taking into account the spatial relationships between the channels. If only a single-channel discriminator is employed, the relationships between the channels cannot be considered. Therefore, in some embodiments, both single-channel and multi-channel discriminators are used to jointly evaluate the quality of individual channels and all channels, respectively.

[0119] interpretation Unless otherwise indicated, as will be apparent from the discussion that follows, throughout the disclosure discussion, terms such as "processing," "computing," "determining," "analyzing," and the like are understood to refer to the actions and / or processes of a computer or computing system, or similar electronic computing device, that manipulates and / or transforms data represented as physical quantities, such as electronic quantities, into other data also represented as physical quantities.

[0120] In a similar manner, the term "processor" can refer to a device or part of a device that processes electronic data, e.g., from registers and / or memory, and converts the electronic data into other electronic data that can be stored, e.g., in registers and / or memory. A "computer," "computing machine," or "computing platform" can include one or more processors.

[0121] In one example embodiment, the metrology described herein can be implemented by one or more processors that accept computer-readable (also called machine-readable) code, including a set of instructions that, when executed by the one or more processors, perform at least one of the methods described herein. Any processor capable of executing a series of instructions (sequential or otherwise) that specify actions to be performed is included. Thus, one example is a typical processing system including one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem including main RAM and / or static RAM and / or ROM. A bus subsystem may be included for communication between components. Furthermore, the processing system may be a distributed processing system with processors coupled by a network. If the processing system requires a display, the display may include, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data entry is required, the processing system may also include input devices, such as one or more alphanumeric input units, such as a keyboard, and a pointing control device, such as a mouse. The processing system may also include a storage system, such as a disk drive unit. The processing system in some configurations may also include a sound output device and a network interface device. The memory subsystem thus includes a computer-readable carrier medium that carries computer-readable code (e.g., software) that includes a set of instructions for performing one or more of the methods described herein when executed by one or more processors. Note that when a method includes multiple elements, e.g., multiple steps, no ordering of such elements is implied unless specifically stated. The software may reside on a hard disk, or may reside completely or at least partially in RAM and / or within the processor during execution by the computer system.The memory and processor therefore also constitute a computer-readable carrier medium carrying computer-readable code. Furthermore, the computer-readable carrier medium may be formed into or included in a computer program product.

[0122] In alternative embodiments, one or more processors may operate as standalone devices or may be networked to other processors, and in a network deployment, one or more processors may operate in the capacity of a server or user machine in a server-user network environment, or as a peer machine in a peer-to-peer or distributed network environment. The one or more processors may form a personal computer (PC), tablet PC, personal digital assistant (PDA), mobile phone, web appliance, network router, switch, or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be performed by that machine.

[0123] The term "machine" is also intended to include a collection of machines that individually or jointly execute a set (or sets) of instructions to perform any one or more of the metrology techniques discussed herein.

[0124] Accordingly, one exemplary embodiment of each method described herein is in the form of a computer-readable carrier medium carrying a set of instructions, e.g., a computer program for execution on one or more processors, e.g., one or more processors that are part of a web server arrangement. Accordingly, as will be appreciated by those skilled in the art, exemplary embodiments of the present disclosure may be embodied as a method, an apparatus such as a dedicated device, an apparatus such as a data processing system, or a computer-readable carrier medium, e.g., a computer program product. The computer-readable carrier medium carries computer-readable code including a set of instructions that, when executed on one or more processors, cause the one or more processors to implement the method. Accordingly, aspects of the present disclosure may take the form of a method, an entirely hardware example embodiment, an entirely software example embodiment, or an example embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.

[0125] The software can further be transmitted and received over a network via a network interface device. While the carrier medium is a single medium in one example embodiment, the term "carrier medium" should be construed to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store one or more sets of instructions. The term "carrier medium" also includes a medium capable of storing, encoding, or carrying sets of instructions for execution by one or more processors, causing the one or more processors to perform any one or more of the metrologies disclosed herein. Carrier media can take various forms, including non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus subsystem. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave or infrared data communications. For example, the term "carrier medium" should be taken to include, but not be limited to, solid-state memory, optically and magnetically implemented computer products, media carrying a propagated signal detectable by at least one processor or one or more processors and representing a set of instructions that when executed implements a method, and transmission media within a network carrying a propagated signal detectable by at least one processor of one or more processors and representing a set of instructions.

[0126] It will be understood that the steps of the methods discussed are, in one example embodiment, performed by a suitable processor(s) (e.g., a computer) of a processing system executing (computer-readable) instructions stored in storage. It will also be understood that this disclosure is not limited to any particular implementation or programming technique, but may be implemented using any suitable technique for implementing the functionality described herein. This disclosure is not limited to any particular programming language or operating system.

[0127] References throughout this disclosure to "one embodiment," "some embodiments," or "an example embodiment" or "exemplary embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the disclosure. Thus, the appearances of the phrases "in one embodiment," "some embodiments," or "in an example embodiment" in various places throughout this disclosure do not necessarily all refer to the same example embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this disclosure, in one or more example embodiments.

[0128] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe a common object merely indicates that different instances of similar objects are being referred to and does not imply that the objects so described must be in a given sequence, either temporally, spatially, ranked, or otherwise.

[0129] In the following claims and the description herein, any of the terms "comprising," "comprised of," or "which comprises" are open terms meaning the inclusion of at least the following element / feature, but not the exclusion of others. Therefore, when used in a claim, the terms "comprising" or "having" should not be interpreted as being limited to the means, elements, or steps listed thereafter. For example, the scope of the expression "a device comprising A and B" should not be limited to a device consisting of only elements A and B. Any of the terms "including" or "which includes" used herein are open terms meaning the inclusion of at least the following element / feature, but not the exclusion of others. Therefore, including is synonymous with having and comprising.

[0130] In the foregoing description of exemplary embodiments of the disclosure, it should be recognized that various features of the disclosure may be grouped together in a single exemplary embodiment, drawing, or description for the purpose of streamlining the disclosure and facilitating understanding of one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention requiring more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in fewer than all features of a single disclosed embodiment as described above. Accordingly, the claims that follow this specification are expressly incorporated herein, with each claim standing on its own as a separate exemplary embodiment of the disclosure.

[0131] Furthermore, although some exemplary embodiments described herein do not include other features included in other exemplary embodiments, a combination of features from different exemplary embodiments is meant to fall within the scope of the disclosure and form different exemplary embodiments, as would be understood by one of ordinary skill in the art. For example, in the following claims, any exemplary embodiment described in the claims may be used in any combination.

[0132] In the description provided herein, numerous specific details are set forth. However, it will be understood that example embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.

[0133] Thus, while what is believed to be the best mode of the disclosure has been described, those skilled in the art will recognize that other and further modifications may be made without departing from the spirit of the disclosure, and it is intended that all such modifications and variations be claimed within the scope of the disclosure. For example, the above formulas merely represent procedures that may be used. Functions may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to methods described within the scope of the disclosure.

[0134] Various aspects and implementations of the present disclosure can be understood from the following enumerated exemplary embodiments (EEE) which are not claims.

[0135] EEE1. 1. A method for generating an enhanced multi-channel audio signal from an audio bitstream containing a multi-channel audio signal in a dynamic range reduced domain, the method comprising: (a) receiving an audio bitstream; (b) core-decoding an audio bitstream and obtaining a dynamic range reduced raw multi-channel audio signal based on the received audio bitstream, where the dynamic range reduced raw multi-channel audio signal includes two or more channels; (c) inputting the reduced dynamic range raw multi-channel audio signals into a multi-channel generator for jointly processing the reduced dynamic range raw multi-channel audio signals; (d) jointly enhancing two or more channels of the reduced dynamic range raw multi-channel audio signal by a multi-channel generator in a reduced dynamic range domain; (e) obtaining an enhanced dynamic range reduced multi-channel audio signal as output from a multi-channel generator for subsequent expansion of the dynamic range, the enhanced dynamic range reduced multi-channel audio signal having two or more channels; Includes.

[0136] EEE2. In the method according to EEE1, step (b) further includes, after core decoding the audio bitstream, performing a dynamic range reduction operation to obtain a dynamic range reduced raw multi-channel audio signal.

[0137] EEE3. The method is according to EEE1, and the audio bitstream is in AC-4 format.

[0138] EEE4. A method according to any of EEE1 to EEE3, further comprising the step (f) of expanding the enhanced dynamic range reduced multi-channel audio signal into an extended dynamic range region by performing an expansion operation on two or more channels.

[0139] EEE5. In the method according to EEE4, the expansion operation is a companding operation based on the p-norm of the spectral magnitude to calculate the respective gain values.

[0140] EEE6. A method according to any one of EEE1 to EEE3, wherein the received audio bitstream includes metadata, and step (a) includes demultiplexing the received audio bitstream.

[0141] EEE7. The method according to EEE6, wherein in step (d), the step of co-emphasizing two or more channels of the reduced dynamic range raw multi-channel audio signal by the multi-channel generator is based on metadata.

[0142] EEE8. A method according to EEE7, wherein the metadata includes one or more items of companding control data.

[0143] EEE9. A method according to EEE8, wherein the companding control data comprises information about a companding mode among one or more companding modes that has been used for encoding the multi-channel audio signal.

[0144] EEE10. In the method according to EEE9, the companding modes include a companding-on companding mode, a companding-off companding mode, and an average companding companding mode.

[0145] EEE11. The method according to EEE9 or 10, wherein in step (d), the step of jointly enhancing two or more channels of the dynamic range reduced raw multi-channel audio signal by the multi-channel generator depends on the companding mode indicated by the companding control data.

[0146] EEE12. In the method according to EEE11 which is dependent on EEE10, if the companding mode is companding off, no co-enhancement is performed by the multi-channel generator.

[0147] EEE13. A method according to any of EEE1 to EEE12, wherein the multi-channel generator is a generator trained in a reduced dynamic range domain in a generative adversarial network setting.

[0148] EEE14. A method according to any of EEE1 to 13, wherein the multi-channel generator includes an encoder stage and a decoder stage arranged in a mirror symmetric manner, the encoder stage and the decoder stage each including L layers with N filters in each layer, L being a natural number greater than 1, and N being a natural number greater than 1, the N filters in each layer of the encoder stage and the decoder stage having the same size, and each of the N filters in the encoder stage and the decoder stage operating with a stride greater than 1.

[0149] EEE15. A method according to EEE14, wherein non-linear operations including one or more of ReLU, PReLU, LReLU, eLU and SeL are performed in at least one layer of the encoder stage and at least one layer of the decoder stage.

[0150] EEE16. The method according to EEE14 or 15, wherein the multi-channel generator further comprises a non-strided convolutional layer as an input layer preceding the encoder stage.

[0151] EEE17. The method according to any of EEE14 to EEE16, wherein the multi-channel generator further comprises a non-strided transposed convolutional layer as an output layer subsequent to the decoder stage.

[0152] EEE18. A method according to any of EEE14 to 17, wherein there are one or more skip connections between each of the homologous layers of the multi-channel generator.

[0153] EEE19. A method according to any of EEE14 to 18, wherein the multi-channel generator includes a stage between the encoder stage and the decoder stage for modifying the multi-channel audio in a reduced dynamic range domain based at least on the reduced dynamic range coded multi-channel audio feature space.

[0154] EEE20. A method according to EEE19, in which a random noise vector z is used in a reduced dynamic range coded multi-channel audio feature space to modify the multi-channel audio in a reduced dynamic range domain.

[0155] EEE21. In the method according to EEE20, the use of the random noise vector z is conditional on the bit rate of the audio bitstream and the number of channels of the multi-channel audio signal.

[0156] EEE22. The method according to any one of EEE1 to EEE21, further comprising the following steps to be performed before step (a): (i) inputting a reduced-dynamic-range raw multi-channel audio training signal (reduced-dynamic-range raw multi-channel audio training signal) into a multi-channel generator, the reduced-dynamic-range raw multi-channel audio training signal including two or more channels; (ii) jointly generating, by a multi-channel generator, an enhanced dynamic range reduced multi-channel audio training signal (enhanced dynamic range reduced multi-channel audio training signal) based on the dynamic range reduced raw multi-channel audio training signal; (iii) inputting each of the two or more channels of the enhanced dynamic range reduced multi-channel audio training signal and a corresponding channel of the original dynamic range reduced multi-channel audio signal from which the dynamic range reduced raw multi-channel audio training signal is derived into one out of a group of one or more single channel discriminators; (iv) further inputting the enhanced dynamic range reduced multi-channel audio training signal and the corresponding original dynamic range reduced multi-channel audio signal into a multi-channel discriminator one at a time; (v) determining, by a single-channel discriminator and a multi-channel discriminator, whether the input reduced-dynamic-range multi-channel audio signal is an enhanced reduced-dynamic-range multi-channel audio training signal or an original reduced-dynamic-range multi-channel audio signal; (vi) tuning parameters of the multi-channel generator until the single-channel discriminator and the multi-channel discriminator can no longer distinguish the enhanced reduced dynamic range multi-channel audio training signal from the original reduced dynamic range multi-channel audio signal.

[0157] EEE23. A method according to EEE22, wherein the set of one or more single-channel discriminators is selected based on a type of original reduced-dynamic-range multi-channel audio signal, the original reduced-dynamic-range multi-channel audio signal comprising a stereo type multi-channel audio signal, a 5.1 type multi-channel audio signal, a 7.1 type multi-channel audio signal, or a 9.1 type multi-channel audio signal.

[0158] EEE24. The method according to EEE22 or 23, wherein in step (i), additionally, a random noise vector z is an input to a multi-channel generator, and wherein the step of jointly generating an enhanced dynamic range reduced multi-channel audio training signal by the multi-channel generator in step (ii) is additionally based on the random noise vector z.

[0159] EEE25. A method according to any of EEE22 to 24, wherein in step (i) the additional metadata is input to the multi-channel generator, and wherein in step (ii) the step of jointly generating an enhanced dynamic range reduced multi-channel audio training signal by the multi-channel generator is additionally based on the metadata.

[0160] EEE26. A method according to EEE25, wherein the metadata includes one or more items of companding control data.

[0161] EEE27. A method according to EEE26, wherein the companding control data comprises information about a companding mode among one or more companding modes used in encoding the original multi-channel audio signal.

[0162] EEE28. In the method according to EEE27, the companding modes include a companding-on companding mode, a companding-off companding mode, and an average companding companding mode.

[0163] EEE29. In the method according to EEE 27 or 28, the step of jointly generating the enhanced reduced dynamic range multi-channel audio training signals by the multi-channel generator in step (ii) depends on the companding mode indicated by the companding control data.

[0164] EEE30. In the method according to EEE29 which is dependent on EEE28, when the companding mode is companding off, no co-enhancement is performed by the multi-channel generator.

[0165] EEE31. 1. A method for training a multi-channel generator in a reduced dynamic range domain in a generative adversarial network setting having a multi-channel generator, a group of one or more single-channel discriminators, and a multi-channel discriminator, the method comprising: (a) inputting a reduced dynamic range raw multi-channel audio training signal into a multi-channel generator, the reduced dynamic range raw multi-channel audio training signal comprising two or more channels; (b) jointly generating, by a multi-channel generator, an enhanced reduced-dynamic-range multi-channel audio training signal based on the reduced-dynamic-range raw multi-channel audio training signal; (c) inputting, one at a time, each of the two or more channels of the enhanced dynamic range reduced multi-channel audio training signal and the corresponding channel of the original dynamic range reduced multi-channel audio signal from which the dynamic range reduced raw multi-channel audio training signal is derived into one single-channel discriminator of a group of one or more single-channel discriminators; (d) further inputting the enhanced dynamic range reduced multi-channel audio training signal and the corresponding original dynamic range reduced multi-channel audio signal one at a time into a multi-channel discriminator; (e) determining, by a single-channel discriminator and a multi-channel discriminator, whether the input reduced-dynamic-range multi-channel audio signal is an enhanced reduced-dynamic-range multi-channel audio training signal or an original reduced-dynamic-range multi-channel audio signal; (f) tuning parameters of the multi-channel generator until the single-channel discriminator and the multi-channel discriminator can no longer distinguish the enhanced reduced dynamic range multi-channel audio training signal from the original reduced dynamic range multi-channel audio signal.

[0166] EEE32. A method according to EEE36, wherein the set of one or more single-channel discriminators is selected based on a type of original reduced-dynamic-range multi-channel audio signal, the original reduced-dynamic-range multi-channel audio signal comprising a stereo type multi-channel audio signal, a 5.1 type multi-channel audio signal, a 7.1 type multi-channel audio signal, or a 9.1 type multi-channel audio signal.

[0167] EEE33. The method according to EEE31 or 32, wherein in step (i), additionally, a random noise vector z is an input to the multi-channel generator, and wherein in step (ii) the step of jointly generating an enhanced dynamic range reduced multi-channel audio training signal by the multi-channel generator is additionally based on the random noise vector z.

[0168] EEE34. A method according to any of EEE31 to EEE33, wherein in step (i) the additional metadata is input to the multi-channel generator, and in step (ii) the step of jointly generating an enhanced dynamic range reduced multi-channel audio training signal by the multi-channel generator is additionally based on the metadata.

[0169] EEE35. A method according to EEE34, wherein the metadata includes one or more items of companding control data.

[0170] EEE36. A method according to EEE35, wherein the companding control data comprises information about a companding mode among one or more companding modes used in encoding the original multi-channel audio signal.

[0171] EEE37. In the method according to EEE36, the companding modes include a companding-on companding mode, a companding-off companding mode, and an average companding companding mode.

[0172] EEE38. In the method according to EEE 36 or 37, the step of jointly generating the enhanced reduced dynamic range multi-channel audio training signals by the multi-channel generator in step (ii) depends on the companding mode indicated by the companding control data.

[0173] EEE39. In the method according to EEE38 which is subordinate to EEE37, if the companding mode is companding off, no co-enhancement is performed by the multi-channel generator.

[0174] EEE40. 1. An apparatus for generating an enhanced multi-channel audio signal in a reduced dynamic range domain from an audio bitstream comprising the multi-channel audio signal, the apparatus comprising: (a) a receiver for receiving an audio bitstream; (b) a core decoder for core-decoding an audio bitstream to obtain a raw multi-channel audio signal having a reduced dynamic range based on the received audio bitstream, the reduced dynamic range raw multi-channel audio signal including two or more channels; (c) a multi-channel generator for jointly enhancing two or more channels of the reduced-dynamic-range raw multi-channel audio signal in a reduced-dynamic-range domain to obtain an enhanced reduced-dynamic-range multi-channel audio signal, wherein the enhanced reduced-dynamic-range multi-channel audio signal has two or more channels.

[0175] EEE41. The apparatus according to EEE40 further includes a demultiplexer for demultiplexing a received audio bitstream, the received audio bitstream including metadata.

[0176] EEE42. An apparatus according to EEE41, wherein the metadata includes one or more items of companding control data.

[0177] EEE43. A device according to EEE42, The companding control data includes information regarding the companding mode, among one or more companding modes, that was used in encoding the multi-channel audio signal.

[0178] EEE44. In the device according to EEE43, the companding modes include a companding on companding mode, a companding off companding mode and an average companding companding mode.

[0179] EEE45. The apparatus according to EEE43 or EEE44, a multi-channel generator, is configured to jointly emphasize two or more channels of a reduced dynamic range raw multi-channel audio signal in a reduced dynamic range region depending on a companding mode indicated by companding control data.

[0180] EEE46. In an apparatus according to EEE45 which is subordinate to EEE44, the multi-channel generator is configured not to perform co-enhancement when the companding mode is companding off.

[0181] EEE47. The apparatus according to any of EEE40 to 46, further comprising an expansion unit configured to perform an expansion operation on two or more channels to expand the enhanced dynamic range reduced multi-channel audio signal into an extended dynamic range region.

[0182] EEE48. The apparatus according to any of EEE40 to 46, further comprising a dynamic range reduction unit configured to perform a domain range reduction operation after core decoding of the audio bitstream to obtain a dynamic range reduced raw multi-channel audio signal.

[0183] EEE49. A computer program product comprising a computer-readable storage medium having instructions adapted, when executed by a device having processing capability, to cause the device to perform a method according to any of EEE1 to EEE30.

[0184] EEE50. A computer program product comprising a computer readable storage medium having instructions adapted, when executed by a device having processing capability, to cause the device to perform a method according to any of EEE31 to EEE39.

[0185] EEE51. 1. A system of an apparatus for generating an enhanced multi-channel audio signal from an audio bitstream in a reduced dynamic range domain and a generative adversarial network, the generative adversarial network having a multi-channel generator, a group of one or more single-channel discriminators, and a multi-channel discriminator, the system being configured to perform a method according to any of EEE1 to EEE30.

[0186] EEE52. A system comprising an apparatus for applying dynamic range reduction to an input multi-channel audio signal and encoding the reduced dynamic range multi-channel audio signal in an audio bitstream, and an apparatus according to any of EEE40 to EEE48 for generating an enhanced multi-channel audio signal from an audio bitstream comprising the multi-channel audio signal in a reduced dynamic range domain.

[0187] [Appendix 1] 1. A method for generating an enhanced multi-channel audio signal in a reduced dynamic range domain from an audio bitstream containing the multi-channel audio signal, comprising: The method comprises: receiving an audio bitstream; core-decoding the audio bitstream and obtaining a dynamic range-reduced raw multi-channel audio signal based on the received audio bitstream, wherein the dynamic range-reduced raw multi-channel audio signal includes two or more channels; inputting the reduced dynamic range raw multi-channel audio signals to a multi-channel generator for jointly processing the reduced dynamic range raw multi-channel audio signals; co-emphasizing the two or more channels of the reduced dynamic range raw multi-channel audio signal by the multi-channel generator in the reduced dynamic range domain; obtaining an enhanced dynamic range reduced multi-channel audio signal as output from the multi-channel generator for subsequent dynamic range expansion, the enhanced dynamic range reduced multi-channel audio signal having two or more channels; A method comprising: [Appendix 2] After core-decoding the audio bitstream, further comprising: performing a dynamic range reduction operation to obtain the dynamic range reduced raw multi-channel audio signal. The method described in Appendix 1. [Appendix 3] the audio bitstream is in AC-4 format; The method described in Appendix 1. [Appendix 4] the method further comprising the step of expanding the enhanced dynamic range reduced multi-channel audio signal to an extended dynamic range region by performing an expansion operation on the two or more channels. The method described in Appendix 1. [Appendix 5] the expansion operation is a companding operation based on the p-norm of the spectral magnitude to calculate respective gain values; The method described in Appendix 4. [Appendix 6] the received audio bitstream includes metadata; receiving the audio bitstream includes demultiplexing the received audio bitstream; 6. The method of any one of appendices 1 to 5. [Appendix 7] and co-emphasizing the two or more channels of the reduced dynamic range raw multi-channel audio signal by the multi-channel generator based on the metadata. The method described in Appendix 6. [Appendix 8] the metadata includes one or more items of companding control data; The method described in Appendix 7. [Appendix 9] the companding control data includes information regarding a companding mode, among one or more companding modes, that was used in encoding the multi-channel audio signal; The method described in Appendix 8. [Appendix 10] the companding mode includes the companding mode of companding on, the companding mode of companding off, and the companding mode of average companding; The method described in Appendix 9. [Appendix 11] co-emphasizing the two or more channels of the dynamic range reduced raw multi-channel audio signal by the multi-channel generator depends on the companding mode indicated by the companding control data. The method described in Appendix 9 or 10. [Appendix 12] When the companding mode is companding off, no co-enhancement is performed by the multi-channel generator. The method of claim 11, which is dependent on claim 10. [Appendix 13] The multi-channel generator is a generator trained in a reduced dynamic range domain in a generative adversarial network setting. 13. The method of any one of appendices 1 to 12. [Appendix 14] the multi-channel generator includes an encoder stage and a decoder stage arranged in mirror symmetry; the encoder stage and the decoder stage each include L layers with N filters within each layer; L is a natural number greater than 1, N is a natural number greater than 1, the N filters in each layer of the encoder stage and the decoder stage have the same size; each of the N filters in the encoder stage and the decoder stage operates with a stride greater than 1; 14. The method of any one of appendices 1 to 13. [Appendix 15] the multi-channel generator further includes a non-strided convolutional layer as an input layer preceding the encoder stage. The method described in Appendix 14. [Appendix 16] a nonlinear operation including one or more of ReLU, PReLU, LReLU, eLU, and SeL is performed in at least one layer of the encoder stage and at least one layer of the decoder stage; 16. The method of claim 14 or 15. [Appendix 17] the multi-channel generator further includes a non-stride permuted convolutional layer as an output layer subsequent to the decoder stage; 17. The method of any one of appendices 14 to 16. [Appendix 18] There are one or more skip connections between each homogeneous layer of the multi-channel generator; 18. The method of any one of appendices 14 to 17. [Appendix 19] the multi-channel generator including a stage between the encoder stage and the decoder stage for modifying the multi-channel audio in the reduced dynamic range domain based at least on a reduced dynamic range coded multi-channel audio feature space; 19. The method of any one of appendices 14 to 18. [Appendix 20] a random noise vector z is used within the reduced dynamic range coded multi-channel audio feature space to modify the multi-channel audio in the reduced dynamic range domain; The method described in Appendix 19. [Appendix 21] the use of the random noise vector z is conditional on the bitrate of the audio bitstream and the number of channels of the multi-channel audio signal; The method described in Appendix 20. [Appendix 22] The method further includes the following steps to be performed before the step of receiving the audio bitstream: inputting a dynamic range reduced raw multi-channel audio training signal (dynamic range reduced raw multi-channel audio training signal) into the multi-channel generator, the dynamic range reduced raw multi-channel audio training signal having two or more channels; jointly generating, by the multi-channel generator, an enhanced reduced-dynamic-range multi-channel audio training signal (enhanced reduced-dynamic-range multi-channel audio training signal) based on the reduced-dynamic-range raw multi-channel audio training signal; inputting, one at a time, each of the two or more channels of the enhanced dynamic range reduced multi-channel audio training signal and a corresponding channel of the original dynamic range reduced multi-channel audio signal from which the dynamic range reduced raw multi-channel audio training signal is derived into one single-channel discriminator of a group of one or more single-channel discriminators; further inputting the enhanced dynamic range reduced multi-channel audio training signal and the corresponding original dynamic range reduced multi-channel audio signal one at a time into a multi-channel discriminator; determining, by the single-channel discriminator and the multi-channel discriminator, whether the input reduced-dynamic-range multi-channel audio signal is the enhanced reduced-dynamic-range multi-channel audio training signal or the original reduced-dynamic-range multi-channel audio signal; tuning parameters of the multi-channel generator until the single-channel discriminator and the multi-channel discriminator can no longer distinguish the enhanced reduced dynamic range multi-channel audio training signal from the original reduced dynamic range multi-channel audio signal; 22. The method of any one of appendices 1 to 21, comprising: [Appendix 23] the set of one or more single-channel discriminators is selected based on a type of the original reduced dynamic range multi-channel audio signal; the original dynamic range reduced multi-channel audio signal comprises a stereo type multi-channel audio signal, a 5.1 type multi-channel audio signal, a 7.1 type multi-channel audio signal, or a 9.1 type multi-channel audio signal; The method described in Appendix 22. [Appendix 24] Additionally, a random noise vector z is input to the multi-channel generator; and co-generating the enhanced dynamic range reduced multi-channel audio training signals by the multi-channel generator additionally based on the random noise vector z. 24. The method of claim 22 or 23. [Appendix 25] additional metadata is input to said multi-channel generator; and co-generating the enhanced dynamic range reduced multi-channel audio training signals by the multi-channel generator is additionally based on the metadata. 25. The method of any one of appendices 22 to 24. [Appendix 26] the metadata includes one or more items of companding control data; The method described in Appendix 25. [Appendix 27] the companding control data includes information about a companding mode, among one or more companding modes, used in encoding the original multi-channel audio signal; The method described in Appendix 26. [Appendix 28] The companding mode includes a companding-on companding mode, a companding-off companding mode, and an average companding companding mode. The method described in Appendix 27. [Appendix 29] the step of jointly generating the enhanced dynamic range reduced multi-channel audio training signals by the multi-channel generator depends on the companding mode indicated by the companding control data. 29. The method of claim 27 or 28. [Appendix 30] When the companding mode is companding off, no co-enhancement is performed by the multi-channel generator. The method of claim 29, which references claim 28. [Appendix 31] 1. A method for training a multi-channel generator in a reduced dynamic range domain in a generative adversarial network setting having a multi-channel generator, a group of one or more single-channel discriminators, and a multi-channel discriminator, the method comprising: inputting a reduced dynamic range raw multi-channel audio training signal into the multi-channel generator, the reduced dynamic range raw multi-channel audio training signal comprising two or more channels; co-generating, by the multi-channel generator, an enhanced reduced dynamic range multi-channel audio training signal based on the reduced dynamic range raw multi-channel audio training signal; inputting, one at a time, each of the two or more channels of the enhanced dynamic range reduced multi-channel audio training signal and a corresponding channel of the original dynamic range reduced multi-channel audio signal from which the dynamic range reduced raw multi-channel audio training signal is derived into one single channel discriminator of the group of one or more single channel discriminators; further inputting the enhanced reduced dynamic range multi-channel audio training signal and the corresponding original reduced dynamic range multi-channel audio signal into the multi-channel discriminator one at a time; determining, by the single-channel discriminator and the multi-channel discriminator, whether the input reduced-dynamic-range multi-channel audio signal is the enhanced reduced-dynamic-range multi-channel audio training signal or the original reduced-dynamic-range multi-channel audio signal; tuning parameters of the multi-channel generator until the single-channel discriminator and the multi-channel discriminator can no longer distinguish the enhanced reduced dynamic range multi-channel audio training signal from the original reduced dynamic range multi-channel audio signal; A method comprising: [Appendix 32] the set of one or more single-channel discriminators is selected based on a type of the original reduced dynamic range multi-channel audio signal; the original dynamic range reduced multi-channel audio signal comprises a stereo type multi-channel audio signal, a 5.1 type multi-channel audio signal, a 7.1 type multi-channel audio signal, or a 9.1 type multi-channel audio signal; The method described in Appendix 31. [Appendix 33] Additionally, a random noise vector z is input to the multi-channel generator; and co-generating the enhanced dynamic range reduced multi-channel audio training signals by the multi-channel generator additionally based on the random noise vector z. 33. The method of claim 31 or 32. [Appendix 34] additional metadata is input to said multi-channel generator; and co-generating the enhanced dynamic range reduced multi-channel audio training signals by the multi-channel generator is additionally based on the metadata. 34. The method of any one of appendices 31 to 33. [Appendix 35] the metadata includes one or more items of companding control data; The method described in Appendix 34. [Appendix 36] the companding control data includes information about a companding mode, among one or more companding modes, used in encoding the original multi-channel audio signal; The method described in Appendix 35. [Appendix 37] The companding mode includes a companding-on companding mode, a companding-off companding mode, and an average companding companding mode. The method described in Appendix 36. [Appendix 38] the step of jointly generating the enhanced dynamic range reduced multi-channel audio training signals by the multi-channel generator is dependent on the companding mode indicated by the companding control data. 38. The method of claim 36 or 37. [Appendix 39] When the companding mode is companding off, no co-enhancement is performed by the multi-channel generator. The method of claim 38, which is dependent on claim 37. [Appendix 40] 1. An apparatus for generating an enhanced multi-channel audio signal in a reduced dynamic range domain from an audio bitstream comprising the multi-channel audio signal, the apparatus comprising: a receiver for receiving the audio bitstream; a core decoder for core-decoding the audio bitstream to obtain a raw multi-channel audio signal with reduced dynamic range (reduced dynamic range raw multi-channel audio signal) based on the received audio bitstream (received audio bitstream), wherein the reduced dynamic range raw multi-channel audio signal includes two or more channels; a multi-channel generator for jointly enhancing the two or more channels of the reduced-dynamic-range raw multi-channel audio signal in the reduced-dynamic-range domain to obtain an enhanced reduced-dynamic-range multi-channel audio signal, wherein the enhanced reduced-dynamic-range multi-channel audio signal has two or more channels; An apparatus comprising: [Appendix 41] a demultiplexer for demultiplexing the received audio bitstream, the received audio bitstream including metadata; 41. The apparatus of claim 40. [Appendix 42] the metadata includes one or more items of companding control data; 42. The apparatus of claim 41. [Appendix 43] the companding control data includes information regarding a companding mode, among one or more companding modes, that was used in encoding the multi-channel audio signal; 43. The apparatus of claim 42. [Appendix 44] The companding mode includes a companding-on companding mode, a companding-off companding mode, and an average companding companding mode. 44. The apparatus of claim 43. [Appendix 45] the multi-channel generator is configured to jointly emphasize the two or more channels of the reduced dynamic range raw multi-channel audio signal in the reduced dynamic range region depending on the companding mode indicated by the companding control data. 45. The apparatus of claim 43 or 44. [Appendix 46] When the companding mode is companding off, the multi-channel generator is configured not to perform co-enhancement. 46. ​​The device of claim 45 dependent on claim 44. [Appendix 47] The device comprises: an expansion unit configured to perform an expansion operation on the two or more channels to expand the enhanced dynamic range reduced multi-channel audio signal to an extended dynamic range region. 47. The apparatus of any one of clauses 40 to 46. [Appendix 48] The device comprises: a dynamic range reduction unit configured to perform a domain range reduction operation after core decoding the audio bitstream to obtain the dynamic range reduced raw multi-channel audio signal; 48. The apparatus of any one of clauses 40 to 47. [Appendix 49] 31. A computer program comprising instructions adapted, when executed by a device having processing capability, to cause the device to carry out the method of any one of clauses 1 to 30. [Appendix 50] 40. A computer program comprising instructions adapted, when executed by a device having processing capability, to cause the device to carry out the method of any one of clauses 31 to 39. [Appendix 51] an apparatus for generating an enhanced multi-channel audio signal from an audio bitstream containing the multi-channel audio signal in a reduced dynamic range domain; 1. A system of a multi-channel generator, a set of one or more single-channel discriminators, and a generative adversarial network having a multi-channel discriminator, 31. The system, wherein the system is configured to perform the method of any one of claims 1 to 30. [Appendix 52] an apparatus for applying dynamic range reduction to an input multi-channel audio signal and encoding the dynamic range reduced multi-channel audio signal in an audio bitstream; 49. An apparatus according to any one of clauses 40 to 48 for generating an enhanced multi-channel audio signal from an audio bitstream comprising a multi-channel audio signal in a reduced dynamic range domain; system.

Claims

[Claim 1] 1. A method for generating an enhanced multi-channel audio signal in a reduced dynamic range domain from an audio bitstream containing the multi-channel audio signal, comprising: The method comprises: receiving an audio bitstream; core-decoding the audio bitstream and obtaining a dynamic range-reduced raw multi-channel audio signal based on the received audio bitstream, wherein the dynamic range-reduced raw multi-channel audio signal includes two or more channels; inputting the reduced dynamic range raw multi-channel audio signals into a multi-channel generator for jointly processing the reduced dynamic range raw multi-channel audio signals; co-emphasizing the two or more channels of the reduced dynamic range raw multi-channel audio signal by the multi-channel generator in the reduced dynamic range domain; obtaining an enhanced dynamic range reduced multi-channel audio signal as output from the multi-channel generator for subsequent dynamic range expansion, the enhanced dynamic range reduced multi-channel audio signal having two or more channels; A method comprising:

Citation Information

Patent Citations

  • Digital voice signal processor

    JP2008089982A

  • Signal processing device, signal processing program, signal processing method, and sound collection device

    JP2020012980A

  • Signal-Dependent Companding System and Method to Reduce Quantization Noise

    US20180358028A1