Apparatus, method and computer program for generating an output downmix representation
The multi-channel decoder harmonizes stereo-to-mono conversion in mobile devices by employing different downmix schemes in the spectral domain, addressing quality and complexity issues in stereo decoding to mono, and optimizing battery usage.
Patent Information
- Application Number
- JP2023144908
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-29
- Filing Date
- 2023-09-07
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-04-22
AI Technical Summary
Existing methods for decoding stereo signals to mono in mobile devices face issues such as phase cancellation effects, energy loss, and spectral imbalances due to passive and active downmixing techniques, which degrade audio quality and increase complexity and latency.
A multi-channel decoder that employs different downmix schemes for various spectral regions of a stereo signal, using active and passive downmixing methods to harmonize the signal in the spectral domain without additional delay or complexity, and generates a high-quality mono output.
The solution provides a high-quality mono output by harmonizing different downmixing techniques in the spectral domain, avoiding phase cancellation and energy imbalances, while reducing processing resources and battery consumption in mobile devices.
Smart Images

Figure 0007757360000016 
Figure 0007757360000017 
Figure 0007757360000018
Abstract
Description
[Technical Field]
[0001] The present application relates to multi-channel processing, in particular multi-channel processing that offers the possibility for a mono output. [Background technology]
[0002] Stereo-encoded bitstreams are typically decoded for playback on stereo systems, but not all devices capable of receiving stereo bitstreams are always capable of outputting a stereo signal. For example, a stereo signal may be played on a mobile phone with only a mono speaker. Therefore, with the emergence of multi-channel mobile communication scenarios, such as those supported by the 3GPP IVAS standard, there is a need for a stereo-to-mono downmixing method that provides the best perceptual quality possible without additional delay and is as efficient as possible in terms of complexity, while still providing the best possible quality that cannot be achieved with simple passive downmixing.
[0003] TIFF0007757360000001.tif27169
[0004] Furthermore, more sophisticated (i.e., active) time-domain-based downmixing methods include energy scaling to preserve the overall energy of the signal [2], [3], phase adjustment to avoid cancellation effects [4], and coherence suppression to prevent comb-filtering effects [5].
[0005] Another approach is to perform energy correction in a frequency-dependent manner by calculating separate weighting coefficients for multiple spectral bands. For example, this is done as part of the MPEG-H format converter [6], which downmixes using a hybrid QMF subband representation of the signal, with additional channel phase adjustment beforehand. A similar band-wise downmix (including both phase and time adjustments) has already been used in [7] for DFT stereo in a parametric low-bitrate mode, where weighting and mixing are performed in the DFT domain. Summary of the Invention [Problem to be solved by the invention]
[0006] The solution of decoding the stereo signal and then passively downmixing from stereo to mono in the time domain is not ideal, since it is well known that purely passive downmixing has drawbacks such as phase cancellation effects and general energy loss, which can significantly degrade the quality of some items.
[0007] Other active downmixing techniques based purely on the time domain alleviate some of the problems of passive downmixing, but are still suboptimal due to the lack of frequency-dependent weighting.
[0008] IVAS (Immersive Voice and Audio Services) Due to the implicit constraints in terms of latency and complexity for mobile codecs such as MPEG-H, having a dedicated post-processing stage to apply band-by-band downmixing, as in the MPEG-H format converter, is not an option, as it requires conversion to the frequency domain and back, which inevitably increases both complexity and latency.
[0009] DFT-based stereo systems, such as those described in [8], which use only parameter-based residual prediction to reconstruct the stereo signal at the decoder and generate an intermediate signal via active downmixing, as described in [7], produce a sufficiently good mono signal at the decoder. However, if the spectral portion of the signal depends on the coded residual signal for stereo reconstruction, generated by M / S conversion, the mono signal obtained before the stereo upmixing is no longer adequate. In this case, the mono signal consists of part of the intermediate signal from the M / S conversion (residual coding section), which is spectrally equivalent to a passive downmix, and part of the active downmix (residual prediction section). This mixing of two different downmixing techniques can result in artifacts and energy imbalances in the signal.
[0010] It is an object of the present invention to provide an improved concept for generating an output downmix representation for multi-channel decoding. [Means for solving the problem]
[0011] This object is achieved by an apparatus for generating an output downmix representation according to claim 1, a multi-channel decoder according to claim 19, a method for generating an output downmix representation according to claim 24, a multi-channel decoding method according to claim 27 or an associated computer program according to claim 28.
[0012] An apparatus for generating an output downmix representation from an input downmix representation, at least a portion of the input downmix representation according to a first downmix scheme, the apparatus comprising: an upmixer for upmixing at least a portion of the input downmix representation using an upmix scheme corresponding to the first downmix scheme to obtain at least one upmixed part, and further comprising a downmixer for downmixing the at least one upmixed part according to a second downmix scheme different from the first downmix scheme.
[0013] In another embodiment, a portion of the input downmix representation complies with a downmix scheme, and a second portion of the input downmix representation complies with a second downmix scheme different from the first downmix scheme. In this embodiment, the downmixer is configured to downmix the upmix portion according to the second downmix scheme or according to a third downmix scheme different from the downmix scheme and the second downmix scheme to obtain a first downmixed portion. Here, the situation regarding the downmixed portion is that the first downmixed portion and the second downmixed portion are related and can be said to be in the same downmix scheme domain, so that the first downmixed portion and the second downmixed portion, or a downmixed portion derived from the second downmixed portion, can be combined by a combiner to obtain an output downmix representation including an output representation for the first portion and an output representation for the second portion. The output representation for the first portion and the output representation for the second portion are based on the same downmix scheme, i.e., are located in one and the same downmix domain, and therefore "match" each other.
[0014] In further embodiments, the whole band or only a part of the input downmix representation is downmixed depending on the parameters and the residual signal or only on the residual signal without parameters. In this situation, the input downmix representation is based on a downmix scheme. In this situation, the input downmix representation consists of a core signal, a residual signal, or a residual signal and parameters. This signal is upmixed using side information, i.e., using parameters and the residual signal, or using only the residual signal. The upmix includes all available information, including the residual signal. The downmix is performed with a second downmix scheme different from the first downmix scheme, i.e., preferably an active downmix having means for addressing energy calculations, or in other words, a downmix scheme that does not generate a residual signal, and preferably does not generate a residual signal and any parameters. Such a downmix offers the possibility of a good, comfortable, and high-quality audio mono rendering, but the core signal of the input downmix representation when used without an upmix and subsequent downmix does not allow any comfortable, high-quality audio reproduction when rendered without taking the residual signal and parameters into account.
[0015] According to this embodiment, the device for generating an output downmix representation performs a conversion from a residual-type downmix scheme to a non-residual-type downmix scheme. This conversion can be performed full-band or partial-band. Typically, and in a preferred embodiment, the low band of a multi-channel encoded signal contains a core signal, a residual signal, and preferably parameters. However, for higher bands, the precision is reduced due to the lower bit rate. Therefore, for such higher bands, active downmixing is sufficient without additional side information such as residual data or parameters. In such a situation, the low band in the residual downmix domain is converted to the non-residual downmix domain, and the result is combined with the higher band already in the "correct" non-residual downmix domain.
[0016] In further embodiments, it is not required that the first part be transformed from the first downmix region to the same downmix region in which the second part is located. Instead, in further embodiments, if the first part is in the first downmix region and the second part of the input representation is in the second downmix region, both of these parts are transformed into a different, third downmix region by upmixing the first part according to a first upmix scheme corresponding to the first downmix scheme. Furthermore, the second part is upmixed according to a second upmix scheme corresponding to the second downmix scheme, and both upmixes are downmixed to a third downmix scheme different from the first and second downmix schemes, preferably by active downmixing without residual data or parametric data.
[0017] Further embodiments may utilize two or more portions, in particular spectral portions or spectral bands, in different downmix representations. According to the present invention, preferably, if the upmix and subsequent downmix are performed in the spectral domain, then individual processing for each band can be performed without interference from one spectral band to another. At the output of the downmixer, all bands are in the same "downmix" domain, and thus the spectrum for the mono output downmix representation is present. This spectrum can be converted to a time-domain representation by a spectral-to-time transformer, such as a synthesis filter bank, an inverse discrete Fourier transform, or an inverse MDCT domain. The combination of the individual bands and the transformation to the time domain can be performed using such a synthesis filter bank. In particular, it does not matter whether the combination is performed before the actual transformation, i.e., in the spectral domain. In such a situation, the combination is performed before the spectral-to-time transformation, i.e., at the input to the synthesis filter bank, and only a single transformation is performed to obtain a single time-domain signal. However, an equivalent implementation would consist of an implementation in which the combiner performs a spectral-to-time transformation for each band separately. The time-domain output of such an individual transform therefore represents a time-domain representation in a specific bandwidth, The individual time domain outputs are combined sample by sample, preferably after some type of upsampling if a critically sampled transform is implemented.
[0018] In a further embodiment, the invention is applied to a multi-channel decoder that can operate in two different modes: a multi-channel output mode as a "normal" mode and a second mode as an "exceptional" mode that is a mono output mode. This mono output mode is particularly useful when the multi-channel decoder is implemented in a device that only has a mono speaker output capability, such as a mobile phone with one speaker, or in a device that is in some kind of power saving mode, and that in principle has the possibility of a multi-channel or stereo output mode, but only provides a mono output mode to save battery and processing resources.
[0019] In such an embodiment, the multi-channel decoder comprises a first time-spectral transform function for the decoded core signal and a second time-spectral transform function for the decoder residual signal, and two different upmix functions in the spectral domain for two different spectral portions in two different downmix domains are provided, where corresponding left channel spectral lines are combined by a combiner such as a synthesis filter bank or an IDFT block, and corresponding other channel spectral lines are combined by an additional or second synthesis filter bank or IDFT (Inverse Discrete Fourier Transform) block.
[0020] To enhance such a multi-channel decoder, a downmixer is provided for downmixing at least one upmixed portion according to a second downmixing scheme different from the first downmixing scheme, preferably implemented as an active downmixer. In an embodiment, two switches and a controller are also provided. The controller controls the first switch to bypass the upmixer for the high-band portion, and the second switch is implemented to provide the output of the upmixer to the downmixer. In such a mono output mode, the second combiner or synthesis filter bank is inactive and the upmixer for the high-band portion is also inactive to save processing power. However, in a stereo output mode, to obtain a left stereo output signal and a right stereo output signal, the first switch provides the upmix for the high-band portion, the second switch bypasses the (active) downmixer, and both output synthesis filter banks are active.
[0021] Because the mono output is calculated in the spectral domain, such as the DFT domain, generating a mono output incurs no additional delay compared to generating a stereo output. This is because no additional time-frequency transform is required compared to the stereo processing mode. Instead, one of the two stereo mode synthesis filterbanks is also used in the mono mode. Furthermore, compared to the stereo output, which typically provides an enhanced audio experience compared to the mono output, the mono processing mode saves complexity, especially processing resources, and ultimately battery power in low-power modes, which is particularly useful for battery-powered mobile devices. This is because the high-band upmixer typically required in the stereo mode can be deactivated, and the second output filterbank, also required in the stereo output mode, can be deactivated as well. Instead, compared to the stereo mode, the only additional processing block required is a low-complexity, low-latency active downmix block that operates entirely in the spectral domain. However, the additional processing resources required by this active downmix block are significantly less than the processing resources saved by deactivating the high-band upmixer and the second synthesis filterbank or IDFT block.
[0022] This embodiment aims to generate a harmonized mono output signal from a mono input signal created by downmixing a stereo signal, where the downmixing is performed using different methods (e.g., active and passive) for at least two different spectral regions of the stereo signal. Harmonization is achieved by selecting one downmix method as the preferred method for the harmonized signal and converting all spectral portions downmixed using different methods to the desired method. This is achieved by first upmixing these spectral portions using all side parameters required for the upmix and recovering LR representations in each spectral region. Next, the preferred method is applied to the stereo representation using all parameters required for the preferred downmix method to convert the spectral portions to mono representations. A harmonized mono output signal is generated, avoiding the problem of uneven downmixing without additional delay or complexity.
[0023] Preferred embodiments will now be described with reference to the accompanying drawings. [Brief explanation of the drawings]
[0024] [Figure 1] FIG. 1 illustrates an apparatus for generating an output downmix representation in one embodiment. [Figure 2] FIG. 2 shows an apparatus for generating an output downmix representation in a further embodiment, where the downmix scheme is based on the residual signal or on the residual signal and parameters. [Figure 3] FIG. 3 illustrates a further embodiment in which different downmix schemes are performed for different parts, such as spectral parts, of the input downmix representation. [Figure 4]FIG. 4 is a diagram illustrating a further embodiment illustrating the use of different downmix schemes in different spectral parts for an input downmix representation, where a first downmix scheme is based on residual data and a second downmix scheme is an active downmix scheme or a downmix scheme without residual or parametric data. [Figure 5] FIG. 5 is a diagram showing a preferred example of an upmix scheme corresponding to the first downmix scheme in the embodiment. [Figure 6] FIG. 6 shows a multi-channel decoder operating in stereo output mode. [Figure 7] FIG. 7 illustrates a multi-channel encoder according to an embodiment that is switchable between a multi-channel output mode or a mono output mode. [Figure 8a] FIG. 8a illustrates a preferred embodiment of the second downmix scheme. [Figure 8b] FIG. 8b shows a further embodiment of the second downmix scheme. [Figure 9] FIG. 9 illustrates the separation of an input downmix representation into a part of the input downmix representation of a first downmix scheme, denoted as a first part, and a second part of the input downmix representation that depends on a downmix scheme with weights. DETAILED DESCRIPTION OF THE INVENTION
[0025] 1 shows an apparatus for generating an output downmix representation from an input downmix representation, at least a portion of the input downmix representation according to a first downmix scheme. The apparatus comprises an upmixer 200 for upmixing at least a portion of the input downmix representation using an upmix scheme corresponding to the first downmix scheme to obtain at least one upmixed portion at the output of block 200. The apparatus further comprises a downmixer 200 for downmixing the at least one upmixed portion according to a second downmix scheme different from the first downmix scheme. The audio signal processing system comprises a downmixer 300 for mixing the output downmix representation with a monophonic playback signal. Preferably, the output of the downmixer 300 is forwarded to an output stage 500 for generating a monophonic output. The output stage is for example an output interface for outputting the output downmix representation to a rendering device, or alternatively the output stage 500 actually constitutes a rendering device for rendering the output downmix representation as a monophonic playback signal.
[0026] The device shown in Fig. 1 provides a conversion from a downmix representation in a first "downmix domain" to another, second downmix domain. As will be explained in other figures, this conversion can be performed, for example, for the lowest three bands b1, b2, b3 exemplarily given in Fig. 9. The downmix conversion may be effective only for a limited portion of the spectrum, such as the first portion shown in the figure. Alternatively, the device may perform the conversion from one downmix domain to another for the full band, i.e., for all bands b1 to b6 exemplarily shown in Figure 9. This portion may be any portion of the signal, such as a spectral portion, a time portion, such as a time block or frame, or any other portion of the signal. It may be a time portion, such as a block or frame, or any portion of the signal.
[0027] FIG. 2 illustrates an embodiment in which a first downmix scheme relies on a residual signal only or on a residual signal and parametric information. FIG. 2 includes an input interface 10, which receives an encoded multichannel signal including an encoded core signal and an encoded side information part. The core signal is decoded by a core decoder 20 to provide an input downmix representation without side information. Furthermore, the side information part from the encoded multichannel signal is provided and processed by a side information decoder 30 in the input interface, which provides a residual signal or a residual signal and parameters, as shown at 210 in FIG. 2. Both the residual data and the input downmix corresponding to the decoded core signal are input to an upmixer 200, which generates an upmix signal having a first channel and a second channel, where the data of the first channel and the second channel are high-quality audio data. 2. This is because high-quality audio data is generated not only by a core signal and some kind of passive upmixing, but also using residual data or residual data and parameters, i.e., all data available from the encoded multi-channel signal. The output of the upmixer 200 is downmixed by the downmixer 300, for example using an active downmix, or in general using a downmix scheme that does not generate a residual signal or does not generate parameters but generates an energy-compensated downmix or a mono signal, i.e., a downmix scheme that does not suffer from energy fluctuations that are usually significant when only passive downmixing is performed, as is the case for the core signal generated by the core decoder 20 of FIG. 2. The output of the downmixer 300 is forwarded, for example, to a renderer for rendering a mono signal, or, for example, to the output stage 500 illustrated in FIG. 1.
[0028] Figure 3 shows a further embodiment, in which, with reference again to Figure 9, a first portion is available with a first downmix scheme, such as a downmix scheme with residual data, and there is a second spectral portion available with a second downmix scheme, for example without residual data, i.e. generated by an active downmix using downmix weights derived based on energy considerations, to counteract fluctuations that would occur if a passive downmix were applied.
[0029] The first part of the downmix representation is up-sampled according to a first downmix scheme. The first part is input to the upmixer 200, which performs the downmixing, and the first part is forwarded to the downmixer 300, which in turn performs the downmixing with a second downmix scheme, as described with reference to FIG. 1 or 2. The second part shown in FIG. 3 can be in, for example, the second downmix scheme from the downmix scheme of the part input to the upmixer 200 or the second downmix scheme output by the downmixer 300, but can also be in a third, i.e., any other, downmix scheme. If the downmix domain of the second part and the output of the downmixer 300 are the same, the second part processor 600 is not required at all. Instead, the second part can be forwarded to the combiner 400, which combines the first and second parts, which currently match in terms of downmix scheme. However, if the second part is in the downmix domain, i.e., if the output of the downmixer 300 has an underlying downmix scheme different from the available downmix schemes, the second part processor 600 is provided. Typically, the second partial processor 600 also comprises an upmixer for upmixing the second part in a third downmix scheme, and the second partial processor 600 further comprises a downmixer for downmixing the upmixer representation into the same downmix domain as that available from the downmixer 300, i.e., using the same downmix scheme. The second partial processor 600 can be implemented using the upmixer 200 followed by the downmixer 300, so as to obtain a perfect match of the data input to the combiner 400. The combiner 400 preferably outputs a spectral representation of the mono output downmix representation that has been converted into the time domain by a spectral-to-time transformer such as a filter bank, IDFT, IMDCT, etc. Alternatively, the combiner 400 is configured to combine the individual inputs into individual time-domain signals, which are combined in the time domain to obtain the time-domain mono output downmix representation.
[0030] Figure 4 includes an input interface that may include a first time-to-spectral transformer 100, such as a DFT block as shown in Figure 4, and a second time-to-spectral transformer 120, such as a second DFT block in Figure 4. The first block 100 outputs a decoded core signal, for example as output by the core decoder 20 in Figure 2. 9. The second time-to-spectral converter 120 is further configured to convert a decoded residual signal, e.g., as output by the side information decoder 30 of FIG. 2, into a spectral representation as illustrated at 210a. Further, the second time-to-spectral converter 120 is further configured to convert a decoded residual signal, e.g., as output by the side information decoder 30 of FIG. 2, into a spectral representation as illustrated at 210a. Furthermore, line 210b illustrates additional optionally provided parametric data, e.g., a side gain, which is also output by the side information decoder 30 of FIG. 2. The upmixer 200 of FIG. 4 outputs a left channel (upmixed 9. The low-band upmix at the output of block 200 is then preferably input to a downmixer 300 which performs an active downmix, providing a low-band representation for the three bands b1, b2, b3 shown exemplarily in FIG. The band downmix is in the same domain as the high-band downmix already generated by the DFT block 100. The high-band output of block 100 is band b4, in the example of FIG. b5, b6 correspond to the downmix representations of b1, b2, b3, b4, b5, b6. Here, at the input to combiner 400, shown in Figure 4 as IDFT 400, the low-band and high-band representations of the downmix are in the same "downmix domain" and have been produced by the same downmix scheme. The low-band and high-band of the harmonic downmix representation can now be combined and preferably transformed into the time domain to provide a mono output signal at the output of block 400.
[0031] Most parametric stereo schemes, such as those described in [8], require a single It is built around the idea of transmitting only the downmixed channels and recreating the stereo image via side parameters. This downmixing at the encoder side is done actively by dynamically calculating weights for both channels in the DFT domain [7]. These weights are calculated band-by-band using the energy of each of the two channels and their cross-correlation. The target energy to be preserved in the downmix is equal to the energy of the phase-rotated intermediate channel.
[0032] TIFF0007757360000002.tif17169
[0033] where L and R represent the left and right channels. Based on this target energy, the channel weights for each band b are calculated as follows:
[0034] TIFF0007757360000003.tif37149
[0035] TIFF0007757360000004.tif28149
[0036] TIFF0007757360000005.tif26148
[0037] TIFF0007757360000006.tif51169TIFF0007757360000007.tif22154
[0038] TIFF0007757360000008.tif24170
[0039] If the stereo processing in such a system is entirely parameter-dependent and the active downmix described is performed on the entire spectrum, then a mono signal that avoids the problems of passive downmix and meets the specified quality requirements is already available after core decoding. This means that in most cases it is sufficient to skip the stereo processing in the decoder altogether and output the signal without entering the DFT domain.
[0040] However, for higher bit rates, this type of system also supports the coding of a residual signal for the lower spectral bands. The residual signal can be seen as a side signal obtained by MS conversion of these lowest bands, while the core signal is a complementary intermediate signal, essentially a passive downmix of left and right. To make the side signal as small as possible, a side gain calculated for each band is used to compensate for the interaural level difference (ILD) between the channels.
[0041] TIFF0007757360000009.tif23168
[0042] TIFF0007757360000010.tif17157
[0043] TIFF0007757360000011.tif23142
[0044] TIFF0007757360000012.tif25150
[0045] The full-band signal input to the core coder is a mixture of a passive downmix of the low band and an active downmix of the high band. Listening tests have shown that there are perceptual problems when such a mixed signal is played back. Therefore, a method is needed to harmonize the different signal parts.
[0046] TIFF0007757360000013.tif52170
[0047] TIFF0007757360000014.tif33169
[0048] Then, an active downmix is applied as described above, but with weights calculated from the upmixed decoded spectra L and R. The low band is combined with the already actively downmixed high band to create a harmonic signal that is brought back to the time domain via an IDFT.
[0049] Figure 6 shows an embodiment of a multi-channel decoder for stereo output. The multi-channel decoder includes elements from Figure 4 that are designated by the same reference numerals. Furthermore, the stereo multi-channel decoder includes, as an embodiment of a multi-channel decoder, a second upmixer 220 for upmixing the high-band downmix, i.e., the second portion, into a second upmix representation, e.g., consisting of a left channel and a right channel, for stereo output. In another implementation of the multi-channel decoder, if there are more than two output channels, e.g., three or more output channels, then upmixer 200 as well as upmixer 220 will generate correspondingly more output channels than just the left and right channels.
[0050] Furthermore, a second combiner 420 is shown in Figure 6 for a multi-channel decoder, i.e., for the stereo decoder shown. In the case of more than two outputs, there would be an additional combiner for the third output channel, another combiner for the fourth output channel, etc. However, in contrast to Figure 6, the downmixer 300 of Figure 4 is not required for a multi-channel output.
[0051] Figure 7 shows a preferred embodiment of a switchable multi-channel decoder, which can be switched between mono and stereo / multi-channel output modes by the action of a controller 700. Furthermore, in contrast to Figure 6, the multi-channel decoder additionally comprises a downmixer 300 as already described with respect to Figure 4 or other Figures. Furthermore, in a switchable implementation, one option is to use two separate switches S1, S2 However, the switching function shown in the lower part of Fig. 7 can also be implemented by other switching means, such as a compound switch or two or more switches. Generally, switch S1 is configured to operate in mono output mode, bypassing the second upmixer 220, also denoted "upmix high". Furthermore, the second switch S2 is controlled by a second control signal CTRL2 to bypass the second upmixer 220, also denoted "upmix low" in Fig. 7. 6 is configured to provide the output of the combiner 200 to the active downmix 300. Furthermore, in mono output mode, only a single combiner 400 is required to generate a single mono output signal, so the upmix high block 220 described with respect to FIG. R The second combiner 42 is marked 0 is also inactive.
[0052] Conversely, in a stereo output mode, or generally a multi-channel output mode, the controller 700 activates the first switch via the control signal CTRL1 to activate the first The output of the time-to-frequency converter 100 is configured to be provided to a second upmixer 220, shown as "upmix high" in FIG. 7. Activation of the switch S1 activates the second combiner 220. Furthermore, the controller 700 is configured to control a second switch S2 720 so that the output of the block 200 is not input to the active downmixer 300, and the downmixer 300 is bypassed. The left channel (low-band) portion of the output of the block 200 is forwarded as the low-band portion for the combiner 400, and the right channel low-band portion at the output of the block 200 is forwarded to the low-band input of the second combiner 420, as illustrated in FIG. 7. Furthermore, in the stereo / multi-channel output mode, the downmix 300 is inactive.
[0053] 8a shows a flowchart of an embodiment used in the downmixer 300 to perform an active downmix. In step 800, weights w are calculated based on the target energy. R and w L is calculated, which is the weight for the right channel, w R and left channel Weight for Nel L This is done band by band so that for each band:
[0054] In block 820, weights are applied to the upmixed signal over the entire band of the signal under consideration, or only in the corresponding part per spectral bin. For this purpose, block 820 receives the signal or bins or spectral values in the spectral domain (complex numbers). Following the application of the weights and in particular the addition of the weighted values to obtain the downmix, a transformation 840 to the time domain is performed. Depending on whether only a part or the entire band is processed in block 820, the transformation to the time domain is performed without the other parts, or in particular together with the other parts, in the case of a harmonized downmix, for example as shown and discussed with reference to FIG. 3 or FIG. 4.
[0055] Figure 8b shows a preferred embodiment of the function performed in block 800 of Figure 8a. In particular, the weights w for each band are R and w L To calculate (b), an amplitude-related measure for L is calculated for the band. For this purpose, the individual spectral lines for the left channel, i.e., the individual spectral lines for the left channel output by block 200 of any of FIGS. 1-7, are input. In block 804, the same procedure is performed for the second or right channel of the same band b. Furthermore, in block 806, another amplitude-related measure is calculated for a linear combination of L and R for band b. In block 806, again, the spectral values of the first channel L and the second channel R are required for the band under consideration. In block 808, a measure of cross-correlation between the left and right channels, or more generally, between the first and second channels, is calculated for the corresponding band b. For this purpose, once again, the spectral values of the first and second channels at measure e are required for the corresponding band.
[0056] TIFF0007757360000015.tif38170
[0057] The same applies to the amplitude-related index calculated in block 804 or the amplitude-related index calculated in block 806 .
[0058] Furthermore, for the cross-correlation index calculated in block 808, the corresponding mathematical equations illustrated previously also rely on calculating the square and square root of the dot product. However, it is also possible to use other exponents different from 2 for the dot product, such as an exponent equal to 3 corresponding to the loudness region, or an exponent greater than 1. At the same time, instead of the square root, other exponents different from 1 / 2 can be used, such as 1 / 3, or in general any exponent between 0 and 1.
[0059] Furthermore, block 810 calculates w based on the three amplitude-related measures and the cross-correlation measure. R and wL It is shown that the target energy is preserved by downmixing. and has been shown to be equal to the energy of the phase-rotated intermediate channel, but w R and w L The calculation of the downmix signal and the calculation of the real downmix signal require rotations with such angles. 4. However, it is not necessary that the L and R cross-correlation measures be actually performed. Instead, if no actual rotation by the rotation angle φ is performed, all that is required is the calculation of the L and R cross-correlation measures in the corresponding band b. While the above embodiments have shown the use of the phase-rotated mid-channel energy as the target energy, other target energies may be used, or no phase rotation may be performed at all. With regard to other target energies, these target energies are those that cause the energy of the downmix signal generated by downmix 300 to vary less with respect to the same signal than the energy of a passive downmix, for example, as the basis for the decoded core signal input to block 100 in FIG. 4.
[0060] 9 shows a general representation of a spectrum, with respect to an input downmix representation, showing a first portion of the low band provided as a downmix containing residual data, and with respect to the input downmix representation, showing a second portion provided by a downmix generated using weights as previously described with respect to FIGS. 8a and 8b. Although FIG. 9 illustrates only six bands, three for the first portion and three for the second portion, and while FIG. 9 illustrates specific bandwidths increasing from the low band to the high band, the specific number, specific bandwidths, and separation of the spectrum into the first and second portions are merely exemplary. In a real scenario, there would be a significantly higher number of bands, and furthermore, the first portion containing the residual signal would be less than 50% of the number of bands b.
[0061] Preferably, the time-to-spectral converters 100, 120 and combiners 400, 420 of Figures 4, 6 and 7 are implemented as DFT or IDFT blocks, preferably implementing FFT or IFFT algorithms. The processing of successive decoded signals input to blocks 100, 120 involves block-wise processing where overlapping blocks are formed, analysis filtered, transformed into the spectral domain, processed, synthesis filtered in combiners 400, 420 and combined again with 50% overlap. The synthesis-side 50% overlap combination is typically performed by an overlap-add operation, preferably with a crossfade from one block to the other, where the crossfade weights are already included in the analysis / synthesis window. However, if this is not the case, the actual crossfade occurs at the output of block 400 (e.g.) or 420 (e.g.) in FIG. 7 or FIG. 6, such that each time-domain output sample of either the mono output signal or the left output signal or the right output signal is generated by the addition of two values from two different blocks. For overlaps of more than 50%, overlaps between three or correspondingly more blocks can be performed as well.
[0062] The overlap process is also used when performing a time-spectral transformation on one side and a spectral-to-time transformation on the other side, for example, using a modified discrete cosine transform. On the spectral-to-time transform side, an overlap-add process is performed, where each output time-domain sample is obtained by summing corresponding time-domain samples from two (or more) different IMDCT blocks.
[0063] Preferably, the harmonization of the downmix scheme is performed entirely in the spectral domain, as shown in Figures 4, 6, and 7. As shown in Figure 7, no additional time-to-spectral or spectral-to-time conversions are required when switching from mono to stereo or from stereo to mono. Only the spectral domain data needs to be manipulated by the downmixer 300 in the case of mono output mode, or by the second upmixer 220 (upmix high) in the case of stereo output mode. The overall processing delay is the same for either mono or stereo output, which is also an important advantage, as subsequent or preceding processing operations do not need to be aware of whether there is a mono or stereo output signal.
[0064] In a preferred embodiment, artifacts and spectral loudness imbalances resulting from different downmix methods in different spectral bands of the system's decoded core signal are removed, as described in [8], without the additional delay and significantly higher complexity introduced by a dedicated post-processing stage.
[0065] In one aspect, the embodiment provides for upmixing of one or more spectral or temporal parts of a mono signal, downmixed using one or more downmix methods to harmonize all spectral or temporal parts of the signal, followed by downmixing at a decoder.
[0066] In one aspect, the present invention provides for decoder-side stereo to mono downmix harmonization.
[0067] In one embodiment, the output downmix is for a playback device that receives the downmix included in the output representation and feeds this downmix of the output representation to a digital-to-analog converter, where the analog downmix signal is rendered by one or more loudspeakers included in the playback device, which may be a mono device such as a mobile phone, a tablet, a digital watch, a Bluetooth speaker, etc.
[0068] It should be mentioned here that all alternatives or aspects as described above and all aspects defined by the following independent claims can be used individually, i.e., without other alternatives or objects other than the contemplated alternatives, objects, or independent claims. However, in other embodiments, two or more alternatives or aspects or independent claims can be combined with each other, and in other embodiments, all aspects or alternatives and all independent claims can be combined with each other.
[0069] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or function of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block, item, or function of a corresponding apparatus.
[0070] Depending on particular implementation requirements, embodiments of the present invention can be implemented in hardware or in software. Implementation can be performed using a digital storage medium, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM or flash memory, having electronically readable control signals stored thereon and cooperating (or capable of cooperating) with a programmable computer system so that the respective methods are performed.
[0071] Some embodiments of the present invention comprise a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.
[0072] Generally, embodiments of the present invention may be implemented as a computer program product with program code operable to perform one of the methods of the present invention when the computer program product runs on a computer. The program code may for example be stored on a machine-readable carrier.
[0073] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier or a non-transitory storage medium.
[0074] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0075] A further embodiment of the inventive method is, therefore, a data carrier (or digital storage medium or computer readable medium) comprising, recorded on it, a computer program for performing one of the methods described herein.
[0076] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be arranged to be transmitted by a data communication connection, for example the Internet.
[0077] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0078] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0079] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.
[0080] The above-described embodiments merely illustrate the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. It is therefore intended that the present invention be limited only by the scope of the impending claims, and not by the specific details expressed by the manner in which the embodiments herein have been described and illustrated.
[0081] reference [1] ITU-R BS.775-2, Multichannel Stereophonic Sound System With And Without Accompanying Picture, 07 / 2006. [2] F. Baumgarte, C. Faller und P. Kroon, "Audio Coder Enhancement using Scalable Binaural Cue Coding with Equalized Mixing," in 116th Convention of the AES, Berlin, 2004. [3] G. Stoll, J. Groh, M. Link, J. Deigmoller, B. Runow, M. Keil, R. Stoll, M. Stoll und C. Stoll, "Method for Generating a Downward-Compatible Sound Format". USA Patent US 2012 / 0 014 526, 2012. [4] M. Kim, E. Oh und H. Shim, "Stereo audio coding improved by phase parameters," in 129th Convention of the AES, San Francisco, 2010. [5] A. Adami, E. Habets und J. Herre, "Down-mixing using coherence suppression," in IEEE International Conference on Acoustics, Speech and Signal Processing, Florence, 2014. [6] ISO / IEC 23008-3:, Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio, 2019. [7] S. Bayer, C. Bors, J. Buthe, S. Disch, B. Edler, G. Fuchs, F. Ghido und M. Multrus, "DOWNMIXER AND METHOD FOR DOWNMIXING AT LEAST TWO CHANNELS AND MULTICHANNEL ENCODER AND MULTICHANNEL DECODER". Patent WO18086946, 17 05 2018. [8] S. Bayer, M. Dietz, S. Dohla, E. Fotopoulou, G. Fuchs, W. Jaegers, G. Markovic, M. Multrus, E. Ravelli und M. Schnell, " APPARATUS AND METHOD FOR ESTIMATING AN INTER-CHANNEL TIME DIFFERENCE". Patent WO17125563, 27 07 2017.
Claims
1. 1. An apparatus for generating an output downmix representation from an input downmix representation, wherein a first part of the input downmix representation is in accordance with a first downmix scheme and a second part of the input downmix representation is in accordance with a second downmix scheme, the apparatus comprising: an upmixer (200) for upmixing the first part of the input downmix representation using an upmix scheme corresponding to the first downmix scheme to obtain a first upmixed part; a downmixer (300) for downmixing the first upmixed part according to the second downmix scheme different from the first downmix scheme to obtain a first downmixed part representing the output downmix representation for the first part of the input downmix representation; An apparatus comprising:
2. a combiner (400) for combining the first downmixed part of the input downmix representation with the second part of the input downmix representation or a downmixed part derived from the second part of the input downmix representation to obtain the output downmix representation comprising a first output representation for the first part of the input downmix representation and a second output representation for the second part of the input downmix representation, wherein the first output representation for the first part of the input downmix representation and the second output representation for the second part of the input downmix representation are based on the second downmix scheme.
10. The apparatus of claim 1.
3. the first part of the input downmix representation is a first frequency band and the first downmix scheme is a downmix scheme that depends on a residual signal, the upmixer (200) is configured to perform an upmix using the residual signal as the upmix scheme, or the second downmix scheme is a fully parametric scheme.
3. The device according to claim 1 or claim 2.
4. the second portion of the input downmix representation is a second frequency band; the combiner (400) is configured to combine the first downmixed part of the input downmix representation and the second part of the input downmix representation to obtain the output downmix representation.
3. The apparatus of claim 2.
5. an audio decoder (10) for generating a decoded core signal for the first portion of the input downmix representation and a residual signal for the first portion of the input downmix representation, the upmixer (200) is configured to use the decoded core signal for the first part of the input downmix representation and the residual signal for the first part of the input downmix representation in the upmix scheme, the downmixer (300) is configured to receive the first upmixed portion, which comprises more channels than the input downmix representation; 10. The apparatus of claim 1.
6. the audio decoder (10) is configured to generate a decoded core signal for the second part of the input downmix representation and a residual signal for the first part of the input downmix representation, and a combiner (400) is configured to combine the first downmixed part and the decoded core signal for the second part of the input downmix representation.
6. The apparatus of claim 5.
7. 2. The apparatus of claim 1, further comprising: a time-to-spectral converter (100) for converting a time-domain input downmix representation of the first part of the input downmix representation into the spectral domain; and a spectral-to-time converter (400) for converting the first downmixed part and the second part of the input downmix representation into the time domain to obtain the output downmix representation, wherein the time-to-spectral converter (100) or the spectral-to-time converter (400) is configured to perform an overlap-add operation or a crossover operation from a previous time block to a subsequent time block.
8. further comprising an output interface (500) for outputting said output downmix representation to a rendering device, or further comprising a rendering device for rendering said output downmix representation as a mono playback signal, or the downmixer (300) is configured to apply as the second downmix scheme an active downmix scheme, an energy-saving downmix scheme, or a downmix scheme in which a target energy of the output downmix representation is a predetermined ratio to an energy of an intermediate channel derived from a first channel and a second channel, wherein at least one of the first channel and the second channel is phase-rotated before being summed to form the input downmix representation; 10. The apparatus of claim 1.
9. the time-to-spectral converter (100) is configured to convert a time-domain input down-mix representation of the second part of the input down-mix representation into the spectral domain, or The downmixer (300) is configured to apply, as an active downmix scheme, a downmix scheme in which a target energy of a downmix signal is a predetermined ratio to an energy of an intermediate channel derived from a first channel and a second channel, and the predetermined ratio indicates that the energy of the first channel and the energy of the second channel are equal, or that there is a deviation in the range of 3 dB with respect to the higher energy of the energy of the first channel and the energy of the second channel.
8. The apparatus of claim 7.
10. the first part of the input downmix representation is according to the first downmix scheme depending on a residual signal or a residual signal and parametric information, the upmixer (200) is configured to upmix the input downmix representation of the first part of the input downmix representation using the upmix scheme corresponding to the first downmix scheme and using the residual signal or the residual signal and the parametric information, respectively, to obtain the first upmixed part; the second downmix scheme is an active downmix scheme or a fully parametric downmix scheme to obtain the output downmix representation comprising at least one downmixed part.
10. The apparatus of claim 1.
11. The apparatus of claim 10, further comprising an output interface (500) for outputting the output downmix representation to a rendering device, or further comprising a rendering device for rendering the output downmix representation as a mono playback signal.
12. the downmixer (300) is configured to apply as the active downmix scheme an energy-saving downmix scheme or a downmix scheme in which the target energy of the output downmix representation is a predetermined ratio to the energy of an intermediate channel derived from a first channel and a second channel, and at least one of the first channel and the second channel is phase-rotated before being summed; 12. Apparatus according to claim 10 or 11.
13. the downmixer (300) is configured to perform the second downmix scheme; The second downmix scheme is Calculating (800) a first weight for a first channel and a second weight for a second channel for a spectral band of the first upmixed portion, the spectral band comprising a plurality of spectral lines; applying the first weights to the spectral lines of the spectral band of the first upmixed portion of the first channel and the second weights to the spectral lines of the spectral band of the first upmixed portion of the second channel, and summing the first weighted lines and the second weighted lines to obtain downmixed spectral lines in the spectral band; the apparatus is configured to transform (840) the downmixed spectral lines into the time domain to obtain time domain samples of the output downmix representation.
10. The apparatus of claim 1.
14. The apparatus of claim 13 , wherein the calculation of the first weights and the second weights is performed for each band using the energies of the first channel and the second channel and a target energy.
15. 15. The apparatus of claim 14, wherein the target energy is equal to the energy of a phase-rotated intermediate channel or is derived from the energies of the first channel and the second channel and from a correlation value between the first channel and the second channel.
16. The calculation of the first weight and the second weight includes, for a spectral band: Calculating (802) an amplitude-related metric for the first channel within the spectral band; Calculating (804) an amplitude-related metric for the second channel within the spectral band; Calculating (806) an amplitude-related metric for a linear combination of the first channel and the second channel within the spectral band; Calculating (808) a measure of cross-correlation between the first channel and the second channel within the spectral band; calculating (810) the first weight and the second weight using the amplitude-related measure for the first channel, the amplitude-related measure for the second channel, the amplitude-related measure for the linear combination, and the cross-correlation measure; 16. The apparatus of any one of claims 13 to 15, comprising:
17. The upmixer (200) is configured to perform the upmix scheme, the upmix scheme comprising: - calculating, from spectral lines of the spectral band of the first portion of the input downmix representation, first channel spectral lines for the spectral band of the first portion of the input downmix representation using prediction parameters for the spectral band and residual signal lines for the spectral band and a first calculation rule; - calculating second channel spectral lines for said spectral bands of said first portion of said input downmix representation from the spectral lines of said spectral bands of said first portion of said input downmix representation using prediction parameters for said spectral bands and residual signal lines for said spectral bands and a second calculation rule; Including, The apparatus of claim 1 , wherein the first calculation rule is different from the second calculation rule.
18. 18. The apparatus of claim 17, wherein the first calculation rule includes one of addition and subtraction, and the second calculation rule includes the other of the addition and the subtraction.
19. an input interface (100, 120) for providing an input downmix representation comprising a first part according to a first downmix scheme, a second part, and parametric data for said second part; a first time-to-spectral converter (100) for generating a first spectral representation of the first part of the input down-mix representation and a second spectral representation of the second part of the input down-mix representation, the second part of the input down-mix representation comprising spectral values for higher frequencies than the first part of the input down-mix representation; a second time-to-spectral transformer (120) for generating a spectral representation of a residual signal for said first portion of said input downmix representation; a first upmixer (200) for upmixing the spectral representation of the first part of the input downmix representation with the spectral representation of the residual signal and with an upmix scheme corresponding to the first downmix scheme to obtain a first upmixed part; a second upmixer (220) for upmixing the second part of the input downmix representation and the parametric data using a second upmix scheme corresponding to a second downmix scheme different from the first downmix scheme to obtain a second upmixed part; a downmixer (300) for downmixing the first upmixed part according to the second downmix scheme to obtain a first downmixed part in the spectral domain, the first downmixed part representing an output downmix representation for the first part of the input downmix representation; a first combiner (400) comprising a spectral-to-time converter for, in a multi-channel output mode, combining a first channel of the first upmixed part and a first channel of the second upmixed part and transforming the result of the combination into the time domain to obtain a first channel of a multi-channel output signal; a second combiner (420) for, in the multi-channel output mode, combining a second channel of the first upmixed portion with a second channel of the second upmixed portion and transforming the combining result into the time domain to obtain a second channel of the multi-channel output signal; a switch (710) connected between the first time-to-spectral converter (100) and the second upmixer (220); a controller (700) configured to control the switch (710) to connect the output of the first time-to-spectral converter (100) to the first combiner (400) in a mono output mode, or to connect the output of the first upmixer (200) to the input of the downmixer (300) while bypassing the second upmixer (220), or to control the switch (710) to connect the output of the first time-to-spectral converter (100) to the input of the second upmixer (220) in the multi-channel output mode; A multi-channel decoder comprising:
20. an input interface (100, 120) for providing an input downmix representation comprising a first part according to a first downmix scheme, a second part, and parametric data for said second part; a first upmixer (200) for upmixing the first part of the input downmix representation using an upmix scheme corresponding to the first downmix scheme to obtain a first upmixed part; a second upmixer (220) for upmixing the second part of the input downmix representation and the parametric data using a second upmix scheme corresponding to a second downmix scheme different from the first downmix scheme to obtain a second upmixed part; a downmixer (300) for downmixing the first upmixed part according to the second downmix scheme to obtain a first downmixed part representing an output downmix representation for the first part of the input downmix representation; a first combiner (400) for, in a multi-channel output mode, combining a first channel of the first upmixed part and a first channel of the second upmixed part and transforming the combining result into the time domain to obtain a first channel of a multi-channel output signal; a second combiner (420) for, in the multi-channel output mode, combining a second channel of the first upmixed portion with a second channel of the second upmixed portion and transforming the combining result into the time domain to obtain a second channel of the multi-channel output signal; a switch (720) connected between the first upmixer (200) and the downmixer (300); a controller (700) configured to control the switch (720) to connect the output of the first upmixer (200) to the input of the downmixer (300) in a mono output mode, and to control the switch (720) to connect the output of the first upmixer (200) to the input of the second combiner (420) or to bypass the downmixer (300) in the multi-channel output mode; A multi-channel decoder comprising:
21. 1. A method for generating an output downmix representation from an input downmix representation, wherein a first part of the input downmix representation is in accordance with a first downmix scheme and a second part of the input downmix representation is in accordance with a second downmix scheme, the method comprising: - upmixing the input downmix representation of the first part of the input downmix representation with an upmix scheme corresponding to the first downmix scheme to obtain a first upmixed part; downmixing the first upmixed part according to the second downmix scheme, which is different from the first downmix scheme, to obtain a first downmixed part representing the output downmix representation of the first part of the input downmix representation; A method comprising:
22. providing an input downmix representation and parametric data for at least a second part of said input downmix representation; 22. The method of claim 21 ; A multi-channel decoding method comprising: The multi-channel decoding method comprises the steps of: upmixing the first part of the input downmix representation according to the upmix scheme corresponding to the first downmix scheme to obtain the first upmixed part; and / or upmixing the second part of the input downmix representation and the parametric data using a second upmix scheme corresponding to the second downmix scheme to obtain a second upmixed part; combining the first upmixed portion and the second upmixed portion to obtain a multi-channel output signal; The multi-channel decoding method further comprises:
23. providing an input downmix representation comprising a first part according to a first downmix scheme, a second part and parametric data for said second part; generating, by a first time-to-spectral converter (100), a first spectral representation of the first portion of the input downmix representation and a second spectral representation of the second portion of the input downmix representation, the second portion of the input downmix representation comprising spectral values for higher frequencies than the first portion of the input downmix representation; generating a spectral representation of a residual signal for the first portion of the input downmix representation; - upmixing, by a first upmixer (200), the spectral representation of the first part of the input downmix representation with the spectral representation of the residual signal and with an upmix scheme corresponding to the first downmix scheme to obtain a first upmixed part; - upmixing, by a second upmixer (220), the second part of the input downmix representation and the parametric data using a second upmix scheme corresponding to a second downmix scheme different from the first downmix scheme to obtain a second upmixed part; - downmixing, by a downmixer (300), the first upmixed part according to the second downmix scheme to obtain a first downmixed part in the spectral domain, which represents an output downmix representation for the first part of the input downmix representation; combining, by a combiner (400), in a multi-channel output mode, a first channel of the first upmixed part and a first channel of the second upmixed part and transforming the result of the combination into the time domain to obtain a first channel of a multi-channel output signal; in the multi-channel output mode, combining a second channel of the first upmixed part and a second channel of the second upmixed part and transforming the combination result into the time domain to obtain a second channel of the multi-channel output signal; using a switch (710) connected between the first time-to-spectral converter (100) and the second upmixer (220); controlling the switch (710) to connect the output of the first time-to-spectral converter (100) to the combiner (400) in a mono output mode, or to connect the output of the first upmixer (200) to the input of the downmixer (300) while bypassing the second upmixer (220), or to connect the output of the first time-to-spectral converter (100) to the input of the second upmixer (220) in the multi-channel output mode; A multi-channel decoding method comprising:
24. providing an input downmix representation comprising a first part according to a first downmix scheme, a second part and parametric data for said second part; - upmixing, by an upmixer (200), the first part of the input downmix representation using an upmix scheme corresponding to the first downmix scheme to obtain a first upmixed part; - upmixing the second part of the input downmix representation and the parametric data using a second upmix scheme corresponding to a second downmix scheme different from the first downmix scheme to obtain a second upmixed part; - downmixing, by a downmixer (300), the first upmixed part according to the second downmix scheme to obtain a first downmixed part representing an output downmix representation for the first part of the input downmix representation; in a multi-channel output mode, combining a first channel of the first upmixed part and a first channel of the second upmixed part and transforming the combination result into the time domain to obtain a first channel of a multi-channel output signal; combining, by a combiner (420), in the multi-channel output mode, a second channel of the first upmixed part and a second channel of the second upmixed part and transforming the combining result into the time domain to obtain a second channel of the multi-channel output signal; using a switch (720) connected between the upmixer (200) and the downmixer (300); controlling the switch (720) to connect the output of the upmixer (200) to the input of the downmixer (300) in a mono output mode, and controlling the switch (720) to connect the output of the upmixer (200) to the input of the combiner (420) or to bypass the downmixer (300) in a multi-channel output mode; A multi-channel decoding method comprising:
25. 25. A computer program for carrying out the method of any one of claims 21 to 24 when the computer program is run on a computer or processor.
Citation Information
Patent Citations
audio channel mixing
JP2001518267A
multichannel encoder
JP2007531913A
Renderer-controlled spatial upmix
JP2016527804A
Audio processing system
JP2017017749A