Apparatus and methods for encoding or decoding multichannel audio signals using frame-controlled synchronization.

CN117238300BActive Publication Date: 2026-08-11FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-01-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,如例如在MPEG USAC中,参数立体声未针对低延迟特定设计,且整个系统示出非常高的算法延迟

Benefits of technology

[0026]因此,本发明的优点是提供一种新的立体声编码方案,其比现有的立体声编码方案更适于立体声语音的转换。本发明的实施例提供了一种新架构,用于实现低延迟立体声编解码器,并在切换式音频编解码器内集成针对语音核心编码器和基于MDCT的核心编码器在频域中执行的共同立体声工具。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117238300B_ABST
    Figure CN117238300B_ABST
Patent Text Reader

Abstract

The multichannel audio signal is encoded using a time-to-spectrum converter for converting a sequence of blocks of sampled values ​​into a sequence of blocks of spectral values, a multichannel processor for applying joint multichannel processing to the blocks of spectral values ​​to obtain at least one result sequence of blocks, a spectrum-to-time converter for converting the result sequence of blocks of spectral values ​​into a time-domain representation of an output sequence of blocks including sampled values, and a core encoder for encoding the output sequence of blocks of sampled values ​​to obtain an encoded multichannel signal, wherein the core encoder operates with first frame control, and wherein the time-to-spectrum converter or the spectrum-to-time converter operates with second frame control synchronized with the first frame control.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Fraunhofer Institute for the Promotion of Applied Research, filed on January 20, 2017, with application number 201780019674.8, entitled "Apparatus and Method for Encoding or Decoding Multichannel Audio Signals Using Frame Control Synchronization". Technical Field

[0002] This application relates to stereo processing or, more generally, to multichannel processing, wherein a multichannel signal has two channels (such as a left channel and a right channel in the case of a stereo signal), or more than two channels (such as three, four, five, or any other number of channels). Background Technology

[0003] Compared to the storage and broadcasting of stereo music, stereo speech, especially conversational stereo speech, has received far less scientific attention. In fact, mono transmission is still the primary method of voice communication. However, with the increase in network bandwidth and capacity, stereo-based communication is expected to become more widespread and deliver a better listening experience.

[0004] For efficient storage or broadcasting, efficient coding of stereo audio materials has been extensively studied in perceptual audio coding of music. At high bit rates where waveform preservation is crucial, sum-difference stereo, known as center / side (M / S) stereo, has long been employed. For low bit rates, intensity stereo and, more recently, parametric stereo coding have been introduced. The latest technologies, such as HeAACv2 and MPEG USAC, are used in various standards. These produce downmixing of two-channel signals and correlate with compact spatial side information.

[0005] Joint stereo coding is typically built on high-frequency resolution (i.e., low temporal resolution, the time-frequency transformation of the signal), and is therefore incompatible with the low latency and time-domain processing performed in most speech encoders. Furthermore, the resulting bit rate is typically high.

[0006] On the other hand, parametric stereo employs an additional filter bank at the front end of the encoder as a preprocessor and an additional filter bank at the back end of the decoder as a postprocessor. Therefore, parametric stereo can be used with conventional speech encoders such as ACELP, as is done in MPEG USAC. Furthermore, the parameterization of auditory scenes can be achieved with minimal side information, which is suitable for low bit rates. However, as in, for example, MPEG USAC, parametric stereo is not specifically designed for low latency and does not deliver consistent quality across different conversational contexts. In the conventional parametric representation of spatial scenes, the width of the stereo image is artificially replicated by the decorrecorder on both synthesized channels and controlled by the inter-channel coherence (IC) parameter calculated and transmitted by the encoder. For most stereo speech, this method of widening the stereo image is unsuitable for recreating the natural environment of speech as fairly direct sound, because fairly direct sound is generated by a single source located in a specific location within space (occasionally with some reverberation from the room). In contrast, musical instruments have a much more natural width than speech, which can be better simulated by decorrecording the channels.

[0007] Problems also arise when recording speech using non-overlapping microphones, such as in A / B configurations where the microphones are far apart or used for binaural recording or rendering. These scenarios are anticipated for capturing speech in teleconferences or for creating virtual auditory scenes with distant speakers in a multipoint control unit (MCU). The arrival times of the signals differ from those of recordings made with overlapping microphones, such as XY (intensity recording) or MS (center-side recording). The coherence calculations of the two untime-aligned channels can then be incorrectly estimated, leading to failures in the synthesis of artificial environments.

[0008] Prior art references for stereo processing are U.S. patents with patent numbers 5,434,948 or 8,811,621.

[0009] Document WO 2006 / 089570 A1 discloses a near-transparent or transparent multichannel encoder / decoder scheme. The multichannel encoder / decoder scheme additionally generates a waveform-type residual signal. This residual signal, along with one or more multichannel parameters, is transmitted to the decoder. In contrast to purely parametric multichannel decoders, the enhanced decoder produces multichannel output signals with improved output quality due to the additional residual signal. On the encoder side, both the left and right channels are filtered by an analysis filter bank. Then, for each sub-band signal, alignment and gain values ​​are calculated for the sub-band. This alignment is then performed before further processing. On the decoder side, dealignment and gain processing are performed, and the corresponding signals are then combined by a synthesis filter to produce the decoded left and right signals.

[0010] On the other hand, parametric stereo employs an additional filter bank, which functions as a preprocessor in the front end of the encoder and as a postprocessor in the back end of the decoder. Therefore, parametric stereo can be used with conventional speech encoders such as ACELP, as is done in MPEG USAC. Furthermore, the parameterization of auditory scenes can be achieved with a minimal amount of side information, which is suitable for low bit rates. However, as in, for example, MPEG USAC, parametric stereo is not specifically designed for low latency, and the entire system exhibits very high algorithmic latency. Summary of the Invention

[0011] The purpose of this invention is to provide an improved concept for multichannel encoding / decoding that is efficient and positioned to achieve low latency.

[0012] This objective is achieved by means of the apparatus for encoding multichannel signals, the method for encoding multichannel signals, the apparatus for decoding encoded multichannel signals, the method for decoding encoded multichannel signals, or the computer program described below.

[0013] This invention is based on the discovery that at least a portion, and preferably all portions, of multichannel processing (i.e., joint multichannel processing) is performed in the spectral domain. Specifically, the downmixing operation of joint multichannel processing is preferably performed in the spectral domain, and additionally, timing and phase alignment operations or even processes for analyzing parameters of joint stereo / joint multichannel processing are performed. Furthermore, frame control for the core encoder and synchronization of stereo processing operating in the spectral domain are performed.

[0014] The core encoder is configured to operate according to the first frame control to provide a frame sequence, wherein the frames are defined by a start frame boundary and an end frame boundary, and the time-to-spectrum converter or spectrum-to-time converter is configured to operate according to the second frame control synchronized with the first frame control, wherein the start frame boundary or end frame boundary of each frame of the frame sequence has a predetermined relationship with the start or end time of the overlapping portion of the window used by the time-to-spectrum converter (1000) for each block of the sequence of sampled values ​​or by the spectrum-to-time converter for each block of the output sequence of sampled values.

[0015] In this invention, the core encoder of the multi-channel encoder is configured to operate according to framing control, and the time-to-spectrum converter, spectrum-to-time converter, and resampler of the stereo post-processor are also configured to operate according to additional framing control synchronized with the framing control of the core encoder. Synchronization is performed in such a manner that the start or end frame boundary of each frame in the frame sequence of the core encoder has a predetermined relationship with the start or end time of the overlapping portion of the window used by the time-to-spectrum converter or spectrum-to-time converter for each block of the sequence of blocks of sampled values ​​or each block of the resampled sequence of blocks of spectral values. Therefore, it is ensured that subsequent framing operations operate synchronously with each other.

[0016] In a further embodiment, the look-ahead operation with a look-ahead portion is performed by the core encoder. In this embodiment, preferably, the look-ahead portion is also used by the analysis window of the time-to-spectrum converter, wherein an overlapping portion of the analysis window is used, the time length of which is less than or equal to the time length of the look-ahead portion.

[0017] Therefore, time-spectrum analysis of the stereo preprocessor cannot be achieved without any additional algorithmic delay by making the overlap between the leading portion of the core encoder and the analysis window equal to each other, or by making the overlap even smaller than the leading portion of the core encoder. To ensure that this leading portion of the window does not excessively affect the leading functionality of the core encoder, it is preferable to correct this portion using the inverse of the analysis window function.

[0018] To ensure good stability, the square root of the sine window shape is used instead of the sine window shape for the analysis window, and a 1.5-power sine synthesis window is used for the purpose of synthesis windowing before performing the overlap operation at the output of the spectrum-time converter. Therefore, it is ensured that the correction function is assumed to be a value that is reduced in magnitude compared to the correction function as the inverse of the sine function.

[0019] Preferably, spectral domain resampling is performed after or even before multichannel processing to provide the output signal from another spectrum-to-time converter, which is already at the output sampling rate required by the subsequently connected core encoder. However, the inventive process of synchronizing frame control of the core encoder with the spectrum-to-time or time-to-spectrum converter can also be applied to scenarios where no spectral domain resampling is performed.

[0020] On the decoder side, at least one operation for generating the first and second channel signals from the downmixed signal in the spectral domain is preferably performed again, and preferably, the entire inverse multichannel processing is even performed in the spectral domain. Furthermore, a time-to-spectrum converter is provided to convert the core-decoded signal into a spectral domain representation, and in the frequency domain, inverse multichannel processing is performed.

[0021] The core decoder is configured to operate according to a first frame control to provide a frame sequence, wherein frames are defined by start and end frame boundaries. A time-to-spectrum converter or a spectrum-to-time converter is configured to operate according to a second frame control synchronized with the first frame control. Specifically, the time-to-spectrum converter or spectrum-to-time converter is configured to operate according to a second frame control synchronized with the first frame control, wherein the start or end frame boundary of each frame in the frame sequence has a predetermined relationship with the start or end time of the overlapping portion of the window used by the time-to-spectrum converter for each block of the sequence of sampled values, or by the spectrum-to-time converter for each block of at least two output sequences of sampled values.

[0022] Of course, it is preferable to use the same analysis and synthesis window shape, as no correction is required. Alternatively, it is preferable to use a time gap on the decoder side, where a time gap exists between the end of the leading overlap portion of the analysis window of the time-to-spectrum converter on the decoder side and the end of the frame output by the core decoder on the multichannel decoder side. Therefore, the core decoder output samples within this time gap are not needed for the immediate analysis windowing performed by the stereo post-processor, but only for the processing / windowing of the next frame. This time gap can be implemented, for example, by using a non-overlapping portion typically in the middle of the analysis window, resulting in a shortening of the overlap portion. However, other alternatives for implementing this time gap can also be used, but implementing the time gap via a non-overlapping portion in the middle is preferred. Therefore, this time gap can be used for other core decoder operations or, preferably, for smoothing operations between switching events when the core decoder switches from a frequency domain to a time domain frame, or for any other smoothing operations that may be useful when parameter changes or coding characteristic variations have occurred.

[0023] In one embodiment, spectral domain resampling is performed before or after multichannel inverse processing, such that the final spectrum-to-time converter converts the spectral resampled signal to the time domain at an output sampling rate intended for use with the time-domain output signal.

[0024] Therefore, the embodiments allow for the complete avoidance of any computationally intensive time-domain resampling operations. Instead, multichannel processing is combined with resampling. In a preferred embodiment, spectral domain resampling is performed by truncating the spectrum in the case of undersampling, or by zero-padding the spectrum in the case of oversampling. These convenient operations (i.e., truncating the spectrum on the one hand or zero-padding the spectrum on the other, along with preferred additional scaling to account for certain normalization operations performed in spectral domain / time domain transformation algorithms such as DFT or FFT) accomplish spectral domain resampling operations in a very efficient and low-latency manner.

[0025] Furthermore, it has been found that at least a portion, or even the entire, of the joint stereo / joint multichannel processing on the encoder side and the corresponding inverse multichannel processing on the decoder side is suitable for execution in the frequency domain. This is effective not only as a minimum of the joint multichannel processing on the encoder side for downmixing operations, or as a minimum of the inverse multichannel processing on the decoder side for upmixing operations. Conversely, it is even possible to perform stereo scene analysis and time / phase alignment on the encoder side or phase and time dealignment on the decoder side in the spectral domain. This also applies to sidechannel encoding preferably performed on the encoder side or sidechannel synthesis and use on the decoder side for generating two decoded output channels.

[0026] Therefore, the advantage of this invention is that it provides a novel stereo coding scheme that is more suitable for stereo speech conversion than existing stereo coding schemes. Embodiments of this invention provide a new architecture for implementing a low-latency stereo codec and integrating common stereo tools for both the speech core encoder and the MDCT-based core encoder, executed in the frequency domain, within a switchable audio codec.

[0027] Embodiments of the present invention relate to a mixing method that blends elements from conventional M / S stereo or parametric stereo. The embodiments utilize aspects and tools from joint stereo coding as well as other aspects and tools from parametric stereo. More specifically, the embodiments employ additional time-frequency analysis and synthesis performed at the front end of the encoder and the back end of the decoder. Time-frequency decomposition and inverse transformation are achieved by employing filter banks or block transforms with complex values. From two-channel or multi-channel inputs, stereo or multi-channel processing combines and modifies the input channels to output an output known as the center and side signals (MS).

[0028] Embodiments of the present invention provide a solution for reducing algorithmic latency introduced by stereo modules, where the latency particularly stems from framing and windowing of their filter banks. It provides a multi-rate inverse transform for feeding switched encoders (such as 3GPP EVS) or encoders switching between speech encoders (such as ACELP) and general-purpose audio encoders (such as TCX) by generating the same stereo processed signal at different sampling rates. Furthermore, it provides different constraints and windowing for stereo processing suitable for low-latency and low-complexity systems. Additionally, embodiments provide a method for combining and resampling different decoded synthesis results in the spectral domain, wherein inverse stereo processing is also applied.

[0029] A preferred embodiment of the invention includes a multifunctional spectral domain resampler that not only generates a single spectral domain resampled block of spectral values, but also additionally generates additional resampled sequences of spectral values ​​corresponding to blocks of spectral values ​​at different higher or lower sampling rates.

[0030] Furthermore, the multichannel encoder is configured to additionally provide an output signal at the output of the spectrum-to-time converter, the output signal having the same sampling rate as the original first and second channel signals input to the time-to-spectrum converter on the encoder side. Therefore, in this embodiment, the multichannel encoder provides at least one output signal at the original input sampling rate, which is preferably used for MDCT-based encoding. Additionally, at least one output signal is provided at an intermediate sampling rate particularly useful for ACELP encoding, and additionally, an additional output signal is provided at a different output sampling rate, which is also useful for ACELP encoding but differs from the other output sampling rates.

[0031] These processes can be performed on the center signal or the side signal, or on two signals derived from the first and second channel signals of a multi-channel signal, wherein in the case of a stereo signal with only two channels (additionally two, e.g., low-frequency enhancement channels), the first signal can also be the left signal and the second signal can be the right signal. Attached Figure Description

[0032] The preferred embodiments of the invention will then be discussed in detail with reference to the accompanying drawings, in which:

[0033] Figure 1 This is a block diagram of an embodiment of a multi-channel encoder;

[0034] Figure 2 The illustration shows an example of spectral domain resampling;

[0035] Figures 3a-3c The diagram illustrates different alternatives for performing time / frequency or frequency / time transformations with different normalizations and corresponding scaling in the spectral domain;

[0036] Figure 3d The illustrations show different frequency resolutions and other frequency-related aspects for some embodiments;

[0037] Figure 4a A block diagram illustrating an embodiment of the encoder;

[0038] Figure 4b A block diagram of a corresponding embodiment of the decoder is shown;

[0039] Figure 5 The illustration shows a preferred embodiment of a multi-channel encoder;

[0040] Figure 6 A block diagram illustrating an embodiment of a multi-channel decoder;

[0041] Figure 7a The illustration shows another embodiment of a multi-channel decoder with a combiner;

[0042] Figure 7b The illustration also includes another embodiment of a multichannel decoder with a combiner (adder);

[0043] Figure 8a The diagram illustrates a table showing the different characteristics of windows used for several sampling rates;

[0044] Figure 8b The illustrations show different proposals / implementations of DFT filter banks as implementations of time-to-spectrum converters and spectrum-to-time converters;

[0045] Figure 8c The figure shows the sequences of two analysis windows of the DFT, with a time resolution of 10 ms;

[0046] Figure 9a The figure illustrates a schematic window opening of the encoder according to the first proposal / embodiment;

[0047] Figure 9b The diagram illustrates a schematic window opening of the decoder according to the first proposal / embodiment;

[0048] Figure 9c The diagram illustrates the windows at the encoder and decoder according to the first proposal / embodiment;

[0049] Figure 9d The illustration shows a preferred flowchart of the modified embodiment;

[0050] Figure 9e The diagram further illustrates the flowchart of the modified embodiment;

[0051] Figure 9f The diagram illustrates a flowchart for explaining an embodiment of the time-slot decoder side;

[0052] Figure 10aThe diagram illustrates a schematic window opening of the encoder according to the fourth proposal / embodiment;

[0053] Figure 10b The figure illustrates a schematic window of the decoder according to the fourth proposal / embodiment;

[0054] Figure 10c The diagram illustrates the windows at the encoder and decoder according to the fourth proposal / embodiment;

[0055] Figure 11a The diagram illustrates a schematic window opening of the encoder according to the fifth proposal / embodiment;

[0056] Figure 11b The diagram illustrates a schematic window opening of the decoder according to the fifth proposal / embodiment;

[0057] Figure 11c The diagram illustrates the windows at the encoder and decoder according to the fifth proposal / embodiment;

[0058] Figure 12 This is a block diagram of a preferred implementation of multi-channel processing with downmixing in a signal processor;

[0059] Figure 13 This is a preferred embodiment of inverse multichannel processing with upmixing operation within the signal processor;

[0060] Figure 14a The diagram illustrates a flowchart of the process performed in the encoding device to align the audio channels;

[0061] Figure 14b The illustration shows a preferred embodiment of the process performed in the frequency domain;

[0062] Figure 14c The illustration shows a preferred embodiment of a process performed in an apparatus for encoding using an analysis window with zero-padding portions and overlap ranges;

[0063] Figure 14d The illustration is a flowchart of a further process performed in an embodiment of the apparatus for encoding;

[0064] Figure 15a The illustration shows the process performed by an embodiment of a device for decoding and encoding multichannel signals;

[0065] Figure 15b The diagram illustrates a preferred implementation of a device for decoding in some aspects; and

[0066] Figure 15c The diagram illustrates the process performed in the context of broadband dealignment within an architecture for decoding encoded multichannel signals. Detailed Implementation

[0067] Figure 1The diagram illustrates an apparatus for encoding a multichannel signal comprising at least two channels 1001 and 1002. In a two-channel stereo scenario, the first channel 1001 is in the left channel, and the second channel 1002 can be the right channel. However, in a multichannel scenario, the first channel 1001 and the second channel 1002 can be any channel of the multichannel signal, such as, for example, one side being the left channel and the other side being the left surround channel, or one side being the right channel and the other side being the right surround channel. However, these channel pairings are merely examples, and other channel pairings may be applied as needed.

[0068] Figure 1 The multichannel encoder includes a time-to-spectrum converter for converting a sequence of blocks of sampled values ​​from at least two channels into a frequency domain representation at the output of the time-to-spectrum converter. Each frequency domain representation has a sequence of blocks of spectral values ​​for one of the at least two channels. Specifically, blocks of sampled values ​​from the first channel 1001 or the second channel 1002 have an associated input sampling rate, and blocks of spectral values ​​from the output sequence of the time-to-spectrum converter have spectral values ​​up to a maximum input frequency associated with the input sampling rate. Figure 1 In the illustrated embodiment, a time-to-spectrum converter is connected to a multichannel processor 1010. This multichannel processor is configured to apply joint multichannel processing to a sequence of blocks of spectral values ​​to obtain at least one resulting sequence of blocks of spectral values ​​that include information relating to at least two channels. A typical multichannel processing operation is downmixing, but preferred multichannel operations include additional processes described later.

[0069] The core encoder 1040 is configured to operate according to first frame control to provide a sequence of frames, wherein the frames are defined by a start frame boundary 1901 and an end frame boundary 1902. The time-to-spectrum converter 1000 or the spectrum-to-time converter 1030 is configured to operate according to second frame control synchronized with the first frame control, wherein the start frame boundary 1901 or end frame boundary 1902 of each frame in the frame sequence has a predetermined relationship with the start or end time of the overlapping portion of the window used by the time-to-spectrum converter 1000 for each block of the sequence of sampled values ​​or by the spectrum-to-time converter 1030 for each block of the output sequence of sampled values.

[0070] like Figure 1As shown, spectral domain resampling is an optional feature. The invention can also be performed without any resampling, or with resampling after or before multichannel processing. In use, the spectral domain resampler 1020 performs a resampling operation in the frequency domain on data input to the spectrum-to-time converter 1030 or on data input to the multichannel processor 1010, wherein the blocks of resampled sequences of spectral values ​​have spectral values ​​up to the maximum output frequencies 1231, 1221 that are different from the maximum input frequency 1211. Embodiments with resampling are subsequently described; however, it should be emphasized that resampling is an optional feature.

[0071] In another embodiment, multichannel processor 1010 is connected to spectral domain resampler 1020, and the output of spectral domain resampler 1020 is input to multichannel processor. This is illustrated by dummy connections 1021, 1022. In this alternative embodiment, multichannel processor is configured to apply joint multichannel processing not to a sequence of blocks of spectral values ​​output by the time-to-spectrum converter, but to a resampled sequence of blocks obtained on connection 1022.

[0072] The spectral domain resampler 1020 is configured to resample the resulting sequence generated by the multichannel processor, or to resample the sequence of blocks output by the time-to-spectrum converter 1000, to obtain a resampled sequence of blocks representing the spectral values ​​of intermediate signals, as shown in line 1025. Preferably, the spectral domain resampler additionally performs resampling on side signals generated by the multichannel processor, and thus also outputs a resampled sequence corresponding to the side signals, as shown at 1026. However, the generation and resampling of side signals are optional and not required for low bit-rate implementations. Preferably, the spectral domain resampler 1020 is configured to truncate blocks of spectral values ​​for downsampling or to zero-padded blocks of spectral values ​​for upsampling. The multichannel encoder additionally includes a spectrum-to-time converter for converting the resampled sequence of blocks of spectral values ​​into a time-domain representation comprising an output sequence of blocks having sampled values ​​with an associated output sampling rate different from the input sampling rate. In an alternative embodiment where spectral domain resampling is performed prior to multichannel processing, the multichannel processor directly provides the resulting sequence to the spectrum-to-time converter 1030 via dashed line 1023. An optional feature of this alternative embodiment is that, additionally, a side signal has already been generated by the multichannel processor in the resampled representation, and then the side signal is also processed by the spectrum-to-time converter.

[0073] Finally, the spectrum-to-time converter preferably provides a time-domain intermediate signal 1031 and an optional time-domain side signal 1032, both of which can be core-encoded by the core encoder 1040. Generally, the core encoder is configured to core-encode the output sequence of blocks of sampled values ​​to obtain an encoded multichannel signal.

[0074] Figure 2 The illustration shows a spectrum diagram that is useful for interpreting spectral domain resampling.

[0075] Figure 2 The upper graph illustrates the spectrum of the channel available at the output of the time-to-spectrum converter 1000. This spectrum 1210 has spectral values ​​up to the maximum input frequency 1211. In the case of upsampling, zero-filling is performed within the zero-fill portion or zero-fill region 1220 extending up to the maximum output frequency 1221. Due to the intention of upsampling, the maximum output frequency 1221 is greater than the maximum input frequency 1211.

[0076] In comparison, Figure 2 The lower diagram illustrates the process caused by downsampling the block sequence. For this purpose, the block is truncated within the truncated region 1230, such that the maximum output frequency of the truncated spectrum at 1231 is lower than the maximum input frequency 1211.

[0077] Usually, with Figure 2 The sampling rate associated with the corresponding spectrum is at least twice the maximum frequency of the spectrum. Therefore, for Figure 2 In the above scenario, the sampling rate will be at least twice the maximum input frequency of 1211.

[0078] exist Figure 2 In the second chart, the sampling rate will be at least twice the maximum output frequency 1221 (i.e., the highest frequency of the zero-fill region 1220). Conversely, in Figure 2 In the bottom chart, the sampling rate will be at least twice the maximum output frequency 1231 (i.e., the highest remaining spectral value after truncation within the truncated region 1230).

[0079] Figures 3a to 3c The diagram illustrates several alternatives that can be used in the context of certain DFT forward or backward transformation algorithms. Figure 3a Consider a case where a DFT of size x is performed, and where no normalization occurs in the forward transform algorithm 1311. At box 1331, a backward transform with a different size y is shown, where a DFT of 1 / N is performed. y Normalization of N. y This is the number of spectral values ​​with an inverse transform of magnitude y. Then, preferably through N... y / Nx Perform scaling, as shown in box 1321.

[0080] In comparison, Figure 3b The illustration shows an implementation where normalization is applied to both the forward transform 1312 and the backward transform 1332. Scaling is then required, as shown in box 1322, where the square root of the relationship between the number of spectral values ​​of the backward transform and the number of spectral values ​​of the forward transform is useful.

[0081] Figure 3c The diagram illustrates another implementation where a global normalization is performed on the forward transformation when performing a forward transformation with a size x. Then, the backward transformation, as shown in box 1333, operates without any normalization, thus requiring no scaling, as... Figure 3c The schematic box 1323 is shown in the diagram. Therefore, depending on the algorithm, some scaling operations may be required, or even none may be required. However, it is preferable to... Figure 3a Perform the operation.

[0082] To keep the total latency low, this invention provides a method on the encoder side to avoid the need for a time-domain resampler and replace it with resampling the signal in the DFT domain. For example, in EVS, it allows for a saving of 0.9375 ms of latency from the time-domain resampler. Resampling in the frequency domain is achieved by zero-padding or truncating the spectrum and scaling it correctly.

[0083] Consider an input windowed signal x sampled at rate fx, which has a magnitude of N. x The spectrum X, and a version y of the same signal resampled at rate fy, having a size of N. y The spectrum. Therefore, the sampling factor is equal to:

[0084] fy / fx = N y / N x

[0085] Sampling N x >N y In this case, downsampling can be easily performed in the frequency domain by directly scaling and truncating the original spectrum X:

[0086] Y[k]=X[k].N y / N x For k = 0..N y

[0087] Sampling N x <N y In this case, upsampling can be easily performed in the frequency domain by directly scaling and zero-padding the original spectrum X:

[0088] Y[k]=X[k].N y / N x For k = 0...N x

[0089] Y[k] = 0, for k = N x ...N y

[0090] The two resampling operations can be summarized as follows:

[0091] Y[k]=X[k].N y / N x For all k = 0...min(N) y N x )

[0092] Y[k] = 0, for all k = min(N) y N x ...N y For if N y >N x

[0093] Once the new spectrum Y is obtained, it can be applied by a size of N. y The associated inverse transform iDFT is used to obtain the time-domain signal y:

[0094] y = iDFT(Y)

[0095] To construct a continuous-time signal across different frames, the output frame y is then windowed and superimposed onto the previously obtained frame.

[0096] The window shape is the same for all sampling rates, but the window has different sizes in the sample and varies depending on the sampling rate. Since the shape is purely analytically defined, the number of samples in the window and its values ​​can be easily derived. Different parts and sizes of the window can be determined in... Figure 8a The ovlp_size coefficient is found to be a function of the target sampling rate. In this case, the sine function in the overlap region (LA) is used for analysis and synthesis windows. For these regions, the increasing ovlp_size coefficient is given by the following equation:

[0097] win_ovlp(k)=sin(pi*(k+0.5) / (2*ovlp_size));, for k=0..ovlp_size-1

[0098] The decreasing ovlp_size coefficient is given by the following formula:

[0099] win_ovlp(k)=sin(pi*(ovip_size-1-k+0.5) / (2*ovlp_size));, for k=0..ovlp_size-1

[0100] Where ovlp_size is a function of the sampling rate and Figure 8a The information is provided in the text.

[0101] The new low-latency stereo coding utilizes joint center / side (M / S) stereo coding with some spatial cues, where the center channel is encoded by the primary mono core encoder (mono core encoder), and the side channels are encoded in the secondary core encoder. The encoder and decoder principles are described in... Figure 4a and 4b Described in the text.

[0102] Stereo processing is primarily performed in the frequency domain (FD). Optionally, some form of stereo processing can be performed in the time domain (TD) prior to frequency analysis. This is in the case of ITD calculations, which can be computed and applied before frequency analysis to align the channels temporally before stereo analysis and processing. Alternatively, ITD processing can be performed directly in the frequency domain. Since commonly used speech encoders such as ACELP do not contain any internal time-frequency decomposition, stereo coding adds an additional complex modulation filter bank by means of an analysis and synthesis filter bank before the core encoder and another analysis and synthesis filter bank after the core decoder. In a preferred embodiment, an oversampled DFT with low overlap regions is employed. However, in other embodiments, any complex-valued time-frequency decomposition with similar temporal resolution can be used. After the stereo filter bank, a filter bank such as QMF or a block transform such as DFT can be referenced.

[0103] Stereo processing involves calculating spatial cues and / or stereo parameters such as inter-channel time difference (ITD), inter-channel phase difference (IPD), inter-channel sound level difference (ILD), and prediction gain for predicting the side signal (S) using the intermediate signal (M). It is important to note that the stereo filter banks at both the encoder and decoder introduce additional delay into the encoding system.

[0104] Figure 4a The diagram illustrates an apparatus for encoding multichannel signals, wherein, in this implementation, a certain joint stereo processing is performed in the time domain using inter-channel time difference (ITD) analysis, and wherein the result of such ITD analysis 1420 is applied in the time domain using a time-shift block 1410 placed before the time-to-spectrum converter 1000.

[0105] Then, in the spectral domain, further stereo processing 1010 is performed, which at least results in downmixing of the left and right sides of the center signal M and optionally in the calculation of the side signals S, and although not in Figure 4a The document clearly states that, however, one of two different alternatives can be applied. Figure 1 The resampling operation performed by the spectral domain resampler 1020 shown is performed either after multichannel processing or before multichannel processing.

[0106] also, Figure 4a The illustration shows further details of the preferred core encoder 1040. Specifically, an EVS encoder is used to encode the intermediate time-domain signal m at the output of the spectrum-to-time converter 1030. Furthermore, for the purpose of side signal encoding, MDCT encoding 1440 and subsequent vector quantization 1450 are performed.

[0107] Encoded or core-encoded intermediate signals and core-encoded side signals are forwarded to multiplexer 1500, which multiplexes these encoded signals together with the side information. One type of side information is the ID parameter output to the multiplexer (and optionally to stereo processing element 1010) at 1421, and other parameters are channel level difference / prediction parameters, inter-channel phase difference (IPD) parameters, or stereo fill parameters, as shown at line 1422. Accordingly, the multi-channel signal represented by bitstream 1510 is used for decoding... Figure 4b The apparatus includes a demultiplexer 1520, which in this embodiment comprises a core decoder consisting of an EVS decoder 1602 for the encoded intermediate signal m, a vector dequantizer 1603, and a subsequently connected inverse MDCT block 1604. Block 1604 provides the core-decoded side signal s. The decoded signals m and s are converted to the spectral domain using a time-to-spectrum converter 1610, and then inverse stereo processing and resampling are performed in the spectral domain. Again, Figure 4bThe illustration depicts a scenario where upmixing is performed from the M signal to the left L and right R, along with additional narrowband alignment using IPD parameters, and further processes are performed to calculate the best possible left and right channels using the inter-channel sound level difference parameter ILD and stereo fill parameter on line 1605. Furthermore, demultiplexer 1520 not only extracts the parameters on line 1605 from bitstream 1510, but also extracts the inter-channel time difference on line 1606 and forwards this information to the block inverse stereo processing / resampler and, additionally, to the inverse time-shift processing in block 1650 performed in the time domain, i.e., after the process performed by the spectrum-to-time converter providing the decoded left and right signals at the output rate, where, for example, the output rate differs from the rate at the output of EVS decoder 1602, or from the rate at the output of IMDCT block 1604.

[0108] The stereo DFT then provides different sampled versions of the signal, which are further passed to the switching core encoder. The signal to be encoded can be the center channel, side channels, left and right channels, or any signal generated from rotation or channel mapping of the two input channels. Since the different core encoders in the switching system accept different sampling rates, the ability of the stereo synthesis filter bank to provide multi-rated signals is an important feature. The principle is as follows... Figure 5 As shown in the image.

[0109] exist Figure 5 In the stereo module, the two input channels I and r are taken as input and transformed into signals M and S in the frequency domain. During stereo processing, the input channels can ultimately be mapped or modified to generate two new signals M and S. M is further encoded using the 3GPP standard EVS mono encoder or a modified version thereof. This encoder is a switchable encoder, switching between the MDCT core (TCX and HQ-Core in the case of EVS) and the speech encoder (ACELP in EVS). It also features preprocessing functions that always run at 12.8 kHz, as well as additional preprocessing functions that run at sampling rates varying depending on the operating mode (12.8, 16, 25.6, or 32 kHz). Furthermore, ACELP runs at 12.8 or 16 kHz, while the MDCT core runs at the input sampling rate. Signal S can be encoded by the standard EVS mono encoder (or a modified version thereof), or by a specific side signal encoder designed specifically for its characteristics. It is also possible to skip encoding the side signal S.

[0110] Figure 5 The illustration shows details of a preferred stereo encoder with a multi-rate synthesis filter bank for signals M and S, featuring stereo processing. Figure 5A time-to-frequency converter 1000 is shown that performs time-frequency conversion at the input rate (i.e., the rate at which signals 1001 and 1002 have this rate). Clearly, Figure 5 The attached map shows the time-domain analysis boxes 1000a and 1000e for each channel. Specifically, although... Figure 5 An explicit time-domain analysis box is shown (i.e., a windower for applying an analysis window to the corresponding channel), but it should be noted that elsewhere in this specification, the windower for applying the time-domain analysis box is considered to be included in boxes indicated by a certain sampling rate as "time-to-spectrum converter" or "DFT". Furthermore, accordingly, references to a spectrum-to-time converter generally include a windower at the output of the actual DFT algorithm for applying the corresponding synthesis window, wherein, in order to finally obtain the output sample, an overlap-addition of blocks of sampled values ​​windowed with the corresponding synthesis window is performed. Therefore, even though, for example, box 1030 only mentions "IDFT", this box generally also indicates subsequent windowing of blocks of time-domain samples using the analysis window, and again, a subsequent overlap-addition operation, to finally obtain the time-domain m-signal.

[0111] also, Figure 5 The diagram illustrates a specific stereo scene analysis box 1011, which executes parameters used in box 1010 to perform stereo processing and downmixing, and these parameters can be, for example, Figure 4a The parameters on line 1422 or 1421. Therefore, box 1011 can be used in the implementation. Figure 4a Box 1420 corresponds to where even the parametric analysis (i.e., stereo scene analysis) is performed in the spectral domain, and specifically utilizes a sequence of blocks of spectral values ​​that are not resampled but are at the maximum frequency corresponding to the input sampling rate.

[0112] Furthermore, the core encoder 1040 includes an MDCT-based encoder branch 1430a and an ACELP encoding branch 1430b. Specifically, the intermediate encoder for the intermediate signal M and the corresponding side encoder for the side signal s perform switching encoding between MDCT-based encoding and ACELP encoding, wherein the core encoder typically additionally has an encoding mode decision unit that typically operates on a certain advance portion to determine whether to use an MDCT-based process or an ACELP-based process to encode a block or frame. Additionally, or alternatively, the core encoder is configured to use an advance portion to determine other characteristics such as LPC parameters, etc.

[0113] In addition, the core encoder includes preprocessors with different sampling rates, such as a first preprocessor 1430c operating at 12.8 kHz and another preprocessor 1430d operating at sampling rates consisting of a group of sampling rates of 16 kHz, 25.6 kHz or 32 kHz.

[0114] Therefore, generally speaking, Figure 5 The embodiment shown is configured to have a spectral domain resampler for resampling from an input rate (which may be 8 kHz, 16 kHz, or 32 kHz) to an output rate different from any of 8, 16, or 32.

[0115] also, Figure 5 The embodiments in the example are additionally configured to have additional branches that are not resampled, namely, branches for intermediate signals and optionally for side signals, represented by "IDFT at input rate".

[0116] also, Figure 5 The encoder preferably includes a resampler that resamples not only to a first output sampling rate but also to a second output sampling rate to have data for both preprocessors 1430c and 1430d. For example, preprocessors 1430c and 1430d are operable to perform some kind of filtering, some kind of LPC calculation, or some kind of signal processing, which preferably has already been implemented. Figure 4a The 3GPP standard for EVS encoders is disclosed in the context of this reference.

[0117] Figure 6 The illustration shows an embodiment of an apparatus for decoding an encoded multichannel signal 1601. The apparatus for decoding includes a core decoder 1600, a time-to-spectrum converter 1610, an optional spectral domain resampler 1620, a multichannel processor 1630, and a spectrum-to-time converter 1640.

[0118] The core decoder 1600 is configured to operate according to a first frame control to provide a frame sequence, wherein frames are defined by a start frame boundary 1901 and an end frame boundary 1902. A time-to-spectrum converter 1610 or a spectrum-to-time converter 1640 is configured to operate according to a second frame control synchronized with the first frame control. The time-to-spectrum converter 1610 or the spectrum-to-time converter 1640 is configured to operate according to a second frame control synchronized with the first frame control, wherein the start frame boundary 1901 or the end frame boundary 1902 of each frame in the frame sequence has a predetermined relationship with the start or end time of the overlapping portion of the window used by the time-to-spectrum converter 1610 for each block of the sequence of sampled values ​​or by the spectrum-to-time converter 1640 for each block of at least two output sequences of sampled values.

[0119] Furthermore, the present invention concerning the apparatus for decoding the encoded multichannel signal 1601 can be implemented in several alternatives. One alternative is to not use a spectral domain resampler at all. Another alternative is to use a resampler and configure it to resample the core-decoded signal in the spectral domain prior to performing multichannel processing. This alternative is provided by... Figure 6 The solid line in the diagram illustrates this. However, an alternative is to perform spectral domain resampling after multichannel processing; that is, multichannel processing is performed at the input sampling rate. This embodiment is shown in... Figure 6 The data is shown in dashed lines. If used, the spectral domain resampler 1620 performs a resampling operation in the frequency domain on data input to the spectrum-to-time converter 1640 or on data input to the multichannel processor 1630, wherein the blocks of the resampled sequence have spectral values ​​up to the maximum output frequency, which is different from the maximum input frequency.

[0120] Specifically, in the first embodiment, i.e., when spectral domain resampling is performed in the spectral domain before multichannel processing, the core-decoded signal representing the block sequence of sampled values ​​is converted into a frequency domain representation of the block sequence having the spectral values ​​of the core-decoded signal at line 1611.

[0121] Furthermore, the core-decoded signal includes not only the M signal at line 1602, but also the side signal at line 1603, which is shown at 1604 in a core-encoded representation.

[0122] Then, the time-to-spectrum converter 1610 additionally generates a sequence of blocks of spectral values ​​for the side signals on line 1612.

[0123] Then, spectral domain resampling is performed by block 1620, and a resampled sequence of blocks of spectral values ​​for the intermediate signal or downmixed channel or first channel is forwarded to the multichannel processor at line 1621, and optionally, a resampled sequence of blocks of spectral values ​​for the side signal is also forwarded from the spectral domain resampler 1620 to the multichannel processor 1630 via line 1622.

[0124] Then, the multichannel processor 1630 performs inverse multichannel processing on the sequence shown at lines 1621 and 1622, which includes sequences from the downmixed signal (and optionally from the side signals), to output at least two resulting sequences of the blocks of spectral values ​​shown at lines 1631 and 1632. These at least two sequences are then converted to the time domain using a spectrum-to-time converter to output time-domain channel signals 1641 and 1642. Alternatively, as shown at line 1615, the time-to-spectrum converter is configured to feed a core-decoded signal, such as an intermediate signal, to the multichannel processor. Furthermore, the time-to-spectrum converter can also feed the decoded side signal 1603, in its spectral domain representation, to the multichannel processor 1630, but... Figure 6 This option is not shown in the diagram. Then, the multi-channel processor performs the inverse processing, and at least two of the output channels are forwarded to the spectral domain resampler via connection line 1635. The spectral domain resampler then forwards the resampled data at these two channels to the spectrum-to-time converter 1640 via line 1625.

[0125] Therefore, with Figure 1 The context discussed is somewhat similar; the apparatus for decoding encoded multichannel signals also includes two alternatives: performing spectral domain resampling before inverse multichannel processing, or alternatively, performing spectral domain resampling at the input sampling rate after multichannel processing. However, the first alternative is preferred because it allows... Figure 7a and Figure 7b Favorable alignment of different signal contributions is shown.

[0126] again, Figure 7a The diagram illustrates a core decoder 1600, but it outputs three distinct output signals: a first output signal 1601 at different sampling rates relative to the output sampling rate; a second core-decoded signal 1602 at the input sampling rate (i.e., the sampling rate at which the core-encoded signal 1601 is sampled); and the core decoder additionally generates an output sampling rate (i.e., at...)... Figure 7a A third output signal 1603 is operable and available at the output of the mid-spectrum time converter 1640 at the final expected sampling rate.

[0127] All three core-decoded signals are input to a time-to-spectrum converter 1610, which generates three different sequences 1613, 1611, and 1612 of blocks of spectral values.

[0128] The sequence 1613 of the block of spectral values ​​has frequencies or spectral values ​​up to the maximum output frequency, and is therefore associated with the output sampling rate.

[0129] The sequence 1611 of blocks of spectral values ​​has spectral values ​​up to different maximum frequencies, therefore this signal does not correspond to the output sampling rate.

[0130] Furthermore, the spectral values ​​of signal 1612 are as high as the maximum input frequency, which is also different from the maximum output frequency.

[0131] Therefore, sequences 1612 and 1611 are forwarded to the spectral domain resampler 1620, while signal 1613 is not forwarded to the spectral domain resampler 1620 because this signal is already associated with the correct output sampling rate.

[0132] The spectral domain resampler 1620 forwards the resampled sequence of spectral values ​​to the combiner 1700, which is configured to perform block-by-block combination on a spectral line-by-spectrum line basis for the corresponding signals in overlapping cases. Therefore, there is typically an overlap region between the switching from an MDCT-based signal to an ACELP signal, and within this overlap range, signal values ​​exist and are combined with each other. However, when this overlap range ends and the signal exists only, for example, in signal 1603 while signal 1602, for example, does not exist, the combiner will not perform block-by-block spectral line addition in this portion. However, when a subsequent switch occurs, block-by-block spectral line addition will occur during this overlap region.

[0133] In addition, such as Figure 7b As shown, it is also possible to perform continuous addition, wherein the bass post-filter output signal shown in block 1600a is executed, which can be generated, for example, from... Figure 7a The interharmonic error signal of signal 1601. Then, after the time-to-spectrum conversion in block 1610 and the subsequent spectral domain resampling 1620, preferably after performing Figure 7b An additional filtering operation 1702 is performed before the addition in box 1700.

[0134] Similarly, the MDCT-based decoding stage 1600d and the temporal bandwidth-extended decoding stage 1600c can be coupled via a cross-fading box 1704 to obtain the core-decoded signal 1603, which is then converted to a spectral domain representation at the output sampling rate. This makes spectral domain resampling unnecessary for this signal 1613, but the signal can be directly forwarded to the combiner 1700. Stereo inverse processing or multichannel processing 1603 then occurs after the combiner 1700.

[0135] Therefore, with Figure 6Compared to the embodiment shown, the multichannel processor 1630 does not operate on the resampled sequence of spectral values, but instead operates on a sequence that includes at least one resampled sequence of spectral values ​​(such as 1622 and 1621), wherein the sequence operated on by the multichannel processor 1630 also includes a sequence 1613 that does not need to be resampled.

[0136] As shown in Figure 7, the different decoded signals from different DFTs operating at different sampling rates are time-aligned because the analysis windows at different sampling rates share the same shape. However, the spectra exhibit different sizes and scaling. To reconcile and make them compatible, all spectra are resampled in the frequency domain at the desired output sampling rate before being added together.

[0137] Therefore, Figure 7 illustrates the combination of different contributions of the synthesized signal in the DFT domain, wherein spectral domain resampling is performed in such a way that all signals to be added by combiner 1700 are finally available and the spectral values ​​extend up to the maximum output frequency corresponding to the output sampling rate (less than or equal to half of the output sampling rate then obtained at the output of spectrum time converter 1640).

[0138] The selection of stereo filter banks is crucial for low-latency systems, and achievable trade-offs are possible. Figure 8b In summary, it can employ DFT (block transform) or pseudo-low-latency QMF called CLDFB (filter bank). Each proposal exhibits different delay, time, and frequency resolutions. For a system, the optimal trade-off between these characteristics must be chosen. Good frequency and time resolution is crucial. This is why using a pseudo-QMF filter bank in Proposal 3 is problematic. The frequency resolution is low. It can be enhanced through hybrid methods like those in MPEG-USAC's MPS212, but this has the drawback of significantly increasing complexity and delay. Another key point is the available delay on the decoder side between the core decoder and the inverse stereo processing. The greater this delay, the better. For example, Proposal 2 cannot provide this delay and is therefore not a valuable solution. For the reasons mentioned above, we will focus on Proposals 1, 4, and 5 in the remainder of the description.

[0139] The analysis and synthesis window of the filter bank is another important aspect. In a preferred embodiment, the same window is used for both DFT analysis and synthesis. This is also true on both the encoder and decoder sides. Special attention should be paid to fulfilling the following constraints:

[0140] • The overlapping region must be equal to or smaller than the overlapping region between the MDCT core and the ACELP prior. In a preferred embodiment,

[0141] All sizes are equal to 8.75ms

[0142] • To allow for linear channel shifting in the DFT domain, zero-padding should be at least approximately 2.5 ms.

[0143] • For different sampling rates: 12.8, 16, 25.6, 32, and 48 kHz, the window size, overlap region size, and zero-filling size must be expressed in integer samples.

[0144] • The complexity of the DFT should be as low as possible, that is, the maximum cardinality of the DFT in the split-radix FFT implementation should be as low as possible.

[0145] • The time resolution is fixed at 10ms.

[0146] Knowing these constraints, in Figure 8c and Figure 8a The window describes proposals 1 and 4.

[0147] Figure 8c The illustration shows a first window consisting of an initial overlapping portion 1801, a subsequent intermediate portion 1803, and a terminating overlapping portion or a second overlapping portion 1802. Furthermore, the first overlapping portion 1801 and the second overlapping portion 1802 additionally have zero-fill portions 1804 and 1805 at their beginning and end, respectively.

[0148] also, Figure 8c The illustration is about Figure 1 Time-to-Spectrum Converter 1000 or alternatively Figure 7a The framing process of 1610 is performed. Another analysis window, consisting of element 1811 (i.e., the first overlapping portion), the intermediate non-overlapping portion 1813, and the second overlapping portion 1812, overlaps the first window by 50%. The second window additionally has zero-padding portions 1814 and 1815 at its beginning and end. These zero-overlapping portions are necessary to position the window for performing wideband time alignment in the frequency domain.

[0149] Furthermore, the first overlapping portion 1811 of the second window begins at the end of the middle portion 1803 (i.e., the non-overlapping portion of the first window), and the overlapping portion of the second window (i.e., the non-overlapping portion 1813) begins at the end of the second overlapping portion 1802 of the first window, as shown in the figure.

[0150] when Figure 8c This is considered to represent the spectrum-to-time converter (such as...) Figure 1During the overlap-addition operation of the spectrum-to-time converter 1030 for the encoder or the spectrum-to-time converter 1640 for the decoder, a first window consisting of blocks 1801, 1802, 1803, 1805, and 1804 corresponds to the synthesis window, and a second window consisting of portions 1811, 1812, 1813, 1814, and 1815 corresponds to the synthesis window for the next block. The overlap between windows is then illustrated, and the overlap portion is shown at 1820. The length of the overlap portion is equal to the current frame divided by two, and in a preferred embodiment, it is equal to 10 ms. Furthermore, in Figure 8c At the bottom, the analytical equations for calculating the increasing window coefficients within the overlapping range 1801 or 1811 are shown as sine functions, and correspondingly, the decreasing overlap size coefficients for the overlapping portions 1802 and 1812 are also shown as sine functions.

[0151] In a preferred embodiment, the same analysis and synthesis window is used only for Figure 6 , Figure 7a , Figure 7b The decoder is shown in the figure. Therefore, the time-to-spectrum converter 1616 and the spectrum-to-time converter 1640 use the exact same window, as shown in the figure. Figure 8c As shown.

[0152] However, in some embodiments, particularly with respect to the subsequent proposal / Embodiment 1, the use of generally consistent with Figure 9c A consistent analysis window is used, but the square root of the sine function is used to calculate the window coefficients for the increasing or decreasing overlapping portions, which have the same characteristics as the sine function. Figure 8c The same independent variable is used in the synthesis window. Accordingly, the synthesis window is calculated using a sine function raised to the power of 1.5, but again has the same independent variable as the sine function.

[0153] Furthermore, it should be noted that, due to the overlap-addition operation, multiplying the sine of the 0.5 power by the sine of the 1.5 power again results in the sine of the 2nd power, which is necessary for achieving energy-saving conditions.

[0154] Proposal 1 has the following key characteristics: the overlapping regions of the DFT are of the same size and aligned with the overlapping regions of the ACELP and MDCT cores. The encoder delay is therefore identical for both the ACELP and MDCT cores, and stereo does not introduce any additional delay at the encoder. In the case of EVS and when using... Figure 5 In the case of the described multi-rate synthesis filter bank method, the stereo encoder delay is as low as 8.75ms.

[0155] Encoder schematic framing in Figure 9a As shown in the diagram, the decoder is... Figure 9e Depicted in [the text]. For the encoder, the window is in [the text]. Figure 9c The center is drawn with a blue dashed line, while the decoder window is drawn with a red solid line.

[0156] A major problem with Proposal 1 is that the advance signal at the encoder is windowed. It can be corrected for subsequent processing, or it can remain windowed if subsequent processing is appropriate to take the windowed advance signal into account. It is possible that if stereo processing performed in the DFT modifies the input channels, and especially when using non-linear operations, then a corrected or windowed signal may not allow for perfect reconstruction when bypassing the core encoder.

[0157] It is worth noting that there is a 1.25ms time gap between the core decoder synthesis and the stereo decoder analysis window, which can be used for core decoder post-processing, bandwidth extension (BWE) (such as temporal BWE used on ACELP), or as a smooth transition between ACELP and MDCT core.

[0158] Since this 1.25ms time gap is lower than the 2.3125ms required by the standard EVS for this operation, the present invention provides a method for combining, resampling, and smoothly switching different synthesis parts of a decoder within the DFT domain of a stereo module.

[0159] like Figure 9a As shown, the core encoder 1040 is configured to operate according to framing control to provide a sequence of frames, where frames are defined by a start frame boundary 1901 and an end frame boundary 1902. Furthermore, the time-to-spectrum converter 1000 and / or the spectrum-to-time converter 1030 are also configured to operate according to a second framing control synchronized with the first framing control. The framing control is shown by two overlapping windows 1903 and 1904 for the time-to-spectrum converter 1000 in the encoder (and specifically for the first channel 1001 and the second channel 1002, which are processed concurrently and fully synchronized). Furthermore, the framing control is also visible on the decoder side, specifically as shown at 1913 and 1914. Figure 6 The time-to-spectrum converter 1610 has two overlapping windows. For example, these windows, 1913 and 1914, are applied to the core decoder signal, which is preferably... Figure 6 A single mono or downmixed signal 1610. Additionally, as from... Figure 9aClearly visible, the framing control of the core encoder 1040 and the synchronization between the time-to-spectrum converter 1000 or the spectrum-to-time converter 1030 ensure that the start frame boundary 1901 or end frame boundary 1902 of each frame in the frame sequence is in a predetermined relationship with the start or end time of the overlapping portion of the window used by the time-to-spectrum converter 1000 or the spectrum-to-time converter 1030 for each block of the sequence of blocks of sampled values ​​or each block of the resampled sequence of blocks of spectral values. Figure 9a In the illustrated embodiment, for example, the predetermined relationship causes the start of the first overlapping portion to coincide with the start time boundary of window 1903, and the start of the overlapping portion of another window 1904 to coincide with the middle portion (such as...) Figure 8c The end of part 1803 coincides. Therefore, when Figure 8c The second window in Figure 9a When window 1904 corresponds to the end frame boundary 1902, it is the same as the end frame boundary. Figure 8c The middle part of 1813 coincides with the end.

[0160] Therefore, it becomes clear that, Figure 9a The second overlapping portion of the second window in 1904 (such as...) Figure 8c (1812) extends beyond the end or termination of frame boundary 1902, and thus extends into the core encoder advance section shown at 1905.

[0161] Therefore, the core encoder 1040 is configured to use a lead portion (such as lead portion 1905) when outputting a block of output sequence of core encoded sample values, wherein the output lead portion is temporally located after the output block. The output block corresponds to a frame defined by frame boundaries 1901, 1904, and the output lead portion 1905 follows this output block used by the core encoder 1040.

[0162] Furthermore, as shown in the figure, the time-to-spectrum converter is configured to use an analysis window (i.e., window 1904, which has an overlapping portion with a time length less than or equal to the time length of the preceding portion 1905), wherein the components located within the overlapping range are... Figure 8c The overlapping portion 1812 is used to generate the leading portion of the windowed section.

[0163] Furthermore, the spectrum-to-time converter 1030 is configured to preferably use a correction function to process the output leading portion corresponding to the windowed leading portion, wherein the correction function is configured to reduce or eliminate the influence of the overlapping portion of the analysis window.

[0164] Therefore, in Figure 9aThe spectrum-to-time converter operating between the core encoder 1040 and the downmixer 1010 / downsampler 1020 blocks is configured to apply corrections in a function to undo the changes made by the core encoder 1040. Figure 9a The window in window 1904 is an application of window opening.

[0165] Therefore, it is ensured that when the core encoder 1040 applies its advance functionality to the part that is as far away from the original part as possible, rather than to the advance part, the part that is closest to the original part.

[0166] However, due to low latency constraints and the synchronization between the stereo preprocessor and the core encoder's framing, there is no original time-domain signal for the lead-in section. However, applying a correction function ensures that any artifacts introduced by this process are minimized as much as possible.

[0167] The series of processes related to this technology Figure 9d , Figure 9e It is shown in more detail below.

[0168] In step 1910, the DFT of block zero is performed. -1 To obtain the zeroth block in the time domain. The zeroth block will be used for... Figure 9a The window to the left of window 1903 in the middle window. However, this zeroth block is not in... Figure 9a It is clearly shown in the text.

[0169] Then, in step 1912, the zeroth block is windowed using a composite window, that is, in Figure 1 A window is shown in the spectrum-to-time converter 1030.

[0170] Then, as shown in box 1911, the DFT of the first block obtained from window 1903 is performed. -1 To obtain the first block in the time domain, and to use a composition window in box 1910 to window this first block again.

[0171] Then, as Figure 9d As indicated in 1918, the second block (i.e., by) Figure 9a The inverse DFT of the block obtained from window 1904 is used to obtain the second block in the time domain, and then the first part of the second block is windowed using a synthesis window, as shown. Figure 9d As shown in 1920. However, importantly, it was by... Figure 9d The second part of the second block obtained in item 1918 is not windowed using a composite window, but is modified, such as... Figure 9d As shown in box 1922, and for the correction function, the inverse of the analysis window function and the corresponding overlapping portion of the analysis window function are used.

[0172] Therefore, if the window used to generate the second block is Figure 8c The sine window shown, then Figure 8c The equation at the bottom of the equation, 1 / sin(), used to decrease the overlap factor, is used as a correction function.

[0173] However, it is preferable to use the square root of the sine window for the analysis window; therefore, the correction function is the window function. This ensures that the modified leading portion obtained through block 1922 is as close as possible to the original signal within the leading portion, but of course not the original left signal or the original right signal, but the original signal obtained by adding the left and right to obtain the intermediate signal.

[0174] Then, in Figure 9d In step 1924, a frame indicated by frame boundaries 1901, 1902 is generated by performing an overlap-add operation in block 1030, such that the encoder has a time-domain signal, and this frame is executed by an overlap-add operation between the block corresponding to window 1903 and the previous sample of the previous block, and by using the first portion of the second block obtained by block 1920. This frame output from block 1924 is then forwarded to core encoder 1040, and additionally, the core encoder receives a modified advance portion of the frame. As shown in step 1926, the core encoder can then use the modified advance portion obtained in step 1922 to determine the characteristics of the core encoder. Then, as shown in step 1928, the core encoder performs core encoding on the frame using the characteristics determined in block 1926 to finally obtain a core-encoded frame corresponding to frame boundaries 1901, 1902, which in a preferred embodiment has a length of 20 ms.

[0175] Preferably, the overlapping portion of the window 1904 extending into the preceding portion 1905 has the same length as the preceding portion, but it may also be shorter than the preceding portion, but preferably not longer than the preceding portion, so that the stereo preprocessor does not introduce any additional delay due to window overlap.

[0176] The process then continues by windowing the second portion of the second block using a compositing window, as shown in box 1930. Thus, the second portion of the second block is corrected on the one hand by box 1922, and on the other hand by being windowed by the compositing window, as shown in box 1930, because this portion is then needed to generate the next frame of the core encoder by overlapping and adding the windowed second portion of the second block, the windowed third block, and the windowed first portion of the fourth block, as shown in box 1932. Naturally, the fourth block, and especially the second portion of the fourth block, will again undergo the process described in relation to… Figure 9dThe correction operation discussed in the second block of item 1922 is then repeated, as discussed above. Furthermore, in step 1934, the core encoder determines its characteristics by correcting the second part of the fourth block, and then encodes the next frame using the determined coding characteristics to ultimately obtain the core-coded next frame in block 1934. Therefore, the alignment of the second overlapping portion of the analysis (corresponding to synthesis) window with the core encoder advance portion 1905 ensures a very low-latency implementation, and this advantage is due to the fact that the advance portion of the windowed section is addressed by performing a correction operation on the one hand and by applying an analysis window that is not equal to the synthesis window but exerts a smaller influence on the other, ensuring a more stable correction function compared to using the same analysis / synthesis window. However, in cases where the core encoder is modified to operate its advance functions, which are typically necessary for determining the core coding characteristics on the windowed section, the correction function need not be performed. However, it has been found that using the correction function is superior to modifying the core encoder.

[0177] Furthermore, as discussed earlier, it is important to note the difference between the end of the window (i.e., analysis window 1914) and the beginning of the end. Figure 9b There is a time gap between the start frame boundary 1901 and the end frame boundary 1902 of the frame defined by the start frame boundary 1901 and the end frame boundary 1902.

[0178] In particular, relative to by Figure 6 The analysis window of the time-to-spectrum converter 1610 application shows the time gap at 1920, and this time gap is also visible at 120 relative to the first output channel 1641 and the second output channel 1642.

[0179] Figure 9f The process of performing steps within the context of a time gap is illustrated, where the core decoder 1600 performs core decoding on a frame, or at least the initial portion of a frame up to time gap 1920. Then, Figure 6 The time-to-spectrum converter 1610 is configured to apply an analysis window 1914 to the initial portion of the frame, which does not extend until the end of the frame (i.e., until time 1902), but only extends to the beginning of time interval 1920.

[0180] Therefore, the core decoder has additional time to perform core decoding and / or post-processing of samples in the time gap, as shown in block 1940. Thus, the time-to-spectrum converter 1610 has output the first block as a result of step 1938, where the core decoder can provide the remaining samples in the time gap, or post-process the samples in the time gap in step 1940.

[0181] Then, in step 1942, the time-to-spectrum converter 1610 is configured to use [the following]... Figure 9b The next analysis window that appears after window 1914 in the previous step windows windows the samples in the time gap together with the samples of the next frame. Then, as shown in step 1944, the core decoder 1600 is configured to decode the next frame or at least the initial portion of the next frame that appears in the next frame up to time gap 1920. Then, in step 1946, the time-to-spectrum converter 1610 is configured to window the samples in the next frame up to time gap 1920 of the next frame, and in step 1948, the core decoder can then perform core decoding and / or post-processing on the remaining samples in the time gap of the next frame.

[0182] Therefore, when considering Figure 9b In some embodiments, such a time gap, for example 1.25ms, can be post-processed by the core decoder, through bandwidth expansion, by temporal bandwidth expansion used, for example in the context of ACELP, or by some kind of smoothing in the case of transmission transition between ACELP and MDCT core signals.

[0183] Therefore, the core decoder 1600 is again configured to operate according to the first framing control to provide a frame sequence, wherein the time-to-spectrum converter 1610 or the spectrum-to-time converter 1640 is configured to operate according to a second framing control synchronized with the first framing control, such that the start or end frame boundary of each frame of the frame sequence is in a predetermined relationship with the start or end time of the overlapping portion of the window used by the time-to-spectrum converter or the spectrum-to-time converter for each block of the sequence of blocks of sampled values ​​or each block of the resampled sequence of blocks of spectral values.

[0184] Furthermore, the time-to-spectrum converter 1610 is configured to use an analysis window to window the frames of a frame sequence having an overlapping range that ends before the end frame boundary 1902, thereby leaving a time gap 1920 between the end of the overlapping portion and the end frame boundary. Therefore, the core decoder 1600 is configured to perform processing on samples in the time gap 1920 in parallel with the windowing of the frames using the analysis window, or to perform further post-processing of the time gap in parallel with the windowing of the frames by the time-to-spectrum converter using the analysis window.

[0185] Furthermore, and preferably, the analysis window for subsequent blocks of the signal decoded by the core is positioned such that the non-overlapping portion of the window is located as follows: Figure 9b Within the time gap shown at point 1920.

[0186] In Proposal 4, the overall system latency is increased compared to Proposal 1. At the encoder, the additional latency comes from the stereo module. Unlike Proposal 1, the issue of perfect reconstruction is no longer relevant in Proposal 4.

[0187] At the decoder, the available latency between the core decoder and the first DFT analysis is 2.5ms, which allows for regular resampling, combining, and smoothing between different core synthesized and extended bandwidth signals, just as is done in standard EVS.

[0188] Encoder schematic framing in Figure 10a As shown in the diagram, the decoder is... Figure 10b Depicted in [the text]. The window is in [the text]. Figure 10c The information is provided in the text.

[0189] In Proposal 5, the time resolution of the DFT is reduced to 5 ms. There are no windowings in the leading and overlapping regions of the core encoder, a common advantage with Proposal 4. On the other hand, the available delay between encoder decoding and stereo analysis is small, requiring the solution proposed in Proposal 1 (Figure 7). The main drawback of this proposal is the low frequency resolution of the time-frequency decomposition and the small overlapping region reduced to 5 ms, which prevents large time shifts in the frequency domain.

[0190] Encoder schematic framing in Figure 11a As shown in the diagram, the decoder is... Figure 11b Depicted in [the text]. The window is in [the text]. Figure 11c The information is provided in the text.

[0191] In view of the above, the preferred embodiment relates to multi-rate time-frequency synthesis on the encoder side, which provides at least one stereo-processed signal to a subsequent processing module at different sampling rates. This module includes, for example, a speech encoder like ACELP, a preprocessing tool, an MDCT-based audio encoder (such as TCX), or a bandwidth-extended encoder (such as a time-domain bandwidth-extended encoder).

[0192] Regarding the decoder, the various contributions from the decoder synthesis are combined by resampling in the stereo frequency domain. These synthesized signals can come from a speech decoder such as the ACELP decoder, an MDCT-based decoder, a bandwidth extension module, or from post-processed interharmonic error signals such as a bass post-filter.

[0193] Furthermore, regarding the encoder and decoder, it is useful to apply complex values ​​for the window used in the DFT, or for zero-filling, low-overlap regions, and hopsize transformation, where the hopsize corresponds to an integer number of samples at different sampling rates (such as 12.9 kHz, 16 kHz, 25.6 kHz, 32 kHz, or 48 kHz).

[0194] The embodiment enables low-bit-rate encoding of stereo audio with low latency. It is specifically designed for efficiently combining low-latency switching audio coding schemes (such as EVS) with filter banks of stereo coding modules.

[0195] Examples can be found in applications such as distributing or broadcasting all types of stereo or multichannel audio content (speech and music-like content with constant perceived quality at a given low bit rate) using digital radio, internet streaming, and audio communication applications.

[0196] Figure 12 The diagram illustrates an apparatus for encoding a multichannel signal having at least two channels. The multichannel signal 10 is input to a parameter determiner 100 and a signal aligner 200. The parameter determiner 100 determines wideband alignment parameters and multiple narrowband alignment parameters from the multichannel signal. These parameters are output via parameter line 12. Furthermore, these parameters are also output to an output interface 500 via another parameter line 14, as shown. On parameter line 14, additional parameters, such as sound level parameters, are forwarded from the parameter determiner 100 to the output interface 500. The signal aligner 200 is configured to align at least two channels of the multichannel signal 10 using the wideband alignment parameters and multiple narrowband alignment parameters received via parameter line 10, to obtain aligned channels 20 at the output of the signal aligner 200. These aligned channels 20 are forwarded to a signal processor 300, which is configured to calculate an intermediate signal 31 and a side signal 32 from the aligned channels received via line 20. The encoding apparatus also includes a signal encoder 400 for encoding the intermediate signal from line 31 and the side signal from line 32 to obtain an encoded intermediate signal on line 41 and an encoded side signal on line 42. Both signals are forwarded to the output interface 500 for generating an encoded multi-channel signal at output line 50. The encoded signal at output line 50 includes the encoded intermediate signal from line 41, the encoded side signal from line 42, narrowband alignment parameters and wideband alignment parameters from line 14, and optionally, sound level parameters from line 14, and also optionally, stereo fill parameters generated by the signal encoder 400 and forwarded to the output interface 500 via parameter line 43.

[0197] Preferably, the signal aligner is configured to align the channels from the multichannel signal using broadband alignment parameters before the parameter determiner 100 actually calculates the narrowband parameters. Therefore, in this embodiment, the signal aligner 200 sends the broadband-aligned channels back to the parameter determiner 100 via connection line 15. The parameter determiner 100 then determines multiple narrowband alignment parameters from the multichannel signal that is already aligned with respect to broadband characteristics. However, in other embodiments, this specific sequence of processes is not required to determine the parameters.

[0198] Figure 14a The diagram illustrates a preferred implementation, in which a specific sequence of steps leading to connector 15 is performed. In step 16, broadband alignment parameters are determined using two channels, and broadband alignment parameters such as inter-channel time difference or ITD parameters are obtained. Then, in step 21, the broadband alignment parameters are used to... Figure 12 The signal aligner 200 aligns the two channels. Then, in step 17, within the parameter determiner 100, narrowband parameters are determined using the aligned channels to determine multiple narrowband alignment parameters, such as inter-channel phase difference parameters for different frequency bands of a multi-channel signal. Then, in step 22, the spectral values ​​in each parameter band are aligned using the corresponding narrowband alignment parameters for this specific frequency band. When this process in step 22 is performed for each frequency band, narrowband alignment parameters are obtained for that frequency band, and then the aligned first and second or left / right channels are obtained for use by... Figure 12 Further signal processing is performed by the signal processor 300.

[0199] Figure 14b Illustration Figure 12 Another implementation of a multi-channel encoder, in which several processes are performed in the frequency domain.

[0200] Specifically, the multichannel encoder also includes a time-to-spectrum converter 150, which is used to convert the time-domain multichannel signal into a spectral representation of at least two channels in the frequency domain.

[0201] Furthermore, as shown at point 152, Figure 12 The parameter determiner, signal aligner, and signal processor shown at positions 100, 200, and 300 all operate in the frequency domain.

[0202] In addition, the multi-channel encoder, and more specifically the signal processor, also includes a spectrum-to-time converter 154 for generating at least a time-domain representation of the intermediate signal.

[0203] Preferably, the spectrum-time converter additionally converts the spectral representation of the side signal, determined by the process also indicated by block 152, into a time-domain representation, and then... Figure 12 The signal encoder 400 is configured to further encode intermediate and / or side signals as time-domain signals, depending on... Figure 12 The specific implementation of the signal encoder 400.

[0204] Preferably, Figure 14b The time-to-spectrum converter 150 is configured to implement Figure 14cSteps 155, 156, and 157. Specifically, step 155 includes providing an analysis window having at least one zero-padding portion at one end, and specifically, zero-padding portions at the initial window portion and at the terminating window portion, for example, as shown later in FIG7. Furthermore, the analysis window additionally has overlapping ranges or overlapping portions at the first and second halves of the window, and additionally preferably, the middle portion is a non-overlapping range, as appropriate.

[0205] In step 156, each channel is windowed using an analysis window with overlapping areas. Specifically, each channel is windowed using an analysis window in a manner that obtains a first block of the channel. Subsequently, a second block of the same channel with a certain overlap with the first block is obtained, and so on, so that after, for example, five blocks of windowed samples for each channel are obtained, and then... Figure 14c As shown in 157, these blocks are individually transformed into spectral representations. The same process is performed for the other channel, such that at the end of step 157, a sequence of blocks with spectral values, and specifically complex spectral values ​​(such as DFT spectral values ​​or complex subband samples), is obtained.

[0206] In the Figure 12 In step 158, performed by parameter determiner 100, broadband alignment parameters are determined, and by... Figure 12 In step 159 of the signal alignment 200, a cyclic shift is performed using wideband alignment parameters. Then, again by... Figure 12 In step 160, the parameter determiner 100 determines narrowband alignment parameters for each frequency band / sub-band, and in step 161, it uses the corresponding narrowband alignment parameters determined for a specific frequency band to rotate the aligned spectral values ​​for each frequency band.

[0207] Figure 14d The diagram illustrates a further process performed by the signal processor 300. Specifically, the signal processor 300 is configured to compute intermediate and side signals, as shown in step 301. In step 302, some further processing of the side signals may be performed. Then, in step 303, each block of the intermediate and side signals is transformed back to the time domain. In step 304, a synthesis window is applied to each block obtained in step 303. And in step 305, an overlap-add operation is performed on both the intermediate and side signals to finally obtain the time-domain intermediate / side signals.

[0208] Specifically, the operations in steps 304 and 305 cause a cross-fading from a block of the intermediate or side signal in the next block of the intermediate or side signal, and execute the side signal so that even if any parameter (such as the inter-channel time difference parameter or the inter-channel phase difference parameter) changes, this will still occur. Figure 14d The intermediate / side signals in the time domain obtained by step 305 are inaudible.

[0209] Figure 13 The figure shows a block diagram of an embodiment of an apparatus for decoding an encoded multichannel signal received at input line 50.

[0210] Specifically, the signal is received by the input interface 600. Connected to the input interface 600 are the signal decoder 700 and the signal dealigner 900. Furthermore, the signal processor 800 is connected to both the signal decoder 700 and the signal dealigner.

[0211] Specifically, the encoded multichannel signal includes an encoded intermediate signal, encoded side signals, information about wideband alignment parameters, and information about multiple narrowband parameters. Therefore, the encoded multichannel signal on line 50 can be [related to] the [other signals]. Figure 12 The output signal from the 500's output interface is exactly the same.

[0212] However, it is important to note here that... Figure 12 Conversely, as shown, the wideband alignment parameters and multiple narrowband alignment parameters included in a certain form in the coded signal can be precisely determined by... Figure 12 The alignment parameter used by the signal aligner 200 can also be its inverse value, that is, a parameter that can be used by the exact same operation performed by the signal aligner 200 but has an inverse value to enable dealignment.

[0213] Therefore, information about alignment parameters can be obtained from Figure 12 The alignment parameters used by the signal aligner 200 in the diagram can be inverse values ​​(i.e., actual "dealignment parameters"). Furthermore, these parameters will typically be quantized in some form, which will be discussed later with reference to Figure 8.

[0214] Figure 13 The input interface 600 extracts information about wideband alignment parameters and multiple narrowband alignment parameters from the encoded intermediate / sideband signals and forwards this information to the signal dealigner 900 via parameter line 610. On the other hand, the encoded intermediate signal is forwarded to the signal decoder 700 via line 601, and the encoded sideband signal is forwarded to the signal decoder 700 via signal line 602.

[0215] A signal decoder is configured to decode the encoded intermediate signal and the encoded side signal to obtain a decoded intermediate signal on line 701 and a decoded side signal on line 702. A signal processor 800 uses these signals to calculate a decoded first channel signal or a decoded left signal and a decoded second channel signal or a decoded right channel signal from the decoded intermediate signal and the decoded side signal, and outputs the decoded first channel and the decoded second channel on lines 801 and 802, respectively. A signal dealigner 900 is configured to dealign the decoded first channel on line 801 and the decoded right channel 802 on line 801 using information about wideband alignment parameters and additionally using information about multiple narrowband alignment parameters to obtain a decoded multi-channel signal (i.e., a decoded signal with at least two decoded and dealigned channels on lines 901 and 902).

[0216] Figure 9a The illustration is by Figure 13 The preferred sequence of steps performed by the signal dealigner 900. Specifically, step 910 receives the signal from the dealigner 900. Figure 13 Aligned left and right channels are available on lines 801 and 802. In step 910, signal dealigner 900 uses information about narrowband alignment parameters to dealign the individual subbands to obtain phase-dealigned decoded first and second or left and right channels at 911a and 911b. In step 912, the channels are dealigned using wideband alignment parameters to obtain phase- and time-dealigned channels at 913a and 913b.

[0217] In step 914, any further processing is performed, including using windowing or any overlap-addition operation, or in general any cross fading operation, to obtain a spoof-reduced or spoof-free decoded signal at 915a or 915b, i.e., a decoded channel without any spoofs, although there are usually time-varying dealignment parameters for broadband on one hand and multiple narrowbands on the other.

[0218] Figure 15b Illustration Figure 13 The preferred implementation of the multi-channel decoder shown is illustrated.

[0219] In particular, Figure 13 The signal processor 800 includes a time-to-spectrum converter 810.

[0220] The signal processor also includes a center / side to left / right converter 820 to calculate the left signal L and the right signal R based on the center signal M and the side signal S.

[0221] However, importantly, in order to calculate L and R via center / side-left / right conversion in block 820, the side signal S is not required. Instead, as discussed later, the left / right signals are initially calculated using only the gain parameters derived from the inter-channel sound level difference parameter ILD. Therefore, in this implementation, the side signal S is used only in the channel updater 830, which operates to provide better left / right signals using the transmitted side signal S, as shown in bypass line 821.

[0222] Therefore, converter 820 operates using the sound level parameters obtained via sound level parameter input 822 and does not actually use the side signal S, but channel updater 830 then operates using side 821 and, depending on the specific implementation, uses the stereo fill parameters received via line 831. Signal aligner 900 then includes a phase dealigner and an energy scaler 910. Energy scaling is controlled by a scaling factor derived by scaling factor calculator 940. Scaling factor calculator 940 is fed by the output of channel updater 830. Phase dealignment is performed based on the narrowband alignment parameters received via input 911, and in block 920, time dealignment is performed based on the wideband alignment parameters received via line 921. Finally, spectrum-time conversion 930 is performed to finally obtain the decoded signal.

[0223] Figure 15c The illustration typically shows, in the preferred embodiment, Figure 15b Another sequence of steps is executed within boxes 920 and 930.

[0224] Specifically, the narrowband de-aligned channel is input to the same channel as... Figure 15b The broadband dealignment function corresponds to block 920. A DFT or any other transformation is performed in block 931. After the actual computation of the time-domain samples, an optional synthesis window is performed using the synthesis window. The synthesis window is preferably identical to or derived from the analysis window, for example, through interpolation or decimation, but depends on the analysis window in some way. This dependency preferably ensures that the multiplication factor defined by the two overlapping windows adds up to one for each point in the overlapping range. Therefore, after the synthesis window in block 932, an overlap operation and subsequent addition operation are performed. Alternatively, instead of the synthesis window and overlap / addition operation, any cross-fading between subsequent blocks of each channel is performed so that, as already done... Figure 15a As discussed in the context, the decoded signal with reduced pseudo-sound is obtained.

[0225] When considering Figure 4b At this point, it becomes clear that the actual decoding operation for the intermediate signal involves, on the one hand, the "EVS decoder," and on the other hand, inverse vector quantization (VQ) for the side signals. -1 and Inverse MDCT operation (IMDCT) and Figure 13 The signal decoder corresponds to 700.

[0226] Furthermore, the DFT operation in box 810 and Figure 15b The component 810 in the middle corresponds to the functions of reverse stereo processing and reverse time shift. Figure 13 The boxes 800 and 900 correspond, and Figure 4b Inverse DFT operation and Figure 15b The corresponding operation is in box 930.

[0227] Then, a more detailed discussion was held. Figure 3d In particular, Figure 3d The diagram shows the DFT spectrum with each spectral line. Preferably, Figure 3d The DFT spectrum or any other spectrum shown is a complex spectrum, and each line is a complex spectral line with magnitude and phase or with real and imaginary parts.

[0228] Furthermore, the spectrum is divided into different parameter bands. Each parameter band has at least one, and preferably more than one, spectral line. Moreover, the parameter bands increase from lower frequencies to higher frequencies. Typically, the broadband alignment parameter is the entire spectrum (i.e., for a spectrum including...). Figure 3d A single broadband alignment parameter for the spectrum of all frequency bands 1 to 6 in the exemplary embodiment.

[0229] Furthermore, multiple narrowband alignment parameters are provided, ensuring that a single alignment parameter exists for each frequency band. This means that the alignment parameter for a frequency band is always applicable to all spectral values ​​within the corresponding frequency band.

[0230] In addition to the narrowband alignment parameters, sound level parameters are provided for each parameter band.

[0231] Rather than providing sound level parameters for each parameter band from band 1 to band 6, it is preferable to provide multiple narrowband alignment parameters for only a limited number of lower frequency bands (such as bands 1, 2, 3, and 4).

[0232] Furthermore, stereo fill parameters are provided for a certain number of frequency bands other than the lower frequency bands (such as frequency bands 4, 5 and 6 in the exemplary embodiment), while side signal spectral values ​​exist for the lower parameter frequency bands 1, 2 and 3. Therefore, there are no stereo fill parameters for these lower frequency bands, wherein waveform matching is obtained using the side signal itself or the predicted residual signal representing the side signal.

[0233] As mentioned above, there are more spectral lines in higher frequency bands, for example, in Figure 3dIn this embodiment, the seven spectral lines in parameter band 6 correspond to only three spectral lines in parameter band 2. However, naturally, the number of parameter bands, the number of spectral lines, the number of spectral lines within a parameter band, and different limits for certain parameters will be different.

[0234] However, Figure 8 illustrates the distribution of parameters and the number of frequency bands providing those parameters in a particular embodiment, which in this embodiment is different from... Figure 3d In comparison, there are actually 12 frequency bands.

[0235] As shown in the figure, a sound level parameter ILD is provided for each of the 12 frequency bands and quantized to a quantization precision of 5 bits per frequency band.

[0236] Furthermore, the narrowband alignment parameter IPD is only provided for the lower frequency bands up to the boundary frequencies of 2.5 kHz. Conversely, the inter-channel time difference or wideband alignment parameter is provided as a single parameter across the entire spectrum, but with very high quantization accuracy, representing the entire frequency band with 8 bits.

[0237] In addition, very coarsely quantized stereo fill parameters are provided, represented by three bits per frequency band, and are not used for lower frequency bands below 1 kHz, because for lower frequency bands, the actual encoded side signal or side signal residual spectrum values ​​are included.

[0238] Subsequently, the optimal processing on the encoder side is summarized. In the first step, DFT analysis of the left and right channels is performed. This process is related to... Figure 14c Steps 155 to 157 correspond to this. Calculate the broadband alignment parameters, particularly the preferred broadband alignment parameter, the inter-channel time difference (ITD). Perform a time shift of L and R in the frequency domain. Alternatively, this time shift can also be performed in the time domain. Then perform an inverse DFT, performing the time shift in the time domain, and perform an additional forward DFT to have a spectral representation again after alignment using the broadband alignment parameters.

[0239] For each parameter band in the shifted L and R representations, calculate the ILD parameters, i.e., the sound level parameters and phase parameters (IPD parameters). For example, this step is related to... Figure 14c Step 160 corresponds to this. The time-shifted L and R represent rotation as a function of the inter-channel phase difference parameter, as shown below. Figure 14c The process is described in step 161. Subsequently, the intermediate and side signals are calculated as shown in step 301, and preferably, additionally, energy session operations as discussed later are performed. Furthermore, a prediction of S is performed, utilizing M as a function of ILD and optionally utilizing past M signals, i.e., the intermediate signals of earlier frames. Subsequently, an inverse DFT of the intermediate and side signals is performed, which, in the preferred embodiment... Figure 14d Steps 303, 304, and 305 correspond to each other.

[0240] In the final step, the intermediate time-domain signal m and the optionally residual signal are encoded. This process is related to... Figure 12 The process performed by the signal encoder 400 in the middle corresponds to that of the signal encoder 400.

[0241] In inverse stereo processing, at the decoder, the side signal is generated in the DFT domain and predicted first from the middle signal, as follows:

[0242]

[0243] Where g is the gain calculated for each parameter band and is a function of the inter-channel sound level difference (ILD) of the transmitted audio.

[0244] The predicted residuals Side-g·Mid can then be refined in two different ways:

[0245] - Through secondary encoding of the residual signal:

[0246]

[0247] Where g cod It is the global gain for the entire spectrum transmission.

[0248] - Residual side spectrum is predicted using a residual prediction termed stereo fill, which utilizes the spectrum of the previously decoded Mid signal from the previous DFT frame:

[0249]

[0250] Where g pred It is the predicted gain for transmission in each parameter frequency band.

[0251] The two types of coding refinement can be mixed within the same DFT spectrum. In a preferred embodiment, residual coding is applied to the lower parameter band, while residual prediction is applied to the remaining band. In a preferred embodiment, such as Figure 12 As described, residual side-end signals are synthesized in the time domain, transformed by MDCT, and then residual coding is performed in the MDCT domain. Unlike DFT, MDCT is key-sampled and more suitable for audio coding. MDCT coefficients are directly vector-quantized via lattice vector quantization, but can alternatively be encoded by a scalar quantizer followed by an entropy encoder. Alternatively, the residual side-end signals can also be encoded in the time domain using speech coding techniques, or directly in the DFT domain.

[0252] Another embodiment of combined stereo / multichannel encoder processing or inverse stereo / multichannel processing is then described.

[0253] 1. Time-Frequency Analysis: DFT

[0254] Importantly, the additional time-frequency decomposition derived from stereo processing using the DFT allows for good auditory scene analysis without significantly increasing the overall latency of the coding system. By default, a time resolution of 10 ms is used (twice the 20 ms framing of the core encoder). The analysis and synthesis windows are identical and symmetrical. This window is represented in Figure 7 at a sampling rate of 16 kHz. It can be observed that the overlapping region is limited to reduce the resulting latency, and zero-padding is also added to balance the cyclic shift when the ITD is applied in the frequency domain. This will be explained below.

[0255] 2. Stereo parameters

[0256] Stereo parameters can be transmitted at a maximum time resolution of the stereo DFT. The minimum resolution can be reduced to the core encoder's framing resolution, i.e., 20 ms. By default, parameters are computed every 20 ms over two DFT windows when no transients are detected. The parameter bands are constructed following a non-uniform and non-overlapping decomposition of the spectrum, approximately 2 or 4 times the equivalent rectangular bandwidth (ERB). By default, a 4x ERB scale is used for a total of 12 bands with a frequency bandwidth of 16 kHz (32 kbps sampling rate, ultra-wideband stereo). Figure 8 summarizes an example configuration where stereo side information is transmitted at approximately 5 kbps.

[0257] 3. Calculation of ITD and channel time alignment

[0258] ITD is calculated by estimating the Time Delay of Arrival (TDOA) using phase transform generalized cross-correlation (GCC-PHAT):

[0259]

[0260] Where L and R are the frequencies of the left and right channels, respectively. Frequency analysis can be performed independently of the DFT used for subsequent stereo processing, or it can be shared. The pseudocode for calculating ITD is as follows:

[0261]

[0262] The ITD calculation can also be summarized as follows. Depending on the spectral flatness measurement, the cross-correlation is calculated in the frequency domain before smoothing. The SFM is defined between 0 and 1. In the case of a noise-like signal, the SFM will be high (i.e., approximately 1) and the smoothing will be weak. In the case of a tone-like signal, the SFM will be low and the smoothing will become stronger. The smoothed cross-correlation is then normalized by its amplitude before being transformed back to the time domain. Normalization corresponds to the phase transformation of the cross-correlation and is known to exhibit better performance than normal cross-correlation in low-noise and relatively high-reverberation environments. The time-domain function thus obtained is first filtered to achieve more robust peaking. The index corresponding to the maximum amplitude corresponds to an estimate of the time difference between the left and right channels (ITD). If the amplitude of the maximum value is below a given threshold, then the ITD estimate is not considered reliable and is set to zero.

[0263] If time alignment is applied in the time domain, then the ITD is calculated in a separate DFT analysis. The shifting is performed as follows:

[0264]

[0265] It requires additional latency at the encoder, which is at most equal to the maximum absolute ITD that can be processed. The ITD is smoothed over time using an analysis window derived from the DFT.

[0266] Alternatively, time alignment can be performed in the frequency domain. In this case, the ITD calculation and cyclic shift are in the same DFT domain, a domain shared with this other stereo processing. The cyclic shift is given by:

[0267]

[0268] A zero-filled DFT window is required to simulate time shifting using cyclic shifting. The size of the zero-fill corresponds to the maximum absolute ITD that can be handled. In a preferred embodiment, the zero-fill is evenly split across the analysis window by adding 3.125 ms of zeros at both ends. Thus, the maximum possible absolute value of the ITD is 6.25 ms. In an AB microphone setup, this corresponds to a worst-case scenario with a maximum distance of approximately 2.15 meters between the two microphones. The variation of the ITD over time is smoothed by the synthesis window and the overlap-addition of the DFT.

[0269] The key difference is that the time shift is followed by windowing of the shifted signal. This is a major difference from existing binaural cue coding (BCC), where the time shift is applied to the windowed signal but not further windowed during the synthesis stage. Therefore, any change in ITD over time will produce pseudo-acoustic transients / clicks in the decoded signal.

[0270] 4. IPD calculation and channel rotation

[0271] IPD is calculated after time alignment of the two channels, and depending on the stereo configuration, this is applied to each parameter band or at least up to a given ipd_max_band.

[0272]

[0273] IPD is then applied to both channels to align their phase:

[0274]

[0275] Where β=atan2(sin(IPD) i [b]), cos(IPD) i [b])+c) And b is the parameter band index to which the frequency index k belongs. The parameter β is responsible for distributing the phase rotation amount between the two channels while aligning their phases. β depends on IPD, but also on the relative amplitude sound level ILD of the channels. If a channel has a higher amplitude, then it will be considered the guiding channel, and the effect of phase rotation on it will be less than on a channel with a lower amplitude.

[0276] 5. Sum-difference and side signal encoding

[0277] A sum-difference transformation is performed on the time- and phase-aligned spectra of the two channels in such a way that energy is preserved in the intermediate signal.

[0278]

[0279] in The limit is defined between 1 / 1.2 and 1.2 (i.e., -1.58 and +1.58 dB). This limit avoids artifacts when adjusting the energies of M and S. It is worth noting that this energy conservation is less important when pre-aligning the time and phase. Alternatively, the boundary can be increased or decreased.

[0280] The lateral signal S is further predicted using M:

[0281] S′(f)=S(f)-g(ILD)M(f)

[0282] in in Alternatively, the optimal prediction gain g can be found by minimizing the residuals and the mean square error (MSE) of the ILD derived from the previous equation.

[0283] The residual signal S′(f) can be modeled in two ways: by predicting it using the delayed spectrum of M or by directly encoding it in the MDCT domain.

[0284] 6. Stereo decoding

[0285] The center signal X and the side signal S are first converted into the left and right channel L and R as follows:

[0286] L i [k]=M i [k]+gM i [k], for band_limits[b]≤k<band_limits[b+1],

[0287] R i [k]=M i [k]-gM i [k], for band_limits[b]≤k<band_limits[b+1],

[0288] The gain g for each frequency band is derived from the ILD parameters:

[0289] in

[0290] For parameter bands below cod_max_band, update both channels using the decoded side signals:

[0291] L i [k]=L i [k]+cod_gain i ·S i [k], for 0 ≤ k < band_limits[cod_max_band], R i [k]=R i [k]-cod_gain i ·S i [k], for 0 ≤ k < band_limits[cod_max_band],

[0292] For higher parameter frequency bands, predict the side signals and update the channels as follows:

[0293] L i [k]=L i [k]+cod_pred i [b]·M i-1 [k], for band_limits[b]≤k<band_limits[b+1],

[0294] R i [k=R i [k]-cod_pred i [b]·M i-1[k], for band_limits[b]≤k<band_limits[b+1],

[0295] Finally, the channels are multiplied by complex values ​​to recover the original energy and inter-channel phase of the stereo signal:

[0296]

[0297]

[0298] in

[0299]

[0300] Where 'a' is defined and delimited as previously defined, and where β = atan2(sin(IPD) i [b]), cos(IPD) i [b])+c), where atan2(x,y) is the arctangent of x with respect to y in all four quadrants.

[0301] Finally, depending on the transmitted ITD, the channels are time-shifted in the time or frequency domain. Time-domain channels are synthesized using inverse DFT and overlap-addition.

[0302] The encoded audio signal of the present invention can be stored on a digital storage medium or a non-transitory storage medium, or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium (the Internet).

[0303] Although some aspects have been described in the context of the apparatus, it is clear that these aspects also represent descriptions of the corresponding methods, where boxes or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent descriptions of corresponding boxes, items, or features of the corresponding apparatus.

[0304] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementation can be performed using a digital storage medium on which electronically readable control signals are stored, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, the electronically readable control signals cooperating with (or being able to cooperate with) a programmable computer system to cause the corresponding methods to be executed.

[0305] Some embodiments of the invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.

[0306] Generally, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operable to perform one of these methods. The program code may, for example, be stored on a machine-readable medium.

[0307] Other embodiments include a computer program for performing one of the methods described herein, stored on a machine-readable carrier or non-transient storage medium.

[0308] In other words, embodiments of the method of the present invention are therefore computer programs having program code that, when run on a computer, performs one of the methods described herein.

[0309] Therefore, another embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein.

[0310] Therefore, another embodiment of the method of the present invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection, such as via the Internet.

[0311] Another embodiment includes a processing means, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.

[0312] Another embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.

[0313] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.

[0314] The above embodiments are merely illustrative of the principles of the invention. It is to be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the intent is limited only by the scope of the following claims, and not by the specific details given in the description and explanation of the embodiments herein.

Claims

1. An apparatus for encoding a multichannel signal comprising at least two channels, wherein the multichannel signal is a multichannel audio or speech signal, the apparatus comprising: A time-to-spectrum converter (1000) is used to convert a sequence of blocks of sampled values ​​of the at least two channels into a frequency domain representation of a sequence of blocks having spectral values ​​for the at least two channels; A multi-channel processor (1010) is used to apply joint multi-channel processing to a sequence of blocks of spectral values ​​to obtain at least one resulting sequence of blocks of spectral values ​​including information related to the at least two channels; A spectrum-to-time converter (1030) is used to convert a result sequence of blocks of spectrum values ​​into a time-domain representation of an output sequence that includes blocks of sampled values; The core encoder (1040) is used to encode the output sequence of blocks of sampled values ​​to obtain an encoded multichannel signal (1510). The core encoder (1040) is configured to operate according to a first frame control to provide a frame sequence, wherein the frames are defined by a start frame boundary (1901) and an end frame boundary (1902), and The time-to-spectrum converter (1000) or the spectrum-to-time converter (1030) is configured to operate according to a second frame control synchronized with the first frame control.

2. The apparatus of claim 1, wherein the analysis window used by the time-to-spectrum converter (1000) or the synthesis window used by the spectrum-to-time converter (1030) each has an increased overlap portion and a decreased overlap portion, wherein the core encoder (1040) comprises a time-domain encoder having a leading portion (1905) or a frequency-domain encoder having the overlap portion of the core window, and The overlapping portion of the analysis window or the synthesis window is less than or equal to the preceding portion (1905) of the core encoder or the overlapping portion of the core window.

3. The apparatus as described in claim 1, The core encoder (1040) is configured to use a lead portion (1905) when performing core encoding on frames derived from the output sequence of blocks with associated output sampling rates, the lead portion (1905) being temporally located after the frame. The time-to-spectrum converter (1000) is configured to use an analysis window (1904) with an overlapping portion, the time length of which is less than or equal to the time length of the preceding portion (1905), wherein the overlapping portion of the analysis window is used to generate the windowed preceding portion (1905).

4. The apparatus as described in claim 3, The spectrum-to-time converter (1030) is configured to process the output leading portion corresponding to the windowed leading portion using a correction function (1922), wherein the correction function is configured to reduce or eliminate the influence of the overlapping portion of the analysis window.

5. The apparatus as described in claim 4, The correction function is the inverse of the function that defines the overlapping portion of the analysis window.

6. The apparatus as described in claim 4 or 5, The overlapping portion is proportional to the square root of the sine function. The correction function is proportional to the reciprocal of the square root of the sine function, and The spectrum-to-time converter (1030) is configured to use an overlap portion proportional to a sine function raised to the power of 1.

5.

7. The apparatus as claimed in claim 1, The spectrum-to-time converter (1030) is configured to generate a first output block using a synthesis window and a second output block using the synthesis window, wherein a second portion of the second output block is an output-ahead portion (1905). The spectrum-to-time converter (1030) is configured to generate sampled values ​​of a frame using an overlap-addition operation between the first output block and another portion of the second output block, excluding the output lead portion (1905). The core encoder (1040) is configured to apply a lookahead operation to the output lookahead portion (1905) to determine encoding information for core encoding of the frame, and The core encoder (1040) is configured to perform core encoding on the frame using the result of the preceding operation.

8. The apparatus as claimed in claim 7, The spectrum-to-time converter (1030) is configured to generate a third output block following the second output block using the synthesis window, wherein the spectrum-to-time converter is configured to overlap a first overlapping portion of the third output block with a second portion of the second output block using the synthesis window to obtain a sample of another frame that follows the frame in time.

9. The apparatus as claimed in claim 8, in, The spectrum-to-time converter (1030) is configured not to window or modify the output-ahead portion (1922) when generating the second output block for the frame, for at least partially undoing the effects of the analysis window used by the time-to-spectrum converter (1000), and The spectrum-to-time converter (1030) is configured to perform an overlap-addition operation (1924) between the second output block and the third output block for the other frame and to window the output lead portion (1920) using the synthesis window.

10. The apparatus as claimed in claim 1, The spectrum-to-time converter (1030) is configured to Use a synthesis window to generate the first block and the second block of the output sample. Overlap and add the second part of the first block and the first part of the second block to generate a portion of the output sample. The core encoder (1040) is configured to apply a look-ahead operation to the portion of the output sample for core encoding of the output sample that is located in time before the portion of the output sample, wherein the look-ahead portion does not include the second portion of the sample of the second block.

11. The apparatus as claimed in claim 1, The spectrum-to-time converter (1030) is configured to use a synthesis window that provides twice the time resolution of the core encoder frame. The spectrum-to-time converter (1030) is configured to use the synthesis window to generate blocks of output samples and perform an overlap-add operation, wherein the overlap-add operation is used to compute all samples in the advance portion of the core encoder, or The spectrum-to-time converter (1030) is configured to apply a look-ahead operation to the output samples for core encoding of the output samples that are located before the portion in time, wherein the look-ahead portion does not include the second portion of the samples of the second block.

12. The apparatus as claimed in claim 1, The blocks of sampled values ​​have an associated input sampling rate, and the blocks of spectral values ​​of the sequence of spectral values ​​have spectral values ​​up to the maximum input frequency (1211) associated with the input sampling rate. The device further includes a spectral domain resampler (1020) for performing a resampling operation in the frequency domain on data input to the spectrum-to-time converter (1030) or on data input to the multichannel processor (1010), wherein the blocks of resampled sequences of spectral values ​​have spectral values ​​up to the maximum output frequencies (1231, 1221) different from the maximum input frequency (1211); The output sequence of the sampled values ​​has an associated output sampling rate that is different from the input sampling rate.

13. The apparatus as claimed in claim 12, The spectral domain resampler (1020) is configured to truncate the block for downsampling or to zero-padded the block for upsampling.

14. The apparatus as claimed in claim 12 or 13, The spectral domain resampler (1020) is configured to scale (1322) the spectral values ​​of the resulting sequence of blocks of blocks using a scaling factor that depends on the maximum input frequency and the maximum output frequency.

15. The apparatus as claimed in claim 14, in, In the case of upsampling, the scaling factor is greater than 1, where the output sampling rate is greater than the input sampling rate; or in the case of downsampling, the scaling factor is less than 1, where the output sampling rate is less than the input sampling rate. The time-to-spectrum converter (1000) is configured to perform a time-to-frequency transformation algorithm (1311) without using normalization of the total number of spectral values ​​of the block of spectral values, and the scaling factor is equal to the quotient between the number of spectral values ​​of the block of the resampled sequence and the number of spectral values ​​of the block of spectral values ​​before resampling, and the spectrum-to-time converter is configured to apply normalization based on the maximum output frequency (1331).

16. The apparatus as claimed in claim 1, The time-to-spectrum converter (1000) is configured to perform a discrete Fourier transform algorithm, or the spectrum-to-time converter (1030) is configured to perform an inverse discrete Fourier transform algorithm.

17. The apparatus as claimed in claim 1, The multi-channel processor (1010) is configured to obtain a further sequence of resulting blocks of spectral values, and The spectrum-to-time converter (1030) is configured to convert the additional resulting sequence of spectrum values ​​into an additional time-domain representation (1032) of an additional output sequence of blocks of sampled values ​​having an associated output sampling rate equal to the input sampling rate.

18. The apparatus of claim 12, The multi-channel processor (1010) is configured to provide a further sequence of resulting blocks of spectral values. The spectral domain resampler (1020) is configured to resample blocks of the additionally resulting sequence in the frequency domain to obtain additional resampled sequences of blocks of spectral values, wherein the additionally resampled sequences of blocks have spectral values ​​up to an additional maximum output frequency different from the maximum input frequency or different from the maximum output frequency. The spectrum-to-time converter (1030) is configured to convert a further resampled sequence of blocks of spectrum values ​​into a further time-domain representation of a further output sequence of blocks of sampled values, the further output sequence of blocks of sampled values ​​having an associated further output sampling rate different from the input sampling rate or the output sampling rate.

19. The apparatus as claimed in claim 1, The multichannel processor (1010) is configured to generate at least one resulting sequence of blocks of intermediate signals as spectral values ​​using only downmixing operations, or to generate another resulting sequence of blocks of additional side signals as spectral values.

20. The apparatus of claim 12, The multi-channel processor (1010) is configured to generate an intermediate signal as the at least one resulting sequence, and the spectral domain resampler (1020) is configured to resample the intermediate signal into two independent sequences having two different maximum output frequencies different from the maximum input frequency. The spectrum-to-time converter (1030) is configured to convert two resampled sequences into two output sequences with different sampling rates. The core encoder (1030) includes a first preprocessor (1430c) for preprocessing the first output sequence at a first sampling rate and a second preprocessor (1430d) for preprocessing the second output sequence at a second sampling rate, and The core encoder is configured to perform core encoding on either the first output sequence or the second output sequence, or The multi-channel processor is configured to generate a side signal as the at least one resulting sequence, and the spectral domain resampler (1020) is configured to resample the side signal into two resampled sequences having two different maximum output frequencies different from the maximum input frequency. The spectrum-to-time converter (1030) is configured to convert two resampled sequences into two output sequences with different sampling rates. The core encoder includes a first preprocessor (1430c) and a second preprocessor (1430d) for preprocessing the first and second output sequences; and The core encoder (1040) is configured to perform core encoding (1430a, 1430b) on either the first preprocessed output sequence or the second preprocessed output sequence.

21. The apparatus as claimed in claim 1, The spectrum-to-time converter (1030) is configured to convert the at least one resulting sequence into a time-domain representation without performing any spectrum-domain resampling. The core encoder (1040) is configured to perform core encoding (1430a) on the non-resampled output sequence to obtain an encoded multi-channel signal, or The spectrum-to-time converter (1030) is configured to convert the at least one resulting sequence into a time-domain representation without any spectral domain resampling in the absence of side signals, and The core encoder (1040) is configured to perform core encoding (1430a) on the non-resampled output sequence for the side signal to obtain an encoded multi-channel signal, or The device also includes a specific spectrum domain side signal encoder (1430e), or The input sampling rate is at least one of the sampling rates in the group consisting of 8kHz, 16kHz, and 32kHz, or The output sampling rate is at least one of the sampling rates in the group consisting of 8kHz, 12.8kHz, 16kHz, 25.6kHz and 32kHz.

22. The apparatus as claimed in claim 1, The time-to-spectrum converter is configured as an application analysis window. The spectrum-to-time converter (1030) is configured to apply a synthesis window. The analysis window's duration is equal to the synthesis window's duration, or is an integer multiple or fraction of the synthesis window's duration, or... The analysis window and the synthesis window each have zero-padding portions at their initial or final portions, or The analysis window and the synthesis window are configured such that the window size, overlap area size, and zero-fill size each include an integer number of samples for at least two sampling rates in a sampling rate group including 12.8 kHz, 16 kHz, 25.6 kHz, 32 kHz, and 48 kHz. The maximum cardinality of the digital Fourier transform in the split cardinality implementation is less than or equal to 7, or the time resolution is fixed to a value less than or equal to the frame rate of the core encoder.

23. The apparatus as claimed in claim 1, The multi-channel processor (1010) is configured to process a sequence of blocks to obtain time alignment using a wideband time alignment parameter (12), and to obtain narrowband phase alignment using multiple narrowband phase alignment parameters (14), and to calculate intermediate and side signals as a result sequence using the aligned sequence.

24. The apparatus of claim 1, wherein the start frame boundary (1901) or end frame boundary (1902) of each frame of the frame sequence is predetermined to have a predetermined relationship with the start or end time of the overlapping portion of the window used by the time-to-spectrum converter (1000) for each block of the sequence of sampled values ​​or used by the spectrum-to-time converter (1030) for each block of the output sequence of sampled values, or The multi-channel processor (1010) is configured to perform downmixing operations.

25. A method for encoding a multichannel signal comprising at least two channels, wherein the multichannel signal is a multichannel audio or speech signal, the method comprising: The sequence of blocks of sampled values ​​from the at least two channels is time-spectral transformed (1000) into a frequency domain representation of the sequence of blocks having spectral values ​​for the at least two channels; (1010) joint multichannel processing is applied to the sequence of blocks of spectral values ​​to obtain at least one resulting sequence of blocks of spectral values ​​including information related to the at least two channels; The result sequence of the block of spectral values ​​is transformed into a time-domain representation of the output sequence including the block of sampled values ​​by a spectral-time transformation (1640). as well as The output sequence of the sampled values ​​is core encoded (1040) to obtain the encoded multi-channel signal (1510). The core coding (1040) operates according to the first frame control to provide a frame sequence, where frames are defined by a start frame boundary (1901) and an end frame boundary (1902), and The time-to-spectrum conversion (1000) or spectrum-to-time conversion (1030) is operated according to the second frame control synchronized with the first frame control.

26. An apparatus for decoding an encoded multichannel signal, wherein the encoded multichannel signal is a multichannel audio or speech signal, the apparatus comprising: The core decoder (1600) is used to generate the core-decoded signal; A time-to-spectrum converter (1610) is used to convert a sequence of blocks of sampled values ​​of the core-decoded signal into a frequency domain representation, the frequency domain representation being a sequence of blocks having spectral values ​​for the core-decoded signal; A multichannel processor (1630) is configured to apply inverse multichannel processing to a sequence (1615) comprising a block, to obtain at least two resulting sequences (1631, 1632, 1635) of the block's spectral values; and A spectrum-to-time converter (1640) is provided for converting at least two result sequences (1631, 1632) of a block of spectrum values ​​into a time-domain representation of at least two output sequences of a block of sampled values. The core decoder (1600) is configured to operate according to the control of the first frame to provide a frame sequence, wherein the frames are defined by a start frame boundary (1901) and an end frame boundary (1902). The time-to-spectrum converter (1610) or the spectrum-to-time converter (1640) is configured to operate according to a second frame control synchronized with the first frame control.

27. The apparatus of claim 26, The core-decoded signal has a frame sequence, with each frame having a start frame boundary (1901) and an end frame boundary (1902). The analysis window (1914) used by the time-to-spectrum converter (1610) for windowing the frames of the frame sequence has an overlapping portion that ends before the end frame boundary (1902), thus leaving a time gap (1920) between the end of the overlapping portion and the end frame boundary (1902), and The core decoder (1600) is configured to perform processing on samples in the time gap (1920) in parallel with the opening of a frame using the analysis window (1914), or to perform core decoder post-processing on samples in the time gap (1920) in parallel with the opening of a frame using the analysis window.

28. The apparatus of claim 26, The core-decoded signal has a frame sequence, with each frame having a start frame boundary (1901) and an end frame boundary (1902). The first overlapping portion of the analysis window (1914) begins at the start of the start frame boundary (1901), and the second overlapping portion of the analysis window (1914) ends before the end frame boundary (1902), such that a time gap (1920) exists between the end of the second overlapping portion and the end frame boundary. The analysis window is positioned for subsequent blocks of the core-decoded signal such that the middle non-overlapping portion of the analysis window is located within the time gap (1920).

29. The apparatus of claim 26, The analysis window used by the time-to-spectrum converter (1610) has the same shape and time length as the synthesis window used by the spectrum-to-time converter (1640).

30. The apparatus of claim 26, The core-decoded signal has a sequence of frames, wherein each frame has a length, and the time-to-spectrum converter (1610) is configured to use a window, wherein the length of the window, excluding any zero-padding portions, is less than or equal to half the length of the frame.

31. The apparatus of claim 26, The spectrum-to-time converter (1640) is configured to A synthesis window is applied to the first output sequence of the at least two output sequences to obtain the first output block of the windowed sample; A synthesis window is applied to the first output sequence of the at least two output sequences to obtain a second output block of the windowed sample. Overlap and add the first output block and the second output block to obtain a first group of output samples for the first output sequence; The spectrum-to-time converter (1640) is configured to A synthesis window is applied to the second output sequence of the at least two output sequences to obtain a first output block of the windowed sample; A synthesis window is applied to the second output sequence of the at least two output sequences to obtain a second output block of the windowed sample; Overlap and add the first output block and the second output block to obtain a second group of output samples for the second output sequence; The first group of output samples for the first output sequence and the second group of output samples for the second output sequence are related to the same time portion of the encoded multichannel signal or to the same frame of the signal decoded by the core.

32. The apparatus of claim 26, The blocks of sampled values ​​have an associated input sampling rate, and the blocks of spectral values ​​have spectral values ​​up to the maximum input frequency associated with the input sampling rate; The device further includes a spectral domain resampler (1620) for performing a resampling operation on data input to the spectrum-to-time converter (1640) or data input to the multichannel processor (1630) in the frequency domain, wherein the blocks of the resampled sequence have spectral values ​​up to a maximum output frequency different from the maximum input frequency; The block of sampled values ​​has at least two output sequences with associated output sampling rates that are different from the input sampling rates.

33. The apparatus as claimed in claim 32, The spectral domain resampler (1020) is configured to truncate the block for downsampling or to zero-padded the block for upsampling.

34. The apparatus as claimed in claim 32, The spectral domain resampler (1020) is configured to scale (1322) the spectral values ​​of the resulting sequence of blocks of blocks using scaling factors that depend on the maximum input frequency and the maximum output frequency.

35. The apparatus as claimed in claim 34, In the case of upsampling, the scaling factor is greater than 1, where the output sampling rate is greater than the input sampling rate; or in the case of downsampling, the scaling factor is less than 1, where the output sampling rate is lower than the input sampling rate. The time-to-spectrum converter (1000) is configured to perform a time-to-frequency transformation algorithm (1311) without using normalization of the total number of spectral values ​​of the block of spectral values, and the scaling factor is equal to the quotient between the number of spectral values ​​of the block of the resampled sequence and the number of spectral values ​​of the block of spectral values ​​before resampling, and the spectrum-to-time converter is configured to apply normalization based on the maximum output frequency (1331).

36. The apparatus of claim 26, The time-to-spectrum converter (1000) is configured to perform a discrete Fourier transform algorithm, or the spectrum-to-time converter (1030) is configured to perform an inverse discrete Fourier transform algorithm.

37. The apparatus of claim 32, The core decoder (1600) is configured to generate an additional core-decoded signal (1601) with a different sampling rate than the input sampling rate. The time-to-spectrum converter (1610) is configured to convert the additional core-decoded signal into a frequency domain representation of another sequence (1611) of blocks having spectral values ​​for the additional core-decoded signal, wherein the blocks of spectral values ​​for the additional core-decoded signal have spectral values ​​up to an additional maximum input frequency that is different from the maximum input frequency and related to the additional sampling rate. The spectral domain resampler (1620) is configured to resample a further sequence of blocks of the additional core-decoded signal in the frequency domain to obtain a further resampled sequence (1621) of blocks of spectral values, wherein the blocks of spectral values ​​of the further resampled sequence have spectral values ​​up to a maximum output frequency different from the additional maximum input frequency; and The device further includes a combiner (1700) for combining the resampled sequence and the additional resampled sequence to obtain a sequence (1701) to be processed by the multichannel processor (1630).

38. The apparatus of claim 26, The core decoder (1600) is configured to generate a signal that is further core-decoded, the further core-decoded signal having an additional sampling rate equal to the output sampling rate (1603). The time-to-spectrum converter (1610) is configured to convert the further core-decoded signal into a frequency domain representation (1613) to obtain a further sequence of blocks with spectral values. The apparatus further includes a combiner (1700) for combining a further sequence of blocks of the spectral values ​​and a resampled sequence of blocks (1622, 1621) in the process of generating a sequence of blocks processed by the multichannel processor (1630).

39. The apparatus of claim 26, The core decoder (1600) includes at least one of the following: an MDCT-based decoding section (1600d), a temporal bandwidth extension decoding section (1600c), an ACELP decoding section (1600b), and a bass post-filter decoding section (1600a). The MDCT-based decoding section (1600d) or the time-domain bandwidth-extended decoding section (1600c) is configured to generate a core-decoded signal with an output sampling rate, or The ACELP decoding section (1600b) or the bass post-filter decoding section (1600a) is configured to generate a core-decoded signal at a sampling rate different from the output sampling rate.

40. The apparatus of claim 26, The time-to-spectrum converter (1610) is configured to apply analysis windows to at least two of a plurality of different core-decoded signals, the analysis windows having the same size in time or the same shape in time. The device further includes a combiner (1700) for combining at least one resampled sequence and any other sequence of blocks having spectral values ​​up to the maximum output frequency on a block-by-block basis to obtain a sequence processed by the multichannel processor (1630).

41. The apparatus of claim 26, The sequence processed by the multi-channel processor (1630) corresponds to the intermediate signal, and The multi-channel processor (1630) is configured to additionally generate side signals using information about side signals included in the encoded multi-channel signal, and The multi-channel processor (1630) is configured to generate at least two result sequences using the intermediate signal and the side signal.

42. The apparatus of claim 26, The multi-channel processor (1630) is configured to convert (820) the sequence into a first sequence for a first output channel and a second sequence for a second output channel using a gain factor for each parameter band; The first and second sequences are updated (830) using decoded side signals, or the first and second sequences are updated using side signals predicted from an earlier block of the sequence for the intermediate signal using stereo fill parameters for the parameter band. (910) phase dealignment and energy scaling are performed using information about multiple narrowband phase alignment parameters; as well as Perform (920) time dealignment using information about the broadband time alignment parameters to obtain at least two result sequences.

43. The apparatus of claim 26, wherein the start frame boundary (1901) or end frame boundary (1902) of each frame of the frame sequence is predetermined to have a predetermined relationship with the start or end time of the overlapping portion of the windows of each block of the sequence of blocks of sampled values ​​used by the time-to-spectrum converter (1610) or each block of at least two output sequences of blocks of sampled values ​​used by the spectrum-to-time converter (1640), or The multi-channel processor (1630) is used to perform upmixing operations.

44. A method for decoding an encoded multichannel signal, wherein the encoded multichannel signal is a multichannel audio or speech signal, the method comprising: Generate (1600) core-decoded signals; The sequence of blocks of sampled values ​​of the core-decoded signal is time-to-spectral transformed (1610) into a frequency domain representation, the frequency domain representing a sequence of blocks having spectral values ​​for the core-decoded signal; Inverse multichannel processing (1630) is applied to a sequence (1615) comprising a block to obtain at least two resulting sequences (1631, 1632, 1635) of the block with spectral values. The spectrum-time transformation (1640) of at least two result sequences (1631, 1632) of the block of spectral values ​​is performed into a time-domain representation of at least two output sequences of the block of sampled values. The generation of the core-decoded signal (1600) is controlled according to the first frame to provide a frame sequence, wherein the frames are defined by a start frame boundary (1901) and an end frame boundary (1902). The time-to-spectrum conversion (1610) or the spectrum-to-time conversion (1640) operates according to the second frame control synchronized with the first frame control.

45. A computer program, when run on a computer or processor, for performing the method of claim 25 or the method of claim 44.

Citation Information

Patent Citations

  • Near-transparent or transparent multi-channel encoder / decoder scheme

    WO2006089570A1

  • Mdct-based complex prediction stereo coding

    CN104851427A

  • High-television system

    CN1226358A