SWITCHING BETWEEN STEREO ENCODING MODES IN A MULTICHANNEL SOUND CODEC
Patent Information
- Application Number
- MX2022009501
- Authority / Receiving Office
- MX · MX
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-03
- Filing Date
- 2022-08-02
- Publication Date
- 2026-02-25
- Estimated Expiration
- 2041-02-01
Smart Images

Figure MX431055B0
Abstract
Description
SWITCHING BETWEEN STEREO ENCODING MODES IN A MULTI-CHANNEL SOUND CODEC TECHNICAL FIELD
[0001] The present invention relates to stereo sound coding, in particular but not exclusively to switching between stereo coding modes (hereinafter also stereo modes) in a multi-channel sound codec capable, in particular but not exclusively, of producing good stereo quality, for example in a complex audio scene at low bit rate and low delay.
[0002] In the present invention and in the appended claims: - The term sound can be related to voice, audio and any other sound; - The term stereo is an abbreviation for stereophonic; and - The term mono is an abbreviation for monophonic. BACKGROUND
[0003] Historically, conversational telephony has been implemented with headsets that had a single transducer to output sound to only one of the user's ears. In the last decade, users have begun to use their handset in conjunction with a headset to receive sound in both of their ears, primarily for listening to music, but also sometimes for listening to voice. However, when a handset is used to transmit and receive conversational voice, the content is still in mono, but is presented to both of the user's ears when using a headset.
[0004] With the latest 3GPP voice coding standard described in reference [1], the entire contents of which are incorporated herein by reference, the quality of encoded sound, e.g., voice and / or audio transmitted and received via a handset, has been significantly improved. The next natural step is to transmit stereophonic information in such a way that the receiver gets as close as possible to a real-life audio scene being picked up at the other end of the communication link.
[0005] In audio codecs, for example those described in Reference [2], the entire contents of which are incorporated herein by reference, stereo information transmission is typically used.
[0006] For conversational speech codecs, a mono signal is the norm. When transmitting a stereo signal, it is often necessary to double the bit rate, since the left and right channels of the stereo signal are encoded using a mono codec. This works well in most cases, but has the drawback of doubling the bit rate and not taking advantage of the possible redundancy between the two channels (left and right of the stereo signal). Furthermore, to keep the overall bit rate at a reasonable level, a very low bit rate is used for each channel, which affects the overall sound quality. To reduce the bit rate, efficient stereo coding techniques have been developed and used. As non-mono examples, LncAnn / zznz / E / YiAi limiting, the following paragraphs discuss the use of three stereo coding techniques that can be efficiently used at low bit rates.
[0007] A first stereo coding technique is called parametric stereo. Parametric stereo coding encodes two channels, left and right, as a mono signal using a common mono codec plus a certain amount of stereo side information (corresponding to the stereo parameters) representing a stereo image. The two input channels, left and right, are down-mixed (adj., down-mixed) into a mono signal, and the stereo parameters are typically computed in the transform domain, for example in the Discrete Fourier Transform (DFT) domain, and related to so-called binaural or inter-channel cues. Binaural cues (Reference [3], the entire contents of which are incorporated herein by reference) comprise interaural level difference (ILD), interaural time difference (ITD), and interaural correlation (IC).Depending on the signal characteristics, the stereo scene configuration, etc., some or all of the binaural cues are encoded and transmitted to the decoder. Information about which binaural cues are encoded and transmitted is sent as signaling information, which is usually part of the stereo side information. A given binaural signal can also be quantized using different coding techniques, resulting in a variable number of bits. In addition to the quantized binaural signals, the stereo side information may contain, typically at medium and high bit rates, a quantized residual signal resulting from down-mixing. The residual signal can be encoded using an entropy coding technique, e.g., an arithmetic coder.Parametric stereo coding with stereo parameters calculated in a transform domain will be referred to herein as DFT stereo coding.
[0008] Another stereo coding technique is one that operates in the time domain (TD). This stereo coding technique mixes the two input channels, left and right, into so-called primary and secondary channels. For example, following the method described in Reference [4], the entire contents of which are incorporated herein by reference, the time domain mixing may be based on a mixing ratio, which determines the respective contributions of the two input channels, left and right, to the production of the primary channel and the secondary channel. The mixing ratio is derived from various metrics, for example, normalized correlations of the input left and right channels with respect to a mono signal version or a long-term correlation difference between the two input left and right channels.The primary channel can be encoded by a common mono codec, while the secondary channel can be encoded by a lower-bit-rate codec. Secondary channel coding can take advantage of coherence between the primary and secondary channels and can reuse some parameters of the primary channel. Time-domain stereo coding will be referred to in the following. LncAnn / zznz / E / YiAi present invention TD stereo coding. In general, TD stereo coding is more efficient at low and medium bit rates for coding speech signals.
[0009] A third stereo coding technique is one that operates in the Modified Discrete Cosine Transform (MDCT) domain. It is based on joint coding of the left and right channels while calculating the overall ILD and mid / side (M / S) processing in the whitened spectral domain.This third stereo coding technique uses several tools adapted from TCX (Transform Coded Excitation) coding in MPEG (Moving Picture Experts Group) codecs as described for example in references [6] and [7], the entire contents of which are incorporated herein by reference; these tools may include TCX kernel coding, TCX LTP (Long Term Prediction) analysis, TCX noise filling, frequency domain noise layering (FDNS), intelligent stereo gap filling (IGF), and / or adaptive channel bit allocation. In general, this third stereo coding technique is effective for encoding all types of audio content at medium and high bit rates. The MDCT stereo coding technique will be referred to herein as MDCT stereo coding.In general, MDCT stereo coding is most efficient at medium and high bit rates for coding general audio signals.
[0010] In recent years, stereo coding has been extended to multi-channel coding. There are several techniques for providing multi-channel coding, but the fundamental core of all these techniques is often based on one or more instances of mono or stereo coding techniques. Thus, the present invention features switching between stereo coding modes that may be part of multi-channel coding techniques, such as Metadata Assisted Spatial Audio (MASA), as described, for example, in reference [8], the entire contents of which are incorporated herein by reference.In the MASA approach, MASA metadata (e.g., direction, energy ratio, propagation coherence, distance, surround coherence, all at multiple time-frequency slots) are generated in a MASA analyzer, quantized, encoded, and passed to the bitstream, while the MASA audio channel(s) are treated as (multi-) mono or (multi-) stereo transport signals encoded by the core encoder(s). In the MASA decoder, the MASA metadata guides the decoding and rendering process to recreate an output spatial sound. LncAnn / zznz / E / YiAi SUMMARY
[0011] The present invention provides devices and methods for encoding stereo sound signals as defined in the appended claims.
[0012] The foregoing and other objects, advantages and features of the stereo encoding and decoding devices and methods will become more apparent upon reading the following non-restrictive description of illustrative embodiments thereof, given by way of example only with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In the attached drawings:
[0014] Figure 1 is a schematic block diagram of a sound processing and communication system representing a possible implementation context of the stereo encoding and decoding devices and methods;
[0015] Figure 2 is a high-level block diagram simultaneously illustrating an Immersive Voice and Audio Services (IVAS) stereo coding device and a corresponding stereo coding method, wherein the IVAS stereo coding device comprises a frequency domain (FD) stereo encoder, a time domain (TD) stereo encoder, and a modified discrete cosine transform (MDCT) stereo encoder, wherein the implementation of the FD stereo encoder is based on the discrete Fourier transform (DFT) (hereinafter DFT stereo encoder) in this illustrative embodiment and in the accompanying drawings;
[0016] Figure 3 is a block diagram simultaneously illustrating the DFT stereo encoder of Figure 2 and a corresponding DFT stereo coding method;
[0017] Figure 4 is a block diagram simultaneously illustrating the TD stereo encoder of Figure 2 and the corresponding TD stereo coding method;
[0018] Figure 5 is a block diagram simultaneously illustrating the MDCT stereo encoder of Figure 2 and the corresponding MDCT stereo coding method;
[0019] Figure 6 is a flowchart illustrating processing operations in the IVAS stereo coding device and method when switching from a TD stereo mode to a DFT stereo mode;
[0020] Figure 7a is a flowchart illustrating processing operations in the IVAS stereo coding device and method when switching from the DFT stereo mode to the TD stereo mode;
[0021] Figure 7b is a flowchart illustrating processing operations related to signals passed in TD stereo when switching from DFT stereo mode to TD stereo mode;
[0022] Figure 8 is a high-level block diagram simultaneously illustrating an IVAS stereo decoding device and a corresponding decoding method, wherein the IVAS stereo decoding device comprises a DFT stereo decoder, a TD stereo decoder, and an MDCT stereo decoder;
[0023] Figure 9 is a flowchart illustrating processing operations in the IVAS stereo decoding device and method when switching from the TD stereo mode to the DFT stereo mode; LncAnn / zznz / E / YiAi
[0024] Figure 10 is a flowchart illustrating a case B) of Figure 9, comprising updating DFT stereo synthesis memories in a TD stereo frame on the decoder side;
[0025] Figure 11 is a flowchart illustrating a case C) of Figure 9, comprising smoothing an output stereo synthesis in the first DFT stereo frame after switching from the TD stereo mode to the DFT stereo mode, on the decoder side;
[0026] Figure 12 is a flowchart illustrating processing operations in the IVAS stereo decoding device and method when switching from the DFT stereo mode to the TD stereo mode;
[0027] Figure 13 is a flowchart illustrating a case A) of Figure 12, comprising updating a TD stereo synchronization memory in a first TD stereo frame after switching from the DFT stereo mode to the TD stereo mode, on the decoder side; and
[0028] Figure 14 is a simplified block diagram of an example configuration of the hardware components implementing each of the IVAS stereo encoding device and method and the IVAS stereo decoding device and method. DETAILED DESCRIPTION
[0029] As mentioned above, the present invention relates to stereo sound coding, in particular, but not exclusively, to switching between stereo coding modes in a sound codec, including speech and / or audio, capable in particular, but not exclusively, of producing good stereo quality, for example, in a complex audio scene at low bit rate and low delay. In the present invention, a complex audio scene includes situations, for example but not exclusively, where (a) the correlation between the sound signals being recorded by the microphones is low, (b) there is significant fluctuation of the background noise, and / or (c) there is an interfering speaker.Non-limiting examples of complex audio scenarios include a large anechoic conference room with an A / B microphone setup, a small echoic room with binaural microphones, and a small echoic room with a mono / side microphone setup. All of these room configurations may include fluctuating background noise and / or interfering speakers.
[0030] Figure 1 is a schematic block diagram of a stereo sound processing and communication system 100 depicting a possible implementation context of the IVAS stereo encoding device and method and the IVAS stereo decoding device and method.
[0031] The stereo sound processing and communication system 100 of Figure 1 supports transmission of a stereo sound signal via a communication link 101. The communication link 101 may comprise, for example, a cable or a fiber optic link. Alternatively, the communication link 101 may comprise, at least in part, a radio frequency link. The radio frequency link often supports multiple, simultaneous communications requiring communication resources. LncRnn / zznz / E / YiAi shared bandwidth, as may be the case with cellular telephony. Although not shown, the communication link 101 may be replaced by a storage device in an implementation of the system 100 that records and stores the encoded stereo sound signal for later playback.
[0032] Continuing with Figure 1, for example, a pair of microphones 102 and 122 produces the left 103 and right 123 channels of an original stereo sound signal. As indicated in the previous description, the sound signal may comprise, in particular, but not exclusively, voice and / or audio.
[0033] The left 103 and right 123 channels of the original analog sound signal are supplied to an analog-to-digital (A / D) converter 104 to convert them to the left 105 and right 125 channels of an original digital stereo sound signal. The left 105 and right 125 channels of the original digital stereo sound signal may also be recorded and supplied from a storage device (not shown).
[0034] A stereo sound encoder 106 encodes the left 105 and right 125 channels of the original digital stereo sound signal, thereby producing a set of encoding parameters that are multiplexed into a bit stream 107 delivered to an optional error correction encoder 108. The optional error correction encoder 108, when present, adds redundancy to the binary representation of the encoding parameters in the bit stream 107 before transmitting the resulting bit stream I1 over the communication link 101.
[0035] At the receiver side, an optional error correction decoder 109 uses the aforementioned redundant information in the received digital bit stream 111 to detect and correct errors that may have occurred during transmission over the communication link 101, producing a bit stream 112 with the received coding parameters. A stereo sound decoder 110 converts the coding parameters received in the bit stream 112 to create synthesized left 113 and right 133 channels of the digital stereo sound signal. The left 113 and right 133 channels of the reconstructed digital stereo sound signal in the stereo sound decoder 110 are converted to the synthesized left 114 and right 134 channels of the analog stereo sound signal in a digital-to-analog (D / A) converter 115.
[0036] The synthesized left 114 and right 134 channels of the analog stereo sound signal are reproduced respectively in a pair of speaker units, or binaural headphones, 116 and 136. Alternatively, the left 113 and right 133 channels of the digital stereo sound signal from the stereo sound decoder 110 may also be supplied and recorded on a storage device (not shown).
[0037] For example, (a) the left channel of Figure 1 may be implemented by the left channel of Figures 2-13, (b) the right channel of Figure 1 may be implemented by the right channel of Figures 2-13, (c) the stereo sound encoder 106 of Figure 1 may be implemented by the LncAnn / zznz / E / YiAi IVAS stereo encoding device of Figures 2-7, and (d) the stereo sound decoder 110 of Figure 1 may be implemented by the IVAS stereo decoding device of Figures 8-13. 1. Switching between stereo modes on the IVAS 200 and Method 250 stereo encoding device
[0038] Figure 2 is a high level block diagram simultaneously illustrating the IVAS stereo encoding device 200 and the corresponding IVAS stereo encoding method 250, Figure 3 is a block diagram simultaneously illustrating the FD stereo encoder 300 of the IVAS stereo encoding device 200 of Figure 2 and the corresponding FD stereo encoding method 350, Figure 4 is a block diagram simultaneously illustrating the TD stereo encoder of the IVAS stereo encoding device 200 of Figure 2 and the corresponding TD stereo encoding method 450, and Figure 5 is a block diagram simultaneously illustrating the MDCT stereo encoder 500 of the IVAS stereo encoding device 200 of Figure 2 and the corresponding MDCT stereo encoding method 550.
[0039] In the illustrative and non-limiting implementation of Figures 2-5, the framework of the IVAS stereo encoding device 200 (and correspondingly the IVAS stereo decoding device 800 of Figure 8) is based on a modified version of the Enhanced Voice Services (EVS) codec (See Reference [1]). Specifically, the EVS codec has been extended to encode (and decode) stereo and multi-channel channels, and is targeted at immersive voice and audio services (IVAS). For this reason, the encoding device 200 and the method 250 are referred to as the IVAS stereo coding device and method herein.In the exemplary implementation described, the IVAS stereo encoding device 200 and method 250 utilize, as a non-limiting example, three stereo coding modes: a DFT (Discrete Fourier Transform) based frequency domain (FD) stereo mode, referred to herein as the DFT stereo mode, a time domain (TD) stereo mode, referred to herein as the TD stereo mode, and a Modified Discrete Cosine Transform (MDCT) based joint stereo coding mode, referred to herein as the MDCT stereo mode. It should be noted that other codec structures may be utilized as the basis for the framework of the IVAS stereo encoding device 200 (and correspondingly the IVAS stereo decoding device 800).
[0040] Stereo mode switching in the IVAS codec (IVAS stereo encoding device 200 and IVAS stereo decoding device 800) refers, in the described, non-limiting implementation, to switching between DFT, TD and MDCT stereo modes. 1.1 Differences between different stereo encoders and coding methods LncAnn / zznz / E / YiAi
[0041] The following nomenclature is used in the present invention and in the accompanying figures: lower case letters indicate signals in the time domain, upper case letters indicate signals in the transform domain, 1 / L means left channel, r / R means right channel, m / M means middle channel, s / S means side channel, PCh means primary channel and SCh means secondary channel. Furthermore, in the figures, unitless numbers correspond to a number of samples at a sampling rate of 16 kHz.
[0042] There are differences between (a) the DFT stereo encoder 300 and the encoding method 350, (b) the TD stereo encoder 400 and the encoding method 450, and (c) the MDCT stereo encoder 500 and the encoding method 550. Some of these differences are summarized in the following paragraphs and at least some of them will be better explained in the following description.
[0043] The IVAS stereo encoding device 200 and encoding method 250 perform operations such as buffering a 20 ms frame (as known in the art, stereo sound signal is processed in successive frames of a given duration containing a given number of sound signal samples) of the stereo input signal (left and right channels), some classification, down-mixing, pre-processing, and actual encoding steps. A look-ahead of 8.75 ms is available which is primarily used for parsing, classification, and OverLap-Add (OLA) operations used in the transform domain, such as in a Transform Coded eXcitation (TCX) kernel, a High Quality (HQ) kernel, and a Frequency-Domain BandWidth-Extension (TDBWE) kernel. These operations are described in reference [1], clauses 5.3 and 5.2.6.2.
[0044] Look-ahead compression is shorter in the IVAS stereo encoding device 200 and coding method 250 compared to the unmodified EVS encoder by 0.9375 ms (corresponding to a resampling delay of the Finite Impulse Response (FIR) filter (see reference [1], clause 5.1.3.1). This impacts the resampling procedure of the down-processed signal (down-mixed signal for TD and DFT stereo modes) in each frame: - DFT 300 stereo encoder and 350 encoding method: Resampling is performed in the DFT domain and therefore does not introduce any additional delay; - TD 400 Stereo Encoder and 450 Coding Method: FIR resampling (decimation) is performed using a 0.9375 ms delay. Since this resampling delay is not available in the IVAS 200 stereo encoder, the resampling delay is compensated for by adding zeros at the end of the downmixed signal. Consequently, the compensated 0.9375 ms portion of the downmixed signal must be recalculated (resampled again) in the next frame. - MDCT 500 Stereo Encoder and 550 Encoding Method: Same as TD 400 Stereo Encoder and 450 Encoding Method. LncAnn / zznz / E / YiAi Resampling in the DFT 300 stereo encoder, the TD 400 stereo encoder, and the MDCT 500 stereo encoder is performed from the input sample rate (typically 16, 32, or 48 kHz) to the internal sample rate(s) (typically 12.8, 16, 25.6, or 32 kHz). The resampled signal(s) are then used in the core preprocessing and encoding.
[0045] Furthermore, the lookahead contains a portion of the down-mixed signal (down-mixed signal for TD and DFT stereo modes) that is not accurate, but extrapolated or estimated, which also impacts the resampling process. The inaccuracy of the down-mixed signal (down-mixed signal for TD and DFT stereo modes) depends on the current stereo encoding mode: - DFT stereo encoder 300 and coding method 350: The 8.75 ms duration of the lookahead corresponds to a windowed overlapping portion of the down-mixed signal related to an OLA portion of the DFT analysis window, respectively an OLA portion of the DFT synthesis window. In order to make the preprocessing into a signal as meaningful as possible, this lookahead portion of the down-mixed signal is repaired (or unwindowed, i.e., the inverse window is applied to the lookahead portion). As a consequence, the repaired down-mixed signal of 8.75 ms duration at the lookahead is not accurately reconstructed in the current frame; - TD 400 stereo encoder and 450 coding method: Before down-mixing in the time domain (TD), an inter-channel alignment (ICA) is performed using an inter-channel time delay (ITD) synchronization between the two input channels 1 yr in the time domain. This is achieved by delaying one of the input channels (1 or) and extrapolating a lost portion of the down-mixed signal corresponding to the duration of the ITD delay; a maximum value of the ITD delay is 7.5 ms. Consequently, the extrapolated down-mixed signal of up to 7.5 ms duration in the lookahead is not accurately reconstructed in the current frame. - MDCT 500 Stereo Encoder and 550 Encoding Method: Normally no down-mixing or time shifting is performed, so the look-ahead portion of the input audio signal is usually accurate.
[0046] The portion of the signal repaired / extrapolated in the lookahead is not subject to the actual coding, but is used for analysis and classification. Therefore, the portion of the signal repaired / extrapolated in the lookahead is recalculated in the next frame, and the resulting processed signal (downmixed signal for TD and DFT stereo modes) is then used for the actual coding. The duration of the recalculated signal depends on the stereo mode and the coding processing: - DFT 300 Stereo Encoder and 350 Encoding Method: The 8.75ms long signal undergoes recalculation at both the input stereo signal sampling rate and the internal sampling rate; LncRnn / zznz / E / YiAi - TD 400 stereo encoder and coding method 450: The signal of 7.5 ms duration is subjected to recalculation at the sampling rate of the input stereo signal, while the signal of 7.5 + 0.9375 = 8.4375 ms duration is subjected to recalculation at the internal sampling rate. - MDCT 500 stereo encoder and coding method 550: Counting is not normally required at the sampling rate of the input stereo signal, while the signal of 0.9375 ms duration is subject to counting at the internal sampling rate. It should be noted that the durations of the repaired signal portion, respectively extrapolated in anticipation, are mentioned here for illustrative purposes, while any other duration can be implemented in general.
[0047] Additional information on the DFT stereo encoder 300 and encoding method 350 can be found in references [2] and [3], Additional information on the TD stereo encoder 400 and encoding method 450 can be found in Reference [4], and additional information on the MDCT stereo encoder 500 and encoding method 550 can be found in references [6] and [7]. 1.2 Structure of the IVAS 200 stereo coding device and processing in the IVAS 250 stereo coding method
[0048] The following Table I lists in a sequential order the processing operations for each frame depending on the current stereo coding mode (See also Figures 2-5). Table 1 - Processing operations on the IVAS 200 stereo encoding device. LncAnn / zznz / E / YiAi DFT Stereo Mode TD Stereo Mode MDCT Stereo Mode Stereo Mode Classification and Selection Memory Allocation / Deallocation Setting TD Stereo Mode Stereo Mode Switching Updates ICA Encoder - Time Alignment and Scaling TD Transient Detectors Stereo Encoder Configuration DFT Analysis TD Analysis Stereo processing and down-mixing in the DFT domain Weighted down-mixing in the TD domain DFT synthesis Front-end preprocessing Core encoder setup TD stereo setup DFT stereo residual coding Additional preprocessing Core coding Joint stereo coding Common stereo updates LncRnn / zznz / E / YiAi
[0049] The IVAS stereo encoding method 250 comprises a switching control operation (not shown) between DFT, TD, and MDCT stereo modes. To perform the switching control operation, the IVAS stereo encoding device 200 comprises a DFT, TD, and MDCT stereo mode switching controller (not shown). Switching between DFT and TD stereo modes in the IVAS stereo encoding device 200 and the encoding method 250 involves using the stereo mode switching controller (not shown) to maintain continuity of the following input signals 1) to 5) to enable proper processing of these signals in the IVAS stereo encoding device 200 and the method 250: 1) the input stereo signal including the left 1 / L and right r / R channels, used, for example, for time-domain transient detection or Inter Channel BWE (IC BWE); 2) The down-mixed processed stereo signal (for TD and DFT stereo modes) at the sampling rate of the input stereo signal: DFT 300 Stereo Encoder and 350 Coding Method: Mid-channel m / M; TD 400 Stereo Encoder and 450 Encoding Method: Primary Channel (PCh) and Secondary Channel (SCh); MDCT 500 Stereo Encoder and 550 Encoding Method: Original left and right channels (no down-mix) 1 yr; 3) Down-processed signal (down-mixed signal for TD and DFT stereo modes) at a sampling rate of 12.8 kHz, used in pre-processing; 4) Down-processed signal (down-mixed signal for TD and DFT stereo modes) at an internal sampling rate, used in core encoding; 5) High band (HB) input signal, used in bandwidth extension (BWE). [0 050] While it is straightforward to maintain continuity for signal 1) above, it is challenging for signals 2) - 5) due to several aspects, e.g. different down-mixing, different duration of the recalculated part of the lookahead, use of inter-channel alignment (ICA) only in TD stereo mode, etc. 1.2.1 Stereo mode classification and selection
[0051] The operation (not shown) of controlling switching between DFT, TD and MDCT stereo modes comprises a stereo sorting and stereo mode selecting operation 255, for example as described in Reference [9], the entire contents of which are incorporated herein by reference. To perform the operation 255, the controller (not shown) of switching between DFT, TD and MDCT stereo modes comprises a stereo sorter and a stereo mode selector 205.
[0052] Switching between TD stereo mode, DFT stereo mode, and MDCT stereo mode is responsive to stereo mode selection. Stereo classification (Reference [9]) is performed in response to the left 1 and right r channels of the input stereo signal, and / or the requested encoded bit rate. Stereo mode selection (Reference [9]) consists of choosing one of DFT, TD, and MDCT stereo modes based on stereo classification.
[0053] The stereo classifier and stereo mode selector 205 produce a stereo mode signaling 270 to identify the selected stereo encoding mode. 1.2.2 Memory allocation / deallocation
[0054] The operation (not shown) of controlling switching between the DFT, TD, and MDCT stereo modes comprises a memory allocation operation (not shown). To perform the memory allocation operation, the DFT, TD, and MDCT stereo mode switching controller (not shown) dynamically allocates / deallocs static memory data structures to / from the DFT, TD, and MDCT stereo modes depending on the current stereo mode. Such memory allocation keeps the static memory impact of the IVAS 200 stereo encoding device as low as possible by maintaining only those data structures that are employed in the current frame.
[0055] For example, in a first DFT stereo frame following a TD stereo frame, data structures related to the TD stereo mode (e.g., the TD stereo data handling, the second encoder-core data structure) are freed (deallocated) and data structures related to the DFT stereo mode (e.g., the DFT stereo data structure) are allocated and initialized in their place. It is noted that the deallocation of unused data structures is performed first, LncAnn / zznz / E / YiAi followed by the allocation of the new data structures used. This order of operations is important to avoid increasing the impact of static memory at any point in the coding.
[0056] Table II shows a summary of the main static memory data structures used in the different stereo modes. Table ll - Data structure assignment in different stereo modes. X means assigned - XX” means twice assigned — - means deallocated and ” means twice deallocated. LncAnn / zznz / E / YiAi Data structures DFT stereo mode Normal TD stereo mode LRTD stereo mode MDCT stereo mode Main structure of XXXX IVAS Stereo classifier XXXX DFT stereo X - - - TD stereo - XX - MDCT stereo - - - XN encoder kernel X XX XX XX ACELP kernel X XX XX — TCX + IGF kernel X X- X- XX TD-BWE XX XX — FD-BWE XX XX — IC-BWE XX - - ICA XXX - Below is an example implementation of the memory allocation / deallocation encoder module in C source code.} void stereo memory ene( CPE_ENC_HANDLE hCPE, / * const int32_t input_Fs, / * const intl6_t max_bwidth, / * float *tdm_last_ratio / * ) { Encoder State *st; i : CPE encoder structure * / i : input sampling rate * / i : máximum audio bandwidth * / o : TD stereo last ratio * / ! x______________________________________________________________* * save parameters from structures that will· be freed x_______________________________________________________________A / if ( hCPE->last_element_mode == IVAS_CPE_TD ) { *tdm last ratio = hCPE->hStereoTD->tdm last ratio; / * note: this must be set to local variable before data structures are allocated / deallocated * / ) if ( hCPE->hStereoTCA != NULL && hCPE->last_element_mode == IVAS_CPE_DFT ) set_s( hCPE->hStereoTCA->prevCorrLagStats, (intl6_t) hCPE->hStereoDft>itd[l], 3 ) ; hCPE->hStereoTCA->prevRefChanlndx = ( hCPE->hStereoDft->itd[1] >= 0 ) ? < L_CH_INDX ) : ( R_CH_INDX ); ) I ** * allocate / deallocate data structures ** / if ( hCPE->element mode !- hCPE->last element mode ) { / ** * switching CPE mode to DFT stereo ** / if ( hCPE->element_mode == IVAS_CPE_DFT ) { / * deallocate data structure of the previous CPE mode * / if ( hCPE->hStereoTD != NULL ) { count_free( hCPE->hStereoTD ); hCPE->hStereoTD = NULL;} if ( hCPE->hStereoMdct != NULL ) { count_free( hCPE->hStereoMdct ) ; hCPE->hStereoMdct = NULL;} / * deallocate CoreCoder secondary channel * / deallocate_CoreCoder_enc( hCPE->hCoreCoder[1] ); / * allocate DFT stereo data structure * / stereo dft ene create ( &( hCPE->hStereoDft ), input Fs, max bwidth ) ; / * allocate ICBWE structure * / if ( hCPE->hStereoICBWE == NULL ) { hCPE->hStereoICBWE = (STEREO_ICBWE_ENC_HANDLE) count_malloc( sizeof ( STEREO_ICBWE_ENC_DATA ) ) ; stereo icBWE init ene ( hCPE->hStereoICBWE ) ;} / * allocate HQ core in M channel * / LncAnn / zznz / E / YiAi st = hCPE->hCoreCoder[0]; if ( st->hHQ_core == NULL ) { st->hHQ core = (HQ ENC HANDLE) count malloc( sizeof( HQ_ENC_DATA ) ); ~ ~ ~ EQ_core_enc_init( st->hHQ_core );} LncAnn / zznz / E / YiAi * switching CPE mode to TD stereo *-------------------------------------------------------------* / if ( hCPE->element_mode == IVAS_CPE_TD ) { / * deallocate data structure of the previous CPE mode * / if ( hCPE->hStereoDft != NULL ) { stereo_dft_enc_destroy( &( hCPE->hStereoDft ) ); hCPE->hStereoDft = NULL; } if ( hCPE->hStereoMdct != NULL ) { count_free( hCPE->hStereoMdct ); hCPE->hStereoMdct = NULL; } / * deallocated TCX / IGF structures for second channel * / deallocate_CoreCoder_TCX_enc( hCPE->hCoreCoder[1] ); / * allocate TD stereo data structure * / hCPE->hStereoTD = (STEREO_TD_ENC_DATA_HANDLE) count_malloc( sizeof ( STEREO_TD_ENC_DATA ) ); stereo_td_init_enc( LCPE->hStereoTD, hCPE->element_brate, hCPE>last_element_mode ) ; / * allocate secondary channel * / allocate_CoreCoder_enc( hCPE->hCoreCoder[1] );} / *-------------------------------------------------------------* * allocate DFT / TD stereo structures after MDCT stereo frame *_____________________________________________________________* / if ( hCPE->last_element_mode == IVAS_CPE_MDCT && ( hCPE->element_mode == IVAS_CPE_DFT || hCPE->element_mode == IVAS_CPE_TD ) ) / * allocate TCA data structure * / hCPE->hStereoTCA = (STEREO_TCA_ENC_HANDLE) count_malloc( sizeof( STEREO_TCA_ENC_DATA ) ) ; stereo tea init ene( hCPE->hStereoTCA, input Fs ); st = hCPE->hCoreCoder[0]; / * allocate primary channel substructures * / allocate_CoreCoder_enc( st ) ; / * allocate CLDFB for primary channel * / if ( st->cldfbAnaEnc == NULL ) { openCldfb( &st->cldfbAnaEnc, CLDFB_ANALYS1S, input_Fs, CLDFB_PROTOTYPE_1_25MS ); } / * allocate BWEs for primary channel * / if ( st->hBWE_TD == NULL ) { st->hBWE_TD = (TD_BWE_ENC_HANDLE) count_malloc( sizeof( TD_BWE_ENC_DATA ) ); if ( st->cldfbSynTd == NULL ) { openCldfb( &st->cldfbSynTd, CLDFB SYNTHESIS, 16000, CLDFB_PROTOTYPE_1_25MS ); } InitSWBencBuffer ( st->hBWE TD ); ResetSHBbúfer_Enc( st->hBWE_TD ); st->hBWE_FD = (FD_BWE_ENC_HANDLE) count_malloc( sizeof( FD_BWE_ENC_DATA ) ) ; fd_bwe_enc_init( st->hBWE_FD );} / *--------------------------------------------------------------* * switching CPE mode to MDCT stereo *_______________________________________________________________* y if ( hCPE->element_mode == IVAS_CPE_MDCT ) intl6_t i; / * deallocate data structure of the previous CPE mode * / if ( hCPE->hStereoDft != NULL ) { stereo dft ene destroy( &( hCPE->hStereoDft ) ); hCPE->hStereoDft = NULL; } if ( hCPE->hStereoTD != NULL ) { count_free( hCPE->hStereoTD ); hCPE->hStereoTD = NULL; } if ( hCPE->hStereoTCA != NULL ) count free( hCPE->hStereoTCA ); hCPE->hStereoTCA = NULL;} if ( hCPE->hStereoICBWE 1= NULL ) { count_free( hCPE->hStereoICBWE ) ; hCPE->hStereoICBWE = NULL;} for ( i = 0; i < CPE_CHANNELS; i++ ) { st = hCPE->hCoreCoder [i]; / * deallocate core channel substructures * / deallocate CoreCoder ene( hCPE->hCoreCoder[i] );} if ( hCPE->last_element_mode == IVAS_CPE_DFT ) { / * allocate secondary channel * / allocate_CoreCoder_enc( hCPE->hCoreCoder[1 ] );} / * allocate TCX / IGF structures for second channel * / st = hCPE->hCoreCoder[1]; st->hTcxEnc = (TCX_ENC_HANDLE) count_malloc( sizeof( TCX_ENC_DATA ) ) ; st->hTcxEnc->spectrum[0] = st->hTcxEnc->spectrum long; st->hTcxEnc->spectrum[1] = st->hTcxEnc->spectrum_long + N_TCX10_MAX; set f( st->hTcxEnc->old out, 0, L ERAME32k ); set f( st->hTcxEnc->spectrum long, 0, N MAX ); if ( hCPE->last_element_mode == IVAS_CPE_DFT ) { st->last core = ACELP CORE; / * needed to set-up TCX core in SetTCXModelnfo() * / } st->hTcxCfg = (TCX_CONFIG_HANDLE) counc_malloc( sizeof ( TCX_config ) ) ; st->hIGFEnc = (IGF_ENC_INSTANCE_HANDLE) count_malloc( sizeof( IGF_ENC_INSTANCE ) ); st->igf = getlgfPresent( st->element_mode, st->total_brate, st>bwidth, st->rf mode ); / * allocate and initialize MDCT stereo structure * / hCPE->hStereoMdct = (STEREO_MDCT_ENC_DATA_HANDLE) count_malloc( sizeof ( STEREO_MDCT_ENC_DATA ) ); initMdctStereoEncData( hCPE->hStereoMdct, hCPE->element_brate, hCPE->hCoreCoder[0]->max_bwidth, SMDCT_MS_DECISION, 0, NULL ); LncRnn / zznz / Ε / γΐΛΐ} } return; } LncAnn / zznz / E / YiAi 1.2.3 Establecer el modo estéreo de TD
[0057] The TD stereo mode may consist of two sub-modes. One is the so-called normal TD stereo sub-mode, for which the TD stereo mixing ratio is greater than 0 and less than 1. The other is the so-called LRTD stereo sub-mode, for which the TD stereo mixing ratio is either 0 or 1; thus, LRTD is an extreme case of the TD stereo mode, in which the TD down-mixing does not actually mix the contents of the time domain left 1 and right r channels to form the primary PCh and secondary SCh channels, but obtains them directly from channels 1 and r.
[0058] When both sub-modes (normal and LRTD) of the TD stereo mode are available, the stereo mode switching operation (not shown) comprises a TD stereo mode setting (not shown). To perform the TD stereo mode setting, as part of the memory allocation, the stereo mode switching controller (not shown) of the IVAS 200 stereo encoding device allocates / de-allocs certain static memory data structures when switching between the normal TD stereo mode and the LRTD stereo mode. For example, an IC-BWE data structure is allocated only in frames using the normal TD stereo mode (see Table II), while several data structures (BWE and Complex Low Delay Filter Bank (CLDFB) for the sub-channel SCh) are allocated only in frames using the LRTD stereo mode (see Table II).A continuación se muestra un ejemplo de implementación del módulo codificador de asignación / desasignación de memoria en el código fuente C:. / * normal TD / LRTD switching * / if ( hCPE->hStereoTD->tdm LRTD flag == 0 ) ( Encoder_State *st; st = hCPE->hCoreCoder[1]; / * deallocate CLDFB ana for secondary channel * / if ( st->cldfbAnaEnc != NULL ) { deleteCldfb( &st->cldfbAnaEnc );} / * deallocate BWEs for secondary channel * / if ( st->hBWE_TD != NULL ) { if ( st->hBWE_TD != NULL ) { count_free( st->hBWE_TD ); st->hBWE TD = NULL; } deleteCldfb( &st->cldfbSynTd ); if ( st->hBWE_FD ! = NULL ) { count_free( st->hBWE_FD ); st->hBWE_FD = NULL;}} / * allocate ICBWE structure * / if ( hCPE->hStereoICBWE == NULL ) { ( hCPE->hStereoICBWE = (STEREO_ICBWE_ENC_HANDLE) count_malloc( sizeof( STEREO_ICBWE_ENC_DATA ) ); stereo_icBWE_init_enc( hCPE->hStereoICBWE ) ;}} else / * tdm LRTD flag == 1 * / { Encoder_State *st; st = hCPE->hCoreCoder[1]; / * deallocate ICBWE structure * / if ( hCPE->hStereoICBWE != NULL ) { / * copy past input signal to be used in BWE * / mvr2r( hCPE->hStereoICBWE->dataChan[1], hCPE->hCoreCoder[1]>old_input_signal, st->input_Fs / 50 ); count free ( hCPE->hStereoICBWE ) ; hCPE->hStereoICBWE = NULL; } / * allocate CLDFB ana for secondary channel * / if ( st->cldfbAnaEnc == NULL ) { openCldfb ( &st->cldfbAnaEnc, CLDFB ANALYSIS, st->input Es, CLDFB_PROTOTYPE_1_25MS ); } / * allocate BWEs for secondary channel * / if ( st->hBWE_TD == NULL ) { st->hBWE_TD = (TD_BWE_ENC_HANDLE) count_malloc( sizeof( TD_BWE_ENC_DATA ) ); openCldfb( &st->cldfbSynTd, CLDFB SYNTHESIS, 16000, CLDFB_PROTOTYPE_1_25MS ); InitSWBencBuffer( st->hBWE TD ); ResetSHBbúfer Ene( st->hBWE TD ); st->hBWE_FD = (FD_BWE_ENC_HANDLE) count_malloc( sizeof( FD_BWE_ENC_DATA ) ) ; LncRnn / zznz / E / YiAi fd_bwe_enc_init( st->hBWE_FD );}}
[0059] Primarily, only the normal TD stereo mode (for simplicity, referred to only as the TD stereo mode) will be described in detail in the present invention. The LRTD stereo mode is mentioned as a possible implementation. 1.2.4 Stereo Mode Switching Updates
[0060] The stereo mode switching control operation (not shown) comprises a stereo switching update operation (not shown). To perform this stereo switching update operation, the stereo mode switching controller (not shown) updates long-term parameters and updates or resets past buffers.
[0061] Upon switching from DFT stereo mode to TD stereo mode, the stereo mode switching controller (not shown) resets the TD and ICA stereo static memory data structures. These data structures store the parameters and memories of the TD stereo analysis and weighted down-mixing (401 in Figure 4), respectively, of the ICA algorithm (201 in Figure 2). The stereo mode switching controller (not shown) then sets a mixing ratio index of past TD frames in accordance with either the normal TD stereo mode or the LRTD stereo mode. As an illustrative, non-limiting example: - The mixing ratio index of the previous frame is set to 15, indicating that the down-mixed middle channel m / M is encoded as the primary channel PCh, where the mixing ratio is 0.5, in normal TD stereo mode; or - The mix ratio index of the previous frame is set to 31, indicating that the left channel 1 is encoded as the primary PCh channel, in LRTD stereo mode.
[0062] Upon switching from TD stereo mode to DFT stereo mode, the stereo mode switching controller (not shown) resets the DFT stereo data structure. This DFT stereo data structure stores parameters and memories related to the DFT stereo processing and down-mixing module (303 in Figure 3).
[0063] Furthermore, the stereo mode switching controller (not shown) transfers some stereo mode related parameters between data structures. For example, parameters related to the time and energy shift between channels 1 and r, namely a side gain (or ILD parameter) and the ITD parameter of the DFT stereo mode are used to update a target gain and correlation offsets (ICA parameters 202) of the TD stereo mode and vice versa. This target gain and correlation offsets are described in more detail in the following Section 1.2.5 of the present invention. LncAnn / zznz / E / YiAi
[0064] Updates / resets related to the core encoders (See Figures 4 and 4) are described later in Section 1.4 of the present invention. An example implementation of handling some memories in the encoder is shown below. void stereo_switching_enc( CPE_ENC_HANDLE hCPE, / * i : CPE encoder structure * / float oíd input signal pri[], / * i : oíd input signal of primary channel * / const intl6_t input_frame / * i : input frame length * / ) { intl6 t i, n, dft ovl, offset; float tmpF; Encoder_State **st; st = hCPE->hCoreCoder; dft_ovl = STEREO_DFT_OVL_MAX * input_frame / L_FRAME48k; / * update DFT analysis overlap memory * / if ( hCPE->element_mode > IVAS_CPE_DFT && hCPE->input_mem[0] != NULL ) for ( n = 0; n < CPE_CHANNELS; n++ ) { mvr2r( st[n]->input + input_frame - dft_ovl, hCPE->input_mem[n], dft_ovl ); } ) / * TD / MDCT -> DFT stereo switching * / if ( hCPE->element mode == IVAS CPE DFT && hCPE->last element mode 1= IVAS_CPE_DFT ) { / * window DFT synthesis overlap memory @input_fs, primary channel * / for ( i = 0; i < dft_ovl; i++ ) { hCPE->hStereoDft->output mem dmx[i] = old_input_signal_pri[input_frame - dft_ovl + i] * hCPE->hStereoDft>win[dft ovl - 1 - i]; / * reset 48kHz BWE overlap memory * / set f( hCPE->hStereoDft->output mem dmx 32k, 0, STEREO DFT OVL 32k ); stereo_dft_enc_reset( hCPE->hStereoDft ); / * update ITD parameters * / if ( hCPE->element mode == IVAS CPE DFT && hCPE->last element mode == IVAS_CPE_TD ) { set f( hCPE->hStereoDft->itd, hCPE->hStereoTCA>prevCorrLagStats[2] , STEREO_DFT_ENC_DFT_NB );} LncAnn / zznz / E / YiAi / * Update the side_gain[] parameters * / if ( hCPE->hStereoTCA != NULL && hCPE->last_element_mode != IVAS_CPE_MDCT ) { tmpF = usdequant( hCPE->hStereoTCA->indx_ica_gD, STEREO_TCA_GDMIN, STEREO_TCA_GDSTEP ) ; ~ ~ ~ for ( i = 0; i < STEREO_DFT_BAND_MAX; i++ ) { b.CPE->hStereoDft->side_gain[S'rEREO_DFl_BAND_MAX + i] = tmpF;}} / * do not allow differential coding of DFT side parameters * / hCPE->hStereoDft->ipd_counter = STEREO_DFT_FEC_THRESHOLD; hCPE->hStereoDft->res_pred_counter = STEREO_DFT_FEC_THRESHOLD; / * update DFT synthesis overlap memory @12.8kHz * / for ( i = 0; i < STEREO_DFT_OVL_12k8; i++ ) { hCPE->hStereoDft->output mem dmx 12k8[i] - st[0]>buf_speech_enc[L_FRAME32k + L_FRAME - STEREO_DFT_OVL_12k8 + i] * hCPE>hStereoDft->win_12k8[STEREO_DFT_OVL_12k8 - 1 - i]; } / * update DFT synthesis overlap memory @16kHz, primary channel only * / lerp( hCPE->hStereoDft->output_mem_dmx, hCPE->hStereoDft>output_mem_dmx_l6k, STEREO_DFT_OVL_16k, dft_ovl ); / * reset DFT synthesis overlap memory @8kHz, secondary channel * / set f( hCPE->hStereoDft->output mem res 8k, 0, STEREO DFT OVL 8k ); hCPE->vad flag[l] = 0;} / * DFT / MDCT -> TD stereo switching * / if ( hCPE->element mode == IVAS CPE TD && hCPE->last element mode != IVAS_CPE_TD ) hCPE->hStereoTD->tdm_last_ratio_idx = LRTD_STEREO_MID_IS_PRIM; hCPE->hStereoTD->tdm_last_ratio_idx_SM = LRTD_STEREO_MID_IS_PRIM; hCPE->hStereoTD->tdm last SM flag = 0; hCPE->hStereoTD->tdm last inst ratio idx = LRTD STEREO MID IS PRIM; / * First frame after DFT frame AND the conoent is uncorrelated or xtalk -> the primary channel is forced to left * / if ( hCPE->hStereoClassif->lrtd mode == 1 ) { hCPE->hStereoTD->tdm_last_ratio = ratio_tabl[LRTD_STEREO_LEFT_IS_PRIM]; hCPE->hStereoTD->tdm_last_ratio_idx = LRTD_STEREO_LEFT_IS_PRIM; if ( hCPE->hStereoTCA->instTargetGain < 0.05f && ( hCPE>vad flag[0] || hCPE->vad flagfl] ) ) / * but if there is no content in the L channel -> the primary channel is forced to right * / { hCPE->hStereoTD->tdm_last_ratio = ratio_tabl[LRTD_STEREO_RIGHT_IS_PRIM]; LncRnn / zznz / E / YiAi hCPE->hStereoTD->tdm last ratio idx = LRTD_STEREO_RIGHT_IS_PRIM; - -}} ) / * DFT -> TD stereo switching * / if ( hCPE->element_mode == IVAS_CPE_TD && hCPE->last_element_mode == 1VAS_CPE_DET ) { offset = st[0]->cldfbAnaEnc->p filter length - st[0]->cldfbAnaEnc>no channels; mvr2r ( oíd input signal pri + input frame - offset - NS2SA( input frame * 50, L MEM RECALO TBE NS ), st[0]->cldfbAnaEnc->cldfb State, offset ); cldfb_reset_memory( st [ 0]->cldfbSynTd ) ; st[0]->currEnergyLookAhead = 6.1e-5f; if ( hCPE->hStereoICBWE == NULL ) { offset = st[1]->cldfbAnaEnc->p_filter_length - st[1]->cldfbAnaEnc>no channels; if ( hCPE->hStereoTD->tdm last ratio idx == LRTD_STEREO_LEFT_IS_PRIM ) { v_multc( hCPE->hCoreCoder[1]->old_input_signal + input_frame offset - NS2SA( input_frame * 50, L_MEM_RECALC_TBE_NS ), -l.Of, st [1]>cldfbAnaEnc->cldfb State, offset ) ;} else { mvr2r( hCPE->hCoreCoder[1]->old input signal + input frame offset - NS2SA( input_frame * 50, L_MEM_RECALC_TBE_NS ), st[1]->cldfbAnaEnc>cldfb state, offset );} cldfb_reset_memory( st[1]->cldfbSynTd ); st[1]-AcurrEnergyLookAhead = 6.1e-5f;} st[1]->last_extl = -1; / * no secondary channel in the previous frame -> memory resets * / set_zero( st[1]->old_inp_12k8, L_INP_MEM ); / *set zero( st[l]->old inp 16k, L INP MEM );* / set_zero( st[1]->mem_decim, 2 * L_FILT_MAX ); / *set_zero( st[1]->mem_deciml6k, 2*L_FILT_MAX );* / st[1]->mem_preemph = 0; / *st[l]->mem preemphl6k = 0;* / set_zero( stΓ1]->buf_speech_enc, L_PAST_MAX_32k + L_FRAME32k + L_NEXT_MAX_32k ); set zero ( st[l]->buf speech ene pe, L PAST MAX 32k + L FRAME32k + L_NEXT_MAX_32k ) ; ~ ----if ( st[1]->hTcxEnc != NULL ) LncAnn / zznz / E / YiAi { set_zero( st[1]->hTcxEnc->buf_speech_ltp, L_PAST_MAX_32k + L_FRAME32k + L_NEXT_MAX_32k ); set_zero( st[1]->buf_wspeech_ene, L_FRAMElok + L_SUBFR + L_FRAME16k + L_NEXT_MAX_16k ); ~ ~ ~ ~ set_zero( st[1]->buf_synth, OLD_SYNTH_SIZE_ENC + L_FRAME32k ); st[1]->mem_wsp = O.Of; st[1]->mem_wsp_enc = O.Of; init_gp_clip( st[1]->clip_var ); set_f( st[1]->Bin_E, 0, L_FFT ); set_f( st[1]->Bin_E_old, 0, L_FFT / 2 ); / * st[1]->hLPDmem reset already done in allocation of handles * / st[1]->last_L_frame = st[0]->last_L_frame; pitch_ol_init( &st[1]->old_thres, &st[1]->old_pitch, &st [1]>delta_pit, Sst[1]->old_corr ); set zero( st[l]->old wsp, L WSP MEM ); set^zero( st[1]->old7wsp2, 7 L_WSP_MEM - L_INTERPOL ) / OPL_DECIM ); set_zero( st[1]->mem_decim2, 3 ); st[l]->Nb ACELP trames = 0; / * popúlate PCh memories into the SCh * / mvr2r ( st[0]->hLPDmem->old_exc, st[1]->hLPDmem->old_exc, L_EXC_MEM ); mvr2r ( st[0]->lsf_old, st[1]->lsf_old, M ); mvr2r( st[0]->lsp oíd, st[l]->lsp oíd, M ); mvr2r( st[0]->lsf_oldl, st[1]->lsf_oldl, M ); mvr2r( st[0]->lsp oldl, st[l]->lsp oldl, M ); st[1]->GSC_noisy_speech = 0; ) else if ( hCPE->element mode == IVAS CPE TD && hCPE->last element mode == IVAS_CPE_MDCT ) set_f( st[0]->hLPDmem->old_exc, O.Of, L_EXC_MEM ); set f( st[1]->hLPDmem->old exc, O.Of, L EXC MEM ); ) LncAnn / zznz / E / YiAi 1.2.5 ICA Encoder
[0065] In TD stereo frames, the stereo mode switching control operation (not shown) comprises an inter-channel time alignment (ICA) operation 251. To perform the operation 251, the stereo mode switching controller (not shown) comprises an ICA encoder 201 for time-aligning channels 1 and r of the input stereo signal and then scaling channel r.
[0066] As described in the previous description, before TD down-mixing, ICA is performed using ITD synchronization between the two input channels 1 yr in the time domain. This is achieved by delaying one of the input channels (1 or r) and extrapolating a missing part of the down-mixed signal corresponding to the duration of the ITD delay; a maximum value of the ITD delay is 7.5 ms. Time warping, i.e., ICA time shifting, is applied first and alters most of the current TD stereo frame. The extrapolated part of the look-ahead down-mixed signal is recalculated and thus time-matched in the next frame based on the ITD estimated in that frame.
[0067] When stereo mode switching is not anticipated, the extrapolated signal of 7.5 ms duration is recalculated in the ICA encoder 201. However, when stereo mode switching can occur, i.e., switching from DFT stereo mode to TD stereo mode, a longer signal is subject to being recalculated. The duration then corresponds to the duration of the repaired DFT stereo signal plus the FIR resampling delay, i.e., 8.75 ms + 0.9375 ms = 9.6875 ms. These features are explained in more detail in section 1.4.
[0068] Another purpose of the ICA encoder 201 is the scaling of the input channel r. The scaling gain, i.e. the mentioned target gain, is estimated as a logarithmic ratio of the energies of the channels 1 and r smoothed with the target gain of the previous frame in each frame, current frame (20 ms) is applied to the last 15 ms of the current input channel r, while the first 5 ms of the current channel r are scaled with a combination of the target gains of the previous and current frames in a fade-in / fade-out manner.
[0069] The ICA encoder 201 produces ICA parameters 202 such as the ITD delay, the target gain, and a target channel index. 1.2.6 Time-domain transient detectors
[0070] The stereo mode switching control operation (not shown) comprises an operation 253 of detecting time domain transients on channel 1 from the ICA encoder 201. To perform the operation 253, the stereo mode switching controller (not shown) comprises a detector 203 for detecting time domain transients on channel 1.
[0071] Likewise, the stereo mode switching control operation (not shown) comprises an operation 254 of detecting time domain transients in channel r from the ICA encoder 201. To perform the operation 254, the stereo mode switching controller (not shown) comprises a detector 204 for detecting time domain transients in channel r.
[0072] Time domain transient detection in the 1 yr channels is a preprocessing step that enables detection and thus proper processing and coding of such transients in the transform domain kernel coding modules (TCX kernel, HQ kernel, FD-BWE). LncAnn / zznz / E / YiAi
[0073] Further information on the time domain transient detectors 203 and 204 and the time domain transient detection operations 253 and 254 can be found, for example, in reference [1], clause 5.1.8. 1.2.7 Stereo Encoder Settings
[0074] To perform stereo encoder settings, the IVAS stereo encoding device 200 sets parameters of the stereo encoders 300, 400, and 500. For example, a nominal bit rate is set for the core encoders. 1.2.8 DFT analysis, stereo processing and down-mixing in the DFT domain, and IDFT synthesis
[0075] Referring to Figure 3, the DFT stereo encoding method 350 comprises an operation 351 for applying a DFT transform to channel 1 from the time domain transient detector 203 of Figure 2. To perform the operation 351, the DFT stereo encoder 300 comprises a calculator 301 of the DFT transform of channel 1 (DFT analysis) to produce a channel L in the DFT domain.
[0076] The DFT stereo encoding method 350 also comprises an operation 352 for applying a DFT transform to channel r from the time domain transient detector 204 of Figure 2. To perform operation 352, the DFT stereo encoder 300 comprises a calculator 302 of the DFT transform of channel r (DFT analysis) to produce a channel R in the DFT domain.
[0077] The DFT stereo encoding method 350 further comprises a stereo processing and down-mixing operation 353 in the DFT domain. To perform the operation 353, the DFT stereo encoder 300 comprises a stereo processor and a down-mixer 303 for producing side information in a side channel S. The down-mixing of the L and R channels also produces a residual signal in the side channel S. The side information and the residual signal of the side channel S are encoded, for example, using an encoding operation 354 and a corresponding encoder 304, and then multiplexed into an output bit stream 310 of the DFT stereo encoder 300. The stereo processor and the down-mixer 303 also down-mix the left L and right R channels from the DFT calculators 301 and 302 to produce the middle channel M in the DFT domain.Further information, for example, on the operation 353 of stereo processing and down-mixing, the stereo processor and down-mixer 303, the mid channel M and side information and the residual signal of the side channel S can be found in Reference [3],.
[0078] In an inverse DFT synthesis (IDFT) operation 355 of the DFT stereo coding method 350, a calculator 305 of the DFT stereo encoder 300 calculates the IDFT transform m of the middle channel M at the sampling frequency of the input stereo signal, for example 12.8 kHz. Similarly, in an inverse DFT synthesis (IDFT) operation 356 of the DFT stereo coding method 350, a calculator 306 of the DFT stereo encoder 300 calculates the IDFT transform m of the channel M at the internal sampling frequency. LncAnn / zznz / E / YiAi 1.2.9 TD analysis and down-mixing in the TD domain
[0079] Referring to Figure 4, the TD stereo encoding method 450 comprises a time domain analysis and weighted down-mixing operation 451 in the TD domain. To perform the operation 451, the TD stereo encoder 400 comprises a time domain analyzer and a downmixer 401 for calculating stereo side parameters 402 such as a submode flag, a mixing ratio index or a linear prediction reuse flag, which are multiplexed into an output bit stream 410 of the TD stereo encoder 400. The time domain analyzer and downmixer 401 also perform a weighted down-mixing of the channels 1 and r of the detectors 203 and 204 (Figure 2) to produce the primary channel PCh and the secondary channel SCh using an estimated mixing ratio, in alignment with the ICA scaling.More information, for example, on the time domain analyzer and down-mixer 401 and operation 451 can be found in Reference [4].
[0080] Down-mixing is performed using the mixing ratio of the current frame, for example, on the last 15 ms of the current frame of the input channels 1 yr, while the first 5 ms of the current frame are mixed using a combination of the mixing ratios of the previous and current frames, such that a fade in / out occurs to smooth the transition from one channel to the other. The two channels (primary channel PCh and secondary channel SCh) sampled at the sampling frequency of the stereo input channel, for example 32 kHz, are resampled using FIR decimation filters to their 12.8 kHz representations, and to the internal sampling frequency.
[0081] In TD stereo mode, it is not only the stereo input signal of the current frame that is downmixed. The stored downmixed signals corresponding to the previous frame are also remixed. The duration of the passed-through signal subjected to this recalculation corresponds to the duration of the time-shifted signal recalculated in the ICA module, i.e., 8.75 ms + 0.9375 ms - 9.6875 ms. 1.2.10 Front-end preprocessing
[0082] In the IVAS codec (IVAS 200 stereo encoding device and IVAS 800 stereo decoding device), there is a restructuring of traditional preprocessing such that some classification decisions are made on the overall bit rate of the codec while other decisions are made depending on the bit rate of the core encoding. Accordingly, traditional preprocessing, as used for example in the EVS codec (Reference [1]), is split into two parts to ensure that the best possible codec configuration is used in each processed frame. In this way, the codec configuration can change from frame to frame, while certain configuration changes can be made as quickly as possible, for example those based on signal activity or signal class.On the other hand, some changes to the codec configuration should not occur too frequently, for example the selection of the encoded audio bandwidth, the selection of the internal sample rate, or the distribution of the bit budget between low-band encoding. LncAnn / zznz / E / YiAi and high-bandwidth; too frequent changes to these codec settings can result in unstable encoded signal quality or even audible artifacts.
[0083] The first part of the preprocessing, the front-end preprocessing, may include preprocessing and classification modules such as resampling to the preprocessing sample rate, spectral analysis, bandwidth detection (BWD), sound activity detection (SAD), linear prediction (LP) analysis, open-loop pitch search, signal classification, speech / music classification. It is to be noted that the decisions of the front-end preprocessing depend exclusively on the overall bitrate of the codec. More information about the operations performed during the preprocessing described above can be found, for example, in reference [1].
[0084] In DFT stereo mode (DFT stereo encoder 300 of Figure 3), the front-end preprocessing is performed by a front-end preprocessor 307 and the corresponding front-end preprocessing operation 357 on the middle channel m in the time domain at the internal sampling rate of the IDFT calculator 306.
[0085] In TD stereo mode, the front-end preprocessing is performed by (a) a front-end preprocessor 403 and corresponding front-end preprocessing operation 453 on the primary channel PCh from the time domain analyzer and down-mixer 401, and (b) a front-end preprocessor 404 and corresponding front-end preprocessing operation 454 on the secondary channel SCh from the time domain analyzer and down-mixer 401.
[0086] In MDCT stereo mode, the front-end preprocessing is performed by a front-end preprocessor 503 and corresponding front-end preprocessing operation 553 on the left input channel 1 of the time-domain transient detector 203 (Figure 2), and (b) a front-end preprocessor 504 and corresponding front-end preprocessing operation 554 on the right input channel r of the time-domain transient detector 204 (Figure 2). 1.2.11 Configuring the core-encoder
[0087] The configuration of the encoder core(s) is done based on the overall bit rate of the codec and the front-end preprocessing.
[0088] Specifically, in the DFT stereo encoder 300 and the corresponding DFT stereo coding method 350 (Figure 3), a core encoder configurator 308 and the corresponding core encoder configuration operation 358 are responsive to the time domain mean channel m of the IDFT calculator 305 and the output of the front end preprocessor 307 to configure the core encoder 311 and the corresponding core encoding operation 361. The core encoder configurator 308 is responsible for, for example, setting the internal sample rate and / or modifying the core encoder type classification. More information about core encoder configuration in the DFT domain can be found, for example, in references [1] and [2]. LncAnn / zznz / E / YiAi
[0089] In the TD stereo encoder 400 and the corresponding TD stereo encoding method 450 (Figure 4), a core encoder configurator 405 and the corresponding core encoder configuration operation 455 are responsive to the primary channel PCh and the secondary channel SCh preprocessed by the front-end preprocessors 403 and 404, respectively, to perform core encoder configuration 406 and the corresponding core encoding operation 456 of the primary channel PCh and the core encoder 407 and the corresponding core encoding operation 457 of the secondary channel SCh. The core encoder configurator 405 is responsible for, for example, setting the internal sampling rate and / or modifying the core encoder type classification. More information about configuring core encoders in the TD domain can be found, for example, in references [1] and [4], 1.2.12 Additional preprocessing.
[0090] The DFT encoding method 350 comprises a further preprocessing operation 362. To perform the operation 362, a so-called front-end preprocessor 312 of the DFT stereo encoder 300 performs a second part of the preprocessing which may include classification, kernel selection, preprocessing at the internal encoding sample rate, etc. The decisions in the front-end preprocessor 307 depend on the kernel encoding bit rate, which typically fluctuates during a session. Additional information about the operations performed during such post preprocessing in the DFT domain can be found, for example, in Reference [1].
[0091] The TD encoding method 450 comprises an additional preprocessing step 458. To perform the step 458, a so-called additional processor 408 of the TD stereo encoder 400 performs, before encoding the core of the primary channel PCh, a second part of the preprocessing which may include classification, core selection, preprocessing to the internal encoding sample rate, etc. The decisions of the preprocessor 408 depend on the encoding bit rate of the core, which usually fluctuates during a session.
[0092] Furthermore, the TD encoding method 450 comprises an additional preprocessing operation 459. To perform the operation 459, the TD stereo encoder 400 comprises a so-called additional preprocessor 409 for performing, before encoding the core of the secondary channel SCh, a second part of the preprocessing which may include classification, core selection, preprocessing at the internal sampling rate of the encoding, etc. The decisions of the preprocessor 409 depend on the encoding bit rate of the core, which usually fluctuates during a session.
[0093] Additional information on this type of preprocessing in the TD domain can be found, for example, in Reference [1]. LncRnn / zznz / E / YiAi
[0094] The MDCT encoding method 550 comprises a further left channel 1 preprocessing operation 555. To perform the operation 555, a so-called further preprocessor 505 of the MDCT stereo encoder 500 performs a second part of the left channel 1 preprocessing which may include sorting, core selection, preprocessing at the encoding internal sample rate, etc., before a joint left channel 1 and right channel r core encoding operation 556 performed by the joint core encoder 506 of the MDCT stereo encoder 500.
[0095] The MDCT encoding method 550 comprises a further right channel r preprocessing operation 557. To perform the operation 557, a so-called further preprocessor 507 of the MDCT stereo encoder 500 performs a second part of the left channel 1 preprocessing which may include sorting, kernel selection, preprocessing at the internal coding sample rate, etc., prior to the joint left channel 1 and right channel r core encoding operation 556 performed by the joint core encoder 506 of the MDCT stereo encoder 500.
[0096] Further information on this type of preprocessing in the field of MDCT can be found, for example, in Reference [1]. 1.2.13 Core coding
[0097] In general, the core encoder 311 in the DFT stereo encoder 300 (performing the core encoding operation 361) and the core encoders 406 (performing the core encoding operation 456) and 407 (performing the core encoding operation 457) in the TD stereo encoder 400 may be any variable bit rate mono codec. In the illustrative implementation of the present invention, the EVS codec (see reference [1]) with fluctuating bit rate capability (see reference [5]) is used. Of course, other suitable codecs may be considered and implemented. In the MDCT stereo encoder 500, the joint core encoder 506 is employed, which may generally be a stereo-enabled stereo encoding module that jointly processes and quantizes the 1 and r channels. 1.2.14 Common Stereo Updates
[0098] Finally, common stereo updates are performed. More information on common stereo updates can be found, for example, in Reference [1]. 1.2.15 Bitstreams
[0099] Referring to Figures 2 and 3, stereo mode signaling 270 from stereo classifier and stereo mode selector 205, a bit stream 313 from side information, residual signal encoder 304, and a bit stream 314 from core encoder 311 are multiplexed to form DFT stereo encoder bit stream 310 (then forming an output bit stream 206 of IVAS stereo encoding device 200 (Figure 2)). LncAnn / zznz / E / YiAi
[00100] Referring to Figures 2 and 4, stereo mode signaling 270 from stereo classifier and stereo mode selector 205, side parameters 402 from time domain analyzer and down-mixer 401, ICA parameters 202 from ICA encoder 201, a bit stream 411 from core encoder 406, and a bit stream 412 from core encoder 407 are multiplexed to form TD stereo encoder bit stream 410 (thus forming output bit stream 206 of IVAS stereo encoding device 200 (Figure 2)).
[00101] Referring to Figures 2 and 5, stereo mode signaling 270 from stereo classifier and stereo mode selector 205, and a bit stream 509 from joint core encoder 506 are multiplexed to form MDCT stereo encoder bit stream 508 (thus forming output bit stream 206 of IVAS stereo encoding device 200 (Figure 2)). 1.3 Switching from TD stereo mode to DFT stereo mode on the IVAS 200 stereo encoding device
[00102] Switching from TD stereo mode (TD stereo encoder 400) to DFT stereo mode (DFT stereo encoder 300) is relatively straightforward as illustrated in Figure 6.
[00103] Specifically, Figure 6 is a flowchart illustrating the processing operations in the IVAS stereo encoding device 200 and method 250 when switching from the TD stereo mode to the DFT stereo mode. As can be seen, Figure 5 shows two stereo input signal frames, i.e., a TD stereo frame 601 followed by a DFT stereo frame 602, with different processing operations and related time instances when switching from the TD stereo mode to the DFT stereo mode.
[00104] A sufficiently long lookahead is available, the processing is performed in the DFT domain (so the FIR decimation filter memory is not handled), and there is a transition from two core encoders 406 and 407 in the last TD stereo frame 501 to one core encoder 311 in the first DFT stereo frame 502.
[00105] The following operations performed when switching from the TD stereo mode (TD stereo encoder 400) to the DFT stereo mode (DFT stereo encoder 300) are performed by the aforementioned stereo mode switching controller (not shown) in response to the selection of the stereo mode.
[00106] Case A) in Figure 6 relates to an update of the DFT analysis memory, specifically the DFT stereo OLA analysis memory as part of the DFT stereo data structure that is windowed prior to DFT computation operations 351 and 352. This update is performed by the stereo mode switching controller (not shown) prior to inter-channel alignment (ICA) (See 251 in Figure 2) and comprises storing samples related to the last 8.75 ms of the current TD stereo frame 601 of channels 1 and r of the input stereo signal. This update is performed every TD stereo frame on both channels 1 and r. More information on the DFT analysis memory can be found, for example, in references [1] and [2]. Lncpnn / zznz / E / YiAi
[00107] Case B) in Figure 6 relates to an update of the DFT synthesis memory, specifically the OLA synthesis memory as part of the DFT stereo data structure resulting from the window after the IDFT calculation operations 355 and 356, when switching from the TD stereo mode to the DFT stereo mode. The stereo mode switching controller (not shown) performs this update in the first DFT stereo frame 602 following the TD stereo frame 601 and uses, for this update, the TD stereo memories as part of the TD stereo data structure and used for the TD stereo processing corresponding to the down-mixed primary PCh channel. More information about the DFT synthesis memory can be found, for example, in References [1] and [2], and more information about the TD stereo memories can be found, for example, in Reference [4].
[00108] Starting with the first DFT stereo frame 602, certain data structures related to TD stereo, for example the TD stereo data structure (as used in TD stereo encoder 400) and a core-encoder data structure 407 related to the secondary channel SCh, are no longer needed and are therefore deallocated, i.e., released by the stereo mode switching controller (not shown).
[00109] In the DFT stereo frame 602 following the TD stereo frame 601, the stereo mode switching controller (not shown) continues the encoding operation of the core 361 in the core encoder 311 of the DFT stereo encoder 300 with core encoder memories of the primary PCh channel 406 (e.g., synthesis memory, pre-emphasis memory, passed signals and parameters, etc.) in the preceding TD stereo frame 601 while controlling for time instance differences between the TD and DFT stereo modes to ensure continuity of various core-encoder buffer memories, e.g., pre-emphasized input signal buffer memories, HB input buffer memories, etc. that are subsequently used in the low-band encoder, resp. the high-band FD-BWE encoder.For further information on the encoding operation of the core 361, the PCh channel core encoder memories 406, the pre-emphasized input signal buffer memories, the HB input buffers, etc., can be found, for example, in Reference [1]. 1.4 Switching from DFT stereo mode to TD stereo mode on the IVAS 200 stereo encoding device
[00110] Switching from the DFT stereo mode to the TD stereo mode is more complicated than switching from the TD stereo mode to the DFT stereo mode, due to the more complex structure of the TD stereo encoder 400. The following operations performed when switching from the DFT stereo mode (DFT stereo encoder 300) to the TD stereo mode (TD stereo encoder 400) are performed by the stereo mode switching controller (not shown) in response to the stereo mode selection.
[00111] Figure 7a is a flowchart illustrating processing operations in the IVAS stereo encoding device 200 and method 250 when switching from the DFT stereo mode to the TD stereo mode. LncAnn / zznz / E / YiAi In particular, Figure 7a shows two frames of the stereo input signal, i.e. a DFT stereo frame 701 followed by a TD stereo frame 702, in different processing operations with time instances related to switching from DFT stereo mode to TD stereo mode.
[00112] Case A) in Figure 7a refers to the update of the FIR resampling filter memory (as employed in FIR resampling the input stereo signal sampling rate to the 12.8 kHz sampling rate and to the inner core encoder sampling rate) used in the primary PCh channel of the TD stereo coding mode. The stereo mode switching controller (not shown) performs this update every DFT stereo frame using the down-mixed mid-m channel and corresponds to a 2 x 0.9375 ms long segment 703 before the last 7.5 ms long segment in the DFT stereo frame 701 (See 704), thus ensuring continuity of the FIR resampling memory for the primary PCh channel.
[00113] Since the s side channel (Figure 3) of the DFT stereo coding method 350 is not available even if used at, for example, the 12.8 kHz sampling rate, the input stereo signal sampling rate, and the internal sampling rate, the stereo mode switching controller (not shown) fills the FIR resampling filter memory of the down-mixed SCh sub-channel differently. To reconstruct the full duration of the down-mixed signal at the internal sampling rate for the core encoder 407, an 8.75 ms segment (see 705) of the down-mixed signal from the previous frame is recalculated in the TD stereo frame 702. Thus, the update of the FIR resampling filter memory of the down-mixed SCh sub-channel corresponds to a 2 x 0.9375 ms duration segment 708 of the down-mixed middle m channel before the last 8 ms segment.75 ms duration (See 705); this is done in the first TD stereo frame 702 after switching from the preceding DFT stereo frame 701. The memory update of the FIR resampling filter of the secondary channel SCh refers to case C) in Figure 7a. As can be seen, the stereo mode switching controller (not shown) recalculates in the TD stereo frame a duration (See 706) of the down-mixed signal that is longer in the secondary channel SCh with respect to the recalculated duration of the down-mixed signal in the primary channel PCh (See 707).
[00114] Case B) in Figure 7a refers to the updating (recalculation) of the primary PCh and secondary SCh channels in the first TD stereo frame 702 following the DFT stereo frame 701. The operations of case B) performed by the stereo mode switching controller (not shown) are illustrated in more detail in Figure 7b. As mentioned in the previous description, Figure 7b is a flowchart illustrating the processing operations when switching from the DFT stereo mode to the TD stereo mode.
[00115] Referring to Figure 7b, in an operation 710, the stereo mode switching controller (not shown) recalculates the ICA memory as used in ICA analysis and calculation (See operation 251 in Figure 2) and subsequently as an input signal for the preprocessing and encoders of LncAnn / zznz / E / YiAi cores (See operations 453-454 and 456-459) of duration 9.6875 ms (as discussed in sections 1.2.7-1.2.9 of the present invention) of the 1 yr channels corresponding to the previous DFT stereo frame 701.
[00116] Thus, in operations 712 and 713, the stereo mode switching controller (not shown) recalculates the primary PCh and secondary SCh channels of the DFT stereo frame 701 by down-mixing the ICA-processed 1 yr channels using a stereo mix ratio of said frame 701.
[00117] For the secondary channel SCh, the duration (See 714) of the passed segment to be recalculated by the stereo mode switching controller (not shown) in operation 712 is 9.6875 ms although a segment of duration only 7.5 ms (See 715) is recalculated when there is no stereo encoding mode switching. For the primary channel PCh (See operation 713), the duration of the segment to be recalculated by the stereo mode switching controller (not shown) using the stereo mixing ratio TD of the passed frame 701 is always 7.5 ms (See 715). This ensures continuity of the primary PCh and secondary channels SCh.
[00118] A continuous down-mixed signal is employed when switching from the middle channel m of the DFT stereo frame 701 to the primary channel PCh of the TD stereo frame 702. To do this, the stereo mode switching controller (not shown) cross-fades (717) the 7.5 ms long segment (See 715) of the middle channel m of the DFT with the recalculated primary channel PCh (713) of the DFT stereo frame 701 to smooth the transition and equalize the differing downmix signal energy between the DFT stereo mode and the TD stereo mode. The reconstruction of the secondary channel SCh in operation 712 uses the mixing ratio of frame 701 while no other smoothing is applied because the secondary channel SCh of the DFT stereo frame 701 is not available.
[00119] The core coding in the first TD stereo frame 702 following the DFT stereo frame 701 then continues with resampling of the down-mixed signals using the F1R filters, pre-emphasizing these signals, calculating the HB signals, etc. More information on these operations can be found, for example, in Reference [1].
[00120] With respect to the pre-emphasis filter implemented as a first order high-pass filter used to emphasize the higher frequencies of the input signal (See Reference [1], Clause 5.1.4), the stereo mode switching controller (not shown) stores two pre-emphasis filter memory values in each DFT stereo frame. These memory values correspond to time instances based on different recalculation durations of the DFT and TD stereo modes. This mechanism ensures optimal recalculation of the pre-emphasis signal in channel m respectively the primary channel PCh with minimum signal duration. For the secondary channel SCh of the TD stereo mode, the pre-emphasis filter memory is reset to zero before processing the first TD stereo frame. LncAnn / zznz / E / YiAi
[00121] Starting with the first TD stereo frame 702 following the DFT stereo frame 701, certain DFT stereo related data structures (e.g., the aforementioned DFT stereo data structure) are no longer needed and are deallocated / released by the stereo mode switching controller (not shown). On the other hand, a second instance of the core encoder data structure is allocated and initialized for core coding (step 457) of the sub-channel SCh. Most of the core encoder data structures of the sub-channel SCh are reset, although some of them are estimated to achieve smoother switching transitions. For example, the previous excitation buffer (ACELP core adaptive codebook), the LSF parameters, and the previous LSP parameters (see reference [1]) of the sub-channel SCh are populated from their counterparts in the primary channel PCh.The restoration or estimation of previous buffers in the secondary channel SCh can be a source of a number of artifacts. Although many of these artifacts are significantly suppressed by decoder-based smoothing processes, some of them could still be a source of subjective artifacts. 1.5 Switching from TD stereo mode to MDCT stereo mode on the IVAS 200 stereo encoding device
[00122] Switching from TD stereo mode to MDCT stereo mode is relatively straightforward because both stereo modes handle two input channels and employ two core encoder instances. The main hurdle is maintaining the correct phase of the left and right input channels.
[00123] To maintain the correct phase of the left and right input channels of the stereo sound signal, the stereo mode switching controller (not shown) alters the TD stereo down-mixing. In the last TD stereo frame before the first MDCT stereo frame, the TD stereo mixing ratio is set to β - 1.0 and an opposite-phase down-mixing of the left and right channels of the stereo sound signal is implemented using, for example, the following formula for TD stereo down-mixing: PCh(i) = r(i) -(1- / / ) + Z(i) β SCh.(i) = Z(0 (1 - β) + r(i) · β where PCh(i) is the primary channel TD, SCh(i) is the secondary channel TD, l(i) is the left channel, r(i) is the right channel, β is the TD stereo mixing ratio, ei is the discrete time index.
[00124] In turn, this means that the primary stereo TD channel PCh(i) is identical to the last stereo MDCT left channel lpast(i) and the secondary stereo TD channel SCh(i) is identical to the last stereo MDCT right channel rpastfi) where i is the discrete time index. For completeness, it is noted that the stereo mode switching controller (not shown) may use in the last stereo TD frame a default stereo TD down-mixing, by using, for example, the following formula: PCh(i) = r(i) - β) + β LncAnn / zznz / E / YiAi SCh(i) = / ( / ) (1 - β) - r(i) · β
[00125] Next, in usual MDCT stereo processing (without stereo mode switching), the front-end preprocessing (front-end preprocessors 503 and 504 and front-end preprocessing operations 553 and 554) does not recalculate the lookahead of the left 1 and right r channels of the stereo sound signal except for its last segment of 0.9375 ms duration. However, in practice, the lookahead of duration 7.5 + 0.9375 ms is subject to recalculation of the internal sampling frequency (12.8 kHz in this illustrative non-limiting implementation). Therefore, no specific treatment is necessary to maintain the continuity of the input signals at the input sampling frequency.
[00126] Next, in the usual MDCT stereo processing (without stereo mode switching), the post preprocessing (post preprocessors 505 and 507 and front preprocessing operations 555 and 557) does not recalculate the lookahead of the left 1 and right r channels of the stereo sound signal except for their last segment of 0.9375 ms duration. In contrast to the front preprocessing, the input signals (left 1 and right r channels of the stereo sound signal) at the internal sampling frequency (12.8 kHz in this illustrative non-limiting implementation) of a duration of only 0.9375 ms are recalculated in the post preprocessing.
[00127] In other words:
[00128] The MDCT stereo encoder 500 comprises (a) front-end preprocessors 503 and 504 which, in the second MDCT stereo mode, recalculate the first duration lookahead of the left 1 and right r channels of the stereo sound signal at the internal sampling rate, and (b) further preprocessors which, in the second MDCT stereo mode, recalculate a last segment of given duration of the lookahead of the left 1 and right r channels of the stereo sound signal at the internal sampling rate, wherein the first and second durations are different.
[00129] The MDCT stereo encoding operation 550 comprises, in the second MDCT stereo mode, (a) recomputing the first duration lookahead of the left 1 and right r channels of the stereo sound signal at the internal sampling frequency, and (b) recomputing a last segment of given duration of the lookahead of the left 1 and right r channels of the stereo sound signal at the internal sampling frequency, wherein the first and second durations are different. 1.6 Switching from MDCT stereo mode to TD stereo mode on the IVAS 200 stereo encoding device
[00130] Similar to switching from TD stereo mode to MDCT stereo mode, two input channels are always available and two core encoder instances are always employed in this scenario. The main hurdle is, again, to maintain the correct phase of the input left and right channels. Thus, in the first TD stereo frame after the last MDCT stereo frame, the stereo mode switching controller (not shown) sets the TD stereo mixing ratio to β = 1.0 and alters LncRnn / zznz / E / YiAi down-mixing stereo TD using the opposite-phase mixing scheme similar to that described in section 1.5.
[00131] Another specific aspect of switching from MDCT stereo mode to TD stereo mode is that the stereo mode switching controller (not shown) appropriately reconstructs in the first TD frame the past segment of the input channels of the stereo sound signal at the internal sampling rate. Thus, a part of the lookahead corresponding to 8.75 - 7.5 = 1.25 ms is reconstructed (resampled and pre-emphasized) in the first TD stereo frame. 1.7 Switching from DFT stereo mode to MDCT stereo mode on the IVAS 200 stereo encoding device
[00132] In this scenario, a mechanism similar to the switching from DFT stereo mode to TD stereo mode described above is used, where the primary PCh and secondary SCh channels of the TD stereo mode are replaced by the left 1 and right r channels of the MDCT stereo mode. 1.8 Switching from MDCT stereo mode to DFT stereo mode on the IVAS 200 stereo encoding device
[00133] In this scenario, a mechanism similar to the switching from TD stereo mode to DFT stereo mode described above is used, where the primary PCh and secondary SCh channels of the TD stereo mode are replaced by the left 1 and right r channels of the MDCT stereo mode. 2. Switching between stereo modes on the IVAS 800 stereo decoding device and method 850
[00134] Figure 8 is a high level block diagram simultaneously illustrating an IVAS stereo decoding device 800 and a corresponding decoding method 850, where the IVAS stereo decoding device 800 comprises a DFT stereo decoder 801 and a corresponding DFT stereo decoding method 851, a TD stereo decoder 802 and a corresponding TD stereo decoding method 852, and an MDCT stereo decoder 803 and a corresponding MDCT stereo decoding method 853. For simplicity, only DFT, TD, and MDCT stereo modes are shown and described; however, it is within the scope of the present invention to utilize and implement other types of stereo modes.
[00135] The IVAS stereo decoding device 800 and the corresponding decoding method 850 receive a bit stream 830 transmitted from the IVAS stereo encoding device 200. In general, the IVAS stereo decoding device 800 and the corresponding decoding method 850 decode, from the bit stream 830, successive frames of an encoded stereo signal, for example, 20 ms long frames as in the case of the EVS codec, up-mix the decoded frames, and finally produce a stereo output signal including channels 1 and r. 2.1 Differences between different stereo decoders and decoding methods LncAnn / zznz / E / YiAi
[00136] The core decoding, performed at the internal sampling rate, is basically the same regardless of the current stereo mode; however, the core decoding is performed once (middle channel) for a DFT stereo frame and twice for a TD stereo frame (primary channels PCh and secondary channels SCh) or for an MDCT stereo frame (left channels 1 and right channels r). One problem is to maintain (update) the secondary channel SCh memories of a TD stereo frame when moving from a DFT stereo frame to a TD stereo frame, resp. to maintain (update) the channel r memories of an MDCT stereo frame when moving from a DFT stereo frame to an MDCT stereo frame.
[00137] Furthermore, the core's post-decoding decoding operations are highly dependent on the current stereo mode, complicating switching between stereo modes. The most fundamental differences are as follows:
[00138] DFT Stereo Decoder 801 and Decoding Method 851: - Resampling of the decoded core synthesis from the internal sampling rate to the output stereo signal sampling rate is performed in the DFT domain with a DFT analysis and synthesis overlap window duration of 3.125 ms. - The low band (LB) bass post-filtering adjustment (in ACELP frames) is performed in the DFT domain. - Core switching (ACELP core <-> TCX / HQ core) is performed in the DFT domain with an available delay of 3.125 ms. - Synchronization between LB synthesis and HB synthesis (in ACELP frames) does not require any additional delay. - Stereo up-mixing is performed in the DFT domain with an available delay of 3.125 ms. - Time synchronization to adjust to the overall decoder delay (which is 3.25 ms) is applied with a duration of 0.125 ms.
[00139] TD 802 Stereo Decoder and Decoding Method 852: (More information about the TD stereo decoder can be found, for example, in Reference [4]) - Resampling of the decoded core synthesis from the internal sampling rate to the output stereo signal sampling rate is performed using CLDFB filters with a delay of 1.25 ms. - The LB low-frequency post-filtering adjustment (in ACELP frames) is performed in the CLDFB domain. - Core switching (ACELP core <-> TCX / HQ core) is performed in the time domain with an available delay of 1.25 ms. - Synchronization between LB synthesis and HB synthesis (in ACELP frames) introduces an additional delay. LncAnn / zznz / E / YiAi - Stereo up-mixing is performed in the TD domain with zero delay. - Time synchronization to match overall decoder delay is applied with a duration of 2.0 ms.
[00140] MDCT Stereo Decoder 803 and Decoding Method 853: - Only one TCX-based core decoder is used, so only a 1.25 ms delay setting is used to synchronize core synthesis signals between different cores. - LB low-frequency post-filtering is omitted (in ACELP frames). - Core switching (ACELP core <-> TCX / HQ core) is performed in time domain only in the first MDCT stereo frame after the TD or DFT stereo frame with an available delay of 1.25 ms. - The timing between LB synthesis and HB synthesis is irrelevant. - Stereo up-mixing is omitted. - Time synchronization to match a global decoder delay is applied with a duration of 2.0 ms.
[00141] The different operations during decoding, mainly the processing in the DFT vs TD domain, and the different delay schemes between the DFT stereo mode and the TD stereo mode are carefully taken into account in the procedure described here for switching between the DFT and TD stereo modes. 2.2 Processing on the IVAS 800 stereo decoding device and decoding method 850
[00142] The following Table III lists in a sequential order the processing operations in the IVAS 800 stereo decoding device for each frame, depending on the current DFT, TD or MDCT stereo mode (See also Figure 8). Table III - Processing steps in the IVAS 800 stereo decoding device LncAnn / zznz / E / YiAi DFT Stereo Mode TD Stereo Mode MDCT Stereo Mode Reads stereo mode and audio bandwidth information Memory Mapping Stereo Mode Switching Updates Stereo Decoder Configuration Core Decoder Configuration TD Stereo Decoder Configuration Core decoding Joint stereo decoding Core switching in DFT domain Core switching in TD domain DFT stereo mode overlap buffers update Update of TCX stereo overlap buffers MDCT DFT stereo overlap buffers reset / refresh DFT analysis DFT stereo decoding incl. residual decoding Up-mixing in DFT domain Up-mixing in TD domain DFT synthesis IC-BWE synthesis synchronization, HB synthesis addition ICA decoder - time-tuning Common stereo upgrades LncAnn / zznz / E / YiAi
[00143] The IVAS stereo decoding method 850 comprises a switching control operation (not shown) between DFT, TD, and MDCT stereo modes. To perform the switching control operation, the IVAS stereo decoding device 800 comprises a DFT, TD, and MDCT stereo mode switching controller (not shown). Switching between DFT, TD, and MDCT stereo modes in the IVAS stereo decoding device 800 and the decoding method 850 involves using the stereo mode switching controller (not shown) to maintain continuity of the following various decoding signals and memories 1) through 6) to enable proper processing of those signals and use of those memories in the IVAS stereo decoding device 800 and the method 850: 1) Down-mixed core post-filter signals and memories to the internal sampling rate, used in core decoding; - DFT 801 stereo decoder: mid-channel m; - TD 802 Stereo Decoder: Primary channel PCh and Secondary channel SCh: - MDCT 803 stereo decoder: left channel 1 and right channel r (not down-mixed). 2) TCX-LTP (Transform Coded eXcitation - Long Term Prediction) post-filter memories. The TCX-LTP post-filter is used to interpolate between past synthesis samples using polyphase FIR interpolation filters (See Reference [1], Clause 6.9.2); 3) OLA DFT analysis memories at the internal sampling rate and at the output stereo signal sampling rate as used in the OLA portion of the windowing in the previous and current frames before the DFT operation 854; 4) OLA DFT synthesis memories as used in the OLA part of windowing in the previous and current frames after IDFT operations 855 and 856 at the sampling frequency of the output stereo signal; 5) Output stereo signal, including channels 1 and 2; and 6) HB signal memories (see Reference [1], Clause 6.1.5), 1 yr channels - used in BWE and IC-BWE.
[00144] While it is relatively simple to maintain continuity for one channel (the middle channel m in DFT stereo mode, respectively the primary channel PCh in TD stereo mode or channel 1 in MDCT stereo mode) at point 1) above, it is challenging for the secondary channel SCh at point 1) above and also for the signals / memories at points 2) - 6) due to several aspects, e.g. total absence of the passed signal and memories of the secondary channel SCh, different down-mixing, different default delay between DFT stereo mode and TD stereo mode, etc. Furthermore, a shorter decoder delay (3.25 ms) compared to the encoder delay (8.75 ms) further complicates the decoding process. 2.2.1 Reading stereo mode and audio bandwidth information
[00145] The IVAS stereo decoding method 850 begins by reading (not shown) the stereo mode and audio bandwidth information from the transmitted bitstream 830. Based on the currently read stereo mode, related decoding operations are performed for each particular stereo mode (see Table III) while the memories and buffers of the other stereo modes are maintained. 2.2.2 Memory Allocation
[00146] Similar to the IVAS 200 stereo encoding device, in a memory allocation operation (not shown), the stereo mode switching controller (not shown) dynamically allocates / deallocs data structures (static memory) depending on the current stereo mode. The stereo mode switching controller (not shown) keeps the static memory impact of the codec as low as possible by keeping only those portions of the static memory that are used in the current frame. The data structures allocated in a particular stereo mode are summarized in Table II. Lncpnn / zznz / B / YiAi
[00147] Furthermore, an LRTD stereo sub-mode flag is read by the stereo mode switching controller (not shown) to distinguish between the normal TD stereo mode and the LRTD stereo mode. Based on the sub-mode flag, the stereo mode switching controller (not shown) allocates / deallocs the related data structures within the TD stereo mode as shown in Table II. 2.2.3 Stereo Mode Switching Updates
[00148] Similar to the IVAS 200 stereo encoding device, the stereo mode switching controller (not shown) handles the memories in case of switching from one of the DFT, TD and MDCT stereo modes to another stereo mode. This keeps the parameters updated long term and updates or resets the past buffer memories.
[00149] Upon receiving a first DFT stereo frame after a TD stereo frame or an MDCT stereo frame, the stereo mode switching controller (not shown) performs a DFT stereo data structure reset operation (already defined in relation to the DFT stereo encoder 300). Upon receiving a first TD stereo frame after a DFT or MDCT stereo frame, the stereo mode switching controller performs a TD stereo data structure reset operation (already described in relation to the TD stereo decoder 400). Finally, upon receiving a first MDCT stereo frame after a DFT or TD stereo frame, the stereo mode switching controller (not shown) performs an MDCT stereo data structure reset operation.Again, when switching from one of the DFT and TD stereo modes to the other stereo mode, the stereo mode switching controller (not shown) performs a transfer operation of some stereo-related parameters between data structures, as described in relation to the IVAS 200 stereo encoding device (see Section 1.2.4 above).
[00150] Updates / resets related to the core decoding SCh secondary channel are described in Section 2.4.
[00151] Furthermore, more information about the operations of stereo decoder setup, core decoder setup, TD stereo decoder setup, core decoding, core switching in DFT domain, core switching in TD domain can be found in Table III, for example, in References [1] and [2], 2.2.4 Updating DFT Stereo Mode Overlap Memories
[00152] The stereo mode switching controller (not shown) maintains or updates the DFT OLA memories at each TD or MDCT stereo frame (See Updating DFT Stereo Mode Overlap Memories, Updating MDCT Stereo TCX Overlap Memory, and Resetting / Updating DFT Stereo Overlap Memories in Table III). In this way, the updated DFT OLA memories are available for a next DFT stereo frame. The actual mechanism of LncAnn / zznz / E / YiAi maintenance / update and related memory buffers are described later in Section 2.3 of the present invention. An exemplary implementation of updating DFT stereo OLA buffers performed on TD or MDCT stereo frames in C source code is given below. if ( st[n]->element_mode != IVAS_CPE_DFT ) { ivas_post_proc( ... ); / * update OLA búfers - needed for switching to DFT stereo * / stereo td2dft update( hCPE, n, output[n], synthfn], hb synthfn], output_frame ); / * update ovl búfer for possible switching from TD stereo SCh ACELP frame to MDCT stereo TCX frame * / if ( st[n]->element_mode == IVAS_CPE_TD && η == 1 && st[n]-RhTcxDec == NULL ) { mvr2r( outputfn] + st[n]->L frame / 2, hCPE->hStereoTD>TCX oíd syn Overl, st[n]->L frame / 2 ); T} void stereO—td2dft_update( CPE_DEC_HANDLE hCPE, / * i / ο: CPE decoder structure * / const int!6_t n, / * i : channel number* / float output[], / * i / o: synthesis @internal Es * / float synth[], / * i / o: synthesis @outputFs * / float hb_synth[], / * i / o: hb synthesis* / const intl6 t output frame / * i : frame length* / { intl6 t ovl, ovl TCX, dft32ms ovl, hq delay comp; Decoder_State **st; / * initialization * / st = hCPE->hCoreCoder; ovl = NS2SA( st[n]->L_frame * 50, STEREO_DFT32MS_OVL_NS ); dft32ms_ovl = ( STEREO_DFT32MS_OVL_MAX * st[0]->output_Fs ) / 48000; hq delay comp = NS2SA( st[0]->output Fs, DELAY CLDFB NS ); LncAnn / zznz / E / YiAi if ( hCPE->element_mode >= IVAS_CPE_DFT && hCPE->element_mode != IVAS_CPE_MDCT ) { if ( st[n]->core == ACELP_CORE ) { if ( n == 0 ) { / * update DFT analysis overlap memory @internal fs: core synthesis * / mvr2r( output + st[n]->L frame - ovl, hCPE>input mem LB[n], ovl ); / * update DFT analysis overlap memory @internal fs: BPF * / if ( st[n]->p_bpf_noise_buf ) { mvr2r( st[n]->p_bpf_noise_buf + st[n]->L_frame ovl, hCPE->input mem BPF[n], ovl ); T / * update DFT analysis overlap memory @output_fs: BWE * / if ( st[n]->extl != -1 || ( st[n]->bws_cnt > 0 && st[n]->core == ACELP_CORE ) ) { mvr2r( hb_synth + output_frame - dft32ms_ovl, hCPE>input_mem[n], dft32ms_ovl ); }} else { / * update DFT analysis overlap memory @internal fs: core synthesis, secondary channel * / mvr2r( output + st[n]->L_frame - ovl, hCPE>input_mem_LB[n], ovl ); }} else / * TCX core * / { / * LB-TCX synthesis * / mvr2r( output + st[n]->L frame - ovl, hCPE>input mem_LB[n] , ovl ) ; / * BPF * / if ( n == 0 && st[n]->p_bpf_noise_buf ) { mvr2r( st[n]->p bpf noise buf + st[n]->L frame - ovl, hCPE->input mem BPF[n], ovl ); “} / * TCX synthesis (it was already delayed in TD stereo in core_switching_post_dec()) * / if ( st[n]->hTcxDec != NULL ) { ovl TCX = NS2SA( st[n]->hTcxDec->L frameTCX * 50, STEREO_DFT32MS_OVL_NS ); mvr2r( synth + st[n]->hTcxDec->L frameTCX + hq delay comp - ovl TCX, hCPE->input memfn], ovl TCX - hq delay comp ); mvr2r( st[n]->delay buf out, hCPE->input memfn] + ovl TCX - hq delay comp, hq delay comp ); T LncAnn / zznz / E / YiAi}} else if ( hCPE->element mode == IVAS CPE MDCT && hCPE->input mem[0] != NULL ) _ _ _ _ { / * reset DFT stereo OLA memories * / set_zero( hCPE->input_mem[n], NS2SA( st[0]->output_Fs, STEREO_DFT32MS_OVL_NS ) ); set_zero( hCPE->input_mem_LB[n], STEREO_DFT32MS_OVL_16k ); if ( n == 0 ) { set_zero( hCPE->input_mem_BPF[n], STEREO_DFT32MS_OVL_16k );}} return; } 2,2,5 Decodificador estéreo DFT 801 y método de decodificación 851
[00153] The DFT decoding method 851 comprises a mid-channel m core decoding operation 857. To perform the operation 857, a core decoder 807 responsively decodes the mid-channel m in the time domain. The core decoder 807 (performing the core decoding operation 857) in the DFT stereo decoder 801 may be any variable bit rate mono codec. In the exemplary implementation of the present invention, the EVS codec (see Reference [1]) with fluctuating bit rate capability (see Reference [5]) is used. Of course, other suitable codecs may be considered and implemented.
[00154] In a DFT calculation operation 854 of the DFT decoding method 851 (DFT analysis of Table 111), a calculator 804 calculates the DFT of the mean channel m to recover the mean channel M in the DFT domain.
[00155] The DFT decoding method 851 also comprises a stereo side information and residual signal S decoding operation 858 (residual decoding of Table III). To perform the operation 858, a decoder 808 responds to the bit stream 830 to recover the stereo side information and the residual signal S.
[00156] In a DFT stereo decoding 859 (Table III DFT stereo decoding) and up-mixing (Table III DFT domain up-mixing) operation, a DFT stereo decoder and up-mixer 809 produces the L and R channels in the DFT domain in response to the mid-channel M and the side information and the residual signal S. Generally speaking, the DFT stereo decoding and up-mixing operation 859 is the inverse of the DFT stereo processing and down-mixing operation 353 of Figure 3.
[00157] In the IDFT calculation operation 855 (DFT synthesis of Table III), a calculator 805 calculates the IDFT of channel L to recover channel 1 in the time domain. Similarly, in the LncAnn / zznz / E / YiAi IDFT calculation 856 (DFT synthesis of Table III), a calculator 806 calculates the IDFT of the R channel to recover the r channel in the time domain. 2.2.6 TD 802 Stereo Decoder and 852 Decoding Method
[00158] The TD decoding method 852 comprises a primary channel PCh core decoding operation 860. To perform operation 860, a core decoder 810 responsively decodes the primary channel PCh bit stream 830.
[00159] The TD decoding method 852 also comprises a sub-channel SCh core decoding operation 861. To perform operation 861, a core decoder 811 responsively decodes the sub-channel SCh in response to the received bit stream 830.
[00160] Again, the core decoder 810 (which performs the core decoding operation 860 in the TD stereo decoder 802) and the core decoder 811 (which performs the core decoding operation 861 in the TD stereo decoder 802) may be any variable bit rate mono codec. In the illustrative implementation of the present invention, the EVS codec (see Reference [1]) with fluctuating bit rate capability (see Reference [5]) is used. Of course, other suitable codecs may be considered and implemented.
[00161] In a time domain (TD) up-mixing operation 862 (TD domain up-mixing of Table III), an up-mixer 812 receives and up-mixes the primary PCh and secondary SCh channels to recover the 1 yr time domain channels of the stereo signal based on the stereo mixing factor TD. 2.2.7 MDCT 803 Stereo Decoder and 853 Decoding Method
[00162] The MDCT decoding method 853 comprises a joint core decoding operation 863 (joint stereo decoding of Table III) of the left channel 1 and the right channel r. To perform the operation 863, a joint core decoder 813 responsively decodes the left channel 1 and the right channel r to the received bit stream 830. It is noted that no up-mixing operation is performed and no up-mixer is employed in the MDCT stereo mode. 2.2.8 Synchronization of synthesis
[00163] To perform a stereo synthesis time synchronization (Table III synthesis sync) and a stereo switching operation 864, the stereo mode switching controller (not shown) comprises a time synchronizer and a stereo switch 814 for receiving the 1 yr channels from the DFT stereo decoder 801, the TD stereo decoder 802 or the MDCT stereo decoder 803 and for synchronizing the up-mixed output stereo channels 1 and r. The time synchronizer and the stereo switch 814 delay the up-mixed output stereo channels 1 yr to match the overall delay value of the codec and handle the transitions between the DFT stereo output channels, the TD stereo output channels and the MDCT stereo output channels. LncAnn / zznz / E / YiAi
[00164] By default, in DFT stereo mode, the time synchronizer and stereo switch 814 introduce a 3.125 ms delay to the DFT stereo decoder 801. To match the overall codec delay of 32 ms (20 ms frame duration, 8.75 ms encoder delay, 3.25 ms decoder delay), the time synchronizer and stereo switch 814 apply a delay of 0.125 ms. In the case of TD or MDCT stereo mode, the time synchronizer and stereo switch 814 apply a delay consisting of the 1.25 ms resampling delay and the 2 ms delay used for synchronization between LB and HB synthesis and to match the overall codec delay of 32 ms.
[00165] After time synchronization and stereo switching (See Stereo Synthesis Time Synchronization and Switching operation 864 and Stereo Time Synchronizer and Switching 814 of Figure 8), HB synthesis (of BWE or IC-BWE) is added to the core synthesis (IC-BWE, Table III HB synthesis addition; See also Figure 8, BWE or IC-BWE calculation operation 865 and BWE or IC-BWE calculator 815) and ICA decoding is performed (ICA decoder - Table III time adjustment that desynchronizes two 1 yr output channels) before the final stereo synthesis of the 1 yr channels is output from the IVAS stereo decoding device 800 (See ICA time operation 866 and corresponding ICA decoder 816). These operations 865 and 866 are omitted in MDCT stereo mode.
[00166] Finally, as shown in Table III, common stereoscopic updates are performed. 2.3 Switching from TD stereo mode to DFT stereo mode on the IVAS stereo decoding device
[00167] Further information on the elements, operations and signals mentioned in section 2.3 and 2.4 can be found, for example, in References [1] and [2].
[00168] The switching mechanism from TD stereo mode to DFT stereo mode in the IVAS stereo decoding device 800 is complicated by the fact that the decoding steps between these two stereo modes are fundamentally different (for details, see Section 2.1 above), including a transition from two core decoders 810 and 811 in the last TD stereo frame to one core decoder 807 in the first DFT stereo frame.
[00169] Figure 9 is a flowchart illustrating processing operations in the IVAS stereo decoding device 800 and method 850 when switching from the TD stereo mode to the DFT stereo mode. Specifically, Figure 9 shows two frames of the decoded stereo signal in different processing operations with related time instances when switching from a TD stereo frame 901 to a DFT stereo frame 902.
[00170] First, core decoders 810 and 811 of TD stereo decoder 802 are used for primary PCh and secondary SCh channels and each of them outputs the corresponding decoded core synthesis at the internal sampling rate. In TD stereo frame 901, core synthesis The decoded LncAnn / zznz / E / YiAi from the two core decoders 810 and 811 is used to update the DFT stereo OLA memory buffers (one memory buffer per channel, i.e., two OLA memory buffers in total; see DFT Analysis and Synthesis buffers described above). These OLA memory buffers are updated every TD stereo frame to be up to date in case the next frame is a DFT stereo frame.
[00171] Case A) of Figure 9 relates, upon receiving a first DFT stereo frame 902 after a TD stereo frame 901, to an operation (not shown) of updating the DFT stereo analysis memories (these are used in the OLA part of the window in the previous and current frame before the DFT calculation operation 854) at the internal sampling rate, input_mem_LB\\, using the stereo mode switching controller (not shown). For this purpose, a number Lovi of the last samples 903 of the TD stereo synthesis at the internal sampling rate of the primary channel PCh and of the secondary channel SCh in the TD stereo frame 901 is used by the stereo mode switching controller (not shown) to update the DFT stereo analysis memories of the DFT stereo mid-channel m and of the side channel s, respectively. The duration of the overlapping segment 903, Lovi, corresponds to the overlapping part of 3.125 ms DFT 905 analysis window duration, e.g., Lm,i = 40 samples at an internal sampling rate of 12.8 kHz.
[00172] Similarly, the stereo mode switching controller (not shown) updates the DFT Bass Post-Filter (BPF) stereo analysis memory (which is used in the OLA portion of the window in the previous and current frame before the DFT calculation operation 854) of the internal bad sample rate mid channel, input_mem_BPF[], using the last Lmi samples of the BPF error signal (See Reference [1], Clause 6.1.4.2) of the primary channel PCh of the TD. Furthermore, the DFT stereo full band analysis (FB) memory (this memory is used in the OLA part of the window in the previous and current frame before the DFT calculation operation 854) of the mid-channel at the sampling rate of the output stereo signal, input_mem[], is updated using the last 3.125 ms samples of the HB PCh stereo TD synthesis (ACELP core) respectively, PCh TCX synthesis.The DFT stereo analysis memories BPF and FB are not used for the side information channel s, so these memories are not updated using the SCh side channel core synthesis.
[00173] Next, in the TD 901 stereo frame, the decoded ACELP core synthesis (primary PCh and secondary SCh channels) at the internal sampling rate is resampled using CLDFB domain filtering which introduces a delay of 1.25 ms. In the case of the TCX / HQ core frame, a compensation delay of 1.25 ms is used to synchronize the core synthesis between the different cores. The TCX-LTP postfilter is then applied to both the PCh and SCh core channels.
[00174] In the following operation, the primary PCh and secondary SCh channels of the TD stereo synthesis at the sampling frequency of the output stereo signal of the TD stereo frame 901 are subjected to an upscaling. LncAnn / zznz / E / YiAi TD stereo mixing (combining the primary PCh and secondary SCh channels using the TD stereo mixing ratio in the TD up-mixer 812 (See Reference [4]) resulting in 1 yr up-mixed stereo channels in the time domain. Since the up-mixing operation 862 is performed in the time domain, it does not introduce any up-mixing delay.
[00175] Next, the up-mixed left 1 and right r channels of the TD stereo frame 901 from the up-mixer 812 of the TD stereo decoder 802 are used in an operation (not shown) of updating the DFT stereo synthesis memories (these are used in the OLA portion of the window in the previous and current frame after the IDFT calculation operation 855). Again, this update is performed at each TD stereo frame by the stereo mode switching controller (not shown) in case the next frame is a DFT stereo frame. Case B) in Figure 9 shows that the number of last available samples from the synthesis of the left 1 and right r TD stereo channels is insufficient to be used for a direct update of the DFT stereo synthesis memories. Therefore, the 3.125 ms long DFT stereo synthesis memories are reconstructed in two segments using approximations. The first segment corresponds to the (3.125 - 1.25) ms duration that is available (i.e., up-mixed synthesis to the output stereo signal sampling rate), while the second segment corresponds to the remaining 1.25 ms duration signal that is not available due to the resampling delay of the core decoder.
[00176] Specifically, the DFT stereo synthesis memories are updated by the stereo mode switching controller (not shown) using the following sub-operations as illustrated in Figure 10. Figure 10 is a flowchart illustrating case B) of Figure 9, comprising updating the DFT stereo synthesis memories on a TD stereo frame at the decoder side:
[00177] (a) The two channels 1 yr of the DFT stereo analysis memories at the internal sampling rate, input_mem_LB[ / , as previously reconstructed during the decoding method 850 (they are identical to the core synthesis at the internal sampling rate), undergo further processing depending on the current decoding core: - ACELP kernel: The last Lm,i 1001 samples of the LB kernel synthesis of the primary PCh and secondary SCh channels at the internal sampling rate are resampled to the output stereo signal sampling rate using simple zero-delay linear interpolation (See 1003). - TCX / HQ core: The last Lm¡ 1001 samples of the LB core synthesis of the primary PCh and secondary SCh channels at the internal sample rate are resampled similarly to the output stereo signal sample rate using simple zero-delay linear interpolation (See 1003). However, the TCX timing memory (the LncAnn / zznz / E / YiAi last 1.25 ms segment of TCX synthesis from the previous frame) is used to update the last 1.25 ms of the resampled kernel synthesis.
[00178] (b) The linearly resampled LB signals corresponding to the 3.125 ms duration portion of the primary PCh and secondary SCh channels of the TD stereo frame 901 are up-mixed (See 1003) to form the left and right channels 1, using the common TD stereo up-mixing routine while using the TD stereo mixing ratio of the current frame (See TD up-mixing operation 862). The resulting signal is further referred to as reconstructed synthesis 1002.
[00179] (c) The reconstruction of the first part (3.125 - 1.25 ms) of the DFT stereo synthesis memories depends on the current decoding core: - ACELP core: A crossfade 1004 is performed between the CLDFB-based resampled synthesis and TD up-mixed synthesis 1005 at the output stereo signal sampling rate and the reconstructed synthesis 1002 (from the previous sub-operation (b)) for both 1 yr channels during the first part (3.125 - 1.25) ms duration of the TD stereo frame channels 901. - TCX / HQ Core: The first part (3.125 - 1.25) ms long DFT stereo synthesis memories are updated using up-mixed 1005 synthesis.
[00180] (d) The last 1.25 ms long portion of the DFT stereo synthesis memories is filled with the last portion of the reconstructed synthesis 1002.
[00181] (e) The DFT synthesis window (904 in Figure 9) is applied to the OLA DFT synthesis memories (defined above) only in the first DFT stereo frame 902 (if switching from TD to DFT stereo mode occurs). It is noted that the last 1.25 ms part of the OLA DFT synthesis memories is of limited significance since the shape of the DFT synthesis window 904 converges to zero and thus masks the coarse samples of the reconstructed synthesis 1002 resulting from resampling based on simple linear interpolation.
[00182] Finally, the reconstructed up-mixed synthesis 1002 of the TD stereo frame 901 is aligned, i.e., delayed by 2 ms in the time synchronizer and stereo switch 814 to match the overall delay of the codec. Specifically: - In case of a switch from a TD stereo frame to a DFT stereo frame, other DFT stereo memories (other than the overlap memories), i.e., the past frame parameters of the DFT stereo decoder and the buffers, are reset by the stereo mode switching controller (not shown). - Then, DFT stereo decoding (Ver 859), up-mixing (Ver 859) and DFT synthesis (Ver 855 and 856) are performed and the stereo output synthesis (1 yr channels) is aligned, i.e., delayed by 0.125 ms in the time synchronizer and stereo switch 814 to match the overall codec delay. LncAnn / zznz / E / YiAi
[00183] Figure 11 is a flowchart illustrating a case C) of Figure 9, comprising smoothing the output stereo synthesis in the first DFT stereo frame 902 after stereo mode switching, on the decoder side.
[00184] Referring to Figure 11, once the DFT stereo synthesis is aligned and synchronized to the overall codec delay in the first DFT stereo frame 902, the stereo mode switching controller (not shown) performs a crossfade operation 1151 between the aligned and synchronized TD stereo synthesis 1101 (from operation 864) and the aligned and synchronized DFT stereo synthesis 1102 (from operation 864) to smooth the switching transition. The crossfade is performed in a segment 1103 of 1.875 ms duration that begins after a delay 1104 of 0.125 ms at the beginning of both output channels 1 yr (all signals are at the sampling frequency of the output stereo signal). This case corresponds to case C) of Figure 9.
[00185] Decoding then continues regardless of the current stereo mode with the IC-BWE calculator 815, the ICA decoder 816, and the common stereo decoder updates. 2.4 Switching from DFT stereo mode to TD stereo mode on the IVAS stereo decoding device
[00186] The fundamentally different decoding operations between DFT stereo mode and TD stereo mode and the presence of two decoder cores 810 and 811 in TD stereo decoder 802 makes switching from DFT stereo mode to TD stereo mode in IVAS stereo decoding device 800 challenging. Figure 12 is a flowchart illustrating processing operations in IVAS stereo decoding device 800 and method 850 in switching from DFT stereo mode to TD stereo mode. Specifically, Figure 12 shows two frames of decoded stereo signal in different processing operations with related time instances in switching from a DFT stereo frame 1201 to a TD stereo frame 1202.
[00187] Core decoding can use the same processing regardless of the actual stereo mode, with two exceptions.
[00188] First exception: In DFT stereo frames, resampling from the internal sample rate to the output stereo signal sample rate is performed in the DFT domain, but CLDFB resampling is executed in parallel to maintain / update the CLDFB analysis and synthesis memories in case the next frame is a TD stereo frame.
[00189] Second exception: Then, the BPF (Bass Post Filter) (a low frequency tone enhancement procedure, see Reference [1], Clause 6.1.4.2) is applied in the DFT domain on DFT stereo frames while the BPF analysis and error signal calculation is done in the time domain regardless of the stereo mode.
[00190] Otherwise, all internal states and memories of the core decoder are simply continuous and are well maintained when moving from the middle channel m of DFT to the primary channel PCh of TD. LncAnn / zznz / E / YiAi
[00191] At the DFT stereo frame 1201, decoding then continues with kernel decoding (857) of the mean channel m, computation (854) of the DFT transform of the mean channel m in the time domain to obtain the mean channel M in the DFT domain, and stereo decoding and up-mixing (859) of the M and S channels into the L and R channels in the DFT domain, including decoding (858) of the residual signal. Analysis and synthesis in the DFT domain introduces an OLA delay of 3.125 ms. The synthesis transitions are then handled in the time synchronizer and stereo switch 814.
[00192] When switching from the DFT stereo frame 1201 to the TD stereo frame 1202, the fact that there is only one core decoder 807 in the DFT stereo decoder 801 makes core decoding of the TD sub-channel SCh complicated because the internal states and memories of the second core decoder 811 of the TD stereo decoder 802 are not continuously maintained (in contrast, the internal states and memories of the first core decoder 810 are continuously maintained using the internal states and memories of the core decoder 807 of the DFT stereo decoder 801). The memories of the second core decoder 811 are therefore typically reset at stereo mode switch updates (See Table III) by the stereo mode switch controller (not shown).However, there are some exceptions where the primary channel's SCh memory is filled with the memory of certain PCh buffers, for example the previous excitation, the previous LSF parameters, and the previous LSP parameters. In any case, the synthesis at the beginning of the first SCh frame of the secondary channel TD after switching from the DFT stereo frame 1201 to the TD stereo frame 1202 consequently suffers from imperfect reconstruction. Accordingly, while the synthesis of the first decoder core 810 decodes well and smoothly during the stereo mode switch, the limited quality synthesis of the second decoder core 811 introduces discontinuities during stereo up-mixing and final synthesis (862). These discontinuities are suppressed by employing the DFT stereo OLA memories during the reconstruction of the first TD stereo output synthesis, as described below.
[00193] The stereo mode switching controller (not shown) suppresses possible discontinuities and differences between the up-mixed DFT stereo and TD stereo channels by simple equalization of the signal energy. If the target gain ICA, gicA, is less than LO, channel 1, y^i), after up-mixing (862) and before time synchronization (864) is altered in the first TD stereo frame 1202 after stereo mode switching, using the following relationship: y'L(0 = a yL(i) for i = 0,..., Leq— 1 where Leq is the duration of the signals to be equalized, which corresponds in the stereo decoding device IVAS 800 to a segment of 8.75 ms duration (which corresponds for example to Leq= 140 samples at a sampling frequency of the output stereo signal of 16 kHz). The value of the gain factor is then obtained by means of the following relation: LncRnn / zznz / E / YiAi ,1Sica „ . . a = gICA+ i-------- for t = 0,..., Leq— 1 ^eq LncAnn / zznz / E / YiAi
[00194] Referring to Figure 12, case A) refers to a missing portion 1203 of the TD stereo up-mixed synchronized synthesis (from step 864) of TD stereo frame 1202 corresponding to a previous DFT stereo up-mixed synchronized synthesis memory of DFT stereo frame 1201. This (3.25 - 1.25) ms duration memory is not available when moving from DFT stereo frame 1201 to TD stereo frame 1202, except for its first 0.125 ms duration segment 1204.
[00195] Figure 13 is a flowchart illustrating case A) of Figure 12, comprising updating the TD stereo up-mixed synchronization synthesis memory in a first TD stereo frame after switching from the DFT stereo mode to the TD stereo mode, on the decoder side.
[00196] Referring to Figures 12 and 13, the stereo mode switching controller (not shown) reconstructs the 3.25 ms 1205 of the synchronized up-mixed stereo TD synthesis using the following operations (a) to (e) for the left 1 and right r channels:
[00197] (a) The DFT stereo OLA synthesis memories (defined herein) are repaired (i.e., the inverse synthesis window is applied to the OLA synthesis memories; See 1301).
[00198] (b) The first 0.125 ms portion 1302 (See 1204 in Figure 12) of the TD stereo up-mixed synchronized synthesis 1303 is identical to the previous DFT stereo up-mixed synchronized synthesis memory 1304 (last 0.125 ms long segment of the DFT stereo up-mixed synchronized synthesis memory from the previous frame) and is therefore reused to form this first portion of the TD stereo up-mixed synchronized synthesis 1303.
[00199] (c) The second part (See 1203 in Figure 12) of the synchronized up-mixed stereo TD synthesis 1303 having a duration of (3.125 - 1.25) ms is approximated with the repaired stereo DFT OLA synthesis memories 1301.
[00200] (d) The part of the synchronized up-mixed stereo synthesis TD 1303 with a duration of 2 ms from the two previous steps (b) and (c) is then filled with the output stereo synthesis in the first stereo frame TD 1202.
[00201] (e) A smoothing of the transition between the previous DFT stereo OLA synthesis memory 1301 and the up-mixed TD synchronized synthesis 1305 of step 864 of the current TD stereo frame 1202 is performed at the beginning of the up-mixed TD synchronized synthesis 1305. The transition segment has a duration of 1.25 ms (see 1306) and is obtained using a crossfade 1307 between the repaired DFT stereo OLA synthesis memory 1301 and the up-mixed TD stereo synchronized synthesis 1305. 2.5 Switching from TD stereo mode to MDCT stereo mode on the IVAS stereo decoding device
[00202] Switching from TD stereo mode to MDCT stereo mode is relatively straightforward because both stereo modes handle two transport channels and employ two core decoder instances.
[00203] Since an opposite-phase down-mixing scheme was employed in the TD stereo encoder 400, the stereo mode switching controller (not shown) similarly alters the up-mixing of the TD stereo channel to maintain the correct phase of the left and right channels of the stereo sound signal in the last TD stereo frame before the first MDCT stereo frame. Specifically, the stereo mode switching controller (not shown) sets the mixing ratio β=LO and implements opposite-phase up-mixing (inverse of the opposite-phase down-mixing employed in the TD stereo encoder 400) of the TD stereo primary channel PCh(i) and the TD stereo secondary channel SCh(i) to calculate the MDCT stereo left channel and the MDCT stereo right channel.stereo MDCT Consequently, the primary channel PCh(i) stereo TD is identical to the left channel past Ipas^i) stereo MDCT and the secondary channel signal SCh(i) stereo TD is identical to the right channel past rpusl(i) MDCT. 2.6 Switching from MDCT stereo mode to TD stereo mode on the IVAS stereo decoding device
[00204] Similar to the switching from TD stereo mode to MDCT stereo mode, two transport channels are available and two core decoder instances are employed in this scenario. In order to maintain the correct phase of the left and right channels of the stereo sound signal, the TD stereo mixing ratio is set to LO and the opposite phase up-mixing scheme is again used by the stereo mode switching controller (not shown) at the first TD stereo frame after the last MDCT stereo frame. 2.7 Switching from DFT stereo mode to MDCT stereo mode on the IVAS stereo decoding device
[00205] In this scenario, a mechanism similar to the decoder-side switching from DFT stereo mode to TD stereo mode is used, where the primary PCh and secondary SCh channels of the TD stereo mode are replaced by the left 1 and right r channels of the MDCT stereo mode. 2.8 Switching from MDCT stereo mode to DFT stereo mode on the IVAS stereo decoding device
[00206] In this scenario, a mechanism similar to the decoder-side switching from TD stereo mode to DFT stereo mode is used, where the primary PCh and secondary SCh channels of the TD stereo mode are replaced by the left 1 and right r channels of the MDCT stereo mode.
[00207] Finally, decoding continues regardless of the current stereo mode with IC-BWE 865 decoding (omitted in MDCT stereo mode), the addition of HB synthesis (omitted in Lncpnn / zznz / B / YiAi MDCT stereo mode), ICA 866 time alignment (omitted in MDCT stereo mode), and common stereo decoder updates. 2.9 Hardware Implementation
[00208] Figure 14 is a simplified block diagram of an example configuration of the hardware components that make up each of the IVAS 200 stereo encoding and IVAS 800 stereo decoding devices described above.
[00209] Each of the IVAS stereo encoding device 200 and the IVAS stereo decoding device 800 may be implemented as a part of a mobile terminal, as a part of a portable media player, or in any similar device. Each of the IVAS stereo encoding device 200 and IVAS stereo decoding device 800 (identified as 1400 in Figure 14) comprises an input 1402, an output 1404, a processor 1406, and a memory 1408.
[00210] Input 1402 is configured to receive left 1 and right r channels of the input stereo sound signal in digital or analog form in the case of IVAS stereo encoding device 200, or bit stream 803 in the case of IVAS stereo decoding device 800. Output 1404 is configured to supply multiplexed bit stream 206 in the case of IVAS stereo encoding device 200 or decoded left 1 channel and right r channel in the case of IVAS stereo decoding device 800. Input 1402 and output 1404 may be implemented in a common module, for example a serial input / output device.
[00211] Processor 1406 is operatively connected to input 1402, output 1404, and memory 1408. Processor 1406 is comprised of one or more processors for executing code instructions in support of the functions of the various elements and operations of the IVAS stereo encoding device 200, IVAS stereo encoding method 250, IVAS stereo decoding device 800, and IVAS stereo decoding method 850 described above, as shown in the accompanying figures, and / or as described herein.
[00212] Memory 1408 may comprise a non-transitory memory for storing code instructions executable by processor 1406, specifically, a processor-readable memory that stores non-transitory instructions that, when executed, cause a processor to implement the elements and operations of IVAS stereo encoding device 200, IVAS stereo encoding method 250, IVAS stereo decoding device 800, and IVAS stereo decoding method 850. Memory 1408 may also comprise a random access memory or buffer memory(ies) for storing intermediate processing data of the various functions performed by processor 1406.
[00213] Those of ordinary skill in the art will realize that the description of the IVAS 200 stereo encoding device, the IVAS 250 stereo encoding method, the IVAS 250 stereo encoding device, The IVAS 800 stereo decoding LncRnn / zznz / E / YiAi and IVAS 850 stereo decoding methods are illustrative only and are not intended to be limiting in any way. Other embodiments will readily suggest themselves to those of ordinary skill in the art having the benefit of the present invention. Furthermore, the disclosed IVAS 200 stereo encoding device, IVAS 250 stereo encoding method, IVAS 800 stereo decoding device, and IVAS 850 stereo decoding method may be customized to provide valuable solutions to existing stereo sound encoding and decoding needs and problems.
[00214] For the sake of clarity, not all routine features of implementations of the IVAS 200 stereo encoding device, the IVAS 250 stereo encoding method, the IVAS 800 stereo decoding device, and the IVAS 850 stereo decoding method are shown or described. It will be appreciated, of course, that in developing any actual implementation of the IVAS 200 stereo encoding device, the IVAS 250 stereo encoding method, the IVAS 800 stereo decoding device, and the IVAS 850 stereo decoding method, numerous implementation-specific decisions may need to be made to achieve specific developer goals, such as compliance with application, system, network, and business-related constraints, and that these specific goals will vary from implementation to implementation and developer to developer.Furthermore, it will be appreciated that a development effort may be complex and time-consuming, but would nonetheless be a routine engineering task for those of ordinary skill in the field of sound processing having the benefit of the present invention.
[00215] In accordance with the present invention, the elements, processing operations, and / or data structures described herein may be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general purpose machines. Furthermore, those skilled in the art will recognize that devices of a less general purpose nature, such as hard-wired devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), or the like, may also be used. When a method comprising a series of operations and sub-operations is implemented by a processor, computer, or machine, and such operations and sub-operations may be stored as a series of non-transitory code instructions readable by the processor, computer, or machine, they may be stored on a tangible and / or non-transitory medium.
[00216] The elements and processing operations of the IVAS stereo encoding device 200, the IVAS stereo encoding method 250, the IVAS stereo decoding device 800, and the IVAS stereo decoding method 850 described herein may comprise software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein. mcRnn / zznz / E / YiAi
[00217] In the IVAS 250 stereo encoding method and IVAS 850 stereo decoding method described herein, the various processing operations and sub-operations may be performed in various orders and some of the processing operations and sub-operations may be optional.
[00218] Although the present invention has been described herein by means of non-restrictive and illustrative embodiments thereof, these embodiments may be modified at will within the scope of the appended claims without departing from the spirit and nature of the present invention.
[00219] The present invention mentions the following references, the entire contents of which are incorporated herein by reference: [1] 3GPPTS 26.445, v. 12.0.0, Enhanced Voice Services (EVS) Code; Detailed Algorithmic Description, Sep 2014. [2] M. Neuendorf, M. Multrus, N. Rettelbach, G. Fuchs, J. Robillard, J. Lecompte, S. Wilde, S. Bayer, S. Disch, C. Helmrich, R. Lefevbre, P. Gournay, et al., ISO / MPEG unified speech and audio coding standard: Consistent high quality for all content types and at all bit rates, J. Audio Eng. Soc., vol. 61, no. 12, pp. 956–977, Dec. 2013. [3] F. Baumgarte, C. Faller, Binaural signal coding - Part I: Psychoacoustic foundations and design principles, IEEE Trans. Speech Audio Processing, vol. 11, PP· 509-519, Nov. 2003. [4] T. Vaillancourt, Method and system utilizing a long-term correlation difference between left and right channels for time-domain down-mixing of a stereo sound signal into primary and secondary channels, PCT application WO2017 / 049397A1. [5] V. Eksler, Method and device for allocating a bit allocation between subframes in a CELP codec, PCT Application WO2019 / 056107A1. [6] M. Neuendorf et al., MPEG Unified Voice and Audio Coding - An ISO / MPEG standard for high-efficiency audio coding of all content types, Journal of the Audio Engineering Society, vol. 61, no. 12, pp. 956-977, December 2013. [7] J. Herre et al., MPEG-H Audio: The New Standard for Universal Coding of Spatial and 3D Audio, in 137th International AES Convention, Paper 9095, Los Angeles, October 9-12, 2014. [8] 3GPP SA4 Contribution S4-180462, On Spatial Metadata for the IVAS Spatial Audio Input Format, SA4 meeting #98, April 9-13, 2018, https: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_98 / Docs / S4-180462.zip [9] V. Malenovsky, T. Vaillancourt, “Method and device for uncorrelated stereo content classification, crosstalk detection, and stereo mode selection in a sound codec,” U.S. provisional patent application 63 / 075,984 filed September 9, 2020. LncAnn / zznz / E / YiAi
Claims
1. A device for encoding a stereo sound signal, comprising: a first stereo encoder of the stereo sound signal using a first stereo mode operating in the time domain (TD), wherein the first stereo mode TD, in the TD frames of the stereo sound signal, (a) produces a first down-mixed signal and (b) uses first data structures and memories; a second stereo encoder of the stereo sound signal using a second stereo mode operating in the frequency domain (FD), wherein the second stereo mode FD, in the FD frames of the stereo sound signal, (a) produces a second down-mixed signal and (b) uses second data structures and memories; a switching controller between (i) the first stereo mode TD and the first stereo encoder, and (ii) the second stereo mode FD and the second stereo encoder for encoding the stereo sound signal in the time domain or in the frequency domain, wherein,When switching from one of the stereo modes first TD and second FD to another of the stereo modes first TD and second FD, the stereo mode switching controller recalculates at least one down-mixed signal duration in a current frame of the stereo sound signal, where the recalculated down-mixed signal duration in the first stereo mode TD is different from the recalculated down-mixed signal duration in the second stereo mode FD.
2. A stereo sound signal encoding device as described in claim 1, wherein the second stereo mode FD is a discrete Fourier transform (DFT) stereo mode.
3. A stereo sound signal encoding device as described in claim 2, wherein, when switching from one of the first TD and second DFT stereo modes to another of the first TD and second DFT stereo modes, the stereo mode switching controller assigns / unassigns data structures to / from the first TD and second DFT stereo modes depending on the current stereo mode, to reduce the impact on memory by keeping only those data structures used in the current frame.
4. A stereo sound signal encoding device as described in claim 3, wherein, when switching from the first stereo mode TD to the second stereo mode DFT, the stereo mode switching controller deals the data structures related to the stereo TD.
5. A stereo sound signal encoding device as described in claim 4, wherein the data structures related to the stereo TD comprise a stereo TD data structure and / or data structures of a core encoder of the first stereo encoder. LncAnn / zznz / E / YiAi 6. A stereo sound signal encoding device as described in any one of claims 2 to 5, wherein, when switching from the first stereo mode TD to the second stereo mode DFT, the second stereo encoder continues a core encoding operation on a stereo DFT frame following a stereo TD frame with memories of a primary channel core encoder PCh.
7. A stereo sound signal encoding device as described in any of claims 2 to 6, wherein the stereo mode switching controller uses stereo parameters related to said stereo mode to update the stereo parameters related to said other stereo mode when switching from said stereo mode to said other stereo mode.
8. A stereo sound signal encoding device as described in claim 7, wherein the related stereo parameters comprise a side gain and an inter-channel time delay (ITD) parameter of the second stereo mode DFT and a target gain and correlation phases of the first stereo mode TD.
9. A stereo sound signal encoding device as described in any one of claims 2 to 8, wherein the stereo mode switching controller updates a DFT analysis memory each TD frame by storing samples related to a last time period of a current TD frame.
10. A stereo sound signal encoding device as described in any one of claims 2 to 9, wherein the stereo mode switching controller retains DFT-related memories during TD frames.
11. A stereo sound signal encoding device as described in any one of claims 2 to 10, wherein the stereo mode switching controller, when switching from the first stereo mode TD to the second stereo mode DFT, updates in a DFT frame following a TD frame a DFT synthesis memory using stereo TD memories corresponding to a primary channel PCh of the TD frame.
12. A stereo sound signal encoding device as described in any one of claims 2 to 11, wherein the stereo mode switching controller maintains a finite impulse response (FIR) resampling filter memory during the DFT frames of the stereo sound signal, and wherein the stereo mode switching controller updates, in each DFT frame, the FIR resampling filter memory used in a primary channel PCh in the first stereo encoder, using a segment of a half-channel m before a last segment of the first half-channel m in the DFT frame. LncAnn / zznz / E / YiAi 13. A stereo sound signal encoding device as described in claim 12, wherein the stereo mode switching controller fills an FIR resampling filter memory used in a secondary channel SCh in the first stereo encoder differently from updating the FIR resampling filter memory used in the primary channel PCh in the first stereo encoder.
14. A stereo sound signal encoding device as described in claim 13, wherein the stereo mode switching controller updates in a current DFT frame the FIR resampling filter memory used in the secondary channel SCh in the first stereo encoder, filling the FIR resampling filter memory using a segment of a medium channel m in the DFT frame before a last segment of the second duration of the medium channel m.
15. A stereo sound signal encoding device as described in any one of claims 2 to 14, wherein, when switching from the second stereo mode DFT to the first stereo mode TD, the stereo mode switching controller recalculates in a current TD frame a down-mixed signal duration that is longer in a secondary channel SCh with respect to a recalculated down-mixed signal duration in a primary channel PCh.
16. A stereo sound signal encoding device as described in any one of claims 2 to 15, wherein, when switching from the second DFT stereo mode to the first TD stereo mode, the stereo mode switching controller cross-fades a recalculated PCh primary channel and a DFT middle channel m from a DFT stereo channel to recalculate a down-mixed PCh primary channel in a first TD frame following a DFT frame.
17. A stereo sound signal encoding device as described in any one of claims 2 to 16, wherein, when switching from the second stereo DFT mode to the first stereo TD mode, the stereo mode switching controller recalculates an ICA memory of the left 1 and right r channels corresponding to a DFT frame preceding a TD frame.
18. A stereo sound signal encoding device as described in claim 17, wherein the stereo mode switching controller recalculates the primary channels PChs and secondary channels SCh of the DFT frame by down-mixing the ICA-processed channels 1 and r using a stereo mix ratio of the DFT frame.
19. A stereo sound signal encoding device as described in claim 18, wherein the stereo mode switching controller shortens the duration of the secondary channel SCh when there is no stereo mode switching. LncAnn / zznz / E / YiAi 20. A stereo sound signal encoding device as described in claim 18 or 19, wherein the stereo mode switching controller recalculates, in the DFT frame preceding the TD frame, a first duration of the primary channel PCh and a second duration of the secondary channel SCh, and wherein the first duration is shorter than the second.
21. A stereo sound signal encoding device as described in any one of claims 2 to 20, wherein the stereo mode switching controller stores two values of a pre-emphasis filter memory in each DFT frame of the stereo sound signal.
22. A stereo sound signal encoding device as described in any one of claims 2 to 21, comprising secondary channel encoder core data structures SCh wherein, when switching from second stereo mode DFT to first stereo mode TD, the stereo mode switching controller resets or estimates the secondary channel encoder core data structures SCh based on the primary channel core data structures PCh.
23. A device for decoding a stereo sound signal, comprising: a first stereo decoder of the stereo sound signal using a first stereo mode operating in the time domain (TD), wherein the first stereo decoder, in the TD frames of the stereo sound signal, (a) decodes a down-mixed signal and (b) uses first data structures and memories;a second stereo decoder of the stereo sound signal using a second stereo mode operating in the frequency domain (FD), wherein the second stereo decoder, in the FD frames of the stereo sound signal, (a) decodes a second down-mixed signal and (b) uses second data structures and memories; a switching controller between (i) the first stereo mode TD and the first stereo decoder and (ii) the second stereo mode FD and the second stereo decoder, wherein, when switching from one of the first stereo mode TD and second stereo mode FD to the other of the first stereo mode TD and second stereo mode FD, the stereo mode switching controller recalculates at least one down-mixed signal duration in a current frame of the stereo sound signal, wherein the recalculated down-mixed signal duration in the first stereo mode TD is different from the recalculated down-mixed signal duration in the second stereo mode FD.
24. A stereo sound signal decoding device as described in claim 23, wherein the second stereo mode FD is a discrete Fourier transform (DFT) stereo mode.
25. A stereo sound signal decoding device as described in claim 24, wherein the first stereo TD mode uses first processing delays, the second LncAnn / zznz / E / YiAi stereo DFT mode uses second processing delays, and the first and second processing delays are different and comprise resampling and up-mixing processing delays.
26. A stereo sound signal decoding device as described in claim 24 or 25, wherein the stereo mode switching controller assigns / unassigns data structures to / from the first mode TD and the second stereo mode DFT depending on the current stereo mode, to reduce a static memory impact by keeping only those data structures that are used in the current frame.
27. A stereo sound signal decoding device as described in any one of claims 24 to 26, wherein, upon receiving a first DFT frame after a TD frame, the stereo mode switching controller restores a stereo DFT data structure.
28. A stereo sound signal decoding device as described in any one of claims 24 to 27, wherein, upon receiving a first TD frame following a DFT frame, the stereo mode switching controller restores a stereo TD data structure.
29. A stereo sound signal decoding device as described in any one of claims 24 to 28, wherein the stereo mode switching controller updates the stereo DFT OLA memory buffers on each stereo TD frame.
30. A stereo sound signal decoding device as described in any one of claims 24 to 29, wherein the stereo mode switching controller updates the DFT stereo analysis memories, and wherein, upon receiving a first DFT frame following a TD frame, the stereo mode switching controller uses a number of last samples from a primary channel PCh and a secondary channel SCh of the TD frame to update in the DFT frame the DFT stereo analysis memories of a DFT stereo middle channel m and a side channel s, respectively.
31. A stereo sound signal decoding device as described in any one of claims 24 to 30, wherein the stereo mode switching controller updates the DFT stereo synthesis memories in each TD stereo frame.
32. A stereo sound signal decoding device as described in claim 31, wherein, for updating the DFT stereo synthesis memories and for an ACELP core, the stereo mode switching controller reconstructs in each TD frame a first part of the DFT stereo synthesis memories by crossfading (a) a resampled and upmixed TD-based left and right channel synthesis and (b) a reconstructed resampled and upmixed left and right channel synthesis. LncAnn / zznz / E / YiAi 33. A stereo sound signal decoding device as described in any one of claims 24 to 32, wherein the stereo mode switching controller crosses a TD aligned and synchronized synthesis with a DFT stereo aligned and synchronized synthesis to smooth the transition when switching from a TD frame to a DFT frame.
34. A stereo sound signal decoding device as described in any one of claims 24 to 33, wherein the encoding mode switching controller updates the TD stereo synthesis memories during DFT frames in case a subsequent frame is a TD frame.
35. A stereo sound signal decoding device as described in any one of claims 24 to 34, wherein, when switching from a DFT frame to a TD frame, the stereo mode switching controller resets the memories of a secondary channel SCh core decoder in the first stereo decoder.
36. A stereo sound signal decoding device as described in any one of claims 24 to 35, wherein, when switching from a DFT frame to a TD frame, the stereo mode switching controller suppresses discontinuities and differences between the up mixed stereo DFT and TD channels by using signal energy equalization.
37. A stereo sound signal decoding device as described in any one of claims 24 to 36, wherein the stereo mode switching controller reconstructs a stereo up-synchronized TD synthesis, and wherein the stereo mode switching controller uses the following operations (a) to (e) for a left channel and a right channel to reconstruct the stereo up-synchronized TD synthesis: (a) repairing a stereo DFT OLA synthesis memory; (b) reusing a stereo up-synchronized DFT synthesis memory as a first part of the stereo up-synchronized TD synthesis; (c) approximating a second part of the stereo up-synchronized TD synthesis, using the repaired stereo DFT OLA synthesis memory;and (d) smoothing a transition between the up-mixed stereo DFT sync synthesis memory and an up-mixed stereo TD sync synthesis at the beginning of the up-mixed stereo TD sync synthesis by crossfading the repaired OLA stereo DFT synthesis memory with the up-mixed stereo TD sync synthesis. LncAnn / zznz / E / YiAi; 38. A method for encoding a stereo sound signal, comprising: providing a first stereo encoder of the stereo sound signal using a first stereo mode operating in the time domain (TD), wherein the first stereo mode TD, in the TD frames of the stereo sound signal, (a) produces a first down-mixed signal and (b) utilizes first data structures and memories;providing a second stereo encoder of the stereo sound signal using a second stereo mode operating in the frequency domain (FD), wherein the second stereo mode FD, in the FD frames of the stereo sound signal, (a) produces a second down-mixed signal and (b) uses second data structures and memories controlling the switching between (i) the first stereo mode TD and the first stereo encoder, and (ii) the second stereo mode FD and the second stereo encoder to encode the stereo sound signal in the time domain or in the frequency domain;Wherein, when switching from one of the stereo modes first TD and second FD to the other of the stereo modes first TD and second FD, the stereo mode switching control comprises recalculating at least one down-mixed signal duration in a current frame of the stereo sound signal, wherein the recalculated down-mixed signal duration in the first stereo mode TD is different from the recalculated down-mixed signal duration in the second stereo mode FD.
39. A method for encoding stereo sound signals as stated in claim 38, wherein the second stereo mode FD is a discrete Fourier transform (DFT) stereo mode.
40. A method for encoding stereo sound signals as described in claim 39, wherein, when switching from one of the first TD and second DFT stereo modes to the other of the first TD and second DFT stereo modes, the stereo mode switching control comprises maintaining the continuity of at least one of the following signals: - an input stereo signal including the left and right channels; - a middle channel used in the second DFT stereo mode; - a primary channel and a secondary channel used in the first TD stereo mode; - a down-mixed signal used in preprocessing; and - a down-mixed signal used in core encoding.
41. A method for encoding stereo sound signals as described in claim 39 or 40, wherein, when switching from one of the first TD and second DFT stereo modes to the other of the first TD and second DFT stereo modes, the stereo mode switching control comprises the allocation / deallocation of data structures to / from the first TD and second DFT stereo modes based on the current stereo mode, to reduce the impact on memory by retaining only those data structures used in the current frame. LncAnn / zznz / E / YiAi 42. A method of encoding stereo sound signals as described in claim 41, wherein, when switching from the first stereo mode TD to the second stereo mode DFT, the stereo mode switching control comprises reassigning data structures related to the stereo mode TD, and wherein the data structures related to the stereo mode TD comprise a stereo TD data structure and / or data structures of a core encoder of the first stereo encoder.
43. A method of encoding stereo sound signals as described in any one of claims 39 to 42, wherein, when switching from the first stereo mode TD to the second stereo mode DFT, the second stereo encoder continues a core encoding operation on a DFT frame following a TD frame with memories of a primary channel core encoder PCh.
44. A method of encoding stereo sound signals as described in any one of claims 39 to 43, wherein the stereo mode switching control comprises using stereo-related parameters of said stereo mode to update the stereo-related parameters of said other stereo mode when switching from said stereo mode to said other stereo mode.
45. A method of encoding stereo sound signals as described in claim 44, wherein the control of the stereo mode switching comprises the transfer of stereo-related parameters between data structures, and wherein the stereo-related parameters comprise a side gain and an inter-channel time delay (ITD) parameter of the second stereo mode DFT and a target gain and correlation phases of the first stereo mode TD.
46. A method of encoding stereo sound signals as described in any one of claims 39 to 45, wherein the stereo mode switching control comprises updating a DFT analysis memory every stereo TD frame by storing samples related to a last time period of a current stereo TD frame.
47. A method of encoding stereo sound signals as described in any one of claims 39 to 46, wherein the stereo mode switching control comprises maintaining DFT-related memories during TD stereo frames.
48. A method for encoding stereo sound signals as described in any one of claims 39 to 47, wherein the stereo mode switching control comprises, when switching from the first stereo mode TD to the second stereo mode DFT, updating in a DFT frame following a TD frame of a DFT synthesis memory using stereo TD memories corresponding to a primary channel PCh of the TD frame. LncAnn / zznz / E / YiAi 49. A method of encoding stereo sound signals as described in any one of claims 39 to 48, wherein the stereo mode switching control comprises maintaining a finite impulse response (FIR) resampling filter memory during DFT frames.
50. A method of encoding stereo sound signals as stated in claim 49, wherein the stereo mode switching control comprises updating in each DFT frame the memory of the FIR resampling filter used in a primary channel PCh in the first stereo encoder, using a segment of a medium channel m before a last segment of first duration of the medium channel m in the DFT frame.
51. A method of encoding stereo sound signals as stated in claim 49 or 50, wherein the switching control comprises filling an FIR resampling filter memory used in a secondary channel SCh in the first stereo encoder, differently from updating the FIR resampling filter memory used in the primary channel PCh in the first stereo encoder.
52. A method of encoding stereo sound signals as stated in claim 51, wherein the stereo mode switching control comprises updating in a current TD frame the memory of the FIR resampling filter used in the secondary channel SCh in the first stereo encoder, filling the FIR resampling filter memory using a segment of a medium channel m in the DFT frame before a last segment of the second duration of the medium channel m.
53. A method of encoding stereo sound signals as described in any one of claims 39 to 52, wherein, when switching from the second stereo mode DFT to the first stereo mode TD, the stereo mode switching control comprises recalculating in a current TD frame a down-mixed signal duration that is longer in a secondary channel SCh with respect to a recalculated down-mixed signal duration in a primary channel PCh.
54. A method of encoding stereo sound signals as described in any one of claims 39 to 53, wherein, when switching from the second stereo DFT mode to the first stereo TD mode, the stereo mode switching control comprises crossfading a recalculated primary channel PCh and a DFT middle channel m from a DFT channel to recalculate a down-mixed primary channel PCh in a first TD frame following a DFT frame.
55. A method for encoding stereo sound signals as described in any one of claims 39 to 54, wherein, when switching from the second stereo mode DFT to the first stereo mode TD, the stereo mode switching control comprises recalculating an ICA memory of the left 1 and right r channels corresponding to a DFT frame preceding a TD frame. LncAnn / zznz / E / YiAi 56. A method of encoding stereo sound signals as stated in claim 55, wherein the stereo mode switching control comprises recalculating the primary channels PCh and secondary channels SCh of the DFT frame by down-mixing the ICA-processed channels 1 and r using a stereo mixing ratio of the DFT frame.
57. A method of encoding stereo sound signals as stated in claim 56, wherein the stereo mode switching control comprises recalculating a shorter duration of the secondary channel SCh when there is no stereo encoding mode switching.
58. A method of encoding stereo sound signals as described in claim 56 or 57, wherein the stereo mode switching control comprises recalculating, in the DFT frame preceding the TD frame, a first duration of the primary channel PCh and a second duration of the secondary channel SCh, and wherein the first duration is shorter than the second.
59. A method of encoding stereo sound signals as described in any one of claims 39 to 58, wherein the stereo mode switching control comprises storing two values of a pre-emphasis filter memory in each DFT frame.
60. A method for encoding stereo sound signals as described in any one of claims 39 to 59, comprising secondary channel encoder core data structures SCh wherein, when switching from second stereo mode DFT to first stereo mode TD, the stereo mode switching control comprises resetting or estimating the secondary channel encoder core data structures SCh based on the primary channel core data structures PCh.
61. A method for decoding a stereo sound signal, comprising providing a first stereo decoder of the stereo sound signal using a first stereo mode operating in the time domain (TD), wherein the first stereo decoder, in TD frames of the stereo sound signal, (a) decodes a down-mixed signal and (b) utilizes first data structures and memories;providing a second stereo decoder of the stereo sound signal using a second stereo mode operating in the frequency domain (FD), wherein the second stereo decoder, in the FD frames of the stereo sound signal, (a) decodes a second down-mixed signal and (b) uses second data structures and memories to control switching between (i) the first stereo mode TD and the first stereo decoder and (ii) the second stereo mode FD and the second stereo decoder LncRnn / zznz / E / YiAi, wherein, when switching from one of the first stereo mode TD and second stereo mode FD to the other of the first stereo mode TD and second stereo mode FD, the stereo mode switching control comprises recalculating at least one down-mixed signal duration in a current frame of the stereo sound signal, wherein the recalculated down-mixed signal duration in the first stereo mode is different from the recalculated down-mixed signal duration in the second stereo mode.
62. A method for decoding stereo sound signals as stated in claim 61, wherein the second stereo mode FD is a discrete Fourier transform (DFT) stereo mode.
63. A method for decoding stereo sound signals as stated in claim 62, wherein the first stereo mode uses first processing delays, the second stereo mode uses second processing delays, and the first and second processing delays are different and comprise resampling and up-mixing processing delays.
64. A method for decoding stereo sound signals as described in claim 62 or 63, wherein, when switching from one of the first TD and second DFT stereo modes to the other of the first FD and second DFT stereo modes, the stereo mode switching control comprises maintaining the continuity of at least one of the following signals and memories: - a middle channel m used in the second DFT stereo mode; - a primary channel PCh and a secondary channel SCh used in the first TD stereo mode; - TCX-LTP post-filtering memories; - OLA DFT analysis memories at an internal sampling rate and at a sampling rate of the output stereo signal; - OLA DFT synthesis memories at the sampling rate of the output stereo signal; - an output stereo signal, including channels 1 yr; and - HB signal memories, and channels 1 yr used in BWE and IC-BWE.
65. A method for decoding stereo sound signals as described in any one of claims 62 to 64, wherein the stereo mode switching control comprises assigning / unassigning data structures to / from the first TD and second DFT stereo modes based on the current stereo mode, to reduce a static memory impact by keeping only those data structures used in the current frame.
66. A method for decoding stereo sound signals as described in any one of claims 62 to 65, wherein, upon receiving a first DFT frame following a TD frame, the stereo mode switching control comprises restoring a stereo DFT data structure. LncRnn / zznz / E / YiAi 67. A method for decoding stereo sound signals as described in any one of claims 62 to 66, wherein, upon receiving a first TD frame following a DFT frame, the switching control comprises restoring a stereo TD data structure.
68. A method for decoding stereo sound signals as described in any one of claims 62 to 67, wherein the control of the stereo mode switching comprises updating the stereo DFT OLA memory buffers in each TD frame.
69. A method for decoding stereo sound signals as described in any one of claims 62 to 68, wherein the stereo mode switching control comprises updating the DFT stereo analysis memories.
70. A method for decoding stereo sound signals as described in claim 69, wherein, upon receiving a first DFT frame following a TD frame, the stereo mode switching control comprises using a number of last samples from a primary channel PCh and a secondary channel SCh of the TD frame to update in the DFT frame the stereo DFT analysis memories of a middle channel m and a stereo DFT side channel s, respectively.
71. A method for decoding stereo sound signals as described in any one of claims 62 to 70, wherein the stereo mode switching control comprises updating the DFT stereo synthesis memories in each TD frame, and wherein, for updating the DFT stereo synthesis memories and for an ACEUP core, the stereo mode switching control comprises reconstructing in each TD frame a first part of the DFT stereo synthesis memories by crossfading (a) a resampled and upmixed TD-based left and right channel synthesis and (b) a reconstructed, resampled, and upmixed left and right channel synthesis.
72. A method for decoding stereo sound signals as described in any one of claims 62 to 71, wherein the stereo mode switching control comprises crossfading a TD aligned and synchronized synthesis with a DFT aligned and synchronized synthesis to smooth the transition when switching from a TD frame to a DFT frame.
73. A method for decoding stereo sound signals as described in any one of claims 62 to 72, wherein the stereo mode switching control comprises updating the TD stereo synthesis memories during DFT frames if a subsequent frame is a TD frame. LncAnn / zznz / E / YiAi 74. A method of decoding stereo sound signals as described in any one of claims 62 to 73, wherein, when switching from a DFT frame to a TD frame, the switching control comprises resetting the memories of a secondary channel SCh core decoder in the first stereo decoder.
75. A method for decoding stereo sound signals as described in any one of claims 62 to 74, wherein, when switching from a DFT frame to a TD frame, the stereo mode switching control comprises suppressing discontinuities and differences between the upmixed DFT and TD stereo channels using signal energy equalization, and wherein, to suppress discontinuities and differences between the upmixed DFT and TD stereo channels, the stereo mode switching control comprises, if a target gain ICA, gicA, is less than 1.0, altering the left channel 1, i), after upmixing and before time synchronization in the TD frame using the following relationship: y'L(i) = a yL(i) for i = 0,, Leq — 1 where Leq is a signal duration to be equalized, and U is a value of a gain factor obtained by the following relationship: . . 1_ Sica n, . a = gICA + i-------- for i = 0,..., Leq — 1 ^eq 76. A method for decoding stereo sound signals as described in any one of claims 62 to 75, wherein the control of the stereo mode switching comprises the reconstruction of a stereo up-synchronized TD synthesis, and wherein the switching control comprises the use of the following operations (a) to (e) for both a left channel and a right channel to reconstruct the stereo up-synchronized TD synthesis: (a) repairing a stereo DFT OLA synthesis memory; (b) reusing a stereo up-synchronized DFT synthesis memory as a first part of the stereo up-synchronized TD synthesis; (c) approximating a second part of the stereo up-synchronized TD synthesis using the repaired stereo DFT OLA synthesis memory;and (d) smoothing a transition between the DFT stereo up-mixed synchronized synthesis memory and a TD stereo up-mixed synchronized synthesis at the beginning of the TD stereo up-mixed synchronized synthesis by crossfading the repaired DFT stereo OLA synthesis memory with the TD stereo up-mixed synchronized synthesis.;