Switching between stereo coding modes in a multi-channel audio codec
The adaptive stereo encoding system addresses the inefficiencies of existing stereo coding by dynamically switching between DFT, TD, and MDCT modes, ensuring high-quality stereo audio at low bit rates in complex environments.
Patent Information
- Application Number
- JP2022547128
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-03
- Filing Date
- 2021-02-01
- Publication Date
- 2025-12-04
- Estimated Expiration
- 2041-02-01
AI Technical Summary
Existing stereo coding techniques in multi-channel audio codecs often double the bit rate when transmitting stereo signals, fail to exploit channel redundancy, and compromise sound quality at low bit rates, especially in complex audio scenes with low correlation and varying noise or multiple speakers.
A stereo sound encoding device and method that switches between frequency-domain (DFT), time-domain (TD), and modified discrete cosine transform (MDCT) stereo modes, optimizing coding efficiency and maintaining stereo quality at low bit rates by dynamically selecting the most suitable mode based on audio characteristics and channel redundancy.
The system maintains high stereo quality and low bit rates by adaptively switching coding modes, effectively handling complex audio scenes with varying noise and multiple speakers, reducing bit rate demands while preserving audio fidelity.
Smart Images

Figure 0007780441000008 
Figure 0007780441000009 
Figure 0007780441000010
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to stereo sound encoding, and in particular, but not exclusively, to switching between "stereo coding modes" (hereinafter also "stereo modes") in multi-channel sound codecs that are capable of producing good stereo quality, for example in complex audio scenes, at low bitrates and low delay.
[0002] In this disclosure and the accompanying claims: The term "sound" may relate to speech, audio, and any other sound; - The term "stereo" is short for "stereophonic" -The term "mono" is short for "monophonic." [Background technology]
[0003] Historically, interactive telephony has been implemented with handsets that have only one transducer for outputting sound to only one ear of the user. Over the past decade, users have begun to use their mobile handsets with headphones to receive sound in both ears, primarily for listening to music, but also occasionally for listening to speech. Nevertheless, when a mobile handset is used to transmit and receive conversational audio, the content is still mono, but when headphones are used, it is delivered to both ears of the user.
[0004] The latest 3GPP® speech coding standards, as described in Non-Patent Document 1, the entire contents of which are incorporated herein by reference, have greatly improved the quality of coded sounds, e.g., speech and / or audio, transmitted and received through mobile handsets. The next natural step is to transmit stereo information, so that the receiver gets as close as possible to the real-world audio scene as seen on the other side of the communication link.
[0005] The transmission of stereo information is commonly used in audio codecs, such as those described in Non-Patent Document 2, the entire contents of which are incorporated herein by reference.
[0006] In speech codecs, mono signals are the norm. When stereo signals are transmitted, both the left and right channels of the stereo signal are coded using a mono codec, often doubling the bit rate. While this works well in most scenarios, it has the drawback of doubling the bit rate and not exploiting any redundancy that may exist between the two channels (the left and right channels of a stereo signal). Furthermore, to keep the overall bit rate reasonable, very low bit rates are used for each channel, which impacts the overall sound quality. To reduce the bit rate, efficient stereo coding techniques have been developed and are being used. As non-limiting examples, the use of three stereo coding techniques that can be used efficiently at low bit rates is discussed in the following paragraphs.
[0007] The first stereo coding technique is called parametric stereo. Parametric stereo coding encodes two channels, the left and right channels, as a mono signal using a common mono codec and a certain amount of stereo side information (corresponding to stereo parameters) that represents the stereo sound image. The two input left and right channels are downmixed to a mono signal, and the stereo parameters are usually calculated in the transform domain, for example, the discrete Fourier transform (DFT) domain, and are related to so-called binaural or inter-channel cues. Binaural cues (see Non-Patent Document 3, the entire contents of which are incorporated herein by reference) include interaural level difference (ILD), interaural time difference (ITD), and interaural correlation (IC). Depending on the signal characteristics, stereo scene configuration, etc., some or all binaural cues are coded and transmitted to the decoder. Information about the binaural cues is coded and transmitted as signaling information, which is usually part of the stereo side information. A particular binaural cue can also be quantized using different coding techniques, resulting in variations in the number of bits used. In addition to the quantized binaural cues, the stereo side information may also include a quantized residual signal resulting from downmixing, usually at intermediate and higher bit rates. The residual signal may be coded using an entropy coding technique, e.g., an arithmetic coder. Parametric stereo coding using stereo parameters calculated in the transform domain is referred to in this disclosure as "DFT stereo" coding.
[0008] Another stereo coding technique operates in the time domain (TD). This stereo coding technique mixes two inputs, a left channel and a right channel, into so-called primary and secondary channels. For example, according to a method described in U.S. Patent Application Publication No. 2009 / 0109990, the entire contents of which are incorporated herein by reference, time-domain mixing may be based on a mixing ratio, which determines the respective contributions of the two inputs, the left channel and the right channel, in generating the primary and secondary channels. The mixing ratio is derived from several metrics, such as the normalized correlation of the input left and right channels with respect to a mono signal version or the difference in the long-term correlation between the two inputs, the left and right channels. The primary channel may be coded with a common mono codec, while the secondary channel may be coded with a lower bitrate codec. Coding of the secondary channel may exploit the coherence between the primary and secondary channels and may reuse some parameters from the primary channel. Time-domain stereo coding is referred to as "TD stereo" coding in this disclosure. In general, TD stereo coding is most efficient at low and medium bit rates for coding speech signals.
[0009] The third stereo coding technique operates in the modified discrete cosine transform (MDCT) domain. It is based on joint coding of both the left and right channels while calculating the global ILD and performing Mid / Side (M / S) processing in the whitened spectral domain. This third stereo coding technique uses several tools adapted from the Transform Coded eXcitation (TCX) coding of the Moving Picture Experts Group (MPEG) codec, as described, for example, in Non-Patent Documents 4 and 5, the entire contents of which are incorporated herein by reference. These tools may include TCX core coding, TCX long-term prediction (LTP) analysis, TCX noise filling, frequency-domain noise shaping (FDNS), stereophonic intelligent gap filling (IGF), and / or adaptive bit allocation between channels. In general, this third stereo coding technique is efficient for encoding all types of audio content at medium and high bit rates. The MDCT-domain stereo coding technique is referred to in this disclosure as "MDCT stereo coding." In general, MDCT stereo coding is most efficient at medium and high bit rates for coding general audio signals.
[0010] In recent years, stereo coding has been further extended to multi-channel coding. While several techniques exist for providing multi-channel coding, the core of all these techniques is often based on one or more instances of mono- or stereo-coding techniques. Therefore, this disclosure presents a switching between stereo coding modes that can be part of a multi-channel coding technique, such as Metadata-Assisted Spatial Audio (MASA), as described, for example, in U.S. Patent Application Publication No. 2010 / 0149994, the entire contents of which are incorporated herein by reference. In the MASA approach, MASA metadata (e.g., direction, energy ratio, spread coherence, distance, and surround coherence, all within a number of time-frequency slots) is generated, quantized, and coded into a bitstream in a MASA analyzer, while MASA audio channels are treated as (multi-)mono or (multi-)stereo transport signals that are coded by a core coder. In a MASA decoder, the MASA metadata then guides the decoding and rendering processes to reconstruct the output spatial audio. [Prior art documents] [Patent documents]
[0011] [Patent Document 1] International Patent Application Publication No. WO2017 / 049397A1 [Patent Document 2] International Patent Application Publication No. WO2019 / 056107A1 [Patent Document 3] U.S. Provisional Patent Application No. 63 / 075,984 [Non-patent literature]
[0012] [Non-Patent Document 1] 3GPP (registered trademark) TS 26.445, v.12.0.0, "Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description", September 2014 [Non-licensed Document 2] M. Neuendorf, M. Multrus, N. Rettelbach, G. Fuchs, J. Robillard, J. Lecompte, S. Wilde, S. Bayer, S. Disch, C. Helmrich, R. Lefevbre, P. Gournay, "The ISO / MPEG Unified Speech and Audio Coding Standard - Consistent High Quality for All Content Types and at All Bit Rates", J. Audio Eng. Soc., vol. 61, no. 12, pp. 956-977, December 2013 [Non-licensed Document 3] F. Baumgarte, C. Faller, "Binaural cue coding - Part I: Psychoacoustic fundamentals and design principles", IEEE Trans. Speech Audio Processing, vol. 11, pp. 509-519, November 2003 [Non-licensed Document 4] M. Neuendorf, "MPEG Unified Speech and Audio Coding - The ISO / MPEG Standard for High-Efficiency Audio Coding of all Content Types", Journal of the Audio Engineering Society, vol. 61, n°12, pp. 956-977, December 2013 [Non-licensed Document 5] J. Herre et al., "MPEG-H Audio - The New Standard for Universal Spatial / 3D Audio Coding," 137th International AES Conference, Paper 9095, Los Angeles, October 9-12, 2014 [Non-patent document 6] 3GPP® SA4 contribution S4-180462, “On spatial metadata for IVAS spatial audio input format,” 98th SA4 Meeting, April 9-13, 2018, https: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_98 / Docs / S4-180462.zip Summary of the Invention [Means for solving the problem]
[0013] The present disclosure provides a stereo sound signal encoding device and method as defined in the appended claims.
[0014] The foregoing and other objects, advantages, and features of the stereo encoding and decoding device and method will become more apparent upon reading the following non-limiting description of illustrative embodiments thereof, given by way of example only with reference to the accompanying drawings. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a schematic block diagram of a sound processing and communication system illustrating a possible implementation of a stereo encoding and decoding device and method. [Figure 2]FIG. 1 is a high-level block diagram illustrating simultaneously an Immersive Voice and Audio Service (IVAS) stereo encoding device and a corresponding stereo encoding method, the IVAS stereo encoding device comprising a frequency domain (FD) stereo encoder, a time domain (TD) stereo encoder, and a modified discrete cosine transform (MDCT) stereo encoder, the implementation of which is based on the discrete Fourier transform (DFT) in this exemplary embodiment and in the accompanying drawings (hereinafter "DFT stereo encoder"). [Figure 3] FIG. 3 is a block diagram illustrating the DFT stereo encoder of FIG. 2 and the corresponding DFT stereo encoding method simultaneously. [Figure 4] FIG. 3 is a block diagram illustrating the TD stereo encoder of FIG. 2 and the corresponding TD stereo encoding method together. [Figure 5] FIG. 3 is a block diagram illustrating the MDCT stereo encoder of FIG. 2 and the corresponding MDCT stereo encoding method simultaneously; [Figure 6] 1 is a flowchart illustrating processing operations in an IVAS stereo encoding device and method when switching from TD stereo mode to DFT stereo mode. [Figure 7a] 1 is a flowchart illustrating processing operations in an IVAS stereo encoding device and method when switching from DFT stereo mode to TD stereo mode. [Figure 7b] 10 is a flowchart showing the processing operation for the TD stereo past signal when switching from the DFT stereo mode to the TD stereo mode. [Figure 8] 1 is a high-level block diagram illustrating an IVAS stereo decoding device and corresponding decoding method simultaneously, the IVAS stereo decoding device comprising a DFT stereo decoder, a TD stereo decoder, and an MDCT stereo decoder. [Figure 9] 1 is a flowchart illustrating processing operations in an IVAS stereo decoding device and method when switching from TD stereo mode to DFT stereo mode. [Figure 10]10 is a flowchart illustrating instance B) of FIG. 9, which comprises updating the DFT stereo synthesis memory in TD stereo frames at the decoder side. [Figure 11] 10 is a flowchart illustrating instance C) of FIG. 9, which comprises smoothing the output stereo synthesis at the first DFT stereo frame after switching from TD stereo mode to DFT stereo mode at the decoder side. [Figure 12] 1 is a flowchart illustrating processing operations in an IVAS stereo decoding device and method when switching from DFT stereo mode to TD stereo mode. [Figure 13] 13 is a flow chart illustrating instance A) of FIG. 12 comprising updating the TD stereo synchronization memory in the first TD stereo frame after switching from DFT stereo mode to TD stereo mode at the decoder side. [Figure 14] 1 is a simplified block diagram of an exemplary configuration of hardware components implementing an IVAS stereo encoding device and method and an IVAS stereo decoding device and method, respectively. DETAILED DESCRIPTION OF THE INVENTION
[0016] As mentioned above, the present disclosure relates to encoding stereo sound, and particularly, but not limited to, switching stereo coding modes in codecs for sound including speech and / or audio, capable of producing good stereo quality at low bit rates and low delays, for example, in complex audio scenes. In the present disclosure, complex audio scenes include, for example, but not limited to, situations where (a) correlation between sound signals recorded by microphones is low, (b) there is significant variation in background noise, and / or (c) there are interfering speakers. Non-limiting examples of complex audio scenes include a large, non-reverberant conference room with an A / B microphone configuration, a small, reverberant room with binaural microphones, and a small, reverberant room with a mono / side microphone setup. All these room configurations may include varying background noise and / or interfering speakers.
[0017] FIG. 1 is a schematic block diagram of a stereo sound processing and communication system 100 illustrating possible implementations of IVAS stereo encoding and decoding devices and methods.
[0018] The stereo sound processing and communication system 100 of Figure 1 supports the transmission of stereo sound signals over a communication link 101. The communication link 101 may comprise, for example, a wire or fiber optic link. Alternatively, the communication link 101 may comprise, at least in part, a radio frequency link. Radio frequency links often support multiple simultaneous communications requiring shared bandwidth resources, such as those found in mobile phones. Although not shown, the communication link 101 may be replaced by a storage device in a single device implementation of the system 100 that records and stores the coded stereo sound signals for later playback.
[0019] 1, for example, a pair of microphones 102 and 122 produce left channel 103 and right channel 123 of an original analog stereo sound signal. As indicated in the preceding description, the sound signal may specifically comprise, but is not limited to, speech and / or audio.
[0020] The left channel 103 and right channel 123 of the original analog sound signal are provided to an analog-to-digital (A / D) converter 104 to convert the left channel 103 and right channel 123 of the original analog sound signal into the left channel 105 and right channel 125 of the original digital stereo sound signal. The left channel 105 and right channel 125 of the original digital stereo sound signal may also be recorded and provided from a storage device (not shown).
[0021] The stereo sound encoder 106 codes the left channel 105 and the right channel 125 of the original digital stereo sound signal, thereby producing a set of coding parameters that are multiplexed in the form of a bitstream 107 that is passed to the optional error correction encoder 108. When present, the optional error correction encoder 108 adds redundancy to the binary representation of the coding parameters in the bitstream 107 before transmitting the resulting bitstream 111 over the communication link 101.
[0022] At the receiver side, an optional error correction decoder 109 utilizes the above-mentioned redundant information in the received digital bitstream 111 to detect and correct errors that may have occurred during transmission over the communication link 101, resulting in a bitstream 112 with the received coding parameters. A stereo sound decoder 110 converts the received coding parameters in the bitstream 112 to create a synthesized left channel 113 and a right channel 133 of a digital stereo sound signal. The reconstructed left channel 113 and right channel 133 of the digital stereo sound signal in the stereo sound decoder 110 are converted to a synthesized left channel 114 and a right channel 134 of an analog stereo sound signal in a digital-to-analog (D / A) converter 115.
[0023] The combined left channel 114 and right channel 134 of the analog stereo sound signal are reproduced on a pair of loudspeaker units or binaural headphones 116 and 136, respectively. Alternatively, the left channel 113 and right channel 133 of the digital stereo sound signal from the stereo sound decoder 110 may also be provided and recorded on a storage device (not shown).
[0024] For example, (a) the left channel of FIG. 1 may be implemented by the left channel of FIGS. 2-13, (b) the right channel of FIG. 1 may be implemented by the right channel of FIGS. 2-13, (c) the stereo sound encoder 106 of FIG. 1 may be implemented by the IVAS stereo encoding device of FIGS. 2-7, and (d) the stereo sound decoder 110 of FIG. 1 may be implemented by the IVAS stereo decoding device of FIGS. 8-13.
[0025] 1. Stereo Mode Switching in IVAS Stereo Encoding Device 200 and Method 250 2 is a high-level block diagram illustrating simultaneously an IVAS stereo encoding device 200 and a corresponding IVAS stereo encoding method 250; FIG. 3 is a block diagram illustrating simultaneously an FD stereo encoder 300 and a corresponding FD stereo encoding method 350 of the IVAS stereo encoding device 200 of FIG. 2; FIG. 4 is a block diagram illustrating simultaneously a TD stereo encoder 400 and a corresponding TD stereo encoding method 450 of the IVAS stereo encoding device 200 of FIG. 2; and FIG. 5 is a block diagram illustrating simultaneously an MDCT stereo encoder 500 and a corresponding MDCT stereo encoding method 550 of the IVAS stereo encoding device 200 of FIG. 2.
[0026] In the exemplary, non-limiting implementations of FIGS. 2-5, the framework of the IVAS stereo encoding device 200 (and correspondingly, the IVAS stereo decoding device 800 of FIG. 8) is based on a modified version of the Enhanced Voice Services (EVS) codec (see Non-Patent Document 1). Specifically, the EVS codec is extended to code (and decode) stereo and multi-channels and to address Immersive Voice and Audio Services (IVAS). For that reason, the encoding device 200 and method 250 are referred to in this disclosure as the IVAS stereo encoding device and method. In the described exemplary implementation, the IVAS stereo encoding device 200 and method 250 use, by way of non-limiting examples, three stereo coding modes: a frequency-domain (FD) stereo mode based on a DFT (Discrete Fourier Transform), referred to in this disclosure as the “DFT stereo mode,” a time-domain (TD) stereo mode, referred to in this disclosure as the “TD stereo mode,” and a joint stereo coding mode based on a modified discrete cosine transform (MDCT) stereo mode, referred to in this disclosure as the “MDCT stereo mode.” It should be noted that other codec structures may be used as the basis for the framework of the IVAS stereo encoding device 200 (and correspondingly, the IVAS stereo decoding device 800).
[0027] Stereo mode switching in the IVAS codec (IVAS stereo encoding device 200 and IVAS stereo decoding device 800) refers to switching between DFT stereo mode, TD stereo mode, and MDCT stereo mode in the described non-limiting implementation.
[0028] 1.1 Differences between various stereo encoders and encoding methods In this disclosure and the accompanying drawings, the following nomenclature is used: lowercase letters indicate time-domain signals, uppercase letters indicate transform-domain signals, l / L indicates left channel, r / R indicates right channel, m / M indicates middle channel, s / S indicates side channel, PCh indicates primary channel, SCh indicates secondary channel, and in the drawings, unitless numbers correspond to the number of samples at a sampling rate of 16 kHz.
[0029] Differences exist between (a) the DFT stereo encoder 300 and encoding method 350, (b) the TD stereo encoder 400 and encoding method 450, and (c) the MDCT stereo encoder 500 and encoding method 550. Some of these differences are summarized in the following paragraphs, and at least some of them are further explained in the following description.
[0030] The IVAS stereo encoding device 200 and encoding method 250 perform operations such as buffering one 20 ms frame of the stereo input signal (left and right channels) (as is known in the art, a stereo audio signal is processed in successive frames of a given time length containing a given number of audio signal samples), a few classification steps, downmixing, preprocessing, and the actual coding. An 8.75 ms look-ahead is available and is used primarily for analysis, classification, and Overlap-Add (OLA) operations used in the transform domain, such as in the Transform Coded eXcitation (TCX) core, the High Quality (HQ) core, and Frequency Domain Bandwidth Extension (FD-BWE). These operations are described in Non-Patent Document 1, Sections 5.3 and 5.2.6.2.
[0031] The look-ahead is 0.9375 ms shorter in the IVAS stereo encoding device 200 and encoding method 250 compared to the unmodified EVS encoder (corresponding to the finite impulse response (FIR) filter resampling delay (see Non-Patent Document 1, Section 5.1.3.1)). This affects the procedure for resampling the down-processed signal (down-mixed signal in TD stereo mode and DFT stereo mode) at every frame. DFT stereo encoder 300 and encoding method 350: The resampling is performed in the DFT domain and therefore does not introduce any additional delay. TD stereo encoder 400 and encoding method 450: FIR resampling (decimation) is performed using a delay of 0.9375 ms. As this resampling delay is not available in the IVAS stereo encoding device 200, it is compensated by adding zeros to the end of the downmixed signal. The compensated part of the downmixed signal, which is 0.9375 ms long, then needs to be recalculated (resampled again) in the next frame. MDCT stereo encoder 500 and encoding method 550: same as TD stereo encoder 400 and encoding method 4500. Resampling is performed from the input sampling rate (usually 16, 32, or 48 kHz) to the internal sampling rate (usually 12.8, 16, 25.6, or 32 kHz) in the DFT stereo encoder 300, the TD stereo encoder 400, and the MDCT stereo encoder 500. The resampled signals are then used in pre-processing and core encoding.
[0032] Also, the look-ahead includes a part of the down-processed signal (down-mixed signal in TD and DFT stereo modes) that is not exact but rather extrapolated or estimated, which also affects the resampling process. The inaccuracy of the down-processed look-ahead signal (down-mixed signal in TD and DFT stereo modes) depends on the current stereo coding mode. DFT stereo encoder 300 and encoding method 350: The 8.75 ms length of the look-ahead corresponds to the windowed overlap of the downmixed signal with respect to the OLA portion of the DFT analysis window, respectively the OLA portion of the DFT synthesis window. To perform pre-processing on the most useful signal possible, the look-ahead portion of the downmixed signal is rectified (or de-windowed, i.e., an inverse window is applied to the look-ahead portion). As a result, the rectified downmixed signal of the 8.75 ms length in the look-ahead is not correctly reconstructed in the current frame. TD stereo encoder 400 and encoding method 450: Before time-domain (TD) downmixing, inter-channel alignment (ICA) is performed using inter-channel time delay (ITD) synchronization between two input channels l and r in the time domain. This is achieved by delaying one of the input channels (l or r) and extrapolating the missing part of the downmixed signal corresponding to the length of the ITD delay, with the maximum value of the ITD delay being 7.5 ms. As a result, the extrapolated downmixed signal with a length of up to 7.5 ms in the look-ahead will not be correctly reconstructed in the current frame. - MDCT stereo encoder 500 and encoding method 550: The look-ahead part of the input audio signal is usually accurate, since no downmixing or time shifting is usually performed.
[0033] The corrected / extrapolated signal portion in the look-ahead section is not subjected to actual coding but is used for analysis and classification. As a result, the corrected / extrapolated signal portion in the look-ahead section is recalculated in the next frame, and the resulting down-processed signal (down-mixed signal in TD stereo mode and DFT stereo mode) is then used for actual coding. The length of the recalculated signal depends on the stereo mode and the coding process. - DFT stereo encoder 300 and encoding method 350: A signal of length 8.75 ms is recalculated at both the input stereo signal sampling rate and the internal sampling rate. TD stereo encoder 400 and encoding method 450: A signal with a length of 7.5 ms is recalculated at the input stereo signal sampling rate, while a signal with a length of 7.5+0.9375=8.4375 ms is recalculated at the internal sampling rate. MDCT stereo encoder 500 and encoding method 550: Although no normal recalculation is necessary at the input stereo signal sampling rate, the 0.9375 ms long signal undergoes recalculation at the internal sampling rate. It should be noted that the lengths of the corrected, respectively extrapolated signal portions in the look-ahead are mentioned here by way of example, but in general any other lengths can be implemented.
[0034] Additional information regarding the DFT stereo encoder 300 and encoding method 350 can be found in Non-Patent Documents 2 and 3. Additional information regarding the TD stereo encoder 400 and encoding method 450 can be found in Patent Document 1. And, additional information regarding the MDCT stereo encoder 500 and encoding method 550 can be found in Non-Patent Documents 4 and 5.
[0035] 1.2 Structure of the IVAS stereo encoding device 200 and processing in the IVAS stereo encoding method 250 Table I below lists the processing operations for each frame in sequential order according to the current stereo coding mode (see also Figures 2-5).
[0036] [Table 1]
[0037] The IVAS stereo coding method 250 includes an operation (not shown) for controlling switching between the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode. To perform the switching control operation, the IVAS stereo coding device 200 includes a controller (not shown) for switching between the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode. Switching between the DFT stereo mode and the TD stereo mode in the IVAS stereo coding device 200 and coding method 250 involves maintaining continuity of the following input signals 1) to 5) using a stereo mode switching controller (not shown) to enable proper processing of these signals in the IVAS stereo coding device 200 and method 250: 1) An input stereo signal including a left l / L channel and a right r / R channel, e.g., used for time-domain transient detection or Inter-Channel BWE (IC-BWE). 2) Stereo down-processed signal at the input stereo signal sampling rate (down-mixed signal in TD stereo mode and DFT stereo mode) DFT stereo encoder 300 and encoding method 350: middle channel m / M TD stereo encoder 400 and encoding method 450: primary channel (PCh) and secondary channel (SCh) MDCT stereo encoder 500 and encoding method 550: original (no downmix) left channel l and right channel r 3) The down-processed signal at a sampling rate of 12.8 kHz used in pre-processing (down-mixed signal in TD stereo mode and DFT stereo mode). 4) The down-processed signal at the internal sampling rate used in the core coding (the down-mixed signal in TD stereo mode and DFT stereo mode). 5) High-bandwidth (HB) input signal used in bandwidth extension (BWE)
[0038] Maintaining continuity for signal 1) above is straightforward, but for signals 2) to 5) it is difficult due to several aspects, such as different downmixing, different lengths of the recalculated parts of the look-ahead, and the use of Inter-Channel Alignment (ICA) in TD stereo mode only.
[0039] 1.2.1 Stereo Classification and Stereo Mode Selection An operation (not shown) for controlling switching between DFT, TD, and MDCT stereo modes comprises a stereo classification and stereo mode selection operation 255, as described, for example, in U.S. Patent Application Publication No. 2007 / 0129994, the entire contents of which are incorporated herein by reference. To perform operation 255, a controller (not shown) for switching between DFT, TD, and MDCT stereo modes comprises a stereo classifier and stereo mode selector 205.
[0040] Switching between the TD stereo mode, the DFT stereo mode, and the MDCT stereo mode is responsive to a stereo mode selection. The stereo classification (Patent Document 3) is performed in response to the left channel l and the right channel r of the input stereo signal and / or the requested coded bit rate. The stereo mode selection (Patent Document 3) consists of choosing one of the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode based on the stereo classification.
[0041] The stereo classifier and stereo mode selector 205 produces stereo mode signaling 270 to identify the selected stereo coding mode.
[0042] 1.2.2 Memory Allocation / Deallocation The operation (not shown) controlling the switching between the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode includes a memory allocation (not shown) operation. To perform the memory allocation operation, the controller (not shown) for switching between the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode dynamically allocates / deallocates static memory data structures to / from the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode depending on the current stereo mode. Such memory allocation keeps the static memory impact of the IVAS stereo encoding device 200 as low as possible by maintaining only data structures utilized in the current frame.
[0043] For example, in the first DFT stereo frame after a TD stereo frame, data structures related to the TD stereo mode (e.g., handling TD stereo data, second core encoder data structures) are freed (deallocated) and data structures related to the DFT stereo mode (e.g., DFT stereo data structures) are allocated and initialized instead. Note that deallocation of any additional unused data structures occurs first, followed by allocation of newly used data structures. This order of operations is important to avoid increasing the static memory impact at any point in the encoding.
[0044] A summary of the data structure of the main static memory as used in the various stereo modes is shown in Table II.
[0045] [Table 2]
[0046] An exemplary implementation of the memory allocation / deallocation encoder module in C source code is shown below: void stereo_memory_enc( CPE_ENC_HANDLE hCPE, / * i : CPE encoder structure * / const int32_t input_Fs, / * i : Input sampling rate * / const int16_t max_bwidth, / * i : Maximum audio bandwidth * / float *tdm_last_ratio / * o : TD stereo final ratio * / ) { Encoder_State *st; / *--------------------------------------------------------------* * Save the parameters from the structure being released *---------------------------------------------------------------* / if ( hCPE->last_element_mode == IVAS_CPE_TD ) { *tdm_last_ratio = hCPE->hStereoTD->tdm_last_ratio; / * NOTE: This must be set to a local variable before the data structure is allocated / deallocated * / } if ( hCPE->hStereoTCA != NULL && hCPE->last_element_mode == IVAS_CPE_DFT ) { set_s( hCPE->hStereoTCA->prevCorrLagStats, (int16_t) hCPE->hStereoDft->itd[1], 3 ); hCPE->hStereoTCA->prevRefChanIndx = ( hCPE->hStereoDft->itd[1] >= 0 ) ? ( L_CH_INDX ) : ( R_CH_INDX ); } / *--------------------------------------------------------------* * Allocate / deallocate data structures *---------------------------------------------------------------* / if ( hCPE->element_mode != hCPE->last_element_mode ) { / *-------------------------------------------------------------* * Switch CPE mode to DFT stereo *-------------------------------------------------------------* / if ( hCPE->element_mode == IVAS_CPE_DFT ) { / * Deallocate previous CPE mode data structures * / if ( hCPE->hStereoTD != NULL ) { count_free( hCPE->hStereoTD ); hCPE->hStereoTD = NULL; } if ( hCPE->hStereoMdct != NULL ) { count_free( hCPE->hStereoMdct ); hCPE->hStereoMdct = NULL; } / * Deallocate the CoreCoder secondary channel * / deallocate_CoreCoder_enc( hCPE->hCoreCoder[1] ); / * Allocate DFT stereo data structure * / stereo_dft_enc_create( &( hCPE->hStereoDft ), input_Fs, max_bwidth ); / * Allocate an ICBWE structure * / if ( hCPE->hStereoICBWE == NULL ) { hCPE->hStereoICBWE = (STEREO_ICBWE_ENC_HANDLE) count_malloc( sizeof( STEREO_ICBWE_ENC_DATA ) ); stereo_icBWE_init_enc( hCPE->hStereoICBWE ); } / * Allocate HQ cores on M channel * / st = hCPE->hCoreCoder[0]; if ( st->hHQ_core == NULL ) { st->hHQ_core = (HQ_ENC_HANDLE) count_malloc( sizeof( HQ_ENC_DATA ) ); HQ_core_enc_init( st->hHQ_core ); } } / *-------------------------------------------------------------* * Switch CPE mode to TD Stereo *-------------------------------------------------------------* / if ( hCPE->element_mode == IVAS_CPE_TD ) { / * Deallocate previous CPE mode data structures * / if ( hCPE->hStereoDft != NULL ) { stereo_dft_enc_destroy( &( hCPE->hStereoDft ) ); hCPE->hStereoDft = NULL; } if ( hCPE->hStereoMdct != NULL ) { count_free( hCPE->hStereoMdct ); hCPE->hStereoMdct = NULL; } / * Deallocated TCX / IGF structure for second channel * / deallocate_CoreCoder_TCX_enc( hCPE->hCoreCoder[1] ); / * Allocate a TD stereo data structure * / hCPE->hStereoTD = (STEREO_TD_ENC_DATA_HANDLE) count_malloc( sizeof( STEREO_TD_ENC_DATA ) ); stereo_td_init_enc( hCPE->hStereoTD, hCPE->element_brate, hCPE->last_element_mode ); / * Allocate secondary channel * / allocate_CoreCoder_enc( hCPE->hCoreCoder[1] ); } / *-------------------------------------------------------------* * Allocate DFT / TD stereo structure after MDCT stereo frame *-------------------------------------------------------------* / if ( hCPE->last_element_mode == IVAS_CPE_MDCT && ( hCPE->element_mode == IVAS_CPE_DFT || hCPE->element_mode == IVAS_CPE_TD ) ) { / * Allocate TCA data structure * / hCPE->hStereoTCA = (STEREO_TCA_ENC_HANDLE) count_malloc( sizeof( STEREO_TCA_ENC_DATA ) ); stereo_tca_init_enc( hCPE->hStereoTCA, input_Fs ); st = hCPE->hCoreCoder[0]; / * Allocate primary channel structure * / allocate_CoreCoder_enc( st ); / * Allocate a CLDFB for the primary channel * / if ( st->cldfbAnaEnc == NULL ) { openCldfb( &st->cldfbAnaEnc, CLDFB_ANALYSIS, input_Fs, CLDFB_PROTOTYPE_1_25MS ); } / * Allocate BWE for primary channel * / if ( st->hBWE_TD == NULL ) { st->hBWE_TD = (TD_BWE_ENC_HANDLE) count_malloc( sizeof( TD_BWE_ENC_DATA ) ); if ( st->cldfbSynTd == NULL ) { openCldfb( &st->cldfbSynTd, CLDFB_SYNTHESIS, 16000, CLDFB_PROTOTYPE_1_25MS ); } InitSWBencBuffer( st->hBWE_TD ); ResetSHBbuffer_Enc( st->hBWE_TD ); st->hBWE_FD = (FD_BWE_ENC_HANDLE) count_malloc( sizeof( FD_BWE_ENC_DATA ) ); fd_bwe_enc_init( st->hBWE_FD ); } } / *--------------------------------------------------------------* * Switch CPE mode to MDCT stereo *---------------------------------------------------------------* / if ( hCPE->element_mode == IVAS_CPE_MDCT ) { int16_t i; / * Deallocate previous CPE mode data structures * / if ( hCPE->hStereoDft != NULL ) { stereo_dft_enc_destroy( &( hCPE->hStereoDft ) ); hCPE->hStereoDft = NULL; } if ( hCPE->hStereoTD != NULL ) { count_free( hCPE->hStereoTD ); hCPE->hStereoTD = NULL; } if ( hCPE->hStereoTCA != NULL ) { count_free( hCPE->hStereoTCA ); hCPE->hStereoTCA = NULL; } if ( hCPE->hStereoICBWE != NULL ) { count_free( hCPE->hStereoICBWE ); hCPE->hStereoICBWE = NULL; } for ( i = 0; i < CPE_CHANNELS; i++ ) { st = hCPE->hCoreCoder[i]; / * Deallocate the core channel substructure * / deallocate_CoreCoder_enc( hCPE->hCoreCoder[i] ); } if ( hCPE->last_element_mode == IVAS_CPE_DFT ) { / * Allocate secondary channel * / allocate_CoreCoder_enc( hCPE->hCoreCoder[1] ); } / * Allocate TCX / IGF structure for second channel * / st = hCPE->hCoreCoder[1]; st->hTcxEnc = (TCX_ENC_HANDLE) count_malloc( sizeof( TCX_ENC_DATA ) ); st->hTcxEnc->spectrum[0] = st->hTcxEnc->spectrum_long; st->hTcxEnc->spectrum[1] = st->hTcxEnc->spectrum_long + N_TCX10_MAX; set_f( st->hTcxEnc->old_out, 0, L_FRAME32k ); set_f( st->hTcxEnc->spectrum_long, 0, N_MAX ); if ( hCPE->last_element_mode == IVAS_CPE_DFT ) { st->last_core = ACELP_CORE; / * Required to set up the TCX core in SetTCXModeInfo() * / } st->hTcxCfg = (TCX_CONFIG_HANDLE) count_malloc( sizeof( TCX_config ) ); st->hIGFEnc = (IGF_ENC_INSTANCE_HANDLE) count_malloc( sizeof( IGF_ENC_INSTANCE ) ); st->igf = getIgfPresent( st->element_mode, st->total_brate, st->bwidth, st->rf_mode ); / * Allocate and initialize the MDCT stereo structure * / hCPE->hStereoMdct = (STEREO_MDCT_ENC_DATA_HANDLE) count_malloc( sizeof( STEREO_MDCT_ENC_DATA ) ); initMdctStereoEncData( hCPE->hStereoMdct, hCPE->element_brate, hCPE->hCoreCoder[0]->max_bwidth, SMDCT_MS_DECISION, 0, NULL ); } } return; }
[0047] 1.2.3 Setting TD Stereo Mode The TD stereo mode can consist of two sub-modes: the so-called normal TD stereo sub-mode, in which the TD stereo mixing ratio is greater than 0 and less than 1, and the so-called LRTD stereo sub-mode, in which the TD stereo mixing ratio is either 0 or 1. LRTD is therefore an extreme case of the TD stereo mode, in which the TD downmixing does not actually mix the contents of the time-domain left channel l and right channel r to form the primary channel PCh and secondary channel SCh, but derives them directly from channels l and r.
[0048] When two sub-modes of the TD stereo mode (normal and LRTD) are available, a stereo mode switching operation (not shown) comprises a TD stereo mode setting (not shown). To perform the TD stereo mode setting, forming part of the memory allocation, a stereo mode switching controller (not shown) of the IVAS stereo encoding device 200 allocates / deallocates certain static memory data structures when switching between the normal TD stereo mode and the LRTD stereo mode. For example, the IC-BWE data structure is allocated only in frames using the normal TD stereo mode (see Table II), while some data structures (BWE and Complex Low Delay Filter Bank (CLDFB) for the secondary channel SCh) are allocated only in frames using the LRTD stereo mode (see Table II). An exemplary implementation of the memory allocation / deallocation encoder module in C source code is shown below. / * Normal TD / LRTD switching * / if ( hCPE->hStereoTD->tdm_LRTD_flag == 0 ) { Encoder_State *st; st = hCPE->hCoreCoder[1]; / * Deallocate CLDFB ana for secondary channel * / if ( st->cldfbAnaEnc != NULL ) { deleteCldfb( &st->cldfbAnaEnc ); } / * Deallocate BWE for secondary channel * / if ( st->hBWE_TD != NULL ) { if ( st->hBWE_TD != NULL ) { count_free( st->hBWE_TD ); st->hBWE_TD = NULL; } deleteCldfb( &st->cldfbSynTd ); if ( st->hBWE_FD != NULL ) { count_free( st->hBWE_FD ); st->hBWE_FD = NULL; } } / * Allocate an ICBWE structure * / if ( hCPE->hStereoICBWE == NULL ) { ( hCPE->hStereoICBWE = (STEREO_ICBWE_ENC_HANDLE) count_malloc( sizeof( STEREO_ICBWE_ENC_DATA ) ); stereo_icBWE_init_enc( hCPE->hStereoICBWE ); } } else / * tdm_LRTD_flag == 1 * / { Encoder_State *st; st = hCPE->hCoreCoder[1]; / * Deallocate the ICBWE structure * / if ( hCPE->hStereoICBWE != NULL ) { / * Copy the previous input signal to be used in BWE * / mvr2r( hCPE->hStereoICBWE->dataChan[1], hCPE->hCoreCoder[1]->old_input_signal, st->input_Fs / 50 ); count_free( hCPE->hStereoICBWE ); hCPE->hStereoICBWE = NULL; } / * Allocate CLDFB ana for secondary channel * / if ( st->cldfbAnaEnc == NULL ) { openCldfb( &st->cldfbAnaEnc, CLDFB_ANALYSIS, st->input_Fs, CLDFB_PROTOTYPE_1_25MS ); } / * Allocate BWE for secondary channel * / if ( st->hBWE_TD == NULL ) { st->hBWE_TD = (TD_BWE_ENC_HANDLE) count_malloc( sizeof( TD_BWE_ENC_DATA ) ); openCldfb( &st->cldfbSynTd, CLDFB_SYNTHESIS, 16000, CLDFB_PROTOTYPE_1_25MS ); InitSWBencBuffer( st->hBWE_TD ); ResetSHBbuffer_Enc( st->hBWE_TD ); st->hBWE_FD = (FD_BWE_ENC_HANDLE) count_malloc( sizeof( FD_BWE_ENC_DATA ) ); fd_bwe_enc_init( st->hBWE_FD ); } }
[0049] For the most part, only the regular TD stereo mode (further referred to simply as TD stereo mode for brevity) is described in detail in this disclosure, and the LRTD stereo mode is mentioned as a possible implementation.
[0050] 1.2.4 Stereo mode switching update The stereo mode switch control operation (not shown) comprises a stereo switch update operation (not shown), to perform which the stereo mode switch controller (not shown) updates long-term parameters and updates or resets the historical buffer memory.
[0051] When switching from DFT stereo mode to TD stereo mode, the stereo mode switch controller (not shown) resets the TD stereo and ICA static memory data structures. These data structures store the parameters and memory of the TD stereo analysis and weighted downmixing (401 in FIG. 4) of the ICA algorithm (201 in FIG. 2), respectively. The stereo mode switch controller (not shown) sets the TD stereo past frame mixing ratio index according to the normal TD stereo mode or the LRTD stereo mode. As a non-limiting illustrative example, In normal TD stereo mode, the previous frame mixing ratio index is set to 15, which indicates that the downmixed intermediate channel m / M is coded as the primary channel PCh, and the mixing ratio is 0.5, or In -LRTD stereo mode, the previous frame mixing ratio index is set to 31, which indicates that the left channel l is coded as the primary channel PCh.
[0052] When switching from TD stereo mode to DFT stereo mode, a stereo mode switch controller (not shown) resets the DFT stereo data structure, which stores parameters and memory for the DFT stereo processing and downmixing module (303 in Figure 3).
[0053] The stereo mode switching controller (not shown) also transfers some stereo-related parameters between data structures. For example, parameters related to the time shift and energy between channels l and r, i.e., the side gains (or ILD parameters) and ITD parameters of the DFT stereo mode, are used to update the target gains and correlation lags (ICA parameters 202) of the TD stereo mode, and vice versa. These target gains and correlation lags are further explained in the next section 1.2.5 of this disclosure.
[0054] The update / reset for the core encoder (see Figures 3 and 4) is explained later in Section 1.4 of this disclosure. An example implementation of some memory handling in the encoder is shown below. void stereo_switching_enc( CPE_ENC_HANDLE hCPE, / * i : CPE encoder structure * / float old_input_signal_pri[], / * i : old input signal of the primary channel * / const int16_t input_frame / * i : input frame length * / ) { int16_t i, n, dft_ovl, offset; float tmpF; Encoder_State **st; st = hCPE->hCoreCoder; dft_ovl = STEREO_DFT_OVL_MAX * input_frame / L_FRAME48k; / * Update DFT analysis overlap memory * / if ( hCPE->element_mode > IVAS_CPE_DFT && hCPE->input_mem[0] != NULL ) { for ( n = 0; n < CPE_CHANNELS; n++ ) { mvr2r( st[n]->input + input_frame - dft_ovl, hCPE->input_mem[n], dft_ovl ); } } / * TD / MDCT -> DFT stereo switching * / if ( hCPE->element_mode == IVAS_CPE_DFT && hCPE->last_element_mode != IVAS_CPE_DFT ) { / * input_fs, window DFT synthesis overlap memory in primary channel * / for ( i = 0; i < dft_ovl; i++ ) { hCPE->hStereoDft->output_mem_dmx[i] = old_input_signal_pri[input_frame - dft_ovl + i] * hCPE->hStereoDft->win[dft_ovl - 1 - i]; } / * Reset 48kHz BWE duplicate memory * / set_f( hCPE->hStereoDft->output_mem_dmx_32k, 0, STEREO_DFT_OVL_32k ); stereo_dft_enc_reset( hCPE->hStereoDft ); / * Update ITD parameters * / if ( hCPE->element_mode == IVAS_CPE_DFT && hCPE->last_element_mode == IVAS_CPE_TD ) { set_f( hCPE->hStereoDft->itd, hCPE->hStereoTCA->prevCorrLagStats[2], STEREO_DFT_ENC_DFT_NB ); } / * Update the side_gain[] parameter * / if ( hCPE->hStereoTCA != NULL && hCPE->last_element_mode != IVAS_CPE_MDCT ) { tmpF = usdequant( hCPE->hStereoTCA->indx_ica_gD, STEREO_TCA_GDMIN, STEREO_TCA_GDSTEP ); for ( i = 0; i < STEREO_DFT_BAND_MAX; i++ ) { hCPE->hStereoDft->side_gain[STEREO_DFT_BAND_MAX + i] = tmpF; } } / * Do not allow differential coding of DFT side parameters * / hCPE->hStereoDft->ipd_counter = STEREO_DFT_FEC_THRESHOLD; hCPE->hStereoDft->res_pred_counter = STEREO_DFT_FEC_THRESHOLD; / * Update DFT synthesis overlap memory at 12.8kHz * / for ( i = 0; i < STEREO_DFT_OVL_12k8; i++ ) { hCPE->hStereoDft->output_mem_dmx_12k8[i] = st[0]->buf_speech_enc[L_FRAME32k + L_FRAME - STEREO_DFT_OVL_12k8 + i] * hCPE->hStereoDft->win_12k8[STEREO_DFT_OVL_12k8 - 1 - i]; } / * Update DFT synthesis overlap memory at 16kHz, primary channel only * / lerp( hCPE->hStereoDft->output_mem_dmx, hCPE->hStereoDft->output_mem_dmx_16k, STEREO_DFT_OVL_16k, dft_ovl ); / * Reset DFT synthesis overlap memory at 8kHz, secondary channel * / set_f( hCPE->hStereoDft->output_mem_res_8k, 0, STEREO_DFT_OVL_8k ); hCPE->vad_flag[1] = 0;} / * DFT / MDCT -> TD stereo switching * / if ( hCPE->element_mode == IVAS_CPE_TD && hCPE->last_element_mode != IVAS_CPE_TD ) { hCPE->hStereoTD->tdm_last_ratio_idx = LRTD_STEREO_MID_IS_PRIM; hCPE->hStereoTD->tdm_last_ratio_idx_SM = LRTD_STEREO_MID_IS_PRIM; hCPE->hStereoTD->tdm_last_SM_flag = 0; hCPE->hStereoTD->tdm_last_inst_ratio_idx = LRTD_STEREO_MID_IS_PRIM; / * First frame after DFT frame and content are decorrelated or crosstalked -> primary channel is forced to the left * / if ( hCPE->hStereoClassif->lrtd_mode == 1 ) { hCPE->hStereoTD->tdm_last_ratio = ratio_tabl[LRTD_STEREO_LEFT_IS_PRIM]; hCPE->hStereoTD->tdm_last_ratio_idx = LRTD_STEREO_LEFT_IS_PRIM; if ( hCPE->hStereoTCA->instTargetGain < 0.05f && ( hCPE->vad_flag[0] || hCPE->vad_flag[1] ) ) / * But if there is no content in the L channel -> the primary channel is forced to the right * / { hCPE->hStereoTD->tdm_last_ratio = ratio_tabl[LRTD_STEREO_RIGHT_IS_PRIM]; hCPE->hStereoTD->tdm_last_ratio_idx = LRTD_STEREO_RIGHT_IS_PRIM; } } } / * DFT -> TD Stereo switching * / if ( hCPE->element_mode == IVAS_CPE_TD && hCPE->last_element_mode == IVAS_CPE_DFT ) { offset = st[0]->cldfbAnaEnc->p_filter_length - st[0]->cldfbAnaEnc->no_channels; mvr2r( old_input_signal_pri + input_frame - offset - NS2SA( input_frame * 50, L_MEM_RECALC_TBE_NS ), st[0]->cldfbAnaEnc->cldfb_state, offset ); cldfb_reset_memory( st[?]->cldfbSynTd ); st[0]->currEnergyLookAhead = 6.1e-5f;<0?00677> if ( hCPE->hStereoICBWE == NULL ) ? { offset = st[1]->cldfbAnaEnc->p_filter_length - st[1]->cldfbAnaEnc->no_channels; if ( hCPE->hStereoTD->tdm_last_ratio_idx == LRTD_STEREO_LEFT_IS_PRIM ) { It should be noted that there seems to be a small error in the original text where "st[?]->cldfbSynTd" in line 23 should probably be "st[0]->cldfbSynTd". This has been marked in the translation for clarity.v_multc( hCPE->hCoreCoder[1]->old_input_signal + input_frame - offset - NS2SA( input_frame * 50, L_MEM_RECALC_TBE_NS ), -1.0f, st[1]->cldfbAnaEnc->cldfb_state, offset ); } else { mvr2r( hCPE->hCoreCoder[1]->old_input_signal + input_frame - offset - NS2SA( input_frame * 50, L_MEM_RECALC_TBE_NS ), st[1]->cldfbAnaEnc->cldfb_state, offset ); } cldfb_reset_memory( st[1]->cldfbSynTd ); st[1]->currEnergyLookAhead = 6.1e-5f; } st[1]->last_extl = -1; / * No secondary channel in previous frame -> memory reset * / set_zero( st[1]->old_inp_12k8, L_INP_MEM ); / *set_zero( st[1]->old_inp_16k, L_INP_MEM );* / set_zero( st[1]->mem_decim, 2 * L_FILT_MAX ); / *set_zero( st[1]->mem_decim16k, 2*L_FILT_MAX );* / st[1]->mem_preemph = 0; / *st[1]->mem_preemph16k = 0;* / set_zero( st[1]->buf_speech_enc, L_PAST_MAX_32k + L_FRAME32k + L_NEXT_MAX_32k ); set_zero( st[1]->buf_speech_enc_pe, L_PAST_MAX_32k + L_FRAME32k + L_NEXT_MAX_32k ); if ( st[1]->hTcxEnc != NULL ) { set_zero( st[1]->hTcxEnc->buf_speech_ltp, L_PAST_MAX_32k + L_FRAME32k + L_NEXT_MAX_32k ); } set_zero( st[1]->buf_wspeech_enc, L_FRAME16k + L_SUBFR + L_FRAME16k + L_NEXT_MAX_16k ); set_zero( st[1]->buf_synth, OLD_SYNTH_SIZE_ENC + L_FRAME32k ); st[1]->mem_wsp = 0.0f; st[1]->mem_wsp_enc = 0.0f; init_gp_clip( st[1]->clip_var ); set_f( st[1]->Bin_E, 0, L_FFT ); set_f( st[1]->Bin_E_old, 0, L_FFT / 2 ); / * st[1]->hLPDmem reset is already done in handle allocation * / st[1]->last_L_frame = st[0]->last_L_frame; pitch_ol_init( &st[1]->old_thres, &st[1]->old_pitch, &st[1]->delta_pit, &st[1]->old_corr ); set_zero( st[1]->old_wsp, L_WSP_MEM ); set_zero( st[1]->old_wsp2, ( L_WSP_MEM - L_INTERPOL ) / OPL_DECIM ); set_zero( st[1]->mem_decim2, 3 ); st[1]->Nb_ACELP_frames = 0; / * Fill SCh with PCh memory * / mvr2r( st[0]->hLPDmem->old_exc, st[1]->hLPDmem->old_exc, L_EXC_MEM ); mvr2r( st[0]->lsf_old, st[1]->lsf_old, M ); mvr2r( st[0]->lsp_old, st[1]->lsp_old, M ); mvr2r( st[0]->lsf_old1, st[1]->lsf_old1, M ); mvr2r( st[0]->lsp_old1, st[1]->lsp_old1, M ); st[1]->GSC_noisy_speech = 0; } else if ( hCPE->element_mode == IVAS_CPE_TD && hCPE->last_element_mode == IVAS_CPE_MDCT ) { set_f( st[0]->hLPDmem->old_exc, 0.0f, L_EXC_MEM ); set_f( st[1]->hLPDmem->old_exc, 0.0f, L_EXC_MEM ); }
[0055] 1.2.5 ICA Encoder In the TD stereo frame, the stereo mode switch control operation (not shown) comprises a temporal Inter-Channel Alignment (ICA) operation 251. To perform operation 251, the stereo mode switch controller (not shown) comprises an ICA encoder 201 for time-aligning channels l and r of the input stereo signal and scaling channel r.
[0056] As explained in the previous section, before TD downmixing, ICA is performed using ITD synchronization between the two input channels l and r in the time domain. This is achieved by delaying one of the input channels (l or r) and extrapolating the missing portion of the downmixed signal corresponding to the length of the ITD delay, with a maximum value of 7.5 ms. Time alignment, i.e., ICA time shift, is applied first, shifting the majority portion of the current TD stereo frame. The extrapolated portion of the look-ahead downmixed signal is recalculated and therefore temporally adjusted in the next frame based on the ITD estimated in the next frame.
[0057] When a stereo mode switch is not expected, a 7.5 ms long extrapolated signal is recalculated in the ICA encoder 201. However, when a stereo mode switch, i.e., a switch from DFT stereo mode to TD stereo mode, is possible, a longer signal is recalculated, whose length then corresponds to the length of the DFT stereo rectified signal plus the FIR resampling delay, i.e., 8.75 ms + 0.9375 ms = 9.6875 ms. Section 1.4 explains these features in more detail.
[0058] Another objective of the ICA encoder 201 is scaling of input channel r. The scaling gain, i.e., the target gain mentioned above, is estimated for each frame as the logarithmic ratio of the energy of the l channel smoothed using the target gain of the previous frame to the energy of the r channel, regardless of the DFT or TD stereo mode used. The target gain estimated for the current frame (20 ms) is applied to the last 15 ms of the current input channel r, while the first 5 ms of the current channel r is scaled by a combination of the target gain of the previous frame and the target gain of the current frame in a fade-in / fade-out manner.
[0059] The ICA encoder 201 produces ICA parameters 202 such as ITD delay, target gain, and target channel index.
[0060] 1.2.6 Time Domain Transient Detector A stereo mode switch control operation (not shown) comprises an operation 253 of detecting a time domain transient in channel l from the ICA encoder 201. To perform operation 253, the stereo mode switch controller (not shown) comprises a detector 203 for detecting a time domain transient in channel l.
[0061] In the same manner, a stereo mode switch control operation (not shown) comprises an operation 254 of detecting a time domain transient in channel r from the ICA encoder 201. To perform operation 254, the stereo mode switch controller (not shown) comprises a detector 204 for detecting a time domain transient in channel r.
[0062] The detection of time domain transients in the time domain channels l and r is a pre-processing step that enables the detection of such transients in the transform domain core coding modules (TCX Core, HQ Core, FD-BWE) and therefore their proper processing and coding.
[0063] Further information regarding the time domain transient detectors 203 and 204 and the time domain transient detection operations 253 and 254 can be found, for example, in Non-Patent Document 1, Section 5.1.8.
[0064] 1.2.7 Stereo Encoder Configuration To implement the stereo encoder configuration, the IVAS stereo encoding device 200 sets the parameters of the stereo encoders 300, 400, and 500. For example, the nominal bit rate for the core encoder is set.
[0065] 1.2.8 DFT analysis, stereo processing and downmixing in the DFT domain, and IDFT synthesis Referring to Figure 3, the DFT stereo encoding method 350 comprises an operation 351 for applying the DFT transform to channel l from the time-domain transient detector 203 of Figure 2. To perform operation 351, the DFT stereo encoder 300 comprises a calculator 301 of a DFT transform (DFT analysis) of channel l to produce channel L in the DFT domain.
[0066] The DFT stereo encoding method 350 also includes an operation 352 for applying a DFT transform to channel r from the time-domain transient detector 204 of Figure 2. To perform operation 352, the DFT stereo encoder 300 includes a calculator 302 of a DFT transform (DFT analysis) of channel r to produce channel R in the DFT domain.
[0067] The DFT stereo encoding method 350 further comprises a stereo processing and downmixing operation 353 in the DFT domain. To perform operation 353, the DFT stereo encoder 300 comprises a stereo processor and downmixer 303 for producing side information on a side channel S. The downmixing of channels L and R also produces a residual signal on the side channel S. The side information and residual signal from the side channel S are coded, for example using a coding operation 354 and a corresponding encoder 304, and then multiplexed in an output bitstream 310 of the DFT stereo encoder 300. The stereo processor and downmixer 303 also downmixes the left channel L and the right channel R from the DFT calculators 301 and 302 to produce a middle channel M in the DFT domain. More information regarding the stereo processing and downmixing operation 353, the stereo processor and downmixer 303, the middle channel M, and the side information and residual signal from the side channel S can be found, for example, in Non-Patent Document 3.
[0068] In an inverse DFT (IDFT) synthesis operation 355 of the DFT stereo encoding method 350, the calculator 305 of the DFT stereo encoder 300 calculates the IDFT transform m of the intermediate channel M at the sampling rate of the input stereo signal, e.g., 12.8 kHz. In the same manner, in an inverse DFT (IDFT) synthesis operation 356 of the DFT stereo encoding method 350, the calculator 306 of the DFT stereo encoder 300 calculates the IDFT transform m of the channel M at the internal sampling rate.
[0069] 1.2.9 TD Analysis and Downmixing in the TD Domain Referring to FIG. 4, the TD stereo encoding method 450 comprises an operation 451 of time-domain analysis and weighted downmixing in the TD domain. To perform operation 451, the TD stereo encoder 400 comprises a time-domain analyzer and downmixer 401 for calculating stereo side parameters 402, such as submode flags, mixing ratio indices, or linear prediction reuse flags, which are multiplexed in the output bitstream 410 of the TD stereo encoder 400. The time-domain analyzer and downmixer 401 also performs weighted downmixing of channels l and r from detectors 203 and 204 (FIG. 2) to produce a primary channel PCh and a secondary channel SCh using the estimated mixing ratios, consistent with ICA scaling. More information regarding the time-domain analyzer and downmixer 401 and operation 451 can be found, for example, in U.S. Pat. No. 6,449,399.
[0070] Downmixing using the current frame mixing ratio is performed, for example, on the last 15 ms of the current frame of input channels l and r, while the first 5 ms of the current frame is downmixed using a combination of the previous frame's mixing ratio and the current frame's mixing ratio in a fade-in / fade-out manner to smooth the transition from one channel to the other.Two channels (primary channel PCh and secondary channel SCh) sampled at a stereo input channel sampling rate, for example 32 kHz, are resampled using FIR decimation filters to their representation at 12.8 kHz and at the internal sampling rate.
[0071] In TD stereo mode, not only the stereo input signal of the current frame is downmixed, but also the stored downmixed signal corresponding to the previous frame is downmixed again. The length of the previous signal subject to this recalculation corresponds to the length of the time-shifted signal recalculated in the ICA module, i.e., 8.75 ms + 0.9375 ms = 9.6875 ms.
[0072] 1.2.10 Initial preprocessing In the IVAS codec (IVAS stereo encoding device 200 and IVAS stereo decoding device 800), conventional preprocessing is constrained such that some classification decisions are made relative to the overall codec bitrate, while other decisions are made according to the core coding bitrate. As a result, conventional preprocessing, such as that used in the EVS codec (Non-Patent Document 1), is divided into two parts to ensure that the best possible codec configuration is used for each processed frame. Thus, although the codec configuration may change from frame to frame, some configuration changes, such as changes based on signal activity or signal class, can be made as quickly as possible. On the other hand, some codec configuration changes, such as selection of the coded audio bandwidth, selection of the internal sampling rate, or bit budget allocation between low-band and high-band coding, should not occur too frequently. Such excessively frequent codec configuration changes may lead to instability in the coded signal quality or even audible artifacts.
[0073] The first part of the preprocessing, i.e., initial preprocessing, may include preprocessing and classification modules such as resampling at the preprocessing sampling rate, spectral analysis, bandwidth detection (BWD), sound activity detection (SAD), linear prediction (LP) analysis, open-loop pitch search, signal classification, speech / music classification, etc. It should be noted that the decisions in the initial preprocessing depend only on the overall codec bitrate. More information on the operations performed during the above-described preprocessing can be found, for example, in Non-Patent Document 1.
[0074] In DFT stereo mode (DFT stereo encoder 300 of FIG. 3), initial pre-processing is performed on intermediate channel m in the time domain at the internal sampling rate from IDFT calculator 306 by initial pre-processor 307 and corresponding initial pre-processing operation 357 .
[0075] In TD stereo mode, initial pre-processing is performed by (a) an initial pre-processor 403 and corresponding initial pre-processing operations 453 on the primary channels PCh from the time domain analyzer and downmixer 401, and (b) an initial pre-processor 404 and corresponding initial pre-processing operations 454 on the secondary channels SCh from the time domain analyzer and downmixer 401.
[0076] In MDCT stereo mode, initial pre-processing is performed (a) on the left channel l of the input from time-domain transient detector 203 (FIG. 2) by initial pre-processor 503 and corresponding initial pre-processing operation 553, and (b) on the right channel r of the input from time-domain transient detector 204 (FIG. 2) by initial pre-processor 504 and corresponding initial pre-processing operation 554.
[0077] 1.2.11 Core Encoder Configuration The core encoder configuration is based on the overall codec bitrate and initial preprocessing.
[0078] Specifically, in the DFT stereo encoder 300 and corresponding DFT stereo encoding method 350 (FIG. 3), the core encoder configurator 308 and corresponding core encoder configuration operation 358 configure the core encoder 311 and corresponding core encoding operation 361 in response to the intermediate channel m in the time domain from the IDFT calculator 305 and the output from the initial pre-processor 307. The core encoder configurator 308 is responsible for, for example, setting the internal sampling rate and / or modifying the core encoder type classification. More information on core encoder configuration in the DFT domain can be found, for example, in Non-Patent Documents 1 and 2.
[0079] In the TD stereo encoder 400 and corresponding TD stereo encoding method 450 (FIG. 4), a core encoder configurator 405 and corresponding core encoder configuration operation 455 configure a core encoder 406 and corresponding core encoding operation 456 for the primary channel PCh and a core encoder 407 and corresponding core encoding operation 457 for the secondary channel SCh in response to the initial pre-processed primary channel PCh and secondary channel SCh from the initial pre-processors 403 and 404, respectively. The core encoder configurator 405 is responsible for, for example, setting the internal sampling rate and / or modifying the core encoder type classification. Further information regarding core encoder configuration in the TD domain can be found, for example, in U.S. Pat. No. 6,239,949 and U.S. Pat. No. 6,239,949.
[0080] 1.2.12 Additional Preprocessing The DFT encoding method 350 comprises an additional pre-processing operation 362. To perform operation 362, a so-called additional pre-processor 312 of the DFT stereo encoder 300 performs a second part of the pre-processing, which may include classification, core selection, pre-processing at the encoding internal sampling rate, etc. The decisions in the initial pre-processor 307 depend on the core encoding bitrate, which usually varies during the session. More information on the operations performed during such additional pre-processing in the DFT domain can be found, for example, in [1].
[0081] The TD encoding method 450 comprises an additional pre-processing operation 458. To perform operation 458, a so-called additional pre-processor 408 of the TD stereo encoder 400 performs a second part of pre-processing before the core encoding of the primary channel PCh, which may include classification, core selection, pre-processing at the encoding internal sampling rate, etc. The decision in the additional pre-processor 408 depends on the core encoding bitrate, which usually varies during the session.
[0082] The TD encoding method 450 also comprises an additional pre-processing operation 459. To perform operation 459, the TD stereo encoder 400 comprises a so-called additional pre-processor 409 to perform a second part of pre-processing before the core encoding of the secondary channels SCh, which may include classification, core selection, pre-processing on the encoding internal sampling rate, etc. The decisions in the additional pre-processor 409 depend on the core encoding bitrate, which usually varies during the session.
[0083] Further information on such additional pre-processing in the TD domain can be found, for example, in [1].
[0084] The MDCT encoding method 550 comprises an operation 555 of additional pre-processing of the left channel l. To perform operation 555, a so-called additional pre-processor 505 of the MDCT stereo encoder 500 performs a second part of pre-processing of the left channel l, which may include classification, core selection, pre-processing at the encoding internal sampling rate, etc., before an operation 556 of joint core encoding of the left channel l and the right channel r, which is performed by a joint core encoder 506 of the MDCT stereo encoder 500.
[0085] The MDCT encoding method 550 comprises an operation 557 of additional pre-processing of the right channel r. To perform operation 557, a so-called additional pre-processor 507 of the MDCT stereo encoder 500 performs a second part of pre-processing of the left channel l, which may include classification, core selection, pre-processing at the encoding internal sampling rate, etc., before the operation 556 of joint core encoding of the left channel l and the right channel r, which is performed by the joint core encoder 506 of the MDCT stereo encoder 500.
[0086] Further information regarding such additional pre-processing in the MDCT domain can be found, for example, in [1].
[0087] 1.2.13 Core Encoding In general, the core encoder 311 (performing the core encoding operation 361) in the DFT stereo encoder 300 and the core encoders 406 (performing the core encoding operation 456) and 407 (performing the core encoding operation 457) in the TD stereo encoder 400 can be any variable bitrate mono codec. In an exemplary implementation of the present disclosure, an EVS codec (see Non-Patent Document 1) with variable bitrate capabilities (see Patent Document 2) is used. Of course, other suitable codecs may be considered and implemented in some cases. The MDCT stereo encoder 500 utilizes a joint core encoder 506, which may generally be a stereo coding module with stereophonic tools that jointly process and quantize the l and r channels.
[0088] 1.2.14 Common Stereo Update Finally, a common stereo update is performed. More information on the common stereo update can be found, for example, in [1].
[0089] 1.2.15 Bitstream 2 and 3, the stereo mode signaling 270 from the stereo classifier and stereo mode selector 205, the bitstream 313 from the side information, the residual signal detector 304, and the bitstream 314 from the core encoder 311 are multiplexed to form the DFT stereo encoder bitstream 310 (and thus the output bitstream 206 of the IVAS stereo encoding device 200 (FIG. 2)).
[0090] 2 and 4, the stereo mode signaling 270 from the stereo classifier and stereo mode selector 205, the side parameters 402 from the time domain analyzer and downmixer 401, the ICA parameters 202 from the ICA encoder 201, the bitstream 411 from the core encoder 406, and the bitstream 412 from the core encoder 407 are multiplexed to form the TD stereo encoder bitstream 410 (and thus form the output bitstream 206 of the IVAS stereo encoding device 200 (FIG. 2)).
[0091] 2 and 5, the stereo mode signaling 270 from the stereo classifier and stereo mode selector 205 and the bitstream 509 from the joint core encoder 506 are multiplexed to form the MDCT stereo encoder bitstream 508 (and thus the output bitstream 206 of the IVAS stereo encoding device 200 (FIG. 2)).
[0092] 1.3 Switching from TD Stereo Mode to DFT Stereo Mode in IVAS Stereo Encoding Device 200 Switching from TD stereo mode (TD stereo encoder 400) to DFT stereo mode (DFT stereo encoder 300) is relatively simple, as shown in FIG.
[0093] Specifically, Figure 6 is a flowchart illustrating the processing operations in the IVAS stereo encoding device 200 and method 250 when switching from TD stereo mode to DFT stereo mode. As can be seen, Figure 5 illustrates two frames of a stereo input signal, namely, a TD stereo frame 601 and a subsequent DFT stereo frame 602, along with various processing operations and associated time instances when switching from TD stereo mode to DFT stereo mode.
[0094] A sufficiently long look-ahead is possible, the resampling is done in the DFT domain (so there is no FIR decimation filter memory handling), and there is a transition from two core encoders 406 and 407 in the last TD stereo frame 501 to one core encoder 311 in the first DFT stereo frame 502.
[0095] The following operations performed upon switching from TD stereo mode (TD stereo encoder 400) to DFT stereo mode (DFT stereo encoder 300) are performed by the stereo mode switching controller (not shown) mentioned above in response to the stereo mode selection.
[0096] Instance A) in Figure 6 refers to updating the DFT analysis memory, specifically the DFT stereo OLA analysis memory as part of the DFT stereo data structure that undergoes windowing before the DFT computation operations 351 and 352. This update is performed by the stereo mode switch controller (not shown) (see 251 in Figure 2) before the Inter-Channel Alignment (ICA) and comprises storing samples for the last 8.75 ms of the current TD stereo frame 601 of channels l and r of the input stereo signal. This update is performed for every single TD stereo frame in both channels l and r. More information regarding the DFT analysis memory can be found, for example, in Non-Patent Documents 1 and 2.
[0097] Instance B) in Figure 6 refers to the update of the DFT synthesis memory when switching from TD stereo mode to DFT stereo mode, specifically the update of the OLA synthesis memory as part of the DFT stereo data structure resulting from windowing after the IDFT calculation operations 355 and 356. The stereo mode switch controller (not shown) performs this update in the first DFT stereo frame 602 after the TD stereo frame 601, and for this update uses the TD stereo memory used for TD stereo processing corresponding to the downmixed primary channel PCh as part of the TD stereo data structure. More information on the DFT synthesis memory can be found, for example, in Non-Patent Documents 1 and 2, and more information on the TD stereo memory can be found, for example, in Patent Document 1.
[0098] Starting with the first DFT stereo frame 602, some TD stereo related data structures, e.g., the TD stereo data structures of the core encoder 407 (as used in the TD stereo encoder 400) and data structures relating to the secondary channel SCh, are no longer needed and are therefore deallocated, i.e., released by the stereo mode switching controller (not shown).
[0099] In the DFT stereo frame 602 following the TD stereo frame 601, the stereo mode switching controller (not shown) continues the core encoding operation 361 in the core encoder 311 of the DFT stereo encoder 300 using the memory of the primary PCh channel core encoder 406 in the preceding TD stereo frame 601 (e.g., synthesis memory, pre-emphasis memory, past signals and parameters, etc.), while controlling the difference in time instances between the TD stereo mode and the DFT stereo mode to ensure continuity of some core encoder buffers, such as the pre-emphasized input signal buffer, the HB input buffer, etc., which will be used later in the low-band encoder and the FD-BWE high-band encoder, respectively. More information regarding the core encoding operation 361, the memory of the PCh channel core encoder 406, the pre-emphasized input signal buffer, the HB input buffer, etc. can be found, for example, in Non-Patent Document 1.
[0100] 1.4 Switching from DFT Stereo Mode to TD Stereo Mode in IVAS Stereo Encoding Device 200 Switching from DFT stereo mode to TD stereo mode is more complicated than switching from TD stereo mode to DFT stereo mode due to the more complex structure of the TD stereo encoder 400. The subsequent operations performed when switching from DFT stereo mode (DFT stereo encoder 300) to TD stereo mode (TD stereo encoder 400) are performed by a stereo mode switching controller (not shown) in response to the stereo mode selection.
[0101] 7a is a flowchart illustrating the processing operations in the IVAS stereo encoding device 200 and method 250 when switching from DFT stereo mode to TD stereo mode. Specifically, FIG. 7a illustrates two frames of a stereo input signal, namely a DFT stereo frame 701 and a subsequent TD stereo frame 702, together with associated time instances, in different processing operations when switching from DFT stereo mode to TD stereo mode.
[0102] Instance A) of Figure 7a refers to the update of the FIR resampling filter memory (as utilized in the FIR resampling from the input stereo signal sampling rate to the 12.8 kHz sampling rate and the inner core encoder sampling rate) used in the primary channel PCh of the TD stereo coding mode. The stereo mode switching controller (not shown) performs this update in every DFT stereo frame using the downmixed intermediate channel m, corresponding to the 2 x 0.9375 ms long interval 703 (see 704) before the last 7.5 ms long interval in the DFT stereo frame 701, thereby ensuring the continuity of the FIR resampling memory for the primary channel PCh.
[0103] Since the side channel s (FIG. 3) of the DFT stereo encoding method 350 is not available, but is used at, for example, a 12.8 kHz sampling rate, the sampling rate of the input stereo signal, and at the internal sampling rate, the stereo mode switching controller (not shown) fills the FIR resampling filter memory of the downmixed secondary channel SCh differently. To reconstruct the entire length of the downmixed signal at the internal sampling rate for the core encoder 407, an 8.75 ms interval (see 705) of the downmixed signal of the previous frame is recalculated in the TD stereo frame 702. Thus, the update of the FIR resampling filter memory of the downmixed secondary channel SCh corresponds to the 2×0.9375 ms interval 708 of the downmixed intermediate channel m before the last 8.75 ms interval (see 705). This occurs in the first TD stereo frame 702 after switching from the previous DFT stereo frame 701. The updating of the FIR resampling filter memory of the secondary channel SCh is indicated by instance C) in Fig. 7a. As can be seen, the stereo mode switching controller (not shown) recalculates the length of the downmixed signal in the secondary channel SCh (see 706) in TD stereo frames to be longer than the recalculated length of the downmixed signal in the primary channel PCh (see 707).
[0104] Instance B) of Figure 7a relates to updating (recalculating) the primary channel PCh and the secondary channel SCh in the first TD stereo frame 702 after the DFT stereo frame 701. The operations of instance B) as performed by a stereo mode switching controller (not shown) are shown in more detail in Figure 7b. As mentioned in the previous description, Figure 7b is a flow chart illustrating the processing operations upon switching from DFT stereo mode to TD stereo mode.
[0105] Referring to FIG. 7b, in operation 710, the stereo mode switch controller (not shown) recalculates the ICA memory used in the ICA analysis and calculation (see operation 251 in FIG. 2 ) and later used as input signals for the pre-processing and core encoder (see operations 453-454 and 456-459) of channels l and r of length 9.6875 ms (as discussed in sections 1.2.7-1.2.9 of this disclosure) corresponding to the previous DFT stereo frame 701.
[0106] Thus, in operations 712 and 713, the stereo mode switching controller (not shown) recalculates the primary channel PCh and secondary channel SCh of the DFT stereo frame 701 by downmixing the ICA processed channels l and r using the stereo mixing ratio of that frame 701.
[0107] For the secondary channel SCh, the length of the past interval to be recalculated by the stereo mode switch controller (not shown) in operation 712 (see 714) is 9.6875 ms, but when there is no stereo coding mode switch, an interval of only 7.5 ms length (see 715) is recalculated. For the primary channel PCh (see operation 713), the length of the interval to be recalculated by the stereo mode switch controller (not shown) using the TD stereo mixing ratio of the past frame 701 is always 7.5 ms (see 715). This ensures the continuity of the primary channel PCh and the secondary channel SCh.
[0108] When switching from the intermediate channel m of the DFT stereo frame 701 to the primary channel PCh of the TD stereo frame 702, a continuous downmixed signal is utilized. To that end, a stereo mode switching controller (not shown) crossfades (717) a 7.5 ms long section (see 715) of the DFT intermediate channel m with the recalculated primary channel PCh (713) of the DFT stereo frame 701 in order to smooth the transition between the DFT and TD stereo modes and equalize the different downmix signal energies. The reconstruction of the secondary channel SCh in operation 712 uses the mixing ratio of frame 701, but since the secondary channel SCh from the DFT stereo frame 701 is not available, no further smoothing is applied.
[0109] The core coding in the first TD stereo frame 702 after the DFT stereo frame 701 then continues with resampling the downmixed signals using an FIR filter, pre-emphasizing these signals, calculating the HB signal, etc. More information on these operations can be found, for example, in [1].
[0110] For a pre-emphasis filter implemented as a first-order high-pass filter used to emphasize the higher frequencies of the input signal (see Non-Patent Document 1, Section 5.1.4), the stereo mode switching controller (not shown) stores two pre-emphasis filter memory values for each DFT stereo frame. These memory values correspond to time instances based on the different recalculation lengths for DFT stereo mode and TD stereo mode. This mechanism ensures optimal recalculation of the pre-emphasis signal in channel m, the primary channel PCh, with the shortest signal length, respectively. For the secondary channel SCh in TD stereo mode, the pre-emphasis filter memory is set to 0 before the first TD stereo frame is processed.
[0111] Starting with the first TD stereo frame 702 after the DFT stereo frame 701, some DFT stereo-related data structures (e.g., the DFT stereo data structures mentioned above) are not needed, so they are deallocated / freed by the stereo mode switching controller (not shown). Meanwhile, a second instance of the core encoder data structure is allocated and initialized for the core encoding of the secondary channel SCh (operation 457). Most of the data structures of the secondary channel SCh core encoder are reset, but some of them are estimated for a smoother switching transition. For example, the previous excitation buffer (the adaptive codebook of the ACELP core), previous LSF parameters, and LSP parameters (see Non-Patent Document 1) of the secondary channel SCh are filled with their counterparts in the primary channel PCh. The resetting or estimation of the previous buffers of the secondary channel SCh may be the cause of some artifacts. Although many of these artifacts are largely suppressed in the smoothing-based processing at the decoder, a small number of them may remain as the cause of subjective artifacts.
[0112] 1.5 Switching from TD Stereo Mode to MDCT Stereo Mode in IVAS Stereo Encoding Device 200 Switching from TD stereo mode to MDCT stereo mode is relatively simple, since both of these stereo modes handle two input channels and utilize two instances of the core encoder. The main hurdle is maintaining the correct phase of the input left and right channels.
[0113] To maintain the correct phase of the left and right channels of the input stereo sound signal, a stereo mode switching controller (not shown) changes the TD stereo downmixing. In the last TD stereo frame before the first MDCT stereo frame, the TD stereo mixing ratio is set to β=1.0, and out-of-phase downmixing of the left and right channels of the stereo sound signal is performed using, for example, the following equation for TD stereo downmixing: PCh(i)=r(i)·(1-β)+l(i)·β SCh(i)=l(i)·(1-β)+r(i)·β where PCh(i) is the TD primary channel, SCh(i) is the TD secondary channel, l(i) is the left channel, r(i) is the right channel, β is the TD stereo mixing ratio, and i is the discrete-time index.
[0114] And this means that the TD stereo primary channel PCh(i) is the left channel l of the MDCT stereo past. past (i), and the TD stereo secondary channel SCh(i) is the same as the previous right channel r past (i), where i is the discrete time index. For completeness, note that the stereo mode switching controller (not shown) may use default TD stereo downmixing in the last TD stereo frame, for example using the following equation: PCh(i)=r(i)·(1-β)+l(i)·β SCh(i)=l(i)·(1-β)-r(i)·β
[0115] Next, in normal (without stereo mode switching) MDCT stereo processing, the initial pre-processing (initial pre-processors 503 and 504 and initial pre-processing operations 553 and 554) do not recalculate the look-ahead of the left channel l and the right channel r of the stereo sound signal except for the last 0.9375 ms long interval. However, in practice, the 7.5+0.9375 ms long look-ahead is recalculated at the internal sampling rate (12.8 kHz in this non-limiting exemplary implementation). Therefore, no special handling is required to maintain the continuity of the input signal at the input sampling rate.
[0116] And in normal (without stereo mode switching) MDCT stereo processing, the additional pre-processing (additional pre-processors 505 and 507 and additional pre-processing operations 555 and 557) do not recalculate the look-ahead of the left channel l and the right channel r of the stereo sound signal except for the last 0.9375 ms long interval. In contrast to the initial pre-processing, the input signal (left channel l and right channel r of the stereo sound signal) at the internal sampling rate (12.8 kHz in this non-limiting exemplary implementation) for only 0.9375 ms long is recalculated in the additional pre-processing.
[0117] In other words:
[0118] The MDCT stereo encoder 500 comprises (a) initial pre-processors 503 and 504 that, in a second MDCT stereo mode, recalculate a first time length look-ahead for the left channel l and the right channel r of the stereo sound signal at the internal sampling rate, and (b) an additional pre-processor that, in the second MDCT stereo mode, recalculate a final interval of a given time length look-ahead for the left channel l and the right channel r of the stereo sound signal at the internal sampling rate, wherein the first and second time lengths are different.
[0119] The MDCT stereo coding operation 550 comprises, in a second MDCT stereo mode, (a) recalculating a first time length look-ahead for the left channel l and the right channel r of the stereo sound signal at the internal sampling rate, and (b) recalculating a last interval of a given time length look-ahead for the left channel l and the right channel r of the stereo sound signal at the internal sampling rate, wherein the first and second time lengths are different.
[0120] 1.6 Switching from MDCT Stereo Mode to TD Stereo Mode in IVAS Stereo Encoding Device 200 Similar to switching from TD stereo mode to MDCT stereo mode, two input channels are always available, and two core encoder instances are always utilized in this scenario. The main obstacle is again maintaining the correct phase of the input left and right channels. Therefore, in the first TD stereo frame after the last MDCT stereo frame, the stereo mode switch controller (not shown) sets the TD stereo mixing ratio to β=1.0 and modifies the TD stereo downmixing by using an anti-phase mixing scheme similar to that described in Section 1.5.
[0121] Another detail about switching from MDCT stereo mode to TD stereo mode is that the stereo mode switching controller (not shown) appropriately reconstructs in the first TD frame the past interval of the input channels of the stereo sound signal at the internal sampling rate. Thus, the look-ahead portion corresponding to 8.75-7.5=1.25 ms is reconstructed (resampling and pre-emphasis) in the first TD stereo frame.
[0122] 1.7 Switching from DFT Stereo Mode to MDCT Stereo Mode in IVAS Stereo Encoding Device 200 A similar mechanism for switching from DFT stereo mode to TD stereo mode as described above is used in this scenario, where the primary channel PCh and secondary channel SCh of the TD stereo mode are replaced by the left channel l and right channel r of the MDCT stereo mode.
[0123] 1.8 Switching from MDCT Stereo Mode to DFT Stereo Mode in IVAS Stereo Encoding Device 200 A similar mechanism for switching from TD stereo mode to DFT stereo mode as described above is used in this scenario, where the primary channel PCh and secondary channel SCh of the TD stereo mode are replaced by the left channel l and right channel r of the MDCT stereo mode.
[0124] 2. Stereo Mode Switching in IVAS Stereo Decoding Device 800 and Method 850 8 is a high-level block diagram illustrating an IVAS stereo decoding device 800 and a corresponding decoding method 850, which includes a DFT stereo decoder 801 and a corresponding DFT stereo decoding method 851, a TD stereo decoder 802 and a corresponding TD stereo decoding method 852, and an MDCT stereo decoder 803 and a corresponding MDCT stereo decoding method 853. For simplicity, only the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode are shown and described. However, implementations using other types of stereo modes are within the scope of this disclosure.
[0125] The IVAS stereo decoding device 800 and corresponding decoding method 850 receive a bitstream 830 transmitted from the IVAS stereo encoding device 200. In general, the IVAS stereo decoding device 800 and corresponding decoding method 850 decode successive frames of the coded stereo signal from the bitstream 830, e.g., frames of 20 ms length as in the case of the EVS codec, and perform upmixing of the decoded frames, finally producing a stereo output signal including channels l and r.
[0126] 2.1 Differences between various stereo decoders and decoding methods The core decoding, performed at the internal sampling rate, is essentially the same regardless of the actual stereo mode. However, the core decoding is performed once for the DFT stereo frame (middle channel) and twice for the TD stereo frame (primary channel PCh and secondary channel SCh) or MDCT stereo frame (left channel l and right channel r). The challenge is to maintain (update) the memory of the secondary channel SCh of the TD stereo frame when switching from DFT stereo frame to MDCT stereo frame, respectively.
[0127] Moreover, the further decoding operations after the core decoding strongly depend on the actual stereo mode, which results in complex switching between stereo modes. The most fundamental differences are:
[0128] DFT stereo decoder 801 and decoding method 851: Resampling of the decoded core synthesis from the internal sampling rate to the output stereo signal sampling rate is performed in the DFT domain using DFT analysis and a synthesis overlap window length of 3.125 ms. - The post-filtering (in the ACELP frame) conditioning of the low-band (LB) bus is done in the DFT domain. -Core switching (ACELP core <-> TCX / HQ core) is done in the DFT domain with an available delay of 3.125ms. - The synchronization of LB and HB synthesis (in the ACELP frame) does not require any additional delay. - Stereo upmixing is done in the DFT domain with an available delay of 3.125ms. -Time synchronization is applied with a duration of 0.125ms to match the overall decoder delay (which is 3.25ms).
[0129] TD stereo decoder 802 and decoding method 852: (More information on TD stereo decoders can be found, for example, in US Pat. No. 6,259,999.) Resampling of the decoded core synthesis from the internal sampling rate to the output stereo signal sampling rate is performed using a CLDFB filter with a delay of 1.25 ms. The post-filtering (in ACELP frames) adjustment of the -LB bus is done in the CLDFB domain. -Core switching (ACELP core <-> TCX / HQ core) is done in the time domain with an available delay of 1.25ms. - The synchronization of LB and HB synthesis (in ACELP frames) introduces additional delay. - Stereo upmixing is done in the TD domain without delay. -Time synchronization is applied with a duration of 2.0 ms to match the overall decoder delay.
[0130] MDCT stereo decoder 803 and decoding method 853: Since only the -TCX-based core decoder is utilized, only a delay adjustment of 1.25 ms is used to synchronize the core synthesis signal between different cores. The adjustment (in ACELP frames) after filtering of the -LB bus is skipped. - Core switching (ACELP core <-> TCX / HQ core) is done in the time domain only at the first MDCT stereo frame after a TD stereo frame or DFT stereo frame with an available delay of 1.25 ms. -The synchronization of LB synthesis with HB synthesis is unrelated. -Stereo upmixing is skipped. -Time synchronization is applied with a duration of 2.0 ms to match the overall decoder delay.
[0131] The various operations during decoding, mainly DFT domain processing "vs" TD domain processing, and the different delay schemes between DFT and TD stereo modes are carefully considered in the procedure described herein below for switching between DFT and TD stereo modes.
[0132] 2.2 Processing in the IVAS Stereo Decoding Device 800 and Decoding Method 850 Table III below lists the processing operations in the IVAS stereo decoding device 800 for each frame in sequential order depending on the current DFT, TD, or MDCT stereo mode (see also FIG. 8).
[0133] [Table 3]
[0134] The IVAS stereo decoding method 850 includes an operation (not shown) for controlling switching between the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode. To perform the switching control operation, the IVAS stereo decoding device 800 includes a controller (not shown) for switching between the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode. Switching between the DFT stereo mode, the TD stereo mode, and the MDCT stereo mode in the IVAS stereo decoding device 800 and the decoding method 850 involves using a stereo mode switching controller (not shown) to maintain continuity of several decoder signals and memories 1) through 6) below to enable proper processing of these signals and use of said memories in the IVAS stereo decoding device 800 and the method 850. 1) The core post-filter downmixed signal at the internal sampling rate and memory used in the core decoding. - DFT stereo decoder 801: middle channel m. - TD stereo decoder 802: primary channel PCh and secondary channel SCh. MDCT stereo decoder 803: left channel l and right channel r (not downmixed). 2) TCX-LTP (Transform Coded eXcitation - Long Term Prediction) postfilter memory. The TCX-LTP postfilter is used to interpolate past synthesis samples using a polyphase FIR interpolation filter (see Non-Patent Document 1, Section 6.9.2). 3) DFT OLA analysis memory at the internal sampling rate and output stereo signal sampling rate as used in the OLA portion of the windowing in the previous and current frames before the DFT operation 854. 4) DFT OLA synthesis memory as used in the OLA part of the windowing in the previous and current frames after IDFT operations 855 and 856 at the output stereo signal sampling rate. 5) Output stereo signal containing channels l and r. 6) HB signal memory (see Non-Patent Document 1, Section 6.1.5), channels l and r, used in BWE and IC-BWE.
[0135] While maintaining continuity for one channel (m in DFT stereo mode, PCh in TD stereo mode, or l in MDCT stereo mode) is relatively straightforward in item 1), it is difficult for the secondary channel SCh in item 1) above and for the signal / memory in items 2)-6) above due to several aspects, such as the complete lack of past signal and memory for the secondary channel SCh, different downmixing between DFT and TD stereo modes, different default delays, etc. The shorter decoder delay (3.25 ms) compared to the encoder delay (8.75 ms) further complicates the decoding process.
[0136] 2.2.1 Reading Stereo Mode and Audio Bandwidth Information The IVAS stereo decoding method 850 begins (not shown) by reading stereo mode and audio bandwidth information from the transmitted bitstream 830. Based on the currently read stereo mode, the associated decoding operations for each particular stereo mode are performed (see Table III), while memories and buffers for other stereo modes are maintained.
[0137] 2.2.2 Memory Allocation Similar to the IVAS stereo encoding device 200, in memory allocation operations (not shown), a stereo mode switch controller (not shown) dynamically allocates / deallocates data structures (static memory) depending on the current stereo mode. The stereo mode switch controller (not shown) keeps the impact on the static memory of the codec as low as possible by maintaining only the portion of static memory used in the current frame. See Table II for an overview of the data structures allocated in specific stereo modes.
[0138] In addition, the LRTD stereo sub-mode flag is read by the stereo mode switch controller (not shown) to distinguish between normal TD stereo mode and LRTD stereo mode. Based on the sub-mode flag, the stereo mode switch controller (not shown) allocates / deallocates the relevant data structures in the TD stereo mode as shown in Table II.
[0139] 2.2.3 Stereo mode switching update Similar to the IVAS stereo encoding device 200, a stereo mode switching controller (not shown) handles the memory when switching from one of the DFT, TD, and MDCT stereo modes to another: it maintains updated long-term parameters and updates or resets the past buffer memory.
[0140] Upon receiving the first DFT stereo frame after the TD stereo frame or MDCT stereo frame, the stereo mode switch controller (not shown) performs an operation to reset the DFT stereo data structure (already defined with respect to the DFT stereo encoder 300). Upon receiving the first TD stereo frame after the DFT stereo frame or MDCT stereo frame, the stereo mode switch controller performs an operation to reset the TD stereo data structure (already described with respect to the TD stereo decoder 400). Finally, upon receiving the first MDCT stereo frame after the DFT stereo frame or TD stereo frame, the stereo mode switch controller (not shown) performs an operation to reset the MDCT stereo data structure. Again, when switching from one of the DFT stereo mode and the TD stereo mode to the other, the stereo mode switch controller (not shown) performs an operation to transfer some stereo-related parameters between data structures as described with respect to the IVAS stereo encoding device 200 (see section 1.2.4 above).
[0141] The update / reset for the secondary channel SCh of the core decoding is explained in section 2.4.
[0142] Further information about the operation of the stereo decoder configuration, core decoder configuration, TD stereo decoder configuration, core decoding, core switching in the DFT domain, and core switching in the TD domain in Table III can be found, for example, in Non-Patent Document 1 and Non-Patent Document 2.
[0143] 2.2.4 DFT Stereo Mode Duplicate Memory Update A stereo mode switching controller (not shown) maintains or updates the DFT OLA memory at each TD or MDCT stereo frame (see "Update DFT Stereo Mode Overlap Memory," "Update MDCT Stereo TCX Overlap Buffer," and "Reset / Update DFT Stereo Overlap Memory" in Table III). In this way, the updated DFT OLA memory is available for the next DFT stereo frame. The actual maintenance / update mechanism and associated memory buffers are described later in Section 2.3 of this disclosure. An exemplary implementation in C source code of the DFT stereo OLA memory update performed at TD or MDCT stereo frames is given below. if ( st[n]->element_mode != IVAS_CPE_DFT ) { ivas_post_proc( ... ); / * Update the OLA buffer - needed to switch to DFT stereo * / stereo_td2dft_update( hCPE, n, output[n], synth[n], hb_synth[n], output_frame ); / * Update ovl buffer for possible switch from TD stereo SCh ACELP frames to MDCT stereo TCX frames * / if ( st[n]->element_mode == IVAS_CPE_TD && n == 1 && st[n]->hTcxDec == NULL ) { mvr2r( output[n] + st[n]->L_frame / 2, hCPE->hStereoTD->TCX_old_syn_Overl, st[n]->L_frame / 2 ); } } void stereo_td2dft_update( CPE_DEC_HANDLE hCPE, / * i / o: CPE decoder structure * / const int16_t n, / * i : channel number * / float output[], / * i / o: synthesis at internal frequency * / float synth[], / * i / o: synthesis at output frequency * / float hb_synth[], / * i / o: hb synthesis * / const int16_t output_frame / * i : frame length * / ) { int16_t ovl, ovl_TCX, dft32ms_ovl, hq_delay_comp; Decoder_State **st; / * Initialization * / st = hCPE->hCoreCoder; ovl = NS2SA( st[n]->L_frame * 50, STEREO_DFT32MS_OVL_NS ); dft32ms_ovl = ( STEREO_DFT32MS_OVL_MAX * st[0]->output_Fs ) / 48000; hq_delay_comp = NS2SA( st[0]->output_Fs, DELAY_CLDFB_NS ); if ( hCPE->element_mode >= IVAS_CPE_DFT && hCPE->element_mode != IVAS_CPE_MDCT ) { if ( st[n]->core == ACELP_CORE ) { if ( n == 0 ) { / * Update DFT analysis overlap memory in internal_fs:core synthesis * / mvr2r( output + st[n]->L_frame - ovl, hCPE->input_mem_LB[n], ovl ); / * Update DFT analysis overlap memory in internal_fs: BPF * / if ( st[n]->p_bpf_noise_buf ) { mvr2r( st[n]->p_bpf_noise_buf + st[n]->L_frame - ovl, hCPE->input_mem_BPF[n], ovl ); } / * Update DFT analysis overlap memory in output_fs: BWE * / if ( st[n]->extl != -1 || ( st[n]->bws_cnt > 0 && st[n]->core == ACELP_CORE ) ) { mvr2r( hb_synth + output_frame - dft32ms_ovl, hCPE->input_mem[n], dft32ms_ovl ); } } else { / * Update DFT analysis overlap memory in internal_fs: core synthesis, secondary channels * / mvr2r( output + st[n]->L_frame - ovl, hCPE->input_mem_LB[n], ovl ); } } else / * TCX core * / { / * LB-TCX synthesis * / mvr2r( output + st[n]->L_frame - ovl, hCPE->input_mem_LB[n], ovl ); / * BPF * / if ( n == 0 && st[n]->p_bpf_noise_buf ) { mvr2r( st[n]->p_bpf_noise_buf + st[n]->L_frame - ovl, hCPE->input_mem_BPF[n], ovl ); } / * TCX synthesis (already delayed in TD stereo in core_switching_post_dec()) * / if ( st[n]->hTcxDec != NULL ) { ovl_TCX = NS2SA( st[n]->hTcxDec->L_frameTCX * 50, STEREO_DFT32MS_OVL_NS ); mvr2r( synth + st[n]->hTcxDec->L_frameTCX + hq_delay_comp - ovl_TCX, hCPE->input_mem[n], ovl_TCX - hq_delay_comp ); mvr2r( st[n]->delay_buf_out, hCPE->input_mem[n] + ovl_TCX - hq_delay_comp, hq_delay_comp ); } } } else if ( hCPE->element_mode == IVAS_CPE_MDCT && hCPE->input_mem[0] != NULL ) { / * Reset DFT stereo OLA memory * / set_zero( hCPE->input_mem[n], NS2SA( st[0]->output_Fs, STEREO_DFT32MS_OVL_NS ) ); set_zero( hCPE->input_mem_LB[n], STEREO_DFT32MS_OVL_16k ); if ( n == 0 ) { set_zero( hCPE->input_mem_BPF[n], STEREO_DFT32MS_OVL_16k ); } } return; }
[0144] 2.2.5 DFT Stereo Decoder 801 and Decoding Method 851 The DFT decoding method 851 comprises an operation 857 of core-decoding the intermediate channel m. To perform operation 857, the core decoder 807 decodes the intermediate channel m in the time domain in response to the received bitstream 830. The core decoder 807 in the DFT stereo decoder 801 (which performs the core decoding operation 857) can be any variable bitrate mono codec. In an exemplary implementation of the present disclosure, the EVS codec (see Non-Patent Document 1) with variable bitrate capabilities (see Patent Document 2) is used. Of course, other suitable codecs may be conceived and implemented in some cases.
[0145] In the DFT calculation operation 854 of the DFT decoding method 851 (DFT analysis of Table III), the calculator 804 calculates the DFT of the intermediate channel m to recover the intermediate channel M in the DFT domain.
[0146] The DFT decoding method 851 also comprises an operation 858 of decoding the stereo side information and the residual signal S (Residual Decoding in Table III). To perform operation 858, the decoder 808 recovers the stereo side information and the residual signal S in response to the bitstream 830.
[0147] In a DFT stereo decoding (DFT stereo decoding in Table III) and upmixing (upmixing in the DFT domain in Table III) operation 859, the DFT stereo decoder and upmixer 809 produces channels L and R in the DFT domain in response to the intermediate channel M and the side information and residual signal S. In general, the DFT stereo decoding and upmixing operation 859 is the inverse of the DFT stereo processing and downmixing operation 353 of FIG. 3.
[0148] In IDFT computation operation 855 (DFT synthesis in Table III), calculator 805 computes the IDFT of channel L to reconstruct channel l in the time domain. Similarly, in IDFT computation operation 856 (DFT synthesis in Table III), calculator 806 computes the IDFT of channel R to reconstruct channel r in the time domain.
[0149] 2.2.6 TD Stereo Decoder 802 and Decoding Method 852 The TD decoding method 852 comprises an operation of core decoding the primary channel PCh 860. To perform operation 860, the core decoder 810 decodes the primary channel PCh in response to the received bitstream 830.
[0150] The TD decoding method 852 also comprises an operation of core decoding the secondary channel SCh 861. To perform operation 861, the core decoder 811 decodes the secondary channel SCh in response to the received bitstream 830.
[0151] Again, the core decoder 810 (performing the core decoding operation 860 in the TD stereo decoder 802) and the core decoder 811 (performing the core decoding operation 861 in the TD stereo decoder 802) can be any variable bitrate mono codec. In an exemplary implementation of the present disclosure, the EVS codec (see Non-Patent Document 1) with variable bitrate capabilities (see Patent Document 2) is used. Of course, other suitable codecs may be conceived and implemented in some cases.
[0152] In a time domain (TD) upmixing operation 862 (upmixing in the TD domain in Table III), the upmixer 812 receives and upmixes the primary channels PCh and secondary channels SCh to recover the time domain channels l and r of the stereo signal based on the TD stereo mixing coefficients.
[0153] 2.2.7 MDCT Stereo Decoder 803 and Decoding Method 853 The MDCT decoding method 853 comprises an operation 863 (joint stereo decoding in Table III) of joint core decoding the left channel l and the right channel r. To perform operation 863, the joint core decoder 813 decodes the left channel l and the right channel r in response to the received bitstream 830. Note that in the MDCT stereo mode, no upmixing operation is performed and no upmixer is utilized.
[0154] 2.2.8 Synchronization To perform the stereo synthesis time synchronization (synthesis synchronization in Table III) and stereo switching operation 864, the stereo mode switch controller (not shown) includes a time synchronizer and stereo switch 814 to receive channels l and r from the DFT stereo decoder 801, the TD stereo decoder 802, or the MDCT stereo decoder 803 and synchronize the upmixed output stereo channels l and r. The time synchronizer and stereo switch 814 delays the upmixed output stereo channels l and r to match the overall codec delay value and handles the transition between the DFT, TD, and MDCT stereo output channels.
[0155] By default, in DFT stereo mode, the time synchronizer and stereo switch 814 introduces a 3.125 ms delay in the DFT stereo decoder 801. A delay synchronization of 0.125 ms is applied by the time synchronizer and stereo switch 814 to match the overall 32 ms codec delay (20 ms frame length, 8.75 ms encoder delay, 3.25 ms decoder delay). In TD or MDCT stereo modes, the time synchronizer and stereo switch 814 applies a delay consisting of a 1.25 ms resampling delay and a 2 ms delay used for synchronization between the LB and HB synthesis to match the overall 32 ms codec delay.
[0156] After time synchronization and stereo switching (synthesis time synchronization and stereo switching operation 864 and time synchronizer and stereo switch 814 in FIG. 8) are performed, the HB synthesis (from BWE or IC-BWE) is added to the core synthesis (IC-BWE, add HB synthesis in Table III; see also BWE or IC-BWE calculation operation 865 and BWE or IC-BWE calculator 815 in FIG. 8), and ICA decoding (ICA decoder - time adjustment in Table III, which desynchronizes the two output channels l and r) is performed (see temporal ICA operation 866 and corresponding ICA decoder 816) before the final stereo synthesis of channels l and r is output from the IVAS stereo decoding device 800. These operations 865 and 866 are skipped in MDCT stereo mode.
[0157] Finally, a common stereo update is performed as shown in Table III.
[0158] 2.3 Switching from TD stereo mode to DFT stereo mode in IVAS stereo decoding device Further information regarding the elements, operations, and signals referred to in sections 2.3 and 2.4 can be found, for example, in non-patent documents 1 and 2.
[0159] The mechanism for switching from TD stereo mode to DFT stereo mode in the IVAS stereo decoding device 800 is complicated by the fact that several compound steps between these two stereo modes are fundamentally different, including the transition from two core decoders 810 and 811 in the last TD stereo frame to one core decoder 807 in the first DFT stereo frame (see section 2.1 above for details).
[0160] 9 is a flowchart illustrating the processing operations in the IVAS stereo decoding device 800 and method 850 when switching from TD stereo mode to DFT stereo mode. Specifically, FIG. 9 shows two frames of a decoded stereo signal at different processing operations, together with the associated time instances, when switching from a TD stereo frame 901 to a DFT stereo frame 902.
[0161] First, core decoders 810 and 811 of the TD stereo decoder 802 are used for both the primary channel PCh and the secondary channel SCh, each outputting a corresponding decoded core synthesis at the internal sampling rate. In a TD stereo frame 901, the decoded core synthesis from the two core decoders 810 and 811 is used to update the DFT stereo OLA memory buffers (one memory buffer per channel, i.e., two OLA memory buffers in total; see DFT OLA analysis and synthesis memory described above). These OLA memory buffers are updated in every TD stereo frame to be up to date in case the next frame is a DFT stereo frame.
[0162] Instance A of FIG. 9 refers to an operation (not shown) using a stereo mode switch controller (not shown) to update the DFT stereo analysis memories (which are used in the OLA part of the windowing in the previous and current frames before the DFT calculation operation 854) at the internal sampling rate input_mem_LB[] upon receiving the first DFT stereo frame 902 after the TD stereo frame 901. For that purpose, the number L of last samples 903 of the TD stereo synthesis at the internal sampling rate of the primary channel PCh and the secondary channel SCh in the TD stereo frame 901 is ovl are used by the stereo mode switching controller (not shown) to update the DFT stereo analysis memories of the DFT stereo mid channel m and side channel s, respectively. ovlFor example, at an internal sampling rate of 12.8 kHz, ovl = 40 samples corresponds to a 3.125 ms long overlap of the DFT synthesis window 905.
[0163] Similarly, the stereo mode switching controller (not shown) controls the last L of the bass post-filter (BPF) error signal of the TD primary channel PCh. ovl samples (see Non-Patent Document 1, Section 6.1.4.2) are used to update the DFT stereo BPF analysis memory input_mem_BPF[] of intermediate channel m at the internal sampling rate (which is used in the OLA part of windowing in the previous and current frames before the DFT calculation operation 854). Additionally, the DFT stereo fullband (FB) analysis memory input_mem[] of intermediate channel m at the output stereo signal sampling rate (which is used in the OLA part of windowing in the previous and current frames before the DFT calculation operation 854) is updated using the last sample of 3.125 ms of the HB synthesis (ACELP core) of the TD stereo PCh, respectively the PCh TCX synthesis. Since the DFT stereo BPF and FB analysis memories are not utilized for the side information channel s, these memories are not updated using the secondary channel SCh core synthesis.
[0164] Next, in the TD stereo frame 901, the decoded ACELP core synthesis (primary channel PCh and secondary channel SCh) at the internal sampling rate is resampled using CLDFB domain filtering, which introduces a 1.25 ms delay. For the TCX / HQ core frame, a compensation delay of 1.25 ms is used to synchronize the core synthesis between different cores. A TCX-LTP postfilter is then applied to both core channels PCh and SCh.
[0165] In a next operation, the primary channel PCh and secondary channel SCh of the TD stereo synthesis at the output stereo signal sampling rate from the TD stereo frame 901 undergo TD stereo upmixing (combining the primary channel PCh and secondary channel SCh) using a TD stereo mixing ratio in the TD upmixer 812 (see Patent Document 1), resulting in upmixed stereo channels l and r in the time domain. Since the upmixing operation 862 is performed in the time domain, it does not introduce an upmixing delay.
[0166] The upmixed left channel l and right channel r of the TD stereo frame 901 from the upmixer 812 of the TD stereo decoder 802 are then used in an operation (not shown) to update the DFT stereo synthesis memory (they are used in the OLA portion of the windowing in the previous and current frames after the IDFT calculation operation 855). Again, this update is performed for every TD stereo frame by the stereo mode switch controller (not shown), in case the next frame is a DFT stereo frame. Instance B) of Figure 9 illustrates that the number of available last samples of the TD stereo left channel l and right channel r synthesis is insufficient to be used for a simple update of the DFT stereo synthesis memory. Therefore, the 3.125 ms long DFT stereo synthesis memory is reconstructed in two intervals using an approximation. The first interval corresponds to the (3.125-1.25) ms long signal that is available (it is the upmixed synthesis at the output stereo signal sampling rate), and the second interval corresponds to the remaining 1.25 ms long signal that is not available due to the core decoder resampling delay.
[0167] Specifically, the DFT stereo synthesis memory is updated by the stereo mode switching controller (not shown) using the following partial operations as shown in Figure 10. Figure 10 is a flow chart illustrating instance B) of Figure 9, which comprises updating the DFT stereo synthesis memory in TD stereo frames at the decoder side.
[0168] (a) The two channels l and r of the DFT stereo analysis memory input_mem_LB[] at the internal sampling rate, as reconstructed earlier during the decoding method 850 (which are identical to the core synthesis at the internal sampling rate), undergo further processing depending on the actual decoding core. -ACELP core: the final L of the LB core synthesis of the primary channel PCh and secondary channel SCh at the internal sampling rate ovl The samples 1001 are resampled to the output stereo signal sampling rate using simple linear interpolation with a delay of 0 (see 1003). -TCX / HQ core: the final L of the LB core synthesis of the primary channel PCh and secondary channel SCh at the internal sampling rate ovl The samples 1001 are similarly resampled to the output stereo signal sampling rate using simple linear interpolation with a delay of 0 (see 1003). However, the TCX synthesis memory (the last 1.25 ms of the TCX synthesis from the previous frame) is then used to update the last 1.25 ms of the resampled core synthesis.
[0169] (b) The linearly resampled LB signals corresponding to 3.125 ms long portions of the primary channel PCh and secondary channel SCh of the TD stereo frame 901 are upmixed (see 1003) to form the left channel l and the right channel r using a common TD stereo upmixing routine, while using the TD stereo mixing ratio from the current frame (see TD upmixing operation 862). The resulting signal is further referred to as the "reconstructed composition" 1002.
[0170] (c) The reconstruction of the first part of the DFT stereo synthesis memory (3.125 to 1.25 ms) depends on the actual decoding core. -ACELP core: Crossfading 1004 between the CLDFB-based resampled, TD upmixed synthesis 1005 at the output stereo signal sampling rate and the reconstructed synthesis 1002 (from the previous partial operation (b)) is performed for both channels l and r during the first (3.125-1.25) ms long portion of the channels of the TD stereo frame 901. - TCX / HQ core: The first (3.125-1.25) ms long part of the DFT stereo synthesis memory is updated using the upmixed synthesis 1005.
[0171] (d) The last 1.25 ms long portion of the DFT stereo synthesis memory is filled with the last portion of the reconstructed synthesis 1002.
[0172] (e) The DFT synthesis window (904 in FIG. 9) is applied to the DFT OLA synthesis memory (defined herein above) only in the first DFT stereo frame 902 (when a switch from TD stereo mode to DFT stereo mode occurs). Note that the last 1.25 ms portion of the DFT OLA synthesis memory is of limited importance because the DFT synthesis window shape 904 converges to 0, and therefore it masks the approximated samples of the reconstructed synthesis 1002 that would result from resampling based on simple linear interpolation.
[0173] Finally, the upmixed reconstructed composition 1002 of the TD stereo frame 901 is aligned to match the overall codec delay, i.e. delayed by 2 ms in the time synchronizer and stereo switch 814 . When there is a switch from TD stereo frames to DFT stereo frames, other DFT stereo memories (other than the duplicated memories), ie parameters and buffers of past frames of the DFT stereo decoder, are reset by the stereo mode switch controller (not shown). DFT stereo decoding (see 859), upmixing (see 859) and DFT synthesis (see 855 and 856) are then performed and the stereo output synthesis (channels l and r) is aligned to match the overall codec delay, i.e. delayed by 0.125 ms in the time synchronizer and stereo switch 814.
[0174] FIG. 11 is a flow chart illustrating instance C) of FIG. 9, which comprises smoothing the output stereo synthesis at the decoder side in the first DFT stereo frame 902 after the stereo mode switch.
[0175] 11, once the DFT stereo synthesis is aligned and synchronized with respect to the overall codec delay in the first DFT stereo frame 902, the stereo mode switch controller (not shown) performs a cross-fading operation 1151 between the aligned and synchronized TD stereo synthesis 1101 (from operation 864) and the aligned and synchronized DFT stereo synthesis 1102 (from operation 864) to smooth the switching transition. The cross-fading is performed at the beginning of both output channels l and r, in a 1.875 ms long interval 1103 starting after a 0.125 ms delay 1104 (all signals are at the output stereo signal sampling rate). This instance corresponds to instance C) of FIG. 9.
[0176] Decoding then continues with updating the IC-BWE calculator 815, the ICA decoder 816, and the common stereo decoder, regardless of the current stereo mode.
[0177] 2.4 Switching from DFT to TD stereo mode in IVAS stereo decoding device The fundamentally different decoding operations between DFT stereo mode and TD stereo mode, and the presence of two core decoders 810 and 811 in the TD stereo decoder 802, make it difficult to switch from DFT stereo mode to TD stereo mode in the IVAS stereo decoding device 800. Figure 12 is a flowchart showing the processing operations in the IVAS stereo decoding devices 800 and 850 when switching from DFT stereo mode to TD stereo mode. Specifically, Figure 12 shows two frames of a decoded stereo signal at different processing operations, together with the associated time instances, when switching from a DFT stereo frame 1201 to a TD stereo frame 1202.
[0178] The core decoding may use the same processing regardless of the actual stereo mode, with two exceptions.
[0179] First exception: For DFT stereo frames, the resampling from the internal sampling rate to the output stereo signal sampling rate is performed in the DFT domain, but the CLDFB resampling is done in parallel to maintain / update the CLDFB analysis and synthesis memory in case the next frame is a TD stereo frame.
[0180] Second exception: The BPF (Bass Post Filter) (low-frequency pitch enhancement procedure, see [1], section 6.1.4.2) is then applied in the DFT domain in the DFT stereo frame, but the BPF analysis and calculation of the error signal is done in the time domain, independent of the stereo mode.
[0181] Otherwise, all internal states and memories of the core decoder are simply continuous and well maintained when switching from the DFT intermediate channel m to the TD primary channel PCh.
[0182] In DFT stereo frame 1201, decoding then continues with stereo decoding and upmixing (859) of channels M and S into channels L and R in the DFT domain, including core decoding of intermediate channel m (857), computing the DFT transform of intermediate channel m in the time domain to obtain intermediate channel M in the DFT domain (854), and decoding the residual signal (858). The DFT domain analysis and synthesis results in an OLA delay of 3.125 ms. The synthesis transition is then handled in time synchronizer and stereo switch 814.
[0183] When switching from the DFT stereo frame 1201 to the TD stereo frame 1202, the fact that there is only one core decoder 807 in the DFT stereo decoder 801 complicates core decoding of the TD secondary channel SCh because the internal state and memory of the second core decoder 811 in the TD stereo decoder 802 are not continuously maintained (conversely, the internal state and memory of the first core decoder 810 are continuously maintained using the internal state and memory of the core decoder 807 in the DFT stereo decoder 801). Therefore, the memory of the second core decoder 811 is usually reset at the stereo mode switch update (see Table III) by the stereo mode switch controller (not shown). However, there are a few exceptions where the primary channel SCh memory is filled with the memory of several PCh buffers, e.g., the previous excitation, previous LSF parameters, and previous LSP parameters. In any case, initial synthesis of the first TD secondary channel SCh frame after switching from the DFT stereo frame 1201 to the TD stereo frame 1202 results in an incomplete reconstruction. Thus, while the synthesis from the first core decoder 810 is well and smoothly decoded during stereo mode switching, the limited quality synthesis from the second core decoder 811 introduces discontinuities during stereo upmixing and final synthesis (862). These discontinuities are suppressed by utilizing a DFT stereo OLA memory during reconstruction of the initial TD stereo output synthesis, as will be explained later.
[0184] A stereo mode switching controller (not shown) suppresses possible discontinuities and differences between the DFT and TD stereo upmixed channels by simple equalization of the signal energy. The ICA target gain g ICA If is less than 1.0, then the channel l after upmixing (862) and before time synchronization (864), i.e., y L(i) is modified in the first TD stereo frame 1202 after the stereo mode switch using the following relationship:
[0185]
number
[0186] L eq corresponds to an interval of 8.75 ms in the IVAS stereo decoding device 800 (for example, L at an output stereo signal sampling rate of 16 kHz). eq = 140 samples). The value of the gain factor α is then obtained using the following relation:
[0187]
number
[0188] 12, instance A) relates to the missing portion 1203 of the TD stereo upmixed synchronized synthesis (from operation 864) of TD stereo frame 1202 that corresponds to the previous DFT stereo upmixed synchronized synthesis memory from DFT stereo frame 1201. This memory, which is (3.25-1.25) ms long, is not available when switching from DFT stereo frame 1201 to TD stereo frame 1202, except for the first 0.125 ms long section 1204.
[0189] FIG. 13 is a flow chart illustrating instance A) of FIG. 12, which comprises updating the TD stereo upmixed synchronous synthesis memory at the first TD stereo frame after switching from DFT stereo mode to TD stereo mode at the decoder side.
[0190] Referring to both Figures 12 and 13, a stereo mode switching controller (not shown) reconstructs 3.25 ms of TD stereo upmixed synchronized synthesis (1205) using the following operations (a) through (e) for both the left channel l and the right channel r:
[0191] (a) The DFT stereo OLA synthesis memory (defined herein above) is rectified (i.e., an inverse synthesis window is applied to the OLA synthesis memory, see 1301).
[0192] (b) The first 0.125 ms portion 1302 (see 1204 in Figure 12) of the TD stereo upmixed synchronized synthesis 1303 is identical to the previous DFT stereo upmixed synchronized synthesis memory 1304 (the last 0.125 ms long section of the previous frame's DFT stereo upmixed synchronized synthesis memory) and is therefore reused to form this first portion of the TD stereo upmixed synchronized synthesis 1303.
[0193] (c) The second part of the TD stereo upmixed synchronized synthesis 1303 (see 1203 in FIG. 12) having a length of (3.125-1.25) ms is approximated using the rectified DFT stereo OLA synthesis memory 1301.
[0194] (d) The portion of the TD stereo upmixed synchronized composition 1303 with a length of 2 ms from the previous two steps (b) and (c) is then filled into the output stereo composition in the first TD stereo frame 1202.
[0195] (e) Smoothing of the transition between the previous DFT stereo OLA synthesis memory 1301 and the TD synchronized upmixed synthesis 1305 from operation 864 of the current TD stereo frame 1202 is performed at the beginning of the synchronized upmixed TD stereo synthesis 1305. The transition interval is 1.25 ms long (see 1306) and is obtained using crossfading 1307 between the rectified DFT stereo OLA synthesis memory 1301 and the synchronized upmixed TD stereo synthesis 1305.
[0196] 2.5 Switching from TD stereo mode to MDCT stereo mode in IVAS stereo decoding device Switching from TD stereo mode to MDCT stereo mode is relatively simple, as both of these stereo modes handle two transport channels and utilize two instances of the core decoder.
[0197] Since the anti-phase downmixing scheme was used in the TD stereo encoder 400, the stereo mode switch controller (not shown) similarly modifies the upmixing of the TD stereo channels to maintain the correct phase of the left and right channels of the stereo sound signal in the last TD stereo frame before the first MDCT stereo frame. Specifically, the stereo mode switch controller (not shown) sets the mixing ratio β=1.0 and performs anti-phase upmixing of the TD stereo primary channel PCh(i) and the TD stereo secondary channel SCh(i) (the inverse of the anti-phase downmixing used in the TD stereo encoder 400) to upmix the past left channel l of the MDCT stereo signal. past (i) and the past right channel r of the MDCT stereo past As a result, the TD stereo primary channel PCh(i) is calculated by subtracting the left channel l from the past channel l of the MDCT stereo. past (i), and the TD stereo secondary channel SCh(i) signal is the same as the previous right channel r past Same as (i).
[0198] 2.6 Switching from MDCT stereo mode to TD stereo mode in IVAS stereo decoding device Similar to switching from TD stereo mode to MDCT stereo mode, two transport channels are available and two instances of the core decoder are utilized in this scenario. To maintain the correct phase of the left and right channels of the stereo sound signal, the TD stereo mixing ratio is set to 1.0 and the anti-phase upmixing scheme is used again by the stereo mode switching controller (not shown) in the first TD stereo frame after the last MDCT stereo frame.
[0199] 2.7 Switching from DFT to MDCT Stereo Mode in IVAS Stereo Decoding Device A similar mechanism for decoder-side switching from DFT stereo mode to TD stereo mode is used in this scenario, where the primary and secondary channels PCh and SCh of the TD stereo mode are replaced by the left and right channels l and r of the MDCT stereo mode.
[0200] 2.8 Switching from MDCT to DFT Stereo Mode in IVAS Stereo Decoding Device A similar mechanism to the decoder-side switching from TD stereo mode to DFT stereo mode is used in this scenario, where the primary channel PCh and secondary channel SCh of the TD stereo mode are replaced by the left channel l and right channel r of the MDCT stereo mode.
[0201] Finally, decoding continues, independent of the current stereo mode, with IC-BWE decoding 865 (skipped in MDCT stereo mode), addition of HB synthesis (skipped in MDCT stereo mode), temporal ICA alignment 866 (skipped in MDCT stereo mode), and a common stereo decoder update.
[0202] 2.9 Hardware Implementation FIG. 14 is a simplified block diagram of an exemplary configuration of hardware components forming each of the IVAS stereo encoding device 200 and IVAS stereo decoding device 800 described above.
[0203] Each of the IVAS stereo encoding device 200 and the IVAS stereo decoding device 800 may be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. Each of the IVAS stereo encoding device 200 and the IVAS stereo decoding device 800 (identified as 1400 in FIG. 14) includes an input 1402, an output 1404, a processor 1406, and a memory 1408.
[0204] The input 1402 is configured to receive the left channel l and the right channel r of an input stereo sound signal in digital or analog form in the case of the IVAS stereo encoding device 200, or to receive the bitstream 803 in the case of the IVAS stereo decoding device 800. The output 1404 is configured to provide the multiplexed bitstream 206 in the case of the IVAS stereo encoding device 200, or to provide the decoded left channel l and right channel r in the case of the IVAS stereo decoding device 800. The input 1402 and the output 1404 may be implemented in a common module, for example a serial input / output device.
[0205] The processor 1406 is operatively connected to the input 1402, the output 1404, and the memory 1408. The processor 1406 may be implemented as one or more processors for executing code instructions in support of the functionality of the various elements and operations of the IVAS stereo encoding device 200, the IVAS stereo encoding method 250, the IVAS stereo decoding device 800, and the IVAS stereo decoding method 850 described above, as shown in the accompanying drawings and / or described in this disclosure.
[0206] The memory 1408 may comprise non-transitory memory for storing code instructions executable by the processor 1406, specifically, processor-readable memory that stores non-transitory instructions that, when executed, cause the processor to implement elements and operations of the IVAS stereo encoding device 200, the IVAS stereo encoding method 250, the IVAS stereo decoding device 800, and the IVAS stereo decoding method 850. The memory 1408 may also comprise random access memory or buffers for storing intermediate processed data from various functions performed by the processor 1406.
[0207] Those skilled in the art will recognize that the descriptions of the IVAS stereo encoding device 200, the IVAS stereo encoding method 250, the IVAS stereo decoding device 800, and the IVAS stereo decoding method 850 are illustrative only and are not intended to be limiting in any way. Other embodiments will be readily apparent to those skilled in the art given the benefit of this disclosure. Furthermore, the disclosed IVAS stereo encoding device 200, the IVAS stereo encoding method 250, the IVAS stereo decoding device 800, and the IVAS stereo decoding method 850 can be customized to provide valuable solutions to existing needs and problems for encoding and decoding stereo sound.
[0208] For the sake of clarity, not all routine features of implementations of the IVAS stereo encoding device 200, the IVAS stereo encoding method 250, the IVAS stereo decoding device 800, and the IVAS stereo decoding method 850 have been shown and described. Of course, it will be understood that in developing any such actual implementation of the IVAS stereo encoding device 200, the IVAS stereo encoding method 250, the IVAS stereo decoding device 800, and the IVAS stereo decoding method 850, numerous implementation-specific decisions may have to be made to achieve the developer's specific goals, such as conformance with application-, system-, network-, and business-related constraints, and that these specific goals will vary from implementation to implementation and from developer to developer. Moreover, it will be understood that the development effort may be complex and time-consuming, but would nevertheless be a routine undertaking for those skilled in the art of sound processing having the benefit of this disclosure.
[0209] In accordance with this disclosure, the elements, processing operations, and / or data structures described herein may be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. In addition, those skilled in the art will recognize that devices with less general-purpose nature, such as hardwired devices, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., may also be used. When a method comprising a series of operations and sub-operations is performed by a processor, computer, or machine, the operations and sub-operations may be stored as a series of non-transitory code instructions readable by the processor, computer, or machine, which may be stored on a tangible and / or non-transitory medium.
[0210] The elements and processing operations of the IVAS stereo encoding device 200, IVAS stereo encoding method 250, IVAS stereo decoding device 800, and IVAS stereo decoding method 850 as described herein may comprise software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.
[0211] In the IVAS stereo encoding method 250 and the IVAS stereo decoding method 850 as described herein, various processing operations and sub-operations may be performed in various orders, and some of the processing operations and sub-operations may be optional.
[0212] Although the present disclosure has been described above by way of non-limiting exemplary embodiments of the present disclosure, these embodiments can be modified at will within the scope of the appended claims without departing from the spirit and scope of the present disclosure.
[0213] This disclosure refers to the following references, the contents of which are incorporated herein by reference in their entireties:
[0214] (References) [Explanation of symbols]
[0215] 101 Communication Links 102 microphones 103 Left Channel 104 A / D converter 105 left channel 106 Stereo Sound Encoder 107 Bitstream 108 Error Correction Encoder 109 Error Correction Decoder 110 Stereo Sound Decoder 111 Bitstream 112 bitstream 113 Left Channel 114 Left Channel 115 D / A converter 116 loudspeaker unit 122 microphones 123 Right Channel 125 right channel 133 Right Channel 134 Right Channel 136 Binaural Headphones 200 IVAS stereo encoding device 202 ICA parameters 203 Time Domain Transient Detector 204 Time Domain Transient Detector 205 Stereo Classifier and Stereo Mode Selector 206 bitstream 270 Stereo Mode Signaling 300 DFT Stereo Encoder 301 Calculator 302 Calculator 303 Stereo Processor and Downmixer 304 Residual Signal Encoder 305 Calculator 306 Calculator 307 Initial Preprocessor 308 Core Encoder Configuration 310 bitstream 311 Core Encoder 312 Additional Preprocessors 313 Bitstream 314 bitstream 400 TD Stereo Encoder 401 Time Domain Analyzer and Downmixer 402 Side Parameters 403 Initial Preprocessor 404 Initial Preprocessor 405 Core Encoder Configuration 406 Core Encoder 407 Core Encoder 410 bitstream 500 MDCT Stereo Encoder 503 Initial Preprocessor 504 Initial Preprocessor 506 Joint Core Encoder 508 bitstream 509 bitstream 601 TD Stereo Frame 602 DFT Stereo Frame 800 IVAS stereo decoding device 801 DFT Stereo Decoder 802 TD Stereo Decoder 803 MDCT Stereo Decoder 807 Core Decoder 808 decoder 809 DFT Stereo Decoder and Upmixer 810 Core Decoder 811 Core Decoder 812 Upmixer 813 Joint Core Decoder 814 Time Synchronizer and Stereo Switch 815 IC-BWE Calculator 816 ICA Decoder 830 bitstream 1402 Input 1404 Output 1406 processor 1408 memory
Claims
1. 1. A device for encoding a stereo sound signal, comprising: a first stereo encoder of the stereo audio signal using a first stereo mode operating in the time domain (TD), the first TD stereo mode producing, in TD frames of the stereo audio signal, (a) a first downmixed signal and (b) using a first data structure and memory; a second stereo encoder of the stereo audio signal using a second stereo mode operating in the frequency domain (FD), the second FD stereo mode (a) producing a second downmixed signal in FD frames of the stereo audio signal, and (b) using a second data structure and memory; a controller for switching between (i) the first TD stereo mode and the first stereo encoder, and (ii) the second FD stereo mode and the second stereo encoder, for coding the stereo sound signal in a time domain or a frequency domain; a stereo mode switching controller configured to, when switching from one of the first TD stereo mode and the second FD stereo mode to the other of the first TD stereo mode and the second FD stereo mode, recalculate at least one downmixed signal section of a certain length in the current frame of the stereo sound signal, wherein a length of the recalculated downmixed signal section in the first TD stereo mode is different from a length of the recalculated downmixed signal section in the second FD stereo mode.
2. 2. The stereo sound signal encoding device of claim 1, wherein the second FD stereo mode is a Discrete Fourier Transform (DFT) stereo mode.
3. 3. The stereo sound signal encoding device of claim 2, wherein when switching from the first TD stereo mode to the second DFT stereo mode, the second stereo encoder continues core encoding operations on a DFT stereo frame after the TD stereo frame using a memory of a primary channel PCh core encoder.
4. 4. The stereo sound signal encoding device of claim 2, wherein the stereo mode switching controller uses stereo-related parameters from the one stereo mode to update stereo-related parameters of the other stereo mode when switching from the one stereo mode to the other stereo mode.
5. The stereo sound signal encoding device of claim 4 , wherein the stereo mode switching controller transfers the stereo-related parameters between data structures.
6. 6. The stereo sound signal encoding device of claim 4, wherein the stereo-related parameters comprise side gain and Inter-Channel Time Delay (ITD) parameters of the second DFT stereo mode and target gain and correlation delay of the first TD stereo mode.
7. 7. The stereo sound signal encoding device according to claim 2, wherein when switching from the second DFT stereo mode to the first TD stereo mode, the stereo mode switching controller recalculates in the current TD frame a duration of the downmixed signal in the secondary channel SCh that is longer than a duration of the recalculated downmixed signal in the primary channel PCh.
8. 8. The stereo sound signal encoding device according to claim 2, wherein when switching from the second DFT stereo mode to the first TD stereo mode, the stereo mode switching controller cross-fades the recalculated primary channel PCh and the DFT intermediate channel m of the DFT stereo channel to recalculate the downmixed primary channel PCh in the first TD frame after the DFT frame.
9. 9. The stereo sound signal encoding device according to claim 2, wherein when switching from the second DFT stereo mode to the first TD stereo mode, the stereo mode switching controller recalculates ICA memories of the left channel l and the right channel r corresponding to a DFT frame previous to a TD frame.
10. 10. The stereo sound signal encoding device of claim 9, wherein the stereo mode switching controller recalculates the primary channel PCh and the secondary channel SCh of the DFT frame by downmixing the ICA processed channels l and r using a stereo mixing ratio of the DFT frame.
11. 11. The stereo sound signal encoding device according to claim 10, wherein the stereo mode switching controller recalculates a shorter section of the secondary channel SCh when there is no stereo mode switching.
12. 12. The stereo sound signal encoding device according to claim 10, wherein the stereo mode switching controller recalculates a first interval of the primary channel PCh and a second interval of the secondary channel SCh in the DFT frame before the TD frame, and the first interval is shorter than the second interval.
13. 1. A device for decoding a stereo sound signal, comprising: a first stereo decoder of the stereo audio signal using a first stereo mode operating in the time domain (TD), the first stereo decoder (a) decoding a downmixed signal in TD frames of the stereo audio signal, and (b) using a first data structure and memory; a second stereo decoder of the stereo audio signal using a second stereo mode operating in the frequency domain (FD), the second stereo decoder (a) decoding a second downmixed signal in FD frames of the stereo audio signal, and (b) using a second data structure and memory; a controller for switching between (i) the first TD stereo mode and the first stereo decoder, and (ii) the second FD stereo mode and the second stereo decoder; a stereo mode switching controller that, when switching from one of the first TD stereo mode and the second FD stereo mode to the other of the first TD stereo mode and the second FD stereo mode, recalculates at least one downmixed signal section of a certain length in the current frame of the stereo sound signal, and a length of the recalculated downmixed signal section in the first TD stereo mode is different from a length of the recalculated downmixed signal section in the second FD stereo mode.
14. 14. The stereo sound signal decoding device according to claim 13, wherein the second FD stereo mode is a Discrete Fourier Transform (DFT) stereo mode.
15. 15. The stereo sound signal decoding device of claim 14, wherein the stereo mode switching controller allocates / deallocates data structures to / from the first TD stereo mode and the second DFT stereo mode according to a current stereo mode, and reduces impact on static memory by maintaining only data structures utilized in the current frame.
16. 16. The stereo sound signal decoding device according to claim 14 or 15, wherein the stereo mode switching controller resets the DFT stereo data structure upon receiving the first DFT frame after a TD frame.
17. 17. The stereo sound signal decoding device according to claim 14, wherein the stereo mode switching controller resets the TD stereo data structure upon receiving the first TD frame after a DFT frame.
18. 18. The stereo sound signal decoding device according to claim 14, wherein the stereo mode switching controller updates a DFT stereo synthesis memory at every single TD stereo frame.
19. 20. The stereo sound signal decoding device of claim 18, wherein, to update the DFT stereo synthesis memory and for the ACELP core, the stereo mode switching controller reconstructs a first portion of the DFT stereo synthesis memory at every TD frame by crossfading (a) a CLDFB-based resampled, TD upmixed left and right channel synthesis and (b) a reconstructed, resampled, upmixed left and right channel synthesis.
20. 20. A stereo sound signal decoding device according to any one of claims 14 to 19, wherein the stereo mode switching controller reconstructs an upmixed and synchronized TD stereo synthesis.
21. the stereo mode switching controller, for both the left and right channels, to reconstruct the upmixed synchronized TD stereo synthesis: (a) Correcting DFT stereo OLA synthesis memory, (b) reusing the upmixed DFT stereo synchronous synthesis memory as a first portion of the upmixed synchronized TD stereo synthesis; (c) approximating a second portion of the upmixed synchronized TD stereo synthesis using the rectified DFT stereo OLA synthesis memory; (d) smoothing the transition between the upmixed DFT stereo synchronous synthesis memory and the synchronized upmixed TD stereo synthesis at the beginning of the synchronized upmixed TD stereo synthesis by crossfading the rectified DFT stereo OLA synthesis memory with the synchronized upmixed TD stereo synthesis; 21. A stereo sound signal decoding device according to claim 20, using the operations (a) to (d) of:
22. 1. A method for encoding a stereo sound signal, the method comprising the steps of: implementing at least one processor and a memory coupled to the processor and storing non-transitory instructions for execution by the processor; Implementing a first stereo encoder of the stereo sound signal using a first stereo mode operating in the time domain (TD), the first TD stereo mode producing, in TD frames of the stereo sound signal, (a) a first downmixed signal and (b) using a first data structure and memory; implementing a second stereo encoder of the stereo audio signal using a second stereo mode operating in the frequency domain (FD), the second FD stereo mode (a) producing a second downmixed signal in FD frames of the stereo audio signal, and (b) using a second data structure and memory; controlling switching between (i) the first TD stereo mode and the first stereo encoder, and (ii) the second FD stereo mode and the second stereo encoder, for coding the stereo sound signal in the time domain or the frequency domain; 10. The method of claim 9, wherein when switching from one of the first TD stereo mode and the second FD stereo mode to the other of the first TD stereo mode and the second FD stereo mode, the step of controlling stereo mode switching comprises the step of recalculating at least one downmixed signal section of a certain length in a current frame of the stereo sound signal, wherein a length of the recalculated downmixed signal section in the first TD stereo mode is different from a length of the recalculated downmixed signal section in the second FD stereo mode.
23. 23. The method of claim 22, wherein the second FD stereo mode is a Discrete Fourier Transform (DFT) stereo mode.
24. controlling stereo mode switching when switching from the one of the first TD stereo mode and the second DFT stereo mode to the other of the first TD stereo mode and the second DFT stereo mode, an input stereo signal containing a left channel and a right channel; a middle channel used in the second DFT stereo mode; a primary channel and a secondary channel used in the first TD stereo mode; a downmixed signal used in preprocessing, and Downmixed signal used in core coding 24. The method of claim 23, further comprising the step of maintaining continuity of at least one of the signals:
25. 25. The method for encoding a stereo sound signal according to claim 23, wherein when switching from one of the first TD stereo mode and the second DFT stereo mode to the other of the first TD stereo mode and the second DFT stereo mode, controlling the stereo mode switching comprises allocating / deallocating data structures to / from the first TD stereo mode and the second DFT stereo mode depending on the current stereo mode so as to reduce memory impact by maintaining only data structures utilized in the current frame.
26. 26. The method of claim 25, wherein when switching from the first TD stereo mode to the second DFT stereo mode, controlling stereo mode switching comprises deallocating TD stereo related data structures.
27. 27. The method of claim 26, wherein the TD stereo related data structures comprise TD stereo data structures and / or data structures of a core encoder of the first stereo encoder.
28. 28. A method for encoding a stereo sound signal according to any one of claims 23 to 27, wherein the step of controlling stereo mode switching comprises the step of updating a DFT analysis memory for each TD stereo frame by storing samples related to the last period of a current TD stereo frame.
29. 29. A method for encoding a stereo sound signal according to any one of claims 23 to 28, wherein the step of controlling stereo mode switching comprises the step of maintaining a DFT-related memory during a TD stereo frame.
30. 30. The method of claim 23, wherein controlling the stereo mode switching comprises, when switching from the first TD stereo mode to the second DFT stereo mode, updating a DFT synthesis memory in a DFT frame following the TD frame using a TD stereo memory corresponding to a primary channel PCh of the TD frame.
31. 31. A method for encoding a stereo sound signal according to any one of claims 23 to 30, wherein controlling stereo mode switching comprises maintaining a finite impulse response (FIR) resampling filter memory during a DFT frame.
32. 32. The method of claim 31, wherein controlling stereo mode switching comprises updating the FIR resampling filter memory used in the primary channel PCh in the first stereo encoder at every DFT frame using an interval of intermediate channel m preceding a last interval of a first length of intermediate channel m in the DFT frame.
33. 33. The method of claim 32, wherein controlling the switching comprises filling an FIR resampling filter memory used in a secondary channel SCh in the first stereo encoder differently from the updating of the FIR resampling filter memory used in the primary channel PCh in the first stereo encoder.
34. 34. The method of claim 33, wherein controlling stereo mode switching comprises updating the FIR resampling filter memory used in the secondary channel SCh in the first stereo encoder in the current TD frame by filling the FIR resampling filter memory using an interval of intermediate channel m before a last interval of a second length of intermediate channel m in the DFT frame.
35. 35. A method for encoding a stereo sound signal according to any one of claims 23 to 34, wherein the step of controlling stereo mode switching comprises the step of storing two values of a pre-emphasis filter memory for each DFT frame.
36. 36. The method of claim 23, further comprising: a secondary channel SCh core encoder data structure; and wherein, when switching from the second DFT stereo mode to the first TD stereo mode, controlling the stereo mode switching comprises resetting or estimating the secondary channel SCh core encoder data structure based on a primary channel PCh core encoder data structure.
37. 1. A method for decoding a stereo sound signal, the method comprising the steps of: implementing at least one processor and a memory coupled to the processor and storing non-transitory instructions to be executed by the processor; implementing a first stereo decoder of the stereo audio signal using a first stereo mode operating in the time domain (TD), the first stereo decoder (a) decoding a downmixed signal in TD frames of the stereo audio signal, and (b) using a first data structure and memory; implementing a second stereo decoder of the stereo audio signal using a second stereo mode operating in the frequency domain (FD), the second stereo decoder (a) decoding a second downmixed signal in FD frames of the stereo audio signal, and (b) using a second data structure and memory; (i) controlling switching between the first TD stereo mode and the first stereo decoder, and (ii) controlling switching between the second FD stereo mode and the second stereo decoder; 10. The method of claim 9, wherein when switching from one of the first TD stereo mode and the second FD stereo mode to the other of the first TD stereo mode and the second FD stereo mode, the step of controlling stereo mode switching comprises the step of recalculating at least one downmixed signal section of a certain length in a current frame of the stereo sound signal, wherein a length of the recalculated downmixed signal section in the first stereo mode is different from a length of the recalculated downmixed signal section in the second stereo mode.
38. 38. The stereo sound signal decoding method of claim 37, wherein the second FD stereo mode is a Discrete Fourier Transform (DFT) stereo mode.
39. 39. The stereo sound signal decoding method of claim 38, wherein the first stereo mode uses a first processing delay and the second stereo mode uses a second processing delay, the first processing delay and the second processing delay being different and comprising resampling and upmixing processing delays.
40. controlling stereo mode switching when switching from one of the first TD stereo mode and the second DFT stereo mode to the other of the first TD stereo mode and the second DFT stereo mode, the intermediate channel m used in the second DFT stereo mode, a primary channel PCh and a secondary channel SCh used in the first TD stereo mode; TCX-LTP post-filter memory, DFT OLA analysis memory at the internal sampling rate and the output stereo signal sampling rate, a DFT OLA synthesis memory at the output stereo signal sampling rate; an output stereo signal containing channels l and r, and HB signal memory, channels l and r, used in BWE and IC-BWE 40. A method for decoding a stereo sound signal according to claim 38 or 39, comprising the step of maintaining continuity of at least one of the signal and the memory.
41. 41. A method for decoding a stereo sound signal according to any one of claims 38 to 40, wherein the step of controlling stereo mode switching comprises the step of updating a DFT stereo OLA memory buffer at every single TD frame.
42. 42. A method for decoding a stereo sound signal according to any one of claims 38 to 41, wherein the step of controlling stereo mode switching comprises the step of updating a DFT stereo analysis memory.
43. 43. The stereo sound signal decoding method of claim 42, wherein the step of controlling stereo mode switching upon receiving a first DFT frame after a TD frame comprises the step of updating the DFT stereo analysis memories of DFT stereo middle channel m and side channel s, respectively, in the DFT frame using a certain number of last samples of the primary channel PCh and secondary channel SCh of the TD frame.
44. 44. A method for decoding a stereo sound signal according to any one of claims 38 to 43, wherein the step of controlling stereo mode switching comprises crossfading the aligned and synchronized TD synthesis with the aligned and synchronized DFT stereo synthesis to provide a smooth transition when switching from TD frames to DFT frames.
45. 45. A method for decoding a stereo sound signal according to any one of claims 38 to 44, wherein the step of controlling stereo mode switching comprises the step of updating a TD stereo synthesis memory between DFT frames in case the next frame is a TD frame.
46. 46. A method for decoding a stereo sound signal according to any one of claims 38 to 45, wherein when switching from DFT frames to TD frames, the step of controlling the switching comprises a step of resetting a memory of a core decoder of a secondary channel SCh in the first stereo decoder.
47. 47. A method for decoding a stereo sound signal according to any one of claims 38 to 46, wherein when switching from DFT frames to TD frames, controlling the stereo mode switching comprises using signal energy equalization to suppress discontinuities and differences between the upmixed DFT stereo channels and the TD stereo channels.
48. The step of controlling the stereo mode switching to suppress discontinuities and differences between the upmixed DFT stereo channels and the TD stereo channels includes controlling the ICA target gain g ICA If is less than 1.0, [Equation 1] the left channel l, y after upmixing in the TD frame and before time synchronization using the relationship L (i) modifying L eq is the length of the signal to be equalized, and α is [Equation 2] 48. A method of decoding a stereo sound signal according to claim 47, wherein the gain factor values are obtained using the relationship:
Citation Information
Patent Citations
An audio encoder for encoding a multi-channel signal and an audio decoder for decoding the encoded audio signal
JP2018511825A
Multi-Channel Coding
JP2019512737A
Method and an apparatus for processing an audio signal
US20100070285A1
Audio encoder for encoding a multichannel signal and audio decoder for decoding an encoded audio signal
US20170365263A1
US63/075,984