Method and apparatus for audio bandwidth detection and audio bandwidth switching in an audio codec
By introducing a simplified audio bandwidth detection and switching algorithm into the IVAS coding framework, the problems of high computational complexity and artifacts in MDCT stereo mode are solved, achieving efficient audio bandwidth detection and switching, and improving coding efficiency and quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VOICEAGE CORPORATION
- Filing Date
- 2021-10-14
- Publication Date
- 2026-07-21
AI Technical Summary
Existing audio codecs suffer from high computational complexity and artifacts when detecting and switching audio bandwidth, especially in MDCT stereo mode, resulting in low coding efficiency and poor quality.
A novel audio bandwidth detection and switching algorithm is employed. By simplifying computational complexity in the preprocessing stage, utilizing MDCT spectral energy analysis and long-term count updates, and combining with the final audio bandwidth determination module, seamless bandwidth switching is achieved, making it suitable for MDCT stereo mode in the IVAS coding framework.
It improves coding efficiency, reduces artifacts, ensures the continuity and stability of coding quality, and adapts to audio signal processing under different bit rates and bandwidths.
Smart Images

Figure CN116529814B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to sound coding, and more particularly, but not exclusively, to methods and apparatus for audio bandwidth detection and for audio bandwidth switching in sound codecs.
[0002] In this disclosure and the appended claims:
[0003] - The term "sound" can be related to speech, audio, and any other sound;
[0004] - The term "stereo" is an abbreviation of "stereophonic"; and
[0005] - The term "mono" is an abbreviation of "monophonic". Background Technology
[0006] Historically, conversational telephones were implemented using handheld devices with only one transducer, outputting sound to only one of the user's ears. In the last decade, users have begun using their portable handheld devices in conjunction with headsets to receive sound through both ears, primarily for listening to music, but sometimes for listening to voice. However, when portable handheld devices are used to send and receive conversational voice, the content remains mono, but is presented to both of the user's ears when headsets are used.
[0007] The quality of encoded sound (e.g., speech and / or audio) transmitted and received via portable handheld devices has been significantly improved by utilizing the latest 3GPP (3rd Generation Partnership Project) speech coding standard, namely the codec for enhanced voice services (EVS), as described in its entirety in reference [1] which is incorporated herein by reference. The next natural step is to transmit stereo information so that the receiver is as close as possible to the real-life audio scene captured at the other end of the communication link.
[0008] In audio codecs, stereo information is typically transmitted.
[0009] For conversational speech codecs, mono signals are the norm. When transmitting stereo audio signals, the bit rate typically needs to be doubled because both the left and right channels of the stereo audio signal are encoded using a mono codec. To reduce the bit rate, efficient stereo coding techniques have been developed and are used. The use of stereo coding techniques is discussed in the following paragraphs as a non-limiting example.
[0010] The first stereo coding technique is called parametric stereo. Parametric stereo uses a common mono codec to encode the left and right channels into mono signals plus a certain amount of stereo side information representing the stereo image (corresponding to stereo parameters). The two input left and right channels are downmixed into mono signals, and the stereo parameters are then typically calculated in the transform domain (e.g., in the Discrete Fourier Transform (DFT) domain) and are associated with so-called binaural or interchannel cues (“cues”). Binaural cues (reference [3], the entire contents of which are incorporated herein by reference) include interaural level difference (ILD), interaural time difference (ITD), and interaural correlation (IC). Depending on the signal characteristics, stereo scene configuration, etc., some or all of the binaural cues are encoded and transmitted to the decoder. Information about which binaural cues are encoded and transmitted is sent as signaling information, which is usually part of the stereo side information. Specific binaural cues can also be quantized using different coding techniques, which results in the use of a variable number of bits. Then, in addition to the quantized binaural cues, the stereo side information can typically contain the quantized residual signal generated by undermixing at medium to high bit rates. The residual signal can be encoded using entropy coding techniques (e.g., arithmetic encoders). Generally, parametric stereo coding is most efficient at low to medium bit rates. Parametric stereo with parameters computed in the DFT domain will be referred to herein as DFT stereo.
[0011] Another stereo coding technique is one that operates in the time domain. This stereo coding technique mixes two inputs (i.e., the left and right channels) into so-called master and consonant channels. For example, according to the method described in its entirety by reference [4] incorporated herein by reference, time-domain mixing can be based on a mixing ratio that determines the respective contribution of the two inputs (i.e., the left and right channels) in producing the master and consonant channels. The mixing ratio is derived from several metrics, such as the normalized correlation of the input left and right channels relative to a mono version of the stereo sound signal or the long-term correlation difference between the two inputs (i.e., the left and right channels). The master channel can be encoded with a common mono codec, while the consonant channel can be encoded with a codec at a lower bit rate. Consonant channel encoding can take advantage of the coherence between the master and consonant channels and can reuse some parameters of the master channel. Time-domain stereo will be referred to as TD stereo in this disclosure. Generally speaking, TD stereo is most effective at encoding speech signals at low and medium bit rates.
[0012] The third stereo coding technique is one that operates in the modified discrete cosine transform (MDCT) domain. It is based on joint coding of both the left and right channels, while simultaneously computing global ILD and center / side (M / S) processing in the whitened spectral domain. It uses several tools adapted from TCX (Transform Coding Excitation) coding in the MPEG (Moving Picture Experts Group) codec, such as TCX core coding, TCX LTP (Long-Term Prediction) analysis, TCX noise filling, frequency domain noise shaping (FDNS), stereo intelligent gap filling (IGF), and / or adaptive bit allocation between channels, as described in the examples in the entirety of which are incorporated herein by reference [7] and [8]. In general, this third stereo coding technique can efficiently encode all kinds of audio content at medium and high bit rates. The MDCT domain stereo coding technique will be referred to as MDCT stereo in this disclosure.
[0013] Furthermore, in recent years, the generation, recording, representation, encoding, transmission, and reproduction of audio have been evolving towards enhanced, interactive, and immersive listener experiences. For example, an immersive experience can be described as a state of deep engagement or immersion in a sound scene when sound comes from all directions. In immersive audio (also known as 3D audio), a wide range of sound characteristics, such as timbre, directionality, reverberation, transparency, and (auditory) spaciousness, are considered for accuracy, and the sound image is reproduced in all three dimensions surrounding the listener. Immersive audio is produced for specific sound playback or reproduction systems such as speaker-based systems, integrated reproduction systems (soundboards), or headphones. The interactivity of the sound reproduction system can then include, for example, the ability to adjust sound levels, change sound positions, or select different languages for reproduction.
[0014] There are three basic methods for achieving immersive experiences.
[0015] The first approach to achieving an immersive experience is a channel-based audio method, which uses multiple spaced microphones to capture sound from different directions, with one microphone corresponding to an audio channel in a specific speaker layout. Each recorded channel is then fed to a speaker in a given location. Examples of channel-based audio methods are, for example, stereo, 5.1 surround sound, 5.1+4, etc. Typically, channel-based audio is encoded by multiple core encoders, where the number of core encoders usually corresponds to the number of recorded channels. For example, channels are encoded by multiple stereo encoders using, for example, TD stereo or MDCT stereo coding techniques. Channel-based audio will be referred to herein as a multichannel (MC) format method.
[0016] The second approach to achieving an immersive experience is a scene-based audio approach, which represents the desired sound field in the local space as a function of time through the combination of dimensional components. The sound signal representing scene-based audio (SBA) is independent of the location of the audio source, while the sound field is converted at the presenter to the selected speaker layout. An example of scene-based audio is ambient stereo. Several SBA coding techniques exist, the most well-known of which is probably Directional Audio Coding (DirAC), as described, for example, in its entirety by reference incorporated herein by reference [6]. The DirAC encoder uses analysis of the ambient stereo input signal in the domain of Complex Low Delay Filter Bank (CLDFB) to estimate spatial parameters (metadata) such as direction and spread grouped in time and frequency slots, and mixes the input channels into a smaller number of so-called transmission channels (typically 1, 2, or 4 channels). The DirAC decoder then decodes the spatial metadata, derives the direct and spread signals from the transmission channels, and presents them to the speaker or headphone setup to suit different listening configurations. Another example of SBA encoding techniques primarily used in motion capture devices is the Metadata-Assisted Spatial Audio (MASA) format, as described, for example, in its entirety by reference [9] incorporated herein by reference. In the MASA approach, MASA metadata (e.g., direction, energy ratio, extended coherence, distance, and surround coherence in several time-frequency slots) is generated, quantized, encoded, and passed to the bitstream in a MASA analyzer, while the MASA audio channels are treated as mono or multi-channel transmission signals encoded by the core encoder. At the MASA decoder, the MASA metadata then guides the decoding and rendering processes to recreate the output spatial sound.
[0017] The third approach to achieving immersive experiences is the object-based audio approach, which represents the auditory scene as a collection of individual audio elements (e.g., singer, drums, guitar, etc.) accompanied by information such as their location, allowing them to be rendered (transformed) by the sound reproduction system at their intended positions. This provides great flexibility and interactivity to the object-based audio approach, as each object remains discrete and can be manipulated individually. Each audio object consists of an audio stream (i.e., waveform) with associated metadata and can therefore also be viewed as an independent stream (ISm) with metadata.
[0018] Each of the aforementioned audio methods for achieving an immersive experience presents both advantages and disadvantages. Therefore, it is often not a single audio method, but rather a combination of several audio methods within a complex audio system to create an immersive auditory scene. An example could be an audio system that combines scene-based or channel-based audio with object-based audio (e.g., ambient stereo with some discrete audio objects).
[0019] In recent years, 3GPP (3rd Generation Partnership Project) has begun to work on developing a 3D (three-dimensional) sound codec based on the EVS codec (see reference [5], the entire contents of which are incorporated herein by reference) for immersive services called IVAS (Immersive Voice and Audio Services). Summary of the Invention
[0020] According to a first aspect, this disclosure relates to an apparatus for detecting the audio bandwidth of an audio signal to be encoded in the encoder portion of an audio codec, comprising: an audio signal analyzer; and a final audio bandwidth determination module for generating a final determination regarding the detected audio bandwidth; wherein, in the encoder portion of the audio codec, the final audio bandwidth determination module is located upstream of the audio signal analyzer.
[0021] According to a second aspect, this disclosure provides a method for detecting the audio bandwidth of an audio signal to be encoded in the encoder portion of an audio codec, comprising: analyzing the audio signal; and using the analysis results of the audio signal to finally determine the detected audio bandwidth; wherein, in the encoder portion of the audio codec, the final determination of the detected audio bandwidth is made upstream of the analysis of the audio signal.
[0022] This disclosure also relates to an apparatus for switching a sound signal to be encoded from a first audio bandwidth to a second audio bandwidth, comprising, in the encoder portion of a sound codec: a final audio bandwidth determination module for generating a final determination of the detected audio bandwidth of the sound signal to be encoded; a count value of a frame in which the audio bandwidth switching occurs, the count value of the frame responding to the final determination of the detected audio bandwidth from the final audio bandwidth determination module; and an attenuator, responding to the count value of the frame, for attenuating the sound signal prior to encoding the sound signal.
[0023] According to another aspect, this disclosure provides a method for switching a sound signal to be encoded from a first audio bandwidth to a second audio bandwidth, comprising, in the encoder portion of a sound codec: generating a final determination of the detected audio bandwidth of the sound signal to be encoded; counting frames in response to the final determination of the detected audio bandwidth; and attenuating the sound signal prior to encoding in response to the counting of frames.
[0024] The above and other objects, advantages and features of the methods and apparatus for audio bandwidth detection and for audio bandwidth switching will become more apparent when reading the following non-limiting description of illustrative embodiments thereof, given by way of example only, with reference to the accompanying drawings. Attached Figure Description
[0025] In the attached diagram:
[0026] Figure 1 This is a schematic flowchart illustrating the conditions used to increase or decrease the count value in audio bandwidth detection;
[0027] Figure 2 This is a schematic flowchart illustrating the logic for determining the final audio bandwidth during the encoding of the input audio signal;
[0028] Figure 3a is a schematic block diagram of the encoder section of the EVS audio codec using conventional audio bandwidth detection.
[0029] Figure 3b is a schematic block diagram of the encoder portion of an IVAS sound codec using the audio bandwidth detection method and apparatus according to the present disclosure.
[0030] Figure 4 This is a schematic flowchart illustrating the logic used to encode audio bandwidth information into joint parameters for two MDCT stereo channels.
[0031] Figure 5 This is a schematic block diagram showing a method and apparatus for audio bandwidth switching according to the present disclosure;
[0032] Figure 6 This is a graph showing the actual values of the attenuation factor in frames after audio bandwidth switching in IVAS running in MDCT stereo mode.
[0033] Figure 7 This is a waveform example illustrating the impact of an audio bandwidth switching mechanism on decoding quality in a segment of a speech signal where the audio bandwidth changes from wideband to ultrawideband during the highlighted portion; and
[0034] Figure 8 This is a simplified block diagram of an example configuration of the hardware components for implementing methods and devices for audio bandwidth detection and methods and devices for audio bandwidth switching. Detailed Implementation
[0035] This disclosure describes audio bandwidth detection and audio bandwidth switching techniques.
[0036] The audio bandwidth detection and audio bandwidth switching techniques are described by way of non-limiting example only with reference to the IVAS coding framework, referred to throughout this disclosure as the IVAS codec (or IVAS sound codec). However, incorporating such audio bandwidth detection and audio bandwidth switching techniques into any other sound codec is also within the scope of this disclosure.
[0037] 1. Introduction
[0038] Specifically, this disclosure describes a method and apparatus for audio bandwidth detection using an audio bandwidth detection algorithm implemented in the IVAS codec baseline, and a method and apparatus for audio bandwidth switching using an audio bandwidth switching algorithm also implemented in the IVAS codec baseline.
[0039] The audio bandwidth detection (BWD) algorithm in IVAS is similar to that in EVS and is applied in its original form in ISM, DFT stereo, and TD stereo modes. However, BWD is not applied in MDCT stereo mode. This disclosure describes a new BWD used in MDCT stereo modes, including higher bitrate DirAC, higher bitrate MASA, and multichannel formats. The goal is to introduce BWD into the modes that are missing in IVAS (i.e., to consistently use BWD at all operating points).
[0040] This disclosure further describes an audio bandwidth switching (BWS) algorithm used in the IVAS coding framework while maintaining the lowest possible computational complexity.
[0041] Traditionally, speech and audio codecs (sound codecs) typically expect to receive input audio signals with an effective audio bandwidth close to the Nyquist frequency. When the effective audio bandwidth of the input audio signal is significantly lower than the Nyquist frequency, these traditional codecs often fail to work optimally because they waste a portion of the available bit budget to represent empty frequency bands.
[0042] Today’s codecs are designed to be flexible in encoding miscellaneous audio material at a wide range of bit rates and bandwidths. An example of state-of-the-art speech and audio codecs is the EVS codec standardized in 3GPP[1]. This codec consists of a multi-rate codec capable of efficiently compressing speech, music and mixed content signals. To maintain high subjective quality for all audio material, it contains a number of different coding modes. These modes are selected based on a given bit rate, the characteristics of the input audio signal (e.g., speech / music, sound / silence), signal activity and audio bandwidth. To select the optimal coding mode, the EVS codec uses BWD. The BWD in the EVS codec is designed to detect changes in the effective audio bandwidth of the input audio signal. Thus, the EVS codec can be flexibly reconfigured to encode only perceptibly meaningful frequency content and to allocate the available bit budget in an optimal manner. In this disclosure, the BWD used in the EVS codec is further elaborated in the context of the IVAS coding framework.
[0043] Reconfiguring the codec due to BWD changes can improve codec performance. However, if the reconfiguration and its associated encoding mode switching are not handled carefully and appropriately, this reconfiguration can introduce artifacts. Artifacts are often associated with abrupt changes in high-frequency (HF) content (typically, HF is intended to specify frequency content above 8 kHz). Therefore, the disclosed Bandwidth Switching (BWS) algorithm smooths the switching and ensures that BWD changes are seamless and pleasant rather than annoying.
[0044] 2. Audio Bandwidth Detection (BWD)
[0045] 2.1 Background
[0046] Figure 3a is a schematic block diagram of the encoder portion of the EVS audio codec using audio bandwidth detection, and Figure 3b is a schematic block diagram of the encoder portion of the IVAS audio codec using the audio bandwidth detection method and apparatus according to the present disclosure. Specifically, Figure 3a shows the BWD embedded in the native EVS audio codec, while Figure 3b shows the BWD according to the present disclosure embedded in the MDCT stereo mode of the IVAS audio codec.
[0047] As shown in Figure 3a, the highlighted BWD 301 forms part of the preprocessing stage 302 of the encoder section of the EVS codec 300 to detect the audio bandwidth (BW) of the input audio signal 310. Additional information about EVS audio codecs including BWD can be found, for example, in reference [1].
[0048] Figure 3b also highlights the BWD. It can be seen that the audio bandwidth detection method and apparatus according to this disclosure are integrated into the pre-processing stage 303 and core coding stage 304 of the encoder portion of the IVAS codec 305 to detect the actual audio bandwidth (BW) of the input audio signal 320 to be encoded. This audio bandwidth information is used to operate the IVAS codec 305 in an optimal configuration tailored for a specific audio bandwidth rather than a specific input sampling frequency. Therefore, the available bit budget is allocated optimally, and thus coding efficiency is significantly improved. For example, if the input sampling frequency is 32 kHz, but there is no "energy-meaning" spectral content above 8 kHz, the codec can operate only in wideband mode without wasting a portion of the bit budget on higher frequency bands (above 8 kHz).
[0049] Additional information about the IVAS audio codec can be found, for example, in reference [5].
[0050] The BWD algorithm in the IVAS codec 305 is based on calculating the energy in certain spectral regions and comparing them with certain thresholds. In the IVAS sound codec 305, the audio bandwidth detection method and device calculates either the CLDFB value (ISm, TD stereo) or the DFT value (DFT stereo). In the AMR-WB IO (Adaptive Multi-Rate Wideband Interoperability) mode associated with the EVS codec, as described in reference [1], the audio bandwidth detection method and device use the DCT transform value to determine the audio bandwidth of the input sound signal.
[0051] The BWD algorithm itself includes several operations:
[0052] 1) Calculate the average and maximum energy values in several spectral regions of the input sound signal 320;
[0053] 2) Update long-term parameters and count values; and
[0054] 3) The final decision on the detected and therefore encoded audio bandwidth.
[0055] The first two operations 1) and 2) above are integrated into the BWD analysis operation 306 performed by the BWD analyzer 356 (which is integrated into the audio signal core coding level 304), and the last operation 3) forms the final BWD determination operation 307 performed by the final audio bandwidth determination module (processor) 357 integrated into the audio signal preprocessing level 303. As shown in Figure 3b), the final audio bandwidth determination module 357 is located upstream of the BWD analyzer 356 in the encoder part of the audio codec 305. Although the operation of the EVS native algorithm related to BWD will be mentioned and described below, its detailed description can be found in sections 5.1.6 and 5.1.7 of reference [1].
[0056] In the following description, as a non-limiting example of implementation, the following audio bandwidths / modes are defined: narrowband (NB, 0-4kHz), wideband (WB, 0-8kHz), ultra-wideband (SWB, 0-16kHz), and full-band (FB, 0-24kHz).
[0057] 2.2BWD signal
[0058] To maintain the computational efficiency of the BWD algorithm, the methods and devices used for audio bandwidth detection reuse as much as possible the signal buffers and parameters available from earlier EVS preprocessing stages (see reference [1]). In the main EVS mode, this includes complex modulation low-delay filter bank (CLDFB) values, local VAD parameters (i.e., voice activity determination without release delay) and long-term estimates of total noise energy, as discussed below.
[0059] The IVAS codec's CLDFB (see 308 in Figure 3b) generates a time-frequency matrix from the input audio signal 320. This matrix can, for example, consist of 16 time slots and several frequency subbands, where each subband is 400 Hz wide. The number of frequency subbands depends on the sampling rate of the input audio signal 320.
[0060] On the other hand, the CLDFB module is absent in the EVSAMR-WB IO mode, which calculates the Discrete Cosine Transform (DCT) to determine the audio bandwidth of the input signal in the BWD. In a non-limiting example of the implementation, the DCT values are obtained by first applying a Hanning window to 320 samples of the audio signal 320 sampled at the input sampling rate. The windowed signal is then transformed to the DCT domain and finally decomposed into several frequency subbands according to the input sampling rate. It should be noted that a constant analysis window length is used at all sampling rates to keep the computational complexity at a reasonably low level.
[0061] More details about CLDFB-based BWD can be found in reference [2], the full contents of which are incorporated herein by reference.
[0062] In MDCT stereo mode, computationally intensive CLDFB is not required, making CLDFB-based BWD inefficient. Therefore, this paper discloses a novel BWD algorithm for MDCT stereo that saves significant computational complexity of CLDFB and BWD in preprocessing stage 303.
[0063] The method and apparatus for audio bandwidth detection in MDCT stereo coding mode can lead to higher quality because bits are not allocated to the high-frequency portion of the spectrum if there is no content there or if the audio bandwidth is limited by a command line or another external request. Furthermore, the method and apparatus for audio bandwidth detection operate continuously to simplify bitrate switching involving different stereo coding techniques. Additionally, the method and apparatus for audio bandwidth detection in MDCT stereo mode can apply BWD in higher bitrate DirAC, higher bitrate MASA, and multichannel (MC) formats.
[0064] The following describes a method and apparatus for audio bandwidth detection in MDCT stereo mode.
[0065] 2.3D CT stereo BWD
[0066] In order not to increase the computational complexity associated with BWD (including CLDFB or other transformations), the BWD analyzer 356 in MDCT stereo mode is not applied to the CLDFB value in the preprocessing stage 303, but is applied to the current MDCT value later in the TCX core encoder 358.
[0067] The TCX core encoder 358 performs several operations: TCX transform based on long MDCT (TCX20) / TCX transform based on short MDCT (TCX10) switching decision, core signal analysis (TCX-LTP, MDCT, time noise shaping (TNS), linear prediction coefficient (LPC) analysis, etc.), envelope quantization and FDNS, fine quantization of the core spectrum and IGF (many of these operations are also part of the EVS codec as described in Section 5.3.3.2 of reference [1]). Core signal analysis includes windowing and MDCT calculations based on transform and overlap length applications.
[0068] The method and apparatus for audio bandwidth detection use the MDCT spectrum as input to the BWD algorithm. To simplify the algorithm, the BWD analysis operation 306 is performed only in frames selected as TCX20 frames and not transition frames; this means that the BWD analysis is performed in frames of a given duration and is skipped in frames shorter and longer than that given duration. This ensures that the length of the MDCT spectrum always corresponds to the length of the frame in the sample at the input sampling rate. Furthermore, in MC format mode, BWD is not applied in the Low Frequency Effects (LFE) channel; the LFE channel contains only low frequencies, e.g., 0–120 Hz, and therefore does not require a full-range core encoder. Additionally, as is known in the art, the input audio signals 310 / 320 are sampled at a given sampling rate and processed through groups of these samples referred to as “frames,” which are divided into several “subframes.”
[0069] In the case of MDCT energy vector, there are nine frequency bands of interest, each with a bandwidth of 1500 Hz. Each spectral region defined in Table 1 is assigned one to four frequency bands.
[0070]
[0071] Table 1: MDCT bands used for energy calculation
[0072] In Table 1 above, the lowercase letters nb (narrowband), wb (wideband), swb (ultra-wideband), and fb (full-band) represent the corresponding frequency spectrum regions, i is the index of the frequency band, and idx start It can include a starting index, and idx end It is a band-terminated index.
[0073] 2.3.1 MDCT Spectral Energy Calculation
[0074] Operation 306 of the BWD analysis in this disclosure is slightly modified from the native EVS BWD algorithm (see reference [1]) to take into account the fact that the MDCT spectrum, whose length is equal to the frame length in the samples at the input sampling rate, must be taken into account. Therefore, the DCT-based EVS native BWD algorithm path (as used in EVS AMR-WB IO mode) is adopted, while the length of the first DCT spectrum of 320 samples (which is the same at all input sampling rates in EVS) is scaled proportionally to the input sampling rate in the MDCT stereo mode of IVAS.
[0075] The energy E of the MDCT spectrum of the input audio signal 320 in MDCT stereo mode bin (i) Therefore, the calculations are as follows across the nine frequency bands:
[0076]
[0077] Where i is the frequency band index, S(k) is the MDCT spectrum, and idx start It is the bandgap starting index defined in Table 1, idx end The bandgap end index is defined in Table 1, and the bandgap width is b. width = 60 samples (regardless of the sampling rate, this corresponds to 1500Hz).
[0078] The above calculations are implemented in the source code as follows, where the "###" marker indicates that the IVAS source code used in the method and device for audio bandwidth detection is a new part relative to the EVS source code:
[0079]
[0080]
[0081]
[0082]
[0083] 2.3.2 Average and maximum energy values for each frequency band
[0084] The BWD analyzer 356 uses, for example, the following relationship to determine the energy value E in the frequency band. bin (i) Transform to the logarithmic field:
[0085] E(i) = log 10 [E bin (i)], i=0,…,8, (1)
[0086] Where i is the index of the frequency band.
[0087] The BWD analyzer 356 uses the logarithmic energy E(i) of each frequency band to calculate the average energy value for each spectral region using, for example, the following relationship:
[0088] E nb =E(0),
[0089]
[0090]
[0091]
[0092] Finally, the BWD analyzer 356 uses the logarithmic energy E(i) of each frequency band to calculate the maximum energy value for each spectral region using, for example, the following relationship:
[0093] E nb,max =E(0),
[0094]
[0095]
[0096]
[0097] The spectral regions nb, wb, swb, and fb are defined in Table 1.
[0098] 2.3.3 Length Period count value
[0099] The BWD analyzer 356 updates the long-term values of the average energy values of the spectral regions nb, wb, and swb using, for example, the following relationships:
[0100]
[0101]
[0102]
[0103] Where λ = 0.25 is an example of the update factor, and the superscript... [–1] This indicates the parameter value from the previous frame. Updates are only performed when the local VAD decision indicates that the input audio signal 320 is active or when the long-term background noise level is above 30 dB. This ensures that the parameter is updated only in frames with perceptually meaningful content. Refer to [2] for additional information on parameters / concepts such as local VAD decision, active signal, and long-term background noise.
[0104] The BWD analyzer 356 then compares the long-term average energy value of equation (4) with a specific threshold, while also considering the current maximum value of each spectral region of equation (3). Based on the comparison results, the BWD analyzer 356 increases or decreases the count values of wb, swb, and fb for each spectral region, such as... Figure 1 As shown. Figure 1 This is a schematic flowchart illustrating the conditions for increasing or decreasing count values in BWD analysis operation 306. For example, refer to... Figure 1 :
[0105] -if (see Figure 1 101) and "2.5E" wb,max >E nb,max (See 102), then the count value is cnt. wb Add, for example, "1" (see 103);
[0106] -If the condition (See 101) is not satisfied, and "3.5E" wb <E nb (See 104), then the count value is cnt. wb Reduce, for example, "1" (see 105);
[0107] -if and (See 106) and “2E” swb,max >E wb,max (See 107), then the count value is cnt. swb Add, for example, "1" (see 108);
[0108] -If the condition and (See 106) is not satisfied, and "3E" swb <E wb (See 109), then the count value is cnt. swb Decrease, for example, "1" (see 110);
[0109] -if and (See 111) and “3E” fb,max >E swb,max (See 112), then the count value is cnt. fb Add, for example, "1" (see 113); and
[0110] -If the condition and (See 111) is not satisfied, and "4.1E" fb <E swb(See 114), then the count value is cnt. fb Reduce, for example, "1" (see 115).
[0111] 2.3.4 Ultimately, the audio bandwidth is determined.
[0112] exist Figure 1 In this scenario, if the BWD analyzer 356 performs tests sequentially, the logic may change the decision regarding audio bandwidth multiple times. After each selection of a specific audio bandwidth, certain count values are reset to their minimum (e.g., "0") or their maximum (e.g., "100"). The audio bandwidth count values are constrained between 0 and 100, and the count values are compared to specific thresholds to determine the BW change. These thresholds are chosen so that the BW change (the switch between audio bandwidths) occurs with a certain lag to avoid frequent changes in the switching between detected and subsequently encoded audio bandwidths. If a potential switch from a lower BW to a higher BW is being tested, the lag is shorter (e.g., 10 frames in EVS). This short lag avoids any potential quality degradation due to the loss of HF content (because changes in HF content are typically sudden and subjectively perceptible). On the other hand, if a potential switch from a higher BW to a lower BW is being tested, a longer lag is applied (e.g., 90 frames in EVS). In this case, there is actually no significant HF content in the spectrum, so the change in spectrum content is not unnatural, abrupt, or annoying.
[0113] Figure 2 This is a schematic flowchart illustrating the decision logic for audio bandwidth detection. Figure 2 The logic output is determined by the final audio bandwidth. (See reference) Figure 2 The final audio bandwidth determination module 357 performs the following operations as the final BWD determination module 307:
[0114] - If the previous audio bandwidth BW (the previous audio bandwidth refers to the audio bandwidth determined in the previous frame) is NB (narrowband) and the count value cnt wb If the value is >10 (see 201), then the final audio bandwidth of module 357 is determined to be WB (wideband) (see 202);
[0115] -If the previous audio bandwidth BW is NB (narrowband) and the count value cnt wb >10 (see 201), and the count value cnt swb 10 (see 203), then the final audio bandwidth of module 357 is determined to be SWB (Ultra-Wideband) (see 204);
[0116] -If the previous audio bandwidth BW is NB (narrowband) and the count value cnt wb>10 (see 201), count value cnt swb >10 (see 203), and the count value cnt fb If the value is >10 (see 205), then the final audio bandwidth of module 357 is determined to be FB (full band) (see 206);
[0117] -If the previous audio bandwidth BW is WB (wideband) and the count value cnt swb If the value is >10 (see 207), then the final audio bandwidth of module 357 is determined to be SWB (Ultra-Wideband) (see 208);
[0118] -If the previous audio bandwidth BW is WB (wideband) and the count value cnt swb >10 (see 207), and the count value cnt fb If the value is >10 (see 209), then the final audio bandwidth of module 357 is determined to be FB (full band) (see 210);
[0119] -If the previous audio bandwidth BW is SWB (Ultra-Wideband) and the count value cnt fb If the value is >10 (see 211), then the final audio bandwidth of module 357 is determined to be FB (full band) (see 212);
[0120] -If the previous audio bandwidth BW is FB (full bandwidth) (see 213) and if:
[0121] -Count value cnt fb If the value is less than 10 (see 214), then the final audio bandwidth of module 357 is determined to be SWB (Ultra-Wideband) (see 215).
[0122] -Count value cnt swb If the value is less than 10 (see 216), then the final audio bandwidth of module 357 is determined to be WB (wideband) (see 217);
[0123] -Count value cnt wb If the value is less than 10 (see 218), then the final audio bandwidth of module 357 is determined to be NB (narrowband) (see 219);
[0124] -If the previous audio bandwidth BW is SWB (Ultra-Wideband) (see 220) and if:
[0125] -Count value cnt swb If the value is less than 10 (see 221), then the final audio bandwidth of module 357 is determined to be WB (wideband) (see 222);
[0126] -Count value cnt wbIf the value is less than 10 (see 223), then the final audio bandwidth of module 357 is determined to be NB (narrowband) (see 224);
[0127] -If the previous audio bandwidth BW is WB (wideband) and the count value cnt wb If the value is less than 10 (see 225), then the final audio bandwidth of module 357 is determined to be NB (narrowband) (see 226).
[0128] Figure 2 The final audio bandwidth is used to select the appropriate audio signal encoding mode.
[0129] 2.3.5 Add code
[0130] In the source code, the newly added code (marked with the sequence "###") can be as follows—the following is excerpted from the IVAS audio codec function ivas_mdct_core_whitening_enc():
[0131]
[0132]
[0133] The calculations associated with the BWD analysis operation 306 at the start of TCX core coding in the current frame (see 358) result in the final BWD decision operation 307 being postponed to the preprocessing of the next frame (see 303). Therefore, the previous EVS BWD algorithm was divided into two parts (see 306 and 307); the BWD analysis operation 306 (i.e., calculating the energy value for each band and updating the long-term count value) was completed at the start of the current TCX core coding, and the final BWD decision operation 307 was only completed in the next frame before the start of TCX core coding.
[0134] Figure 3 illustrates the differences between the BWD-related elements in the EVS codec (Figure 3a) and IVAS codec (Figure 3b) discussed above.
[0135] 2.3.6 BWD information in CPE
[0136] In MDCT stereo coding, the final BWD decision regarding the input and thus encoded audio bandwidth from the decision module 357 is not made individually for each of the two channels, but rather as a joint decision for both channels. In other words, in MDCT stereo coding, both channels are always encoded using the same audio bandwidth, and each Channel Pair Element (CPE) transmits information about the encoded audio bandwidth only once (CPE is a coding technique that encodes two channels using stereo coding). If the final BWD decision differs between the two CPE channels, then both CPE channels are encoded using the wider audio bandwidth BW of the two channels. For example, if the detected audio bandwidth BW is the WB bandwidth of the first channel and the SWB bandwidth of the second channel, then the encoded audio bandwidth BW of the first channel is rewritten as the SWB bandwidth, and the SWB bandwidth information is transmitted in the bitstream. The only exception is when one of the MDCT stereo channels corresponds to an LFE channel, in which case the encoded audio bandwidth of the other channel is set to the audio bandwidth of that channel. This is primarily used in MC format mode when several MDCT stereo CPEs are used to encode multiple MC channels.
[0137] The final audio bandwidth determination module 357 can be used Figure 4 The logic is used to encode audio bandwidth information (the audio bandwidth of the detected channel) as a joint parameter for the two MDCT stereo channels.
[0138] refer to Figure 4 If the audio bandwidth of two CPE channels is detected:
[0139] - If MDCT stereo is not used (see 401):
[0140] -Audio bandwidth (BW) used for encoding the first channel coded,ch1 The audio bandwidth BW is detected by the final audio bandwidth determination module 357. detected,ch1 The audio bandwidth BW used to encode the second channel coded,ch2 The audio bandwidth BW is detected by the final audio bandwidth determination module 357. detected,ch2 (See 402), and the audio bandwidth information includes two bitstream parameters (see 404);
[0141] - If using MDCT stereo (see 401):
[0142] - If channel X is an LFE channel (see 403), then the audio bandwidth BW used to encode the other channel Y coded,chY The audio bandwidth BW is detected by the final audio bandwidth determination module 357. detected,chYFurthermore, the audio bandwidth information is a bitstream parameter (see 406);
[0143] -If channel X is not an LFE channel (see 403):
[0144] -If the final audio bandwidth determination module 357 detects the audio bandwidth BW used to encode the first channel. detected,ch1 This is not equal to the audio bandwidth BW used for encoding the second channel detected by the final audio bandwidth determination module 357. detected,ch2 (See 407), audio bandwidth BW used to encode the first channel coded,ch1 Equal to the audio bandwidth BW used to encode the second channel coded,ch2 And equal to BW detected,ch1 and BW detected,ch2 The maximum value (see 408), and the audio bandwidth information is a bitstream parameter (see 409); and
[0145] -If the final audio bandwidth determination module 357 detects the audio bandwidth BW used to encode the first channel. detected,ch1 This is equal to the audio bandwidth BW used for encoding the second channel, detected by the final audio bandwidth determination module 357. detected,ch2 (See 407), audio bandwidth BW used to encode the first channel coded,ch1 Equal to the audio bandwidth BW used to encode the second channel coded,ch2 And equal to BW detected,ch1 (See 410), and the audio bandwidth information is a bitstream parameter (see 411).
[0146] The audio bandwidth information from blocks 405, 408 and 410 is encoded by the MDCT core encoder 358 (Figure 3b) into joint parameters for the two CPE channels.
[0147] In the source code of the IVAS audio codec, the final BW decision logic can be seen as follows, where newly added code is marked with a "###" sequence:
[0148]
[0149]
[0150]
[0151] The above function runs at the core codec configuration block, that is, at the end of preprocessing and before the start of TCX core encoding.
[0152] It should be noted that the same principle of joint audio bandwidth information encoding can be used in other stereo coding techniques, such as using two core encoders to encode two channels in TD stereo.
[0153] 3. Bandwidth Switching (BWS)
[0154] 3.1 Background
[0155] In the EVS codec, changes in audio bandwidth (BW) may occur due to bit rate variations or changes in encoded audio bandwidth. When a change occurs from wideband (WB) to ultra-wideband (SWB), or from SWB to WB, audio bandwidth switching post-processing is performed at the decoder to improve the perceived quality for end users. Smoothing is applied to the switch from WB to SWB, and blind audio bandwidth extension is used for the switch from SWB to WB. An overview of the EVS BWS algorithm is given in the following paragraphs, while more information can be found in section 6.3.7 of reference [1].
[0156] First, in the EVS, the audio bandwidth switching detector receives the transmitted BW information and, in response to such BW information, detects whether an audio bandwidth switch exists (see Section 6.3.7.1 of [1]), and updates some count values accordingly. Then, in the case of switching from SWB to WB, the high-frequency band (HB) portion of the spectrum (HB>8kHz) is estimated in the next frame based on the SWB bandwidth extension (BWE) technique of the previous frame. The HB spectrum fades out over 40 frames, while the time-domain signal at the output sampling rate is used to perform the estimation of the SWB BWE parameters. On the other hand, in the case of switching from WB to SWB, the HB portion of the spectrum fades out over 20 frames.
[0157] 3.2 Problem
[0158] In IVAS, the BWS technology used in EVS can be implemented in the decoder, but it has never been applied due to the bit rate limitations in the native EVS BWS algorithm. Furthermore, the native EVS BWS algorithm does not support BWS in the TCX core. Finally, the native EVS BWS algorithm cannot be applied to DFT stereo CNG (comfort noise generation) frames because the time-domain signal cannot be used for algorithmic estimation.
[0159] 3.3 BWS in IVAS
[0160] Therefore, a new and different BWS algorithm was implemented in the IVAS audio codec.
[0161] First, this BWS algorithm is implemented in the encoder section of the IVAS sound codec. Compared to the native EVS algorithm, this choice has the advantage of a very low complexity footprint for the IVAS BWS algorithm.
[0162] Another design option is that the BWS algorithm in IVAS is implemented only for handovers from lower BW to higher BW (e.g., from WB to SWB). In this direction, the handover is relatively fast (see Section 2.3.4 above), and the resulting abrupt changes in HF content can be annoying. Therefore, new and different BWS algorithms are designed to smooth this handover. On the other hand, no special treatment is implemented for handovers from higher BW to lower BW because there is actually no significant HF content in the spectrum in this direction, so the change in spectral content is not unnaturally abrupt and annoying.
[0163] 3.4 Proposed BWS
[0164] Figure 5 This is a schematic block diagram simultaneously showing a method 500 and a device 550 for audio bandwidth switching according to this disclosure. Figure 5 As shown, the method for audio bandwidth switching includes final audio bandwidth determination operation 307, cnt bwidth_sw Count value update operation 502, comparison operation 503, high-frequency band spectrum fade-in operation 504. Also, Figure 5 As shown, the device for audio bandwidth switching includes: a final audio bandwidth determination module 357 for performing the final BWD determination operation 307, and a module for performing CNTB... width_sw The calculator 552 for the count value update operation 502, the comparator 553 for performing the comparison operation 503, and the attenuator 554 for performing the high-frequency band spectrum fade-in operation 504.
[0165] Depend on Figure 5 The proposed BWS algorithm used in method 500 and device 550 smooths the perceived impact of audio bandwidth switching already present in the encoder portion of the IVAS audio codec, while removing artifacts in the composite. As shown by the final audio bandwidth determination module 357, the high-frequency band (HB>8kHz) portion of the spectrum is attenuated in several consecutive frames following the BWS instance. More specifically, the gain of the HB spectrum is faded in attenuator 554 and is thus subtly controlled in the case of BWS to avoid unpleasant artifacts. The attenuation is applied before the HB spectrum is quantized and encoded in the core encoder 555 and the corresponding core encoding operation 505, so the smooth BW transition is already present in the transmitted bitstream 506 and does not require further processing at the decoder. For example, in the case of audio bandwidth switching from WB to SWB, the HB spectrum corresponding to frequencies above 8kHz is smoothed before further processing. In other words, the audio bandwidth switching is inherent in the encoded audio signal, no additional bits related to the audio bandwidth switching are transmitted to the decoder, and the decoder does not perform any additional processing on the audio bandwidth switching.
[0166] 3.4.1 BWS technology
[0167] Figure 5 The BWS mechanism of the method and device used for audio bandwidth switching works as follows.
[0168] First, calculator 552 updates the count value of frame cntb. width_sw Audio bandwidth switching occurs in these frames, and at the end of preprocessing, based on the final BWD decision, 307 will apply attenuation to each IVAS transmission channel, as follows.
[0169] Calculator 552 initially counted the frame value cnt. bwidth_sw The value is set to the initial value "0". When a BW change from a lower audio bandwidth to a higher audio bandwidth (typically from WB to SWB or FB) is detected—in response to the final BWD decision from the final audio bandwidth determination module 357—the frame count value is incremented by 1. In subsequent frames, the count value is incremented by 1 in each frame until it reaches its maximum value B as defined below. tran When the count reaches its maximum value B tran At that time, the counter value is then reset to 0, and a new BW switch can be detected.
[0170] In the source code, the newly added code (marked with the sequence "###") can be seen as follows. The code excerpt can be found at the end of the IVAS audio codec function `core_switching_pre_enc()`:
[0171]
[0172]
[0173] Next, when the calculator 552 updates or does not update the count value cnt as determined by comparator 553... bwidth_sw When the value is greater than 0, attenuator 554 applies an attenuation factor β to the audio signal in frame i, for example, as defined below. i (507):
[0174]
[0175] Among them, cnt bwidth_sw This is the count value of the audio bandwidth switching frames mentioned above (bwidth_sw_cnt in the source code above), and B tran (The macro BWS_TRAN_PERIOD in the source code above) is the BWS transition period, which corresponds to the number of frames for which attenuation is applied after a BW handover from a lower BW to a higher BW. Constant B tranIt was discovered through experiments and is set to 5 in the IVAS framework.
[0176] Figure 6 This is a graph showing the actual value of the attenuation factor β in a frame after BWD detects a change in BW in IVAS running in MDCT stereo mode. Figure 6 The non-restrictive example assumes that the BW change is detected in the fastest possible time (i.e., a 10-frame lag), the final BWD decision is made in the subsequent frame (n+11), and the BWS is made in the next BW. tran =Applied across 5 frames (frames n+12 to n+16). Ultimately, the attenuation factor β is adjusted according to the coding mode in B. tran The following applies to the frame.
[0177] In the TCX and HQ core frames (HQ stands for High Quality MDCT Encoder in EVS, see Section 5.3.4 of Reference [1]), the spectrum X of length L as defined in Section 5.3.2 of Reference [1] is... M The high-frequency gain of (k) is controlled, and it is exactly in the spectrum X after the time-to-frequency domain transformation. M The high-frequency band (HB) portion of (k) is updated (fade in) by attenuator 554 using, for example, the following relationship:
[0178] X′ M (k+L WB )=β i *X M (k+L WB ), i = 0, ..., B tran -1,
[0179] Among them, L WB It is the spectral length corresponding to the WB audio bandwidth, i.e., L WB = In 320 samples of the IVAS example with a frame length of 20 milliseconds (normal HQ or TCX20 frames), L WB = 80 samples in the transient frame, L WB = 160 samples in the TCX10 frame, and k is the sample index in the range [0, K–LWB–1], where K is the length of the entire spectrum, especially for the transform sub-modes (normal, transient, TCX20, TCX10).
[0180] In the ACELP core with time-domain BWE (TBE) frames, attenuator 554 sets the attenuation factor β before further processing these parameters. iThe SWB gain shape parameter applied to the HB portion of the spectrum. The time gain shape parameter gs(j) is defined in section 5.2.6.1.14.2 of reference [1] and consists of four values. Therefore, in the example of the implementation:
[0181] gs′(j)=β i *gs(j), i=0,...,B tran -1,
[0182] Where j = 0, ..., 3 are the gain shape numbers.
[0183] In the ACELP core with frequency domain BWE (FD-BWE) frames, the transformed original input signal X of length L is defined as in Section 5.2.6.2.1 of reference [1]. M The high-frequency band gain of (k) is controlled, and the HB portion of the MDCT spectrum is updated by attenuator 554 using, for example, the following relationship:
[0184] X′ M (k+L WB )=β i *X M (k+L WB ), i = 0, ..., B tran -1
[0185] Please note that NB encoding is not considered in IVAS, and the SWB to FB switch is not processed because its subjective and objective impact is negligible. However, the same principles described above can be used to cover all BWS scenarios.
[0186] The attenuated audio signal from attenuator 554 is then encoded in core encoder 555. The count value cnt, updated or not updated, is determined by comparator 553. bwidth_sw If the value is not greater than 0, the sound signal is encoded in the core encoder 555 without attenuation.
[0187] 3.4.2 BWS Impact Examples
[0188] Figure 7 This is an example of a waveform illustrating the impact of the BWS mechanism on decoding quality. Specifically, Figure 7 A speech signal (0.3 seconds long in the example) is shown, in which the highlighted part shows a BW change from WB to SWB. Figure 7The following are shown from top to bottom: (1) input signal waveform, (2) BW parameters (value 1 corresponds to WB, and value 2 corresponds to SWB), (3) combined waveform of decoding without BWS, (4) combined spectrum of decoding without BWS, (5) combined waveform of decoding with BWS, and (6) combined spectrum of decoding with BWS. Also shown are... Figure 7 The arrows highlight that, as can be observed, the decoded composites when BWS is applied do not suffer from sudden energy increases in the time domain in the HF frequency domain. Therefore, when using the BWS technique disclosed herein, artifacts (annoying clicking sounds) are removed from the composites.
[0189] 4. Hardware Implementation
[0190] Figure 8 This is a simplified block diagram of an example configuration of the hardware components of the encoder portion of the IVAS audio codec 305, which are formed using an audio bandwidth detection method and device and an audio bandwidth switching method and device.
[0191] The encoder portion of the IVAS audio codec 305, which uses the audio bandwidth detection method and apparatus and the audio bandwidth switching method and apparatus, can be implemented as part of a mobile terminal, a portable media player, or in any similar device. The encoder portion of the IVAS audio codec 305, which uses the audio bandwidth detection method and apparatus and the audio bandwidth switching method and apparatus (… Figure 8 The part number (marked as 800) includes input 802, output 804, processor 806, and memory 808.
[0192] Input 802 is configured to receive the input audio signal 320 in digital or analog form (Figure 3b). Output 804 is configured to supply the encoded audio signal for output. Input 802 and output 804 can be implemented in a common module, such as a serial input / output device.
[0193] Processor 806 is operatively connected to input 802, output 804, and memory 808. Processor 806 is implemented as one or more processors for executing code instructions that support the functionality of various components of the encoder portion of IVAS sound codec 305 using the audio bandwidth detection method and apparatus and the audio bandwidth switching method and apparatus shown in FIG3b).
[0194] Memory 808 may include non-transient memory for storing code instructions executable by processor 806, specifically including / storing processor-readable memory for storing non-transient instructions that, when executed, cause the processor to implement the aforementioned encoder portion of the IVAS sound codec 305 using the audio bandwidth detection method and apparatus and the audio bandwidth switching method and apparatus as described in this disclosure. Memory 808 may also include random access memory or buffers for storing intermediate processing data from various functions performed by processor 806.
[0195] Those skilled in the art will recognize that the description of the encoder portion of the IVAS audio codec 305 using the audio bandwidth detection method and apparatus, and the audio bandwidth switching method and apparatus, is merely illustrative and not intended to be limiting in any way. Other embodiments will readily conceive of those skilled in the art upon receiving this disclosure. Furthermore, the encoder portion of the disclosed IVAS audio codec 305 using the audio bandwidth detection method and apparatus, and the audio bandwidth switching method and apparatus, can be customized to provide valuable solutions to existing needs and problems in encoding and decoding sound.
[0196] For clarity, not all conventional features of the encoder portion of the IVAS audio codec 305 implemented using audio bandwidth detection methods and devices, and audio bandwidth switching methods and devices, are shown and described. It should be understood, of course, that in the development of any such actual implementation of the encoder portion of the IVAS audio codec 305 using audio bandwidth detection methods and devices, and audio bandwidth switching methods and devices, many implementations may require specific decisions to achieve the developer's specific goals, such as complying with constraints related to the application, system, network, and business, and these specific goals will vary from one implementation to another and from one developer to another. Furthermore, it should be understood that the development work may be complex and time-consuming, but it remains routine engineering work for those skilled in the art of sound processing who will benefit from this disclosure.
[0197] According to this disclosure, the components / processors / modules, processing operations, and / or data structures described herein can be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. Furthermore, those skilled in the art will recognize that less general-purpose devices, such as hardwired devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., can also be used. When a method comprising a series of operations and sub-operations is implemented by a processor, computer, or machine, and those operations and sub-operations can be stored as a series of non-transitory code instructions readable by the processor, computer, or machine, they can be stored on a tangible and / or non-transitory medium.
[0198] The encoder portion of the IVAS audio codec 305, which uses the audio bandwidth detection method and apparatus as described herein, as well as the audio bandwidth switching method and apparatus, can use software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.
[0199] In the encoder portion of the IVAS audio codec 305 using the audio bandwidth detection method and apparatus as described herein, and the audio bandwidth switching method and apparatus, various operations and sub-operations can be performed in various orders, and some of the operations and sub-operations can be optional.
[0200] Although this disclosure has been described above by way of its non-limiting, illustrative embodiments, these embodiments may be modified freely within the scope of the appended claims without departing from the spirit and nature of this disclosure.
[0201] 5. References
[0202] This disclosure references the following sources, the entire contents of which are incorporated herein by reference:
[0203] [1] 3GPP TS 26.445, v.16.1.0, "Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description", July 2020.
[0204] [2]V.Eksler, M.Jelínek, and W.Jaegers, "Audio Bandwidth Detection in the EVS Codec," in Proc. IEEE Global Conf.on Signal and Information Processing (GlobalSIP), Orlando, FL, USA, 2015.
[0205] [3] F.Baumgarte, C.Faller, "Binaural cue coding-Part I: Psychoacousticfundamentals and design principles," IEEE Trans.Speech Audio Processing, vol.11, pp.509-519, Nov.2003.
[0206] [4]T.Vaillancourt,“Method and system using a long-term correlationdifference between left and right channels for time domain down mixing astereo sound signal into primary and secondary channels,”PCT ApplicationWO2017 / 049397A1.
[0207] [5]3GPP SA4 contribution S4-170749,“New WID on EVS Codec Extensionfor Immersive Voice and Audio Services”,SA4 meeting#94,June26-30,2017, http: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_94 / Docs / S4-170749.zip
[0208] [6]V.Pulkki,C.Faller,"Directional audio coding:Filterbank and STFT-based design,"in 120 th AES Convention,Paper 6658,Paris,May 2006.
[0209] [7]M.Neuendorf et al.,“MPEG Unified Speech and Audio Coding-The ISO / MPEG Standard for High-Efficiency Audio Coding of all Content Types”,Journalof the Audio Engineering Society,vol.61 n°12,pp.956-977,December 2013.
[0210] [8]J.Herre et al.,“MPEG-H Audio-The New Standard for UniversalSpatial / 3D Audio Coding”,in 137th International AES Convention,Paper 9095,LosAngeles,October 9-12,2014.
[0211] [9]3GPP SA4 contribution S4-180462,“On spatial metadata for IVASspatial audio input format”,SA4 meeting#98,April 9-13,2018, https: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_98 / Docs / S4-180462.zip 。
Claims
1. An audio bandwidth detection device for detecting the audio bandwidth of an audio signal to be encoded in the encoder section of an audio codec, the encoder section of the audio codec including an audio signal preprocessing stage and a subsequent audio signal modified discrete cosine transform (MDCT) core coding stage, the audio bandwidth detection device comprising: A sound signal analyzer, integrated into the MDCT core coding level, is used to analyze the MDCT spectrum of the sound signal; as well as The final audio bandwidth determination module, integrated into the pre-processing stage, is used to generate a final determination about the detected audio bandwidth using the results of analysis of the MDCT spectrum of the audio signal from the audio signal analyzer. In the encoder portion of the audio codec, the final audio bandwidth determination module, integrated into the pre-processing stage, is located upstream of the audio signal analyzer integrated into the MDCT core coding stage. The result of the audio signal analyzer's analysis of the MDCT spectrum of the audio signal in the current frame is used by the final audio bandwidth determination module in the next frame after the current frame to generate a final determination of the audio bandwidth of the detected audio signal.
2. The audio bandwidth detection device according to claim 1, wherein: In MDCT stereo mode, the audio signal analyzer analyzes the MDCT spectrum of the audio signal in the MDCT core coding stage of the encoder section of the audio codec, instead of calculating the complex low-delay filter bank (CLDFB) values, which are not required in the audio signal preprocessing stage of the encoder section of the audio codec, in the MDCT stereo mode.
3. The audio bandwidth detection device according to claim 1, wherein, The sound signal analyzer calculates the average energy of the MDCT spectrum of the sound signal in several spectral regions.
4. The audio bandwidth detection device according to claim 1, wherein, The sound signal analyzer calculates the maximum energy of the MDCT spectrum of the sound signal in several spectral regions.
5. The audio bandwidth detection device according to claim 1, wherein, The sound signal analyzer calculates the average and maximum energy of the MDCT spectrum of the sound signal in several spectral regions.
6. The audio bandwidth detection device according to claim 5, wherein, The audio signal analyzer calculates the energy of the MDCT spectrum of the audio signal in multiple frequency bands, wherein each of the frequency regions is defined by at least one frequency band in the multiple frequency bands, and wherein the audio signal analyzer uses the calculated energy of the MDCT spectrum of the audio signal in the frequency band to calculate the average value and the maximum value of the energy of the MDCT spectrum.
7. The audio bandwidth detection device according to claim 3, wherein, The sound signal analyzer calculates the long-term value of the average energy value of the MDCT spectrum of the sound signal in a region among the several spectral regions.
8. The audio bandwidth detection device according to claim 5, wherein, The sound signal analyzer updates the count values associated with the spectral region.
9. The audio bandwidth detection device according to claim 7, wherein, The audio signal analyzer calculates the maximum energy of the MDCT spectrum of the audio signal in several spectral regions, and wherein the audio signal analyzer increases or decreases a count value associated with a corresponding spectral region in response to the long-term value of the average energy value of the MDCT spectrum of the audio signal and the maximum energy value of the MDCT spectrum of the audio signal.
10. The audio bandwidth detection device according to any one of claims 1 to 9, wherein, The audio signal analyzer performs audio signal analysis in frames of a given duration and skips audio signal analysis in frames that are longer or shorter than the given duration.
11. The audio bandwidth detection device according to claim 8 or 9, wherein, The final audio bandwidth determination module uses decision logic to switch between audio bandwidths in response to a comparison between the count value and a given threshold.
12. The audio bandwidth detection device according to claim 11, wherein, The decision logic of the final audio bandwidth determination module also responds to the previously determined audio bandwidth.
13. The audio bandwidth detection device according to claim 11, wherein, The final audio bandwidth determination module uses lag to avoid frequent switching between audio bandwidths.
14. The audio bandwidth detection device according to claim 13, wherein, The hysteresis used by the final audio bandwidth determination module is shorter in the case of a potential switch from lower audio bandwidth to higher audio bandwidth, and longer in the case of a potential switch from higher audio bandwidth to lower audio bandwidth.
15. The audio bandwidth detection device according to any one of claims 1 to 9, wherein, The sound signal is a multi-channel signal comprising multiple channels, and wherein the final audio bandwidth determination module encodes the detected audio bandwidth of the channels into joint parameters.
16. The audio bandwidth detection device according to any one of claims 1, 3 to 9, wherein, The MDCT spectrum of the audio signal is used in MDCT stereo coding mode.
17. The audio bandwidth detection device according to any one of claims 1 to 9, wherein, The audio signal analyzer performs the analysis of the audio signal only within frames of a given duration.
18. An audio bandwidth detection device for detecting the audio bandwidth of an audio signal to be encoded in the encoder section of an audio codec, the encoder section of the audio codec including an audio signal preprocessing stage and a subsequent audio signal modified discrete cosine transform (MDCT) core coding stage, the audio bandwidth detection device comprising: At least one processor; as well as A memory, coupled to the processor, stores non-transitory instructions that, when executed, cause the processor to implement: A sound signal analyzer, integrated into the MDCT core coding level, is used to analyze the MDCT spectrum of the sound signal; and The final audio bandwidth determination module, integrated into the pre-processing stage, is used to generate a final determination about the detected audio bandwidth using the results of analysis of the MDCT spectrum of the audio signal from the audio signal analyzer. In the encoder portion of the audio codec, the final audio bandwidth determination module, integrated into the pre-processing stage, is located upstream of the audio signal analyzer integrated into the MDCT core coding stage. The result of the audio signal analyzer's analysis of the MDCT spectrum of the audio signal in the current frame is used by the final audio bandwidth determination module in the next frame after the current frame to generate a final determination of the audio bandwidth of the detected audio signal.
19. An audio bandwidth detection device for detecting the audio bandwidth of an audio signal to be encoded in the encoder section of an audio codec, the encoder section of the audio codec including an audio signal preprocessing stage and a subsequent audio signal modified discrete cosine transform (MDCT) core coding stage, the audio bandwidth detection device comprising: At least one processor; as well as A memory, coupled to the processor, stores non-transitory instructions that, when executed, cause the processor to: In the MDCT core coding level, the MDCT spectrum of the audio signal is analyzed; as well as In the preprocessing stage, the results of the analysis of the MDCT spectrum of the audio signal are used to make a final decision about the detected audio bandwidth; In the encoder portion of the sound codec, the final determination of the detected audio bandwidth is made in the pre-processing stage upstream of the analysis of the MDCT spectrum of the sound signal in the MDCT core coding stage, and the result of the analysis of the MDCT spectrum of the sound signal in the current frame in the MDCT core coding stage is used in the pre-processing stage in the next frame after the current frame to make the final determination of the audio bandwidth of the detected sound signal.
20. An audio bandwidth detection method for detecting the audio bandwidth of an audio signal to be encoded in the encoder section of an audio codec, wherein the encoder section of the audio codec includes an audio signal preprocessing stage and a subsequent audio signal modified discrete cosine transform (MDCT) core coding stage, the audio bandwidth detection method comprising: In the MDCT core coding level, the MDCT spectrum of the audio signal is analyzed; as well as In the preprocessing stage, the results of the analysis of the MDCT spectrum of the audio signal are used to make a final decision about the detected audio bandwidth; In the encoder portion of the sound codec, the final determination of the detected audio bandwidth is made in the pre-processing stage upstream of the analysis of the MDCT spectrum of the sound signal in the MDCT core coding stage, and the result of the analysis of the MDCT spectrum of the sound signal in the current frame in the MDCT core coding stage is used in the pre-processing stage in the next frame after the current frame to make the final determination of the audio bandwidth of the detected sound signal.
21. The audio bandwidth detection method according to claim 20, wherein: In MDCT stereo mode, the MDCT spectrum of the audio signal is analyzed in the MDCT core coding stage of the audio signal in the encoder section of the audio codec, instead of calculating the complex low-delay filter bank (CLDFB) values, which are not required in the audio signal pre-processing stage of the encoder section of the audio codec, in the MDCT stereo mode.
22. The audio bandwidth detection method according to claim 20, wherein, The analysis of the MDCT spectrum of the sound signal includes: calculating the average energy of the MDCT spectrum of the sound signal in several spectral regions.
23. The audio bandwidth detection method according to claim 20, wherein, The analysis of the MDCT spectrum of the sound signal includes: calculating the maximum energy of the MDCT spectrum of the sound signal in several spectral regions.
24. The audio bandwidth detection method according to claim 20, wherein, The analysis of the MDCT spectrum of the sound signal includes: calculating the average and maximum values of the energy of the MDCT spectrum of the sound signal in several spectral regions.
25. The audio bandwidth detection method according to claim 24, wherein, The analysis of the MDCT spectrum of the sound signal includes calculating the energy of the MDCT spectrum of the sound signal in multiple frequency bands, wherein each of the frequency bands is defined by at least one frequency band, and wherein the analysis of the MDCT spectrum of the sound signal includes using the calculated energy of the MDCT spectrum of the sound signal in the frequency bands to calculate the average value and the maximum value of the energy of the MDCT spectrum.
26. The audio bandwidth detection method according to claim 22, wherein, The analysis of the MDCT spectrum of the sound signal includes calculating the long-term value of the average energy value of the MDCT spectrum of the sound signal in a region among the several spectral regions.
27. The audio bandwidth detection method according to claim 24, wherein, The analysis of the MDCT spectrum of the sound signal includes updating the count values associated with the spectral region.
28. The audio bandwidth detection method according to claim 26, wherein, The analysis of the MDCT spectrum of the audio signal includes: calculating the maximum value of the energy of the MDCT spectrum of the audio signal in several spectral regions, and wherein the analysis of the MDCT spectrum of the audio signal includes increasing or decreasing the count value associated with the corresponding spectral region in response to the long-term value of the average energy value of the MDCT spectrum of the audio signal and the maximum value of the energy of the MDCT spectrum of the audio signal.
29. The audio bandwidth detection method according to any one of claims 20 to 28, wherein, The analysis of the MDCT spectrum of the audio signal is performed in frames of a given duration and is skipped in frames that are longer or shorter than the given duration.
30. The audio bandwidth detection method according to claim 27 or 28, wherein, The final determination of the detected audio bandwidth includes decision logic that switches between audio bandwidths in response to a comparison between the count value and a given threshold.
31. The audio bandwidth detection method according to claim 30, wherein, The decision logic also responds to the previously determined audio bandwidth.
32. The audio bandwidth detection method according to claim 30, wherein, The final decision regarding the detected audio bandwidth includes using hysteresis to avoid frequent switching between audio bandwidths.
33. The audio bandwidth detection method according to claim 32, wherein, The hysteresis used in the final determination of the detected audio bandwidth is shorter in the case of a potential switch from lower audio bandwidth to higher audio bandwidth, and longer in the case of a potential switch from higher audio bandwidth to lower audio bandwidth.
34. The audio bandwidth detection method according to any one of claims 20 to 28, wherein, The audio signal is a multichannel signal comprising multiple channels, and wherein the final determination regarding the detected audio bandwidth comprises encoding the detected audio bandwidth of the channels as joint parameters.
35. The audio bandwidth detection method according to any one of claims 20, 22 to 28, wherein, The MDCT spectrum of the audio signal is used in MDCT stereo coding mode.
36. The audio bandwidth detection method according to any one of claims 20 to 28, wherein, The analysis of the MDCT spectrum of the audio signal is performed only within frames of a given duration.