Method and device for audio bandwidth detection and audio bandwidth switching in audio codecs

The IVAS codec's audio bandwidth detection and switching algorithm optimizes bit allocation and transitions, addressing inefficiencies in immersive audio scenarios by using a novel BWD algorithm for MDCT stereo mode, enhancing coding efficiency and subjective quality.

JP7850145B2Active Publication Date: 2026-04-22VOICEAGE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
VOICEAGE CORPORATION
Filing Date
2021-10-14
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing audio codecs struggle to efficiently manage audio bandwidth detection and switching, particularly in immersive audio scenarios, leading to inefficient bit allocation and potential artifacts due to abrupt changes in high-frequency components.

Method used

Implementing an audio bandwidth detection and switching algorithm within the IVAS codec framework, which includes a novel BWD algorithm for MDCT stereo mode to optimize bit allocation and smooth transitions, reducing computational complexity and minimizing artifacts.

Benefits of technology

Enhances coding efficiency and subjective quality by optimizing bit distribution based on actual audio bandwidth, ensuring seamless transitions and reducing artifacts in high-frequency components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850145000019
    Figure 0007850145000019
  • Figure 0007850145000020
    Figure 0007850145000020
  • Figure 0007850145000021
    Figure 0007850145000021
Patent Text Reader

Abstract

The method and device detect the audio bandwidth of a sound signal in an encoder portion. The device includes an analyzer for the sound signal and a final audio bandwidth determination module for using the results of the analysis of the sound signal to deliver a final decision regarding the detected audio bandwidth. In the encoder portion, the final audio bandwidth determination module is located upstream of the sound signal analyzer. The method and device switch from a first audio bandwidth of the sound signal to a second audio bandwidth. In the encoder portion, the method and device include a final audio bandwidth determination module for delivering a final decision regarding the detected audio bandwidth of the sound signal, a counter of frames at which the audio bandwidth switch occurs responsive to the final decision of the detected audio bandwidth, and an attenuator responsive to the counter of frames for attenuating the sound signal before encoding the sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to audio coding, and in particular to, but not limited to, methods and devices for audio bandwidth detection, and methods and devices for audio bandwidth switching in audio codecs.

[0002] In this disclosure and the attached claims, The term "sound" may refer to speech, audio, and any other sound. The term "stereo" is an abbreviation of "stereophonic," The term "mono" is an abbreviation of "monophonic." [Background technology]

[0003] Historically, conversational telephone communication has been implemented with handsets having only one transducer to output sound to only one ear of the user. In the last decade, users have begun using their portable handsets in combination with headphones to receive sound through both ears, primarily for listening to music, but sometimes also for listening to speech. Nevertheless, when portable handsets are used to transmit and receive conversational speech, the content is still mono, but when headphones are used, it is presented to both of the user's ears.

[0004] The latest 3GPP® (Third Generation Partnership Project) speech coding standard, codec for Enhanced Voice Services (EVS), described in reference [1], whose entire contents are incorporated herein by reference, has significantly improved the quality of coded sounds, such as speech and / or audio transmitted and received via a portable handset. The next natural step is for the receiver to transmit stereo information so as close as possible to the real audio scene captured on the other side of the communication link.

[0005] In audio codecs, the transmission of stereo information is typically used.

[0006] For conversational speech codecs, mono signals are the standard. When stereo signals are transmitted, the bitrate often needs to be doubled because both the left and right channels of the stereo signal are coded using a mono codec. To reduce the bitrate, efficient stereo coding techniques have been developed and are in use. The use of stereo coding techniques is discussed in the following paragraphs as an unrestricted example.

[0007] The first stereo coding technique is called parametric stereo. Parametric stereo encodes the left and right channels as mono signals using a common mono codec and a specific amount of stereo side information (corresponding to stereo parameters) that represents the stereo image. The two input channels, left and right, are downmixed into mono signals, and then the stereo parameters are calculated, usually in the transformation domain, e.g., the Discrete Fourier Transform (DFT) domain, and are related to so-called binaural cues or inter-channel cues. Binaural cues (reference [3], whose entire content is incorporated herein by reference) include interaural level difference (ILD), interaural time difference (ITD), and interaural correlation (IC). Depending on the signal characteristics, stereo scene configuration, etc., some or all binaural cues are coded and sent to the decoder. Information about which binaural cues are coded and sent is usually sent as signaling information, which is part of the stereo side information. Certain binaural cues may be quantized using various coding techniques, resulting in the use of a variable number of bits. In addition to the quantized binaural cues, stereo side information may then include quantized residual signals resulting from downmixing, typically at medium to high bitrates. These residual signals can be coded using entropy coding techniques, such as arithmetic encoders. Generally, parametric stereo coding is most efficient at low and medium bitrates. Parametric stereo with parameters calculated in the DFT domain is referred to in this disclosure as DFT stereo.

[0008] Another stereo coding technique is one that operates in the time domain. This stereo coding technique mixes two input channels, left and right, into so-called primary and secondary channels. For example, according to the method described in reference [4], the entire content of which is incorporated herein by reference, the mixing in the time domain can be based on a mixing ratio that determines the respective contributions of the two input channels, left and right, to the generation of the primary and secondary channels. The mixing ratio is derived from several metrics, e.g., the normalized correlation of the left and right input channels to the mono version of the stereo sound signal, or the long-term correlation difference between the two input channels, left and right. The primary channel can be coded by a general mono codec, and the secondary channel can be coded by a low bitrate codec. Secondary channel coding may take advantage of the coherence between the primary and secondary channels and may reuse some parameters from the primary channel. Time-domain stereo is referred to herein as TD stereo. Generally, TD stereo is most efficient at low and medium bitrates for coding speech signals.

[0009] The third stereo coding technique is one that operates in the Modified Discrete Cosine Transform (MDCT) domain. This technique is based on combined coding of both the left and right channels while computing global ILD and mid / side (M / S) processing in the whitened spectral domain. This technique uses several tools adapted from Transform Coded eXcitation (TCX) coding in MPEG (Moving Picture Experts Group), as described in references [7] and [8], the full contents of which are incorporated herein by reference, such as TCX core coding, TCX Long-Term Prediction (TCX) analysis, TCX noise filling, Frequency-Domain Noise Shaping (FDNS), Intelligent Gap Filling (IGF), and / or adaptive bit assignment between channels. In general, this third stereo coding technique is efficient for coding all types of audio content at medium and high bitrates. The MDCT region stereo coding technique is referred to as MDCT stereo in this disclosure.

[0010] Furthermore, in recent years, the generation, recording, representation, coding, transmission, and playback of audio have shifted towards enhanced, interactive, and immersive experiences for the listener. An immersive experience can be described, for example, as being deeply involved in or participating in a sound scene while sound is coming from all directions. In immersive audio (also known as 3D audio), the sound image is reproduced in all three dimensions around the listener, taking into account a wide range of sonic characteristics such as timbre, directivity, reverberation, clarity, and accuracy of (auditory) spaciousness. Immersive audio is generated for specific sound playback or reproduction systems, such as loudspeaker-based systems, integrated playback systems (soundbars), or headphones. The interactivity of the sound playback system may then include, for example, the ability to adjust sound levels, change the position of sounds, or select different languages ​​for playback.

[0011] There are three basic methods for achieving an immersive experience.

[0012] The first technique for achieving an immersive experience is a channel-based audio technique that uses multiple spaced microphones to capture sound from different directions, with each microphone corresponding to one audio channel in a particular loudspeaker layout. Each recorded channel is then fed into a loudspeaker at a given location. Examples of channel-based audio techniques include stereo, 5.1 surround, and 5.1+4. Generally, channel-based audio is coded by multiple core coders, the number of core coders usually corresponding to the number of recorded channels. For example, channels are coded by multiple stereo coders, for example, using TD stereo or MDCT stereo coding techniques. Channel-based audio is referred to in this disclosure as a multi-channel (MC) format technique.

[0013] A second approach to achieving an immersive experience is a scene-based audio technique that represents a desired sound field in a localized space as a function of time through a combination of dimensional components. The sound signal representing scene-based audio (SBA) is independent of the location of the audio source while the sound field is transformed into a selected layout of loudspeakers in the renderer. An example of scene-based audio is ambisonics. Several SBA coding techniques exist, but perhaps the most well-known is Directional Audio Coding (DirAC), as described, for example, in reference [6], the full content of which is incorporated herein by reference. The DirAC encoder uses an analysis of the ambisonics input signal in a Complex Low Delay Filter Bank (CLDFB) domain to estimate spatial parameters (metadata) such as direction and diffusivity grouped in time and frequency slots, and downmixes the input channels to a smaller number of so-called transport channels (typically 1, 2, or 4 channels). The DirAC decoder then decodes the spatial metadata, deriving direct and spread signals from the transport channels and rendering them into loudspeaker or headphone settings to accommodate various listening configurations. Another example of an SBA coding technique, primarily aimed at mobile capture devices, is the Metadata-Assisted Spatial Audio (MASA) format, as described, for example, in reference [9], the full content of which is incorporated herein by reference. In the MASA method, MASA metadata (e.g., direction, energy ratio, spread coherence, distance, surround coherence, all within several time-frequency slots) is generated, quantized, coded, and passed to a bitstream in a MASA analyzer, and the MASA audio channels are processed as mono or multi-channel transport signals coded by a core encoder.Next, in the MASA decoder, the MASA metadata proceeds with the decoding and rendering process to recreate the output spatial sound.

[0014] A third method for achieving an immersive experience is an object-based audio approach that represents the auditory scene as a set of individual audio elements (e.g., singer, drums, guitar, etc.), along with information such as their positions, so that these audio elements can be rendered (translated) by the sound playback system at their intended locations. This gives object-based audio approaches very high flexibility and interactivity, as each object can be held and manipulated individually. Each audio object consists of an audio stream, i.e., a waveform, with associated metadata, and can therefore also be considered an Independent Stream with metadata (ISm).

[0015] Each of the audio techniques described above for achieving an immersive experience has its own advantages and disadvantages. Therefore, to create an immersive auditory scene, it is common to combine several audio techniques in a complex audio system, rather than using just one. One example is an audio system that combines scene-based or channel-based audio with object-based audio, such as an audio system that combines ambisonics with a small number of discrete audio objects.

[0016] In recent years, 3GPP (Third Generation Partnership Project) has begun working on developing a 3D audio codec for immersive services called IVAS (Immersive Voice and Audio Services) based on the EVS codec (see reference [5], the full content of which is incorporated herein by reference). [Overview of the project] [Means for solving the problem]

[0017] According to a first aspect, the disclosure relates to a device for detecting the audio bandwidth of an audio signal to be coded in the encoder portion of an audio codec, the device comprising an audio signal analyzer and a final audio bandwidth determination module for delivering a final determination regarding the detected audio bandwidth, wherein in the encoder portion of the audio codec, the final audio bandwidth determination module is located upstream of the audio signal analyzer.

[0018] According to a second aspect, the disclosure provides a method for detecting the audio bandwidth of an audio signal to be coded in the encoder portion of an audio codec, the method comprising the steps of analyzing the audio signal and making a final determination regarding the detected audio bandwidth using the results of the analysis of the audio signal, wherein the final determination regarding the detected audio bandwidth in the encoder portion of the audio codec is made upstream of the analysis of the audio signal.

[0019] The disclosure also relates to a device for switching from a first audio bandwidth to a second audio bandwidth of an audio signal to be coded, the device comprising, in the encoder portion of an audio codec, a final audio bandwidth determination module for delivering a final determination of a detected audio bandwidth of an audio signal to be coded; a frame counter in which the audio bandwidth switching occurs, wherein the frame counter responds to the final determination of the detected audio bandwidth from the final audio bandwidth determination module; and an attenuator that responds to the frame counter for attenuating an audio signal before coding the audio signal.

[0020] In a further embodiment, the Disclosure provides a method for switching from a first audio bandwidth to a second audio bandwidth of an audio signal to be coded, the method comprising the steps of: delivering a final determination regarding a detected audio bandwidth of the audio signal to be coded in an encoder portion of an audio codec; counting frames in which an audio bandwidth switch occurs in response to the final determination of the detected audio bandwidth; and attenuating an audio signal before coding the audio signal in response to the count of frames.

[0021] The aforementioned and other purposes, advantages, and features of the methods and devices for audio bandwidth detection and audio bandwidth switching will become more apparent from the following non-limiting description of exemplary embodiments, which are given only as examples in relation to the accompanying drawings. [Brief explanation of the drawing]

[0022] [Figure 1] This is a schematic flowchart illustrating the conditions for increasing or decreasing the counter in audio bandwidth detection. [Figure 2] This is a schematic flowchart illustrating the logic for determining the final audio bandwidth for switching between audio bandwidths during the coding of the input audio signal. [Figure 3A] This is a schematic block diagram of the encoder portion of an EVS audio codec that uses conventional audio bandwidth detection. [Figure 3B] This is a schematic block diagram of the encoder portion of an IVAS audio codec that uses the audio bandwidth detection method and device described herein. [Figure 4] This is a schematic flowchart illustrating the logic for coding audio bandwidth information as coupling parameters for two MDCT stereo channels. [Figure 5] This schematic block diagram simultaneously illustrates the method and device for audio bandwidth switching according to the present disclosure. [Figure 6] This graph shows the actual values ​​of the attenuation coefficient in frames after audio bandwidth switching in IVAS operating in MDCT stereo mode. [Figure 7] This diagram shows an example waveform illustrating the effect of the audio bandwidth switching mechanism on decoding quality in a speech signal segment where a change in audio bandwidth from broadband to ultra-broadband occurs in the highlighted portion. [Figure 8] This is a simplified block diagram of an exemplary configuration of hardware components implementing a method and device for audio bandwidth detection and a method and device for audio bandwidth switching. [Modes for carrying out the invention]

[0023] This disclosure describes audio bandwidth detection and audio bandwidth switching techniques.

[0024] Audio bandwidth detection and audio bandwidth switching techniques are described only as non-restrictive examples, with reference to the IVAS coding framework referred to throughout this disclosure as the IVAS codec (or IVAS audio codec). However, incorporating such audio bandwidth detection and audio bandwidth switching techniques into any other audio codec is within the scope of this disclosure.

[0025] 1. Introduction Specifically, this disclosure describes a method and device for audio bandwidth detection using an audio bandwidth detection algorithm implemented in the IVAS codec baseline, and a method and device for audio bandwidth switching using an audio bandwidth switching algorithm also implemented in the IVAS codec baseline.

[0026] The Audio Bandwidth Detection (BWD) algorithm in IVAS is similar to the BWD algorithm in EVS and is applied in its original form in ISm mode, DFT stereo mode, and TD stereo mode. However, BWD was not applied in MDCT stereo mode. This disclosure describes a new BWD used in MDCT stereo mode (including higher bitrate DirAC, higher bitrate MASA, and multi-channel formats). The goal is to introduce BWD into modes that were missing in IVAS (i.e., to use BWD consistently at all operating points).

[0027] This disclosure further describes the audio bandwidth switching (BWS) algorithm used in the IVAS coding framework while keeping computational complexity as low as possible.

[0028] Traditionally, speech and audio codecs (sound codecs) generally expect to receive input audio signals with an effective audio bandwidth close to the Nyquist frequency. If the effective audio bandwidth of the input audio signal is significantly lower than the Nyquist frequency, these conventional codecs typically do not function optimally, as they waste some of the available bit budget to represent empty frequency bands.

[0029] Today's codecs are designed to be flexible in coding diverse audio material across a wide range of bitrates and bandwidths. An example of a state-of-the-art speech and audio codec is the EVS codec, standardized in 3GPP[1]. This codec consists of a multirate codec that can efficiently compress speech, music, and mixed content signals. To maintain high subjective quality for all audio material, the codec features several different coding modes. These modes are selected depending on a given bitrate, input audio signal characteristics (e.g., speech / music, voiced / silent), signal activity, and audio bandwidth. To select the optimal coding mode, the EVS codec uses BWD. The BWD in the EVS codec is designed to detect changes in the effective audio bandwidth of the input audio signal. As a result, the EVS codec can be flexibly reconfigured to encode only perceptually meaningful frequency components and distribute the available bit budget in the most optimal way. In this disclosure, the BWD used in the EVS codec is further described in the context of the IVAS coding framework.

[0030] Codec reconfiguration as a result of BWD changes improves codec performance. However, this reconfiguration can introduce artifacts if the reconfiguration and its associated coding mode switching are not handled carefully and properly. Artifacts are typically associated with abrupt changes in high-frequency (HF) components (generally, HF is intended to specify frequency components above 8 kHz). Therefore, the Bandwidth Switching (BWS) algorithm disclosed here smooths the switching, ensuring that BWD changes are seamless, pleasant, and unobtrusive.

[0031] 2. Audio Bandwidth Detection (BWD) 2.1 Background Figure 3A is a schematic block diagram of the encoder portion of an EVS audio codec using audio bandwidth detection, and Figure 3B is a schematic block diagram of the encoder portion of an IVAS audio codec using the audio bandwidth detection method and device according to this disclosure. Specifically, Figure 3A shows the BWD embedded within the native EVS audio codec, and Figure 3B shows the BWD according to this disclosure embedded within the MDCT stereo mode of the IVAS audio codec.

[0032] As shown in Figure 3A, the highlighted BWD 301 forms part of the preprocessing stage 302 of the encoder portion of the EVS codec 300, which detects the audio bandwidth (BW) of the input sound signal 310. Additional information regarding the EVS sound codec, including the BWD, can be found, for example, in reference [1].

[0033] In Figure 3B, BWD is again highlighted. As can be seen, the audio bandwidth detection method and device according to this disclosure are integrated into the front preprocessing stage 303 and core coding stage 304 of the encoder portion of the IVAS codec 305 to detect the actual audio bandwidth (BW) of the input sound signal 320 to be coded. This audio bandwidth information is used to run the IVAS codec 305 in its optimal configuration, which is tuned to a specific audio bandwidth, rather than a specific input sampling frequency. Thus, the available bit budget is distributed in an optimal manner, resulting in a significant improvement in coding efficiency. For example, if the input sampling frequency is 32 kHz but there are no "energetically" meaningful spectral components above 8 kHz, the codec can operate only in broadband mode and will not waste any bit budget on higher bandwidths (above 8 kHz).

[0034] Additional information regarding the IVAS audio codec can be found, for example, in reference [5].

[0035] The BWD algorithm in the IVAS codec 305 is based on calculating the energy in a specific spectral region and comparing it to a specific threshold. In the IVAS audio codec 305, the audio bandwidth detection method and device operate with CLDFB values ​​(ISm, TD stereo) or DFT values ​​(DFT stereo). In the AMR-WB IO (Adaptive MultiRate WideBand InterOperable) mode described in reference [1] in relation to the EVS codec, the audio bandwidth detection method and device use DCT transformed values ​​to determine the audio bandwidth of the input audio signal.

[0036] The BWD algorithm itself involves several operations, namely, 1) Calculation of the average and maximum energy values ​​of the input sound signal 320 in several spectral regions. 2) Update of long-term parameters and counters, and 3) Final determination regarding the detected and therefore coded audio bandwidth Includes.

[0037] The first two operations 1) and 2) described above are integrated into an operation 306 of BWD analysis performed by a BWD analyzer 356 integrated into the audio signal core coding stage 304, and the last operation 3) forms an operation 307 of final BWD determination performed by a final audio bandwidth determination module (processor) 357 integrated into the audio signal preprocessing stage 303. As seen in Figure 3B, the final audio bandwidth determination module 357 is located upstream of the BWD analyzer 356 in the encoder portion of the audio codec 305. The operation of the EVS native algorithm related to BWD will be referenced and introduced later in this specification, but a detailed explanation can be found in sections 5.1.6 and 5.1.7 of reference [1].

[0038] In the following description, the following audio bandwidths / modes are defined as unrestricted examples of implementation: narrow-band (NB, 0-4kHz), wide-band (WB, 0-8kHz), ultra-wideband (SWB, 0-16kHz), and full-band (FB, 0-24kHz).

[0039] 2.2 BWD signal To maintain the BWD algorithm computationally efficient, the methods and devices for audio bandwidth detection reuse as much as possible the signal buffers and parameters available from previous EVS preprocessing stages (see reference [1]). In EVS primary mode, this includes complex modulated low delay filter bank (CLDFB) values, local VAD parameters (i.e., hangover-free audio activity determination), and long-term estimates of total noise energy, which are discussed below.

[0040] The IVAS codec's CLDFB (see 308 in Figure 3B) generates a time-frequency matrix from the input audio signal 320. The matrix may consist, for example, 16 time slots and several frequency subbands, each with a width of 400 Hz. The number of frequency subbands depends on the sampling rate of the input audio signal 320.

[0041] On the other hand, the CLDFB module does not have a discrete cosine transform (DCT) in the EVS AMR-WB IO mode, where the DCT is calculated to determine the audio bandwidth of the input signal in BWD. In a non-restrictive example of implementation, the DCT value is obtained by first applying a Hanning window to 320 samples of the audio signal sampled at the input sampling rate. The windowed signal is then transformed into the DCT domain and finally decomposed into several frequency subbands depending on the input sampling rate. It should be noted that a constant analysis window length is used across all sampling rates to keep the computational complexity reasonably low.

[0042] Further details regarding BWD based on CLDFB can be found in reference [2], the full content of which is incorporated herein by reference.

[0043] In MDCT stereo modes, the computationally intensive CLDFB, which makes CLDFB-based BWD inefficient, is not required. Therefore, this specification discloses a novel BWD algorithm for MDCT stereo that significantly reduces the computational complexity of CLDFB and BWD in the preprocessing stage 303.

[0044] If the high-bandwidth portion of the spectrum contains no content, or if the audio bandwidth is limited by the command line or another external requirement, bits are not allocated to the high-bandwidth portion of the spectrum, thus enabling the method and device for audio bandwidth detection in MDCT stereo coding mode to deliver higher quality. Furthermore, the method and device for audio bandwidth detection is continuously performed to facilitate bitrate switching with switching between different stereo coding techniques. In addition, the method and device for audio bandwidth detection in MDCT stereo mode enables the application of BWD in higher bitrate DirAC, higher bitrate MASA, and multi-channel (MC) formats.

[0045] The following describes a method and device for detecting audio bandwidth in MDCT stereo mode.

[0046] 2.3 BWD in MDCT Stereo To avoid increasing the complexity of calculations related to BWD (including CLDFB or other transformations), the BWD analyzer 356 in MDCT stereo mode is not applied to the CLDFB value in the front preprocessing stage 303, but is later applied to the MDCT value in the TCX core encoder 358.

[0047] The TCX core encoder 358 performs several options, namely, switching decisions between long MDCT-based TCX conversion (TCX20) and short MDCT-based TCX conversion (TCX10), core signal analysis (TCX-LTP, MDCT, Temporal Noise Shaping (TNS), Linear Prediction Coefficient (LPC) analysis, etc.), envelope quantization and FDNS, fine quantization of the core spectrum, and IGF (as described in section 5.3.3.2 of reference [1], many of these operations are also part of the EVS codec). Core signal analysis includes windowing and MDCT calculations applied based on the conversion length and overlap length.

[0048] A method and device for audio bandwidth detection uses an MDCT spectrum as input to a BWD algorithm. To simplify the algorithm, operation 306 of the BWD analysis is performed only on frames selected as TCX20 frames and not transition frames, meaning that the BWD analysis is performed on frames of a given duration and skipped on frames shorter and longer than this given duration. This ensures that the length of the MDCT spectrum always corresponds to the length of the frame in the sample at the input sampling rate. Also, in MC format mode, BWD is not applied to the Low-Frequency Effect (LFE) channel, as the LFE channel contains only low frequencies, e.g., 0-120 Hz, and therefore does not require a full-range core encoder. Also, as is well known in the art, the input audio signal 310 / 320 is sampled at a given sampling rate and processed by groups of these samples called "frames" which are divided into several "subframes".

[0049] In the case of the MDCT energy vector, there are nine frequency bands of interest, each with a bandwidth of 1500 Hz. As defined in Table 1, frequency bands 1 to 4 are assigned to each of the spectral regions.

[0050] [Table 1]

[0051] In Table 1 above, the lowercase letters nb (narrowband), wb (broadband), swb (ultra-broadband), and fb (fullband) represent their respective spectral regions, i is the frequency band index, and idx start This is the start index of the energy band, and idx end This is the end index of the energy band.

[0052] 2.3.1 MDCT Spectral Energy Calculation The operation of the BWD analysis 306 is slightly modified in this disclosure from the EVS native BWD algorithm (see reference [1]) to take into account the fact that an MDCT spectrum of equal length to the frame length of the sample at the input sampling rate must be considered. Thus, the DCT-based pass of the EVS native BWD algorithm (used in EVS AMR-WB IO mode) is employed, and the previous DCT spectral length of 320 samples (which is the same for all input sampling rates in EVS) is scaled proportionally to the input sampling rate in the MDCT stereo mode of IVAS.

[0053] Therefore, the energy E of the MDCT spectrum of the input sound signal 320 in MDCT stereo mode bin (i) is calculated in nine frequency bands as follows:

[0054]

number

[0055] Here, i is the frequency band index, S(k) is the MDCT spectrum, and idx start This is the energy band start index defined in Table 1, and idx end This is the energy band termination index defined in Table 1, and the width of the energy band is b width This corresponds to 60 samples (which is equivalent to 1500Hz regardless of the sampling rate).

[0056] The above calculation is implemented in the source code as follows, where the symbol "###" identifies the portion of the IVAS source code used in the new audio bandwidth detection method and device with respect to the EVS source code. void bw_detect( Encoder_State *st, / * i / o: Encoder State * / const float signal_in[], / * i : input signal * / const int16_t localVAD, / * i : localVAD flag * / const float spectrum[], / * i : MDCT spectrum * / const float enerBuffer[] / * i : CLDFB energy buffer * / ) { #define BWD_TOTAL_WIDTH 320 if ( enerBuffer != NULL ) / * CLDFB-based processing in EVS native mode * / { ... } else { / * set width of a speactral bin (corresponds to 1.5kHz) * / if ( st->input_Fs == 16000 ) { bw_max = WB; bin_width = 60; } else if ( st->input_Fs == 32000 ) { bw_max = SWB; bin_width = 30; } else / * st->input_Fs == 48000 * / { bw_max = FB; bin_width = 20; } ### if ( signal_in != NULL ) / * DCT-based processing in EVS AMR-WB IO * / ### { / * windowing of the input signal * / pt = signal_in; pt1 = hann_window_320; / * 1st half of the window * / for ( i = 0; i < BWD_TOTAL_WIDTH / 2; i++ ) { in_win[i] = *pt++ * *pt1++; } pt1--; / * 2nd half of the window * / for ( ; i < BWD_TOTAL_WIDTH; i++ ) { in_win[i] = *pt++ * *pt1--; } / * tranform into frequency domain * / edct( in_win, spect, BWD_TOTAL_WIDTH, st->element_mode ); ###} ### else / * MDCT-based processing in IVAS * / ### { ### bin_width *= ( st->input_Fs / 50 ) / BWD_TOTAL_WIDTH; ### mvr2r( spectrum, spect, st->input_Fs / 50 ); ###} / * compute energy per spectral bins * / set_f( spect_bin, 0.001f, n_bins ); for ( k = 0; k <= bw_max; k++ ) { for ( i = bwd_start_bin[k]; i <= bwd_end_bin[k]; i++ ) { for ( j = 0; j < bin_width; j++ ) { spect_bin[i] += spect[i * bin_width + j] * spect[i * bin_width + j]; } spect_bin[i] = (float) log10( spect_bin[i] ); } } } ... }

[0057] 2.3.2 Average and Maximum Energy Values for Each Frequency Band The BWD analyzer 356, for example, uses the following relationship E(i)=log 10 [E bin (i)], i = 0,..., 8, (1) to convert the energy value E bin (i) in the frequency band to the logarithmic domain, where i is the index of the frequency band.

[0058] The BWD analyzer 356, for example, uses the following relationship

[0059]

Number

[0060] to calculate the average energy value for each spectral region using the logarithmic energy E(i) for each frequency band.

[0061] Finally, the BWD analyzer 356, for example, uses the following relationship

[0062]

Number

[0063] to calculate the maximum energy value for each spectral region using the logarithmic energy E(i) for each frequency band, where the spectral regions nb, wb, swb, and fb are defined in Table 1.

[0064] 2.3.3 Long-Term Counter The BWD analyzer 356, for example, uses the following relationship

[0065]

Number

[0066] Using this, the long-term average energy values ​​for the spectral regions nb, wb, and swb are updated, where λ=0.25 is an example of an update coefficient, and superscripts are used. [-1] This indicates the parameter value from the previous frame. Updates are performed only if the local VAD determination indicates that the input sound signal 320 is active, or if the long-term background noise level is higher than 30 dB. This ensures that parameters are updated only in frames that have perceptually meaningful components. For additional information on parameters / concepts such as local VAD determination, active signal, and long-term background noise, refer to [2].

[0067] Next, the BWD analyzer 356 compares the long-term energy mean values ​​from equation (4) to a specific threshold, while also considering the current maximum values ​​for each spectral region from equation (3). Depending on the results of the comparison, the BWD analyzer 356 increases or decreases the counters for each spectral region wb, swb, and fb, as shown in Figure 1. Figure 1 is a schematic flowchart showing the conditions for increasing or decreasing the counters in the BWD analysis operation 306. Referring to Figure 1, for example, "

[0068]

number

[0069] (See 101 in Figure 1) and "2.5·E wb,max >E nb,max In the case of (see 102), the counter cnt wb However, for example, if it is increased by only "1" (see 103), "

[0070]

number

[0071] The condition " (see 101) is not met, and "3.5·E wb <E nbIf (see 104), then counter cnt wb However, for example, it is reduced by "1" (see 105), "

[0072]

number

[0073] " and "

[0074]

number

[0075] (See 106) and "2·E swb,max >E wb,max If (see 107), then counter cnt swb However, for example, if it is increased by only "1" (see 108), "

[0076]

number

[0077] " and "

[0078]

number

[0079] The condition " (see 106) is not met, and "3·E swb <E wb If (see 109), then counter cnt swb However, for example, it is reduced by only "1" (see 110), "

[0080]

number

[0081] "

[0082]

number

[0083] ", and "

[0084]

number

[0085] (See 111), and "3·E fb,max >E swb,max If (see 112), then counter cnt fb However, for example, if it is increased by only "1" (see 113), "

[0086]

number

[0087] "

[0088]

number

[0089] ", and "

[0090]

number

[0091] The condition " (see 111) is not met, and "4.1·E fb <E swb If (see 114), then counter cnt fb However, for example, it is reduced by "1" (see 115).

[0092] 2.3.4 Final Audio Bandwidth Determination In Figure 1, if the BWD analyzer 356 performs tests in a sequential order, the decision regarding the audio bandwidth may change several times using this logic. With each selection of a particular audio bandwidth, a specific counter is reset to a minimum value, e.g., "0", or a maximum value, e.g., "100". The audio bandwidth counter is constrained between 0 and 100, and the counter value is compared to a specific threshold to determine the change in BW. These thresholds are chosen so that the change in BW (switching between audio bandwidths) occurs with a specific hysteresis to avoid frequent changes in switching between the detected audio bandwidth and the subsequently coded audio bandwidth. If a potential switch from a lower BW to a higher BW is being tested, the hysteresis will be shorter (e.g., 10 frames in EVS). Since changes in the HF component are usually abrupt and subjectively noticeable, this short hysteresis avoids any potential quality degradation due to loss of the HF component. On the other hand, if a potential switch from a higher BW to a lower BW is being tested, a longer hysteresis (e.g., 90 frames in EVS) is applied. In this case, since there are virtually no significant HF components in the spectrum, the changes in spectral components are not unnaturally abrupt or jarring.

[0093] Figure 2 is a schematic flowchart showing the decision logic for audio bandwidth detection. The output of the logic in Figure 2 is the final audio bandwidth determination. Referring to Figure 2, the final audio bandwidth determination module 357 performs the final BWD determination operation 307 as follows. The last audio bandwidth BW (the last audio bandwidth refers to the audio bandwidth determined in the previous frame) is NB (narrowband), and the counter cnt wb If >10 (see 201), the final audio bandwidth determination by module 357 is WB (Wideband) (see 202). The last audio bandwidth BW is NB (narrowband), and the counter cnt wb >10 (see 201), counter cntswb If >10 (see 203), the final audio bandwidth determination by module 357 is SWB (Super Wideband) (see 204). The last audio bandwidth BW is NB (narrowband), and the counter cnt wb >10 (see 201), counter cnt swb >10 (see 203), counter cnt fb If >10 (see 205), the final audio bandwidth determination by module 357 is FB (full bandwidth) (see 206). The last audio bandwidth BW is WB (wideband), and the counter cnt swb If >10 (see 207), the final audio bandwidth determination by module 357 is SWB (Super Wideband) (see 208). The last audio bandwidth BW is WB (wideband), and the counter cnt swb >10 (see 207), counter cnt fb If >10 (see 209), the final audio bandwidth determination by module 357 is FB (full bandwidth) (see 210). The last audio bandwidth BW is SWB (Super Wideband), and the counter cnt fb If >10 (see 211), the final audio bandwidth determination by module 357 is FB (full bandwidth) (see 212). The last audio bandwidth BW is FB (full bandwidth) (see 213), counter cnt fb If <10 (see 214), the final audio bandwidth determination by module 357 is SWB (Super Wideband) (see 215), counter cnt swb If <10 (see 216), the final audio bandwidth determination by module 357 is WB (Wideband) (see 217), counter cnt wb If <10 (see 218), the final audio bandwidth determination by module 357 is NB (narrowband) (see 219). The last audio bandwidth BW is SWB (Super Wideband) (see 220), counter cnt swb If <10 (see 221), the final audio bandwidth determination by module 357 is WB (Wideband) (see 222), counter cnt wb If <10 (see 223), the final audio bandwidth determination by module 357 is NB (narrowband) (see 224). The last audio bandwidth BW is WB (wideband), and the counter cnt wb If <10 (see 225), the final audio bandwidth determination by module 357 is NB (narrowband) (see 226).

[0094] The final audio bandwidth determination from Figure 2 is used to select the appropriate audio signal coding mode.

[0095] 2.3.5 Newly added code In the source code, newly added code (marked by the "###" sequence) may look like this, and the following excerpt is from the ivas_mdct_core_whitening_enc() function of the IVAS audio codec. for ( ch = 0; ch < CPE_CHANNELS; ch++ ) { SetCurrentPsychParams( ... ); tcx_ltp_encode( ... ); core_signal_analysis_high_bitrate( ... ); ### if ( sts[ch]->hTcxEnc->transform_type[0] == TCX_20 && ### sts[ch]->hTcxCfg->tcx_last_overlap_mode != TRANSITION_OVERLAP ) ### { ### if ( sts[ch]->mct_chan_mode != MCT_CHAN_MODE_LFE ) ### { ### bw_detect( ... ); ###} ###} }

[0096] The calculations associated with the BWD analysis operation 306 at the start of TCX core coding (see 358) in the current frame result in the final BWD determination operation 307 being deferred to the front preprocessing (see 303) of the next frame. Thus, the previous EVS BWD algorithm is divided into two parts (see 306 and 307), where the BWD analysis operation 306 (i.e., calculating energy values ​​for each frequency band and updating the long-run counter) is performed at the start of the current TCX core coding, and the final BWD determination operation 307 is performed only in the next frame before the TCX core coding begins.

[0097] Figure 3 shows the differences between BWD-related elements in the EVS codec (Figure 3A) and the IVAS codec (Figure 3B) as discussed above.

[0098] 2.3.6 BWD Information in CPE In MDCT stereo coding, the final BWD determination from the decision module 357 regarding the input and thus coded audio bandwidth is made as a joint determination for both channels, rather than separately for each of the two channels. In other words, in MDCT stereo coding, both channels are always coded using the same audio bandwidth, and information about the coded audio bandwidth is transmitted only once for each channel pair element (CPE) (CPE is a coding technique that encodes two channels using a stereo coding technique). If the final BWD determination differs between two CPE channels, both CPE channels are coded using the wider audio bandwidth BW of the two channels. For example, if the detected audio bandwidth BW is the WB bandwidth for the first channel and the SWB bandwidth for the second channel, the coded audio bandwidth BW for the first channel is rewritten to the SWB bandwidth, and the SWB bandwidth information is transmitted in the bitstream. The only exception is when one of the MDCT stereo channels corresponds to the LFE channel, in which case the coded audio bandwidth of the other channel is set to the audio bandwidth of that channel. This mainly applies in MC format mode when multiple MC channels are coded using several MDCT stereo CPEs.

[0099] The final audio bandwidth determination module 357 may use the logic shown in Figure 4 to code audio bandwidth information (detected audio bandwidth of the channels) as a joint parameter for the two MDCT stereo channels.

[0100] Referring to Figure 4, if the audio bandwidth for two CPE channels is detected, If MDCT stereo is not used (see 401), Audio bandwidth BW for coding the first channel coded,ch1This is the audio bandwidth BW detected by the final audio bandwidth determination module 357. detected,ch1 The audio bandwidth BW for coding the second channel. coded,ch2 This is the audio bandwidth BW detected by the final audio bandwidth determination module 357. detected,ch2 (See 402), and the audio bandwidth information includes two bitstream parameters (See 404), When MDCT stereo is used (see 401) If channel X is an LFE channel (see 403), then the audio bandwidth BW is used to code the other channel Y. coded,chY This is the audio bandwidth BW detected by the final audio bandwidth determination module 357. detected,chY The audio bandwidth information is a single bitstream parameter (see 406), If channel X is not an LFE channel (see 403), The audio bandwidth BW detected by the final audio bandwidth determination module 357 for coding the first channel detected,ch1 The audio bandwidth BW detected by the final audio bandwidth determination module 357 for coding the second channel detected,ch2 If not equal to (see 407), the audio bandwidth BW for coding the first channel. coded,ch1 This is the audio bandwidth BW for coding the second channel. coded,ch2 Equivalent to BW detected,ch1 and BW detected,ch2 Equal to the maximum value (see 408), audio bandwidth information is one bitstream parameter (see 409), The audio bandwidth BW detected by the final audio bandwidth determination module 357 for coding the first channel detected,ch1 The audio bandwidth BW detected by the final audio bandwidth determination module 357 for coding the second channel detected,ch2If equal to (see 407), the audio bandwidth BW for coding the first channel. coded,ch1 This is the audio bandwidth BW for coding the second channel. coded,ch2 Equivalent to BW detected,ch1 Equally to (see 410), audio bandwidth information is a single bitstream parameter (see 411).

[0101] Audio bandwidth information from blocks 405, 408, and 410 is coded by the MDCT core encoder 358 (Figure 3B) as a joint parameter for the two CPE channels.

[0102] In the source code of the IVAS audio codec, the final BW determination logic may be as follows, where newly added codes are marked with the "###" sequence. ### void set_bw_stereo( ### CPE_ENC_HANDLE hCPE, / * i / o: CPE encoder structures * / ###) ### { ### Encoder_State **st = hCPE->hCoreCoder; ### ### if ( hCPE->element_mode == IVAS_CPE_MDCT ) ### { ### / * do not check band-width in LFE channel * / ### if ( sts[0]->mct_chan_mode == MCT_CHAN_MODE_LFE) ### { ### st[0]->bwidth = st[0]->input_bwidth; ###} ### else if ( sts[1]->mct_chan_mode == MCT_CHAN_MODE_LFE) ### { ### st[1]->bwidth = st[1]->input_bwidth; ###} ### / * ensure that both CPE channels have the same audio band-width * / ### else if ( st[0]->input_bwidth == st[1]->input_bwidth ) ### { ### st[0]->bwidth = st[0]->input_bwidth; ### st[1]->bwidth = st[0]->input_bwidth; ###} ### else if( st[0]->input_bwidth != st[1]->input_bwidth ) ### { ### st[0]->bwidth = max( st[0]->input_bwidth, st[1]->input_bwidth ); ### st[1]->bwidth = max( st[0]->input_bwidth, st[1]->input_bwidth ); ###} ###} ### ### st[0]->bwidth = max( st[0]->bwidth, WB ); ### st[1]->bwidth = max( st[1]->bwidth, WB ); ### ### return; ###}

[0103] The above function is executed in the core codec configuration block, that is, at the end of front-end preprocessing and before TCX core coding begins.

[0104] It should be noted that the same principle of joint audio bandwidth information coding can be used in other stereo coding techniques that code two channels using two core encoders, such as in TD stereo.

[0105] 3. Bandwidth Switching (BWS) 3.1 Background In the EVS codec, changes in audio bandwidth (BW) may occur as a result of changes in bitrate or coding of audio bandwidth. When a change from wideband (WB) to ultra-wideband (SWB) or from SWB to WB occurs, audio bandwidth switching post-processing is performed in the decoder to improve perceived quality for the end user. Smoothing is applied to the WB to SWB switch, and blind audio bandwidth expansion is used for the SWB to WB switch. A summary of the EVS BWS algorithm is given in the following paragraphs, but more information can be found in section 6.3.7 of reference [1].

[0106] First, in EVS, the audio bandwidth switching detector receives transmitted BW information and, in response to such BW information, detects whether an audio bandwidth switching is present (reference [1], section 6.3.7.1), and therefore hardly updates the counter. Next, in the case of a switch from SWB to WB, the high-bandwidth (HB) portion of the spectrum (HB > 8 kHz) is estimated in the next frame based on the SWB Band-Width Extension (BWE) technique of the last frame. The HB spectrum fades out over 40 frames, but the time-domain signal at the output sampling rate is used to perform the estimation of the SWB BWE parameters. On the other hand, in the case of a switch from WB to SWB, the HB portion of the spectrum fades out over 20 frames.

[0107] 3.2 Problem In IVAS, the BWS technique used in EVS can be implemented in the decoder, but it can never be applied due to the bitrate limitations of the EVS native BWS algorithm. Furthermore, the EVS native BWS algorithm does not support BWS in the TCX core. Finally, since time-domain signals cannot be used to perform algorithm estimation, the EVS native BWS algorithm cannot be applied to DFT stereo CNG (Comfort Noise Generation) frames.

[0108] 3.3 BWS in IVAS Therefore, a new and different BWS algorithm will be implemented in the IVAS audio codec.

[0109] First, such a BWS algorithm is implemented in the encoder portion of the IVAS audio codec. This choice has the advantage of the very low footprint complexity of the IVAS BWS algorithm compared to the native EVS one.

[0110] Another design choice is that the BWS algorithm in IVAS is implemented only for switching from lower BW to higher BW (e.g., switching from WB to SWB). In this direction, the switching is relatively fast (see section 2.3.4 above), and the resulting abrupt change in HF components can be bothersome. Therefore, a new, different BWS algorithm is designed to smooth such switching. On the other hand, in this direction, there are virtually no significant HF components in the spectrum, so the change in spectral components is not unnaturally abrupt and bothersome, and therefore no special handling is implemented for switching from higher BW to lower BW.

[0111] 3.4 Proposed BWS Figure 5 is a schematic block diagram showing simultaneously the method 500 and device 550 for audio bandwidth switching according to the present disclosure. As shown in Figure 5, the method for audio bandwidth switching includes a final audio bandwidth determination operation 307 and cnt bwidth_sw The device includes a counter update operation 502, a comparison operation 503, and a high-bandwidth spectral fade-in operation 504. Similarly, as shown in Figure 5, the device for audio bandwidth switching includes a final audio bandwidth determination module 357 for performing the final BWD determination operation 307, and cnt bwidth_sw The system includes a computer 552 for performing a counter update operation 502, a comparator 553 for performing a comparison operation 503, and an attenuator 554 for performing a high-bandwidth spectral fade-in operation 504.

[0112] The proposed BWS algorithm used by Method 500 and Device 550 in Figure 5 smooths the perceptual effects of audio bandwidth switching already present in the encoder portion of the IVAS audio codec while removing artifacts in synthesis. The high-bandwidth (HB > 8kHz) portion of the spectrum is attenuated in several consecutive frames after the BWS instance, as shown by the final audio bandwidth determination module 357. More specifically, the gain of the HB spectrum is faded in attenuator 554 and thus intelligently controlled in the case of BWS to avoid unpleasant artifacts. Since the attenuation is applied before the HB spectrum is quantized and encoded in the core encoder 555 and the corresponding core coding operation 505, the smoothed BW transition is already present in the transmitted bitstream 506, and no further processing is required in the decoder. For example, in the case of audio bandwidth switching from WB to SWB, the HB spectrum corresponding to frequencies above 8kHz is smoothed before further processing. In other words, audio bandwidth switching is inherent to the coded audio signal, no extra bits related to audio bandwidth switching are sent to the decoder, and no additional processing is performed by the decoder regarding audio bandwidth switching.

[0113] 3.4.1 BWS technique The BWS mechanism for the audio bandwidth switching method and device shown in Figure 5 functions as follows:

[0114] First, the computer 552, based on the final BWD determination 307, calculates the counter cnt for each IVAS transport channel at the end of preprocessing when audio bandwidth switching occurs and attenuation is applied to frames, as follows: bwidth_sw Update.

[0115] Calculator 552 counts the frame counter cnt bwidth_sw The value of is initially set to its initial value of "0". In response to the final BWD determination from the final audio bandwidth determination module 357, when a BW change from a lower audio bandwidth to a higher audio bandwidth is detected, typically a BW change from WB to SWB or FB, the value of the frame counter is incremented by 1. In the following frame, the counter is incremented by its maximum value B as defined below. tran The counter is incremented by 1 each frame until it reaches its maximum value B. tran When it reaches this point, the counter is reset to 0, and a new BW switch detection can occur.

[0116] In the source code, newly added code (marked by the "###" sequence) may look like this: An excerpt of the code can be found at the end of the IVAS audio codec function core_switching_pre_enc(). ### / *---------------------------------------------------------------------* ### * band-width switching from WB -> SWB / FB ### *---------------------------------------------------------------------* / ### ### if( st->bwidth_sw_cnt == 0 ) ### { ### if( st->bwidth >= SWB && st->last_bwidth == WB ) ### { ### st->bwidth_sw_cnt++; ###} ###} ### else ### { ### st->bwidth_sw_cnt++; ### ### if ( st->bwidth_sw_cnt == BWS_TRAN_PERIOD ) ### { ### st->bwidth_sw_cnt = 0; ###} ###}

[0117] Next, the counter cnt, which has been updated or not updated by the computer 552. bwidth_sw However, if it is greater than 0 as determined by comparator 553, attenuator 554 modifies the sound signal in frame i, for example, as follows:

[0118]

number

[0119] The damping coefficient β is defined as shown above. i Apply (507), where cnt bwidth_sw This is the audio bandwidth switching frame counter (bwidth_sw_cnt in the source code above), and B tran (The macro BWS_TRAN_PERIOD in the source code above) is the BWS transition period, which corresponds to the number of frames to which attenuation is applied after a BW switch from a lower BW to a higher BW. Constant B tranIt was discovered experimentally and set to 5 in the IVAS framework.

[0120] Figure 6 is a graph showing the actual values ​​of the attenuation coefficient β in the frame after the BW change is detected in the IVAS operating in MDCT stereo mode. In the unrestricted example of Figure 6, the BW change is detected in the fastest possible time (i.e., 10 frames of hysteresis), the final BWD determination is made in the next frame (n+11), and the BWS is in the next B tran =Assuming it is applied over 5 frames (frames n+12 to n+16). Finally, the damping coefficient β is B depending on the coding mode as follows. tran Applied within the frame.

[0121] In TCX and HQ core frames (HQ represents a high-quality MDCT coder in EVS, see section 5.3.4 of reference [1]), spectral X of length L as defined in section 5.3.2 of reference [1] M The high-band gain of (k) is controlled, and the spectral X immediately after the time-domain to frequency-domain conversion is M The high-bandwidth (HB) portion of (k) is related to, for example, the following relationship X' M (k+L WB )=β i *X M (k+L WB ), i=0,...,B tran -1 It is updated (faded in) by attenuator 554, where L WB L is the spectral length corresponding to the WB audio bandwidth, i.e., in an example of an IVAS with a frame length of 20 milliseconds (normal HQ, or TCX20 frame), L WB =320 samples, and in the temporary frame L WB =80 samples, and in TCX10 frames L WB = 160 samples, where k is in the range [0, KL]. WBis the sample index in [-1], where K is the length of the entire spectrum in a specific conversion sub-mode (usually transient, TCX20, TCX10).

[0122] In the ACELP core with a time-domain BWE (TBE) frame, the attenuator 554 applies the attenuation coefficient β to these parameters before the SWB gain shape parameters of the HB part of the spectrum are additionally processed. The time gain shape parameter gs(j) is defined in section 5.2.6.1.14.2 of reference [1] and consists of four values. Thus, in an example implementation, i gs'(j) = β *gs(j), i = 0,..., B i -1 tran where j = 0,..., 3 are the gain shape numbers.

[0123] In the ACELP core with a frequency-domain BWE (FD-BWE) frame, the high-band gain of the original input signal X(k) of length L defined in section 5.2.6.2.1 of reference [1] is controlled, and the HB part of the MDCT spectrum is updated by the attenuator 554 using, for example, the following relationship M M X' (k + L WB ) = β i *X M (k + L WB )、i = 0,..., B tran -1 It should be noted that NB coding is not considered in IVAS, and the switch from SWB to FB is not handled because its subjective and objective effects can be ignored. However, the same principle as above can be used to cover all BWS scenarios.

[0124]

[0125] ​​​Next, the attenuated sound signal from the attenuator 554 is encoded in the core encoder 555. The counter cnt, which has been updated or not, is then processed by the computer 552. bwidth_sw However, if the value is not greater than 0 as determined by the comparator 553, the sound signal is encoded in the core encoder 555 without attenuation.

[0126] 3.4.2 Examples of the effects of BWS Figure 7 shows an example waveform illustrating the effect of the BWS mechanism on decoding quality. Specifically, Figure 7 shows a segment of the audio signal (in this example, 0.3 seconds long) where the BW change from WB to SWB occurs in the highlighted portion. From top to bottom, Figure 7 shows: (1) the input signal waveform, (2) BW parameters (value 1 corresponds to WB, value 2 corresponds to SWB), (3) the decoded-synthesized waveform without BWS, (4) the decoded-synthesized spectrum without BWS, (5) the decoded-synthesized waveform with BWS applied, and (6) the decoded-synthesized spectrum with BWS applied. Furthermore, as highlighted by the arrows in Figure 7, it may be observed that the decoded-synthesized waveform with BWS applied is unaffected by the rapid energy increase in the time domain at HF ​​in the frequency domain. As a result, when the BWS technique disclosed herein is used, artifacts (annoying clicks) are removed from the synthesis.

[0127] 4. Hardware implementation Figure 8 is a simplified block diagram of an exemplary configuration of hardware components forming the encoder portion of the IVAS audio codec 305 described above, which uses an audio bandwidth detection method and device and an audio bandwidth switching method and device.

[0128] The encoder portion of the IVAS audio codec 305 using the audio bandwidth detection method and device and the audio bandwidth switching method and device may be implemented as part of a mobile terminal, as part of a portable media player, or in any similar device. The encoder portion of the IVAS audio codec 305 using the audio bandwidth detection method and device and the audio bandwidth switching method and device (identified as 800 in Figure 8) comprises an input 802, an output 804, a processor 806, and memory 808.

[0129] Input 802 is configured to receive the input audio signal 320 of Figure 3B in digital or analog format. Output 804 is configured to supply the output, coded audio signal. Input 802 and output 804 can be implemented in a common module, such as a serial input / output device.

[0130] Processor 806 is operably coupled to input 802, output 804, and memory 808. Processor 806 is implemented as one or more processors for executing code instructions to support the functionality of various components of the encoder portion of the IVAS audio codec 305, which use audio bandwidth detection methods and devices and audio bandwidth switching methods and devices as shown in Figure 3B.

[0131] Memory 808 may include non-temporary memory for storing code instructions executable by processor 806, specifically, processor-readable memory that, when executed, causes the processor to implement the operation and components of the encoder portion of the IVAS audio codec 305 described above, using the audio bandwidth detection method and device and the audio bandwidth switching method and device described herein. Memory 808 may also include random-access memory or buffers for storing intermediate processing data from various functions performed by processor 806.

[0132] Those skilled in the art will understand that the description of the encoder portion of the IVAS audio codec 305 using the audio bandwidth detection method and device and the audio bandwidth switching method and device is merely illustrative and not intended to limit in any way. Other embodiments will readily suggest themselves to those skilled in the art who are interested in this disclosure. Furthermore, the disclosed encoder portion of the IVAS audio codec 305 using the audio bandwidth detection method and device and the audio bandwidth switching method and device can be customized to provide a valuable solution to existing needs and problems of encoding and decoding sound.

[0133] For clarity, this disclosure does not describe or explain all the routine functions of an implementation of the encoder portion of the IVAS audio codec 305 that uses an audio bandwidth detection method and device and an audio bandwidth switching method and device. Of course, it will be understood that in the development of any actual implementation of the encoder portion of the IVAS audio codec 305 that uses an audio bandwidth detection method and device and an audio bandwidth switching method and device, many implementation-specific decisions may need to be made to achieve the developer's specific goals, such as compliance with application, system, network, and business-related constraints, and that these specific goals will differ from implementation to implementation and from developer to developer. Furthermore, it will be understood that development efforts, while complex and time-consuming, are nevertheless routine engineering work for those skilled in the art in the field of audio processing who are interested in this disclosure.

[0134] According to this disclosure, the components / processors / modules, processing operations, and / or data structures described herein may be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. In addition, those skilled in the art will recognize that less general-purpose devices such as hardwired devices, field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs) may also be used. If a method comprising a series of operations and suboperations is implemented by a processor, computer, or machine, and those operations and suboperations can be stored as a series of non-temporary code instructions readable by the processor, computer, or machine, they may be stored on tangible and / or non-temporary media.

[0135] The encoder portion of the IVAS audio codec 305 using the audio bandwidth detection method and device and the audio bandwidth switching method and device described herein may use software, firmware, hardware, or any combination of software, firmware, or hardware suitable for the purposes described herein.

[0136] In the encoder portion of the IVAS audio codec 305 using the audio bandwidth detection method and device and the audio bandwidth switching method and device described herein, various operations and suboperations may be performed in various orders, and some of the operations and suboperations may be optional.

[0137] While the present disclosure has been described above by its non-limiting and exemplary embodiments, these embodiments may be modified at will within the scope of the appended claims without departing from the spirit and nature of the present disclosure.

[0138] 5.References This disclosure references the following references, the entire contents of which are incorporated herein by reference. (References) [1] 3GPP TS 26.445, v.16.1.0, “Codec for Enhanced Voice Services (EVS); Detailed Algorithmic Description”, July 2020. [2] V. Eksler, M. Jelinek, and W. Jaegers, "Audio Bandwidth Detection in the EVS Codec," in Proc. IEEE Global Conf. on Signal and Information Processing (GlobalSIP), Orlando, FL, USA, 2015. [3] F. Baumgarte, C. Faller, "Binaural cue coding - Part I: Psychoacoustic fundamentals and design principles," IEEE Trans. Speech Audio Processing, vol. 11, pp. 509-519, Nov. 2003. [4] T. Vaillancourt, “Method and system using a long-term correlation difference between left and right channels for time domain down mixing a stereo sound signal into primary and secondary channels,” PCT Application WO2017 / 049397A1. [5] 3GPP SA4 contribution S4-170749, “New WID on EVS Codec Extension for Immersive Voice and Audio Services”, SA4 meeting #94, June 26-30, 2017, http: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_94 / Docs / S4-170749.zip [6] V. Pulkki, C. Faller, "Directional audio coding: Filterbank and STFT-based design," in 120th AES Convention, Paper 6658, Paris, May 2006. [7] M. Neuendorf et al., “MPEG Unified Speech and Audio Coding - The ISO / MPEG Standard for High-Efficiency Audio Coding of all Content Types”, Journal of the Audio Engineering Society, vol. 61 n° 12, pp. 956-977, December 2013. [8] J. Herre et al., “MPEG-H Audio - The New Standard for Universal Spatial / 3D Audio Coding”, in 137th International AES Convention, Paper 9095, Los Angeles, October 9-12, 2014. [9] 3GPP SA4 contribution S4-180462, “On spatial metadata for IVAS spatial audio input format”, SA4 meeting #98, April 9-13, 2018, https: / / www.3gpp.org / ftp / tsg_sa / WG4_CODEC / TSGS4_98 / Docs / S4-180462.zip [Explanation of Symbols]

[0139] 300 EVS codecs 301 BWD 302 Pre-treatment stage 303 Front pre-processing stage, audio signal pre-processing stage 304 Core coding stage, audio signal core coding stage 305 IVAS codec, IVAS audio codec 306 BWD analysis operation, BDW analysis operation 307 Final BWD determination operation, final BWD determination operation, final audio bandwidth determination operation 310 Input audio signal 320 Input audio signal, audio signal 356 BWD Analyzer 357 Final audio bandwidth determination module (processor), final audio bandwidth determination module, module, determination module 358 TCX core encoders 506 Transmit bitstream 550 devices 552 Calculator 553 Comparator 554 Attenuator 555 Core Encoder 802 Input 804 output 806 Processor 808 memory

Claims

1. An audio bandwidth detection device for detecting the audio bandwidth of an audio signal to be coded in the encoder portion of an audio codec, wherein the encoder portion includes an audio signal front preprocessing stage prior to the Modified Discrete Cosine Transform (MDCT) core coding stage. A sound signal analyzer for the sound signal, integrated into the MDCT core coding stage, for analyzing the MDCT spectrum of the sound signal, A final audio bandwidth determination module, integrated into the front preprocessing stage, delivers a final determination regarding the detected audio bandwidth using the results of the analysis of the MDCT spectrum of the sound signal. The audio codec comprises, in the encoder portion, the final audio bandwidth determination module, integrated into the front preprocessing stage, positioned upstream of the sound signal analyzer, integrated into the MDCT core coding stage, wherein the result of the analysis of the MDCT spectrum of the sound signal by the sound signal analyzer in the current frame is used by the final audio bandwidth determination module in the next frame, and the next frame follows the current frame to deliver the final determination of the detected audio bandwidth of the sound signal. Audio bandwidth detection device.

2. The sound signal analyzer analyzes the MDCT spectrum of the sound signal in MDCT stereo mode in the MDCT core coding stage of the encoder portion of the sound codec without calculating Complex Low Delay Filter Bank (CLDBB) values ​​that are not required in MDCT stereo mode in the front preprocessing stage of the encoder portion of the sound codec. The audio bandwidth detection device according to claim 1.

3. The audio bandwidth detection device according to claim 1, wherein the sound signal analyzer calculates the average value of the energy of the MDCT spectrum of the sound signal in several spectral regions.

4. The audio bandwidth detection device according to claim 3, wherein the sound signal analyzer calculates the maximum energy of the MDCT spectrum of the sound signal in several spectral regions.

5. The audio bandwidth detection device according to claim 4, wherein the sound signal analyzer calculates the energy of the MDCT spectrum of the sound signal in a plurality of frequency bands, the spectral region is defined by at least one of the frequency bands, and the sound signal analyzer uses the calculated energy of the MDCT spectrum of the sound signal in the frequency bands to calculate the mean and maximum values ​​of the energy of the MDCT spectrum.

6. The audio bandwidth detection device according to claim 3, wherein the sound signal analyzer calculates the long-term value of the average value of the energy of the MDCT spectrum of the sound signal in a region among several spectral regions.

7. The audio bandwidth detection device according to claim 1, wherein the sound signal analyzer updates counters related to several spectral regions.

8. The audio bandwidth detection device according to claim 6, wherein the sound signal analyzer calculates the maximum energy of the MDCT spectrum of the sound signal in several spectral regions, and the sound signal analyzer increases or decreases a counter associated with each of the spectral regions in response to the long-term value of the average energy of the MDCT spectrum of the sound signal and the maximum energy of the MDCT spectrum of the sound signal.

9. The audio bandwidth detection device according to any one of claims 3 to 8, wherein the sound signal analyzer performs sound signal analysis in frames of a given duration and skips sound signal analysis in frames longer or shorter than the given duration.

10. The audio bandwidth detection device according to claim 7 or 8, wherein the final audio bandwidth determination module uses determination logic for switching between audio bandwidths in response to a comparison between the counter and a given threshold.

11. The audio bandwidth detection device according to claim 10, wherein the determination logic of the final audio bandwidth determination module also responds to previously determined audio bandwidths.

12. The audio bandwidth detection device according to claim 10, wherein the final audio bandwidth determination module uses hysteresis to avoid frequent switching between audio bandwidths.

13. The audio bandwidth detection device according to claim 12, wherein the hysteresis used by the final audio bandwidth determination module is shorter in the case of a potential switch from a lower audio bandwidth to a higher audio bandwidth and longer in the case of a potential switch from a higher audio bandwidth to a lower audio bandwidth.

14. The audio bandwidth detection device according to any one of claims 3 to 8, wherein the sound signal is a multi-channel signal including multiple channels, and the final audio bandwidth determination module codes the detected audio bandwidth of the channels as a common parameter.

15. An audio bandwidth detection method for detecting the audio bandwidth of an audio signal to be coded in the encoder portion of an audio codec, wherein the encoder portion includes an audio signal front preprocessing stage before a modified discrete cosine transform (MDCT) core coding stage, and the audio bandwidth detection method is The MDCT core coding stage includes the step of analyzing the MDCT spectrum of the sound signal, In the front preprocessing stage, the results of the analysis of the MDCT spectrum of the sound signal are used to make a final determination regarding the detected audio bandwidth. The encoder portion of the audio codec includes a step in the front preprocessing stage to make a final determination regarding the detected audio bandwidth, which is performed upstream of the analysis of the MDCT spectrum of the audio signal in the MDCT core coding stage, wherein the result of the analysis of the MDCT spectrum of the audio signal in the current frame is used in the next frame, and the next frame follows the current frame to deliver the final determination regarding the detected audio bandwidth of the audio signal. Audio bandwidth detection method.

16. In the front preprocessing stage of the encoder portion of the audio codec, the Complex Low Delay Filter Bank (CLDBB) values ​​that are not required in MDCT stereo mode are not calculated, and in the MDCT core coding stage of the encoder portion of the audio codec, the MDCT spectrum of the audio signal is analyzed in MDCT stereo mode. The audio bandwidth detection method according to claim 15.

17. The audio bandwidth detection method according to claim 15, wherein the analysis of the MDCT spectrum of the sound signal includes the step of calculating the average energy of the MDCT spectrum of the sound signal in several spectral regions.

18. The audio bandwidth detection method according to claim 17, wherein the analysis of the MDCT spectrum of the sound signal includes the step of calculating the maximum energy of the MDCT spectrum of the sound signal in several spectral regions.

19. The audio bandwidth detection method according to claim 18, wherein the analysis of the MDCT spectrum of the sound signal includes the step of calculating the energy of the MDCT spectrum of the sound signal in a plurality of frequency bands, each of which is defined by at least one of the frequency bands, and the analysis of the MDCT spectrum of the sound signal includes the step of using the calculated energy of the MDCT spectrum of the sound signal in the frequency bands to calculate the mean and maximum values ​​of the energy of the MDCT spectrum.

20. The audio bandwidth detection method according to claim 17, wherein the analysis of the MDCT spectrum of the sound signal includes the step of calculating a long-term value of the average value of the energy of the MDCT spectrum of the sound signal in a region of several spectral regions.

21. The audio bandwidth detection method according to claim 15, wherein the analysis of the MDCT spectrum of the sound signal includes the step of updating counters related to several spectral regions.

22. The audio bandwidth detection method according to claim 20, wherein the analysis of the MDCT spectrum of the sound signal comprises the step of calculating the maximum energy of the MDCT spectrum of the sound signal in several spectral regions, and the analysis of the MDCT spectrum of the sound signal comprises the step of increasing or decreasing a counter associated with each of the spectral regions in response to the long-term value of the mean value of the MDCT spectrum of the sound signal and the maximum energy of the MDCT spectrum of the sound signal.

23. The audio bandwidth detection method according to any one of claims 17 to 22, wherein the analysis of the MDCT spectrum of the sound signal is performed in frames of a given duration and skipped in frames longer or shorter than the given duration.

24. The audio bandwidth detection method according to claim 21 or 22, wherein the step of finally determining the detected audio bandwidth includes using a decision logic for switching between audio bandwidths in response to a comparison between the counter and a given threshold.

25. The audio bandwidth detection method according to claim 24, wherein the decision logic also responds to previously determined audio bandwidths.

26. The audio bandwidth detection method according to claim 24, wherein the step of finally determining the detected audio bandwidth includes the step of using hysteresis to avoid frequent switching between audio bandwidths.

27. ​​The audio bandwidth detection method according to claim 26, wherein the hysteresis is shorter in the case of a potential switch from a lower audio bandwidth to a higher audio bandwidth and longer in the case of a potential switch from a higher audio bandwidth to a lower audio bandwidth.

28. The audio bandwidth detection method according to any one of claims 17 to 22, wherein the sound signal is a multi-channel signal including a plurality of channels, and the step of finally determining the detected audio bandwidth includes the step of coding the detected audio bandwidth of the channels as a common parameter.