Improved transitions in multi-mode audio decoder
By using low-complexity analysis to initialize envelope stability parameters in a multi-mode audio codec, the high computational complexity and memory update problems during mode switching are solved, and more efficient mode conversion and memory update are achieved.
Patent Information
- Application Number
- CN202511119768.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-13
- Filing Date
- 2023-12-12
- Publication Date
- 2025-09-19
AI Technical Summary
When a multi-mode audio codec switches between codec modes, there are problems of high computational complexity and memory reinitialization, resulting in non-smooth mode transitions.
By using low-complexity analysis such as energy analysis and spectral shape analysis when switching between codec modes, the envelope stability parameters of the codec mode are initialized and the memory is kept updated.
The invention realizes smooth conversion between encoding and decoding modes, reduces computational complexity, keeps the memory updated, and improves the efficiency of mode conversion.
Smart Images

Figure CN120673771A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to communications, and more particularly to communication methods and related devices and nodes supporting wireless communications. Background Art
[0002] Modern audio codecs are designed to compress a wide variety of input audio signals. For low bit rate coding, it has been shown that it is beneficial to utilize several audio codecs designed to process different types of audio signals. Each audio codec corresponds to a specific operating mode of the audio encoder, and therefore corresponds to the term multi-mode audio encoder. For example, speech signals are typically encoded using a codec mode based on a speech model, such as algebraic code excited linear prediction (ACELP), while conventional audio signals (e.g., music) are better captured using a transform-based codec mode, such as transform coding residual (TCX) based on modified discrete cosine transform (MDCT). Several examples of audio codecs based on this principle exist, such as 3GPP 26.290 AMR-WB+, ISO / IEC 23003-3 MPEG-D USAC, and 3GPP 26.445 EVS.
[0003] There are certain challenges. A challenge with multi-mode audio codecs is handling the transitions between different codec modes. While a codec mode handles its designated audio signal type very well, it can exhibit specific signatures or types of distortion, and the differences in signatures between codec modes become apparent when switching between modes. Therefore, handling the transitions between modes is a fundamental task in a multi-mode audio codec. Because codec modes are inherently different, they also contain different signal processing tools, which can include memories of past encoded audio. When switching between modes, these memories need to be reinitialized to avoid old or outdated memory being used after the transition. Summary of the Invention
[0004] One solution is to run all encoding modes in parallel, thereby keeping all encoding modes and their memories up to date. However, in most cases, this solution is computationally too complex.
[0005] Another challenge is that the memory or state for one codec mode may not exist in other codec modes. Performing the required analysis to keep the memory updated may also require high computational effort.
[0006] Certain aspects of the present disclosure and embodiments thereof may provide solutions to these and other challenges. Various embodiments initialize envelope stability parameters in one coding mode based on low-complexity analysis performed in another coding mode when switching between coding modes. Low-complexity analysis includes, but is not limited to, energy analysis and spectral shape analysis. When switching to the coding mode, the results of the low-complexity analysis are used to initialize the key memory for the coding mode.
[0007] According to a first aspect, a method for decoding an encoded audio frame in a decoder is provided, the audio frame being encoded using one of at least two modes. The method comprises: receiving information indicating a selected codec mode and determining whether the selected codec mode is a first mode. In response to determining that the selected codec mode is the first mode, determining whether the previous codec mode is the first mode. In response to the selected codec mode being the first mode and the previous codec mode not being the first mode, estimating an envelope stability metric using energy stability and shape stability of a previous frame. Decoding the encoded audio frame using the estimated envelope stability metric, and determining energy stability of a current frame and shape stability of the current frame.
[0008] According to a second aspect, a method for decoding encoded audio in a decoder is provided, the encoded audio being encoded using multiple modes. The method comprises receiving information indicating a selected codec mode and determining whether a current mode is a first mode and whether a previous mode was not the first mode. In response to determining yes, estimating an envelope stability metric using energy stability and shape stability of a previous frame. The encoded audio frame is decoded based on the current mode, and energy stability of a current frame and shape stability of the current frame are determined.
[0009] According to a third aspect, there is provided a decoder for decoding encoded audio, the encoded audio being encoded using a plurality of modes, the decoder being adapted to perform the method according to the first aspect or the second aspect.
[0010] According to a fourth aspect, a decoder is provided for decoding encoded audio, the encoded audio being encoded using multiple modes. The decoder comprises: a processing circuit; and a memory coupled to the processing circuit, wherein the memory comprises instructions that, when executed by the processing circuit, cause the decoder to perform the operations according to the first aspect or the second aspect.
[0011] According to a fifth aspect, there is provided a computer program comprising program code to be executed by processing circuitry of a decoder, wherein execution of the program code causes the decoder to perform operations according to the first or second aspect.
[0012] According to a sixth aspect, a computer program product is proposed, which comprises a non-transitory storage medium comprising program code to be executed by a processing circuit of a decoder, wherein execution of the program code causes the decoder to perform operations according to the first aspect or the second aspect.
[0013] According to a seventh aspect, a decoder is provided for decoding an encoded audio frame, which is encoded using one of at least two modes. The decoder is configured to: receive information indicating a selected codec mode, and determine whether the selected codec mode is a first mode. In response to determining that the selected codec mode is the first mode, the decoder is configured to determine whether the previous codec mode is the first mode. In response to the selected codec mode being the first mode and the previous codec mode not being the first mode, the decoder is configured to estimate an envelope stability metric using energy stability and shape stability of a previous frame. The encoded audio frame is decoded using the estimated envelope stability metric, and the energy stability of a current frame and the shape stability of the current frame are determined.
[0014] Certain embodiments may provide one or more of the following technical advantages. Advantages that may be achieved include improved conversion between modes of a multi-mode decoder because the memory remains updated during conversion between modes. These advantages may be achieved with little impact on computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application, illustrate certain non-limiting embodiments of the inventive concept. In the drawings:
[0016] Figure 1 are examples of operating environments in which various embodiments of the present disclosure may be implemented, according to some embodiments;
[0017] Figure 2 is a graphical representation of a sigmoid function according to some embodiments;
[0018] Figure 3 is a block diagram of a decoder according to some embodiments of the present disclosure;
[0019] Figure 4 is a diagram showing some embodiments of the present disclosure Figure 3 A block diagram of a decoder with additional details;
[0020] Figures 5 and 6 is a flowchart illustrating the operation of a decoder according to some embodiments of the present disclosure;
[0021] Figure 7 is a block diagram of a user equipment according to some embodiments;
[0022] Figure 8 is a block diagram of a host computer in communication with a user device according to some embodiments; and
[0023] Figure 9 is a block diagram of a virtualization environment according to some embodiments. DETAILED DESCRIPTION
[0024] Some embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. The embodiments are provided by way of example to convey the scope of the present subject matter to those skilled in the art, wherein examples of embodiments of the inventive concept are shown. However, the inventive concept can be embodied in many different forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope of the inventive concept to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Components from one embodiment may be appropriately assumed to exist / be used in another embodiment.
[0025] As previously indicated, a challenge with multi-mode audio codecs is handling the transitions between different codec modes. While a codec mode handles its designated audio signal type very well, it can exhibit specific signatures or types of distortion, and the differences in signatures between codec modes become apparent when switching between modes. Therefore, handling transitions between modes is a fundamental task in multi-mode audio codecs. Because codec modes are inherently different, they also contain different signal processing tools, which can include memories of past encoded audio. When switching between modes, these memories need to be reinitialized to avoid old or outdated memory being used after the transition.
[0026] Before describing embodiments that address the challenges, an example of an operating environment will be described. Figure 1 An example of an operating environment in which various embodiments of the present disclosure may be implemented is shown. Figure 1In the example operating environment 100, an encoder 102 receives data to be encoded (such as an audio file) from an entity, a storage device 108, and / or an audio recorder connected to a microphone 105 via a network 104 (such as a host 106). In some embodiments, the host 106 can communicate directly with the encoder 102 and include the audio recorder 105. The encoder 102 encodes the audio file as described herein and stores the encoded audio file in the storage device 108 or sends the encoded audio file to a decoder 112 via a network 110. The decoder 112 decodes the audio file and sends the decoded audio file to an audio player 114 for playback. The audio player 114 can be or be included in a user device, a terminal, a mobile phone, etc. In other embodiments, the host 106 can send the encoded audio file to the decoder 112 via the network 110.
[0027] The audio codec system consists of two main parts, an audio encoder 102 and a decoder 112. The input audio is processed in time segments called frames x(m,n), which consist of samples n=0, 1, 2, ...L-1 in frame m. Frames can be extracted using overlap, making the analysis frame longer than the output synthesis of each frame. Based on the properties of the audio signal in frame x(m,n), a codec mode is selected to encode it. Mode selection can be done by using a signal classifier to decide which codec mode will achieve the best performance, or in a so-called closed-loop manner, where all available codec modes are run and the best performing mode is selected.
[0028] Various embodiments will be described using three codec modes, and the three codec modes will be denoted as Mode A, Mode B, and Mode C.
[0029] Mode A - a codec mode based on Modified Discrete Cosine Transform (MDCT).
[0030] In this mode, the input frame x(m,n) is transformed into the modified discrete cosine transform domain by the following equation: where w a (n) is the analysis window.
[0031] It should be noted that the frame length used in the transform is twice the original frame length. However, the effective length of the input frame is given by w a The length of the non-zero portion of (n) is determined by the MDCT spectrum X(m,k) which now represents the MDCT coefficient k of frame m. The coefficients of the spectrum are divided into groups or bands. These bands are sized unevenly to simulate the frequency resolution of a human listener, using narrower bands for low frequencies and wider bandwidths for higher frequencies. Each band E(m,b) is calculated according to the following formula, with b = 0, 1, ..., Nband -1 Energy: where k start (b),…,k end (b) represents the index of frequency band b, and N band is the number of frequency bands.
[0032] The band energy is transformed into a logarithmic energy index I(m,b) using the following equation: I(m,b)=[34-2log2E(m,b)] Then quantized to be stored or sent to the decoder. Here, [·] represents a rounding operation.
[0033] The logarithmic energy index I(m,b) can be viewed as the inverse (negative) logarithmic energy with a scaling factor applied. The encoder reconstructs the band energy The energy of this frequency band It is also used to normalize the MDCT spectrum according to the following formula:
[0034] The normalized spectrum can be encoded using a suitable coding method such as a vector quantizer (VQ) or a scalar quantizer followed by an entropy codec such as an arithmetic codec. The encoding of the normalized spectrum Y(m,k) is based on a bit allocation R(m,b) that allocates the available bit budget to the frequency bands. The bit allocation is done using a perceptual model that aims to allocate bits to maximize perceptual performance. The perceptual model can use the reconstructed spectral envelope Or equivalently, the logarithmic energy index I(m,b) is used to allocate bits to the frequency band. The coding parameters, including the representation of the logarithmic energy index I(m,b), the normalized spectrum Y(m,k), and information indicating the selected codec mode, are combined into a bitstream to be stored or sent to a decoder.
[0035] Codec mode based on Mode B-ACELP
[0036] In this mode, encoding is performed in the time domain with the aid of the Algebraic Code Excited Linear Prediction (ACELP) codec method. This method relies on linear prediction analysis to obtain coefficients a representing the spectral envelope of frame m. j ,j=0,1,…,M linear predictor filter A m (z). The codec mode is based on A m (z) derive the weighted filter W m (z). It then performs a weighted synthesis by passing the coded excitation signal through a weighted synthesis filter To search for the best matching synthesis in the weighted domain, where is the reconstructed predictor filter. The encoding of the filter coefficients can be done in a domain more suitable for quantization, such as the line spectral frequency (LSF) domain. The encoded parameters of the representation and the information indicating the selected codec mode are combined into a bitstream to be stored or sent to a decoder.
[0037] Mode 2 encoding and decoding based on mode C-MDCT
[0038] This mode is similar to mode A, but has a different structure. For the purposes of this description, it is sufficient to mention that it does not have the corresponding envelope energy that mode A has.
[0039] Envelope stability measure in Mode A decoder
[0040] The decoder of mode A produces the logarithmic energy index I(m,b) and the reconstructed normalized spectrum For bands b that receive zero bits in the bit allocation R(m,b)=0, a noise filling algorithm is used. For low bit rate bands and noise filled bands, adaptive attenuation is applied. This is done based on an envelope stability metric denoted as env_stab(m). The envelope stability metric represents the changes in the spectral envelope, including changes in both energy and shape. It can also be called a spectral stability metric. The envelope stability metric is based on the band energies of the current and previous frames determined according to the following formula:
[0041] This difference is low-pass filtered to form a long-term estimate of the log energy change. D LP (m) = αD(m) + (1-α)D LP (M-1)
[0042] Here, α is the low-pass filter coefficient, where a suitable value may be α = 0.1 or in the range α∈[0.01,0.5]. The envelope stability metric env_stab(m) is obtained by converting D LP (m) is mapped to the range [0,1] to provide a smooth transition. The constants b, c, and d can be set to b=6.11, c=1.91, and d=2.26.
[0043] An alternative expression for this transformation is Where -a / b is the midpoint of the transition, with env_stab(m) = 0.5, and suitable values for constant a and constant b may be a = -15.7 and b = 6.11, which yields -a / b = 2.57.
[0044] This is a sigmoid function, which can be viewed as a soft threshold function with a crossing point at -a / b. LP (m) means strong changes, resulting in low env_stab(m). A graphical representation of this function can be found at Figure 2 Found in.
[0045] The function can be sampled discretely, which allows the transformation to be implemented by lookup in a table. Note that env_stab(m) captures changes in both energy and spectral shape.
[0046] Shape stability measures in mode B
[0047] Mode B is a linear predictor-based codec mode where the spectral shape is modeled by a linear predictor (LP) filter. The LP filter is represented in the line spectrum frequency (LSF) domain, which is suitable for quantization and interpolation of the LP filter. The shape stability factor stab_fac(m) is calculated according to the following formula: Among them L ModeB is the frame length of the LP coding band, and D LSF (m) is the Euclidean distance between the current frame LSF vector and the previous frame LSF vector. Since stab_fac(m) is based on the difference in LP filters between the current and previous frames, it captures changes in spectral shape but excludes changes in energy.
[0048] Multi-mode decoder - Mode A
[0049] Decoder 112 Figure 3 is shown in the block diagram and is Figure 4 As shown in more detail in FIG. 1 , in some embodiments, the decoder performs Figure 5 The decoder has a multi-mode decoder 310 in communication with a codec mode memory 320 that stores variables of past decoded frames of previous frames during decoding of the encoded audio frame, such as the logarithmic energy index I(m-1,b), the past decoded predictor filter A(z), the energy stability E Δ,LP (m-1) 410 and shape stability stab_fac_lt(m-1) 420. When using the energy stability estimator 430, the shape stability estimator 440 and the shape stability factor estimator 450, D LP When (m-1) is outdated or does not exist, the estimator 330 calculates the energy stability E of the previous frame based on the energy stability E of the previous frame. Δ,LP (m-1) and shape stability stab_fac_lt(m-1) to estimate.
[0050] In box 501, the multi-mode decoder 310 receives a packet from a bitstream representing an encoded audio frame, the packet including information indicating the selected codec mode CURRENT_MODE and the encoded parameters required for the multi-mode decoder to perform reconstruction of the encoded audio frame. The codec mode of the current frame is required to select an appropriate decoding method for the frame, and the codec mode of the current frame is determined in box 503. When processing of the current frame is completed, CURRENT_MODE is stored in the variable PREVIOUS_MODE for use in subsequent frames. In response to the current mode CURRENT_MODE=FIRST (i.e., mode A), the previous mode is checked in box 505. If the previous mode is also the first mode (e.g., PREVIOUS_MODE=FIRST), then in box 507, the decoder 112 continues to decode the current frame. When decoding a first mode frame, env_stab(m) is based on the logarithmic energy index I(m,b) from D LP (m) is calculated and used for decoding. If the previous mode is different from the first mode PREVIOUS_MODE≠FIRST, then D LP (m) cannot be calculated because D LP (m) and I(m-1,b) are outdated or do not exist because D LP (m) has not been updated for one or more frames. In this case, in block 509, the estimator 330 calculates the energy stability E of the previous frame according to the following formula: Δ,LP (m-1) and shape stability stab_fac_lt(m-1) to estimate: D LP (m) = D LP,est (m)=P1+P2stab_fac_lt(m-1)+P3E Δ,LP (m-1) where P1, P2 and P3 are constants, stab_fac_lt(m-1) is the shape stability implemented as the long-term estimate of the shape stability factor from frame m-1, and E Δ,LP (m-1) is the energy stability, implemented as a long-term estimate of the absolute logarithmic energy difference between the synthesized frames estimated in frame m-1.
[0051] Note that the shape stability and energy stability of the previous frame m-1 need to be used because the updated values need to be decoded for the current frame m. The constants P1, P2 and P3 can be set experimentally, for example, based on stab_fac_lt(m) and E for a test database running the first mode (i.e., mode A). Δ,LP (m), using the least squares approximation to match D LP (m)(in D LP(m) is available). Another approach is to use machine learning techniques, such as a linear regression model using a representative database with cross validation. The coefficients from this model are P1 = 2.93, P2 = -2.20, and P3 = 0.741. More sophisticated mapping functions can also be used, but D is usually estimated LP,est (m) is the energy stability E Δ,LP (m-1) and the function of shape stability stab_fac_lt(m-1), that is, D LP,est (m) = f(E Δ,LP (m-1),stab_fac_lt(m-1))
[0052] Then, as mentioned before, env_stab(m) is based on D LP (m) to be determined.
[0053] The determined energy stability and shape stability are stored in the memory 320 together with other memories of the multi-mode decoder. Then, in block 507, decoding of the current first mode is performed using the estimated env_stab(m). In block 511, the energy stability E Δ,LP (m) is determined. Here, it is defined as the long-term estimate of the absolute logarithmic energy difference according to the following formula: E Δ,LP (m) = βE Δ (m)+(1-β)E Δ,LP (m-1) E Δ (m)=|E log (m)-E log (m-1)| where β is the low-pass filter coefficient, is the output composite of frame m and L out Indicates the output composite frame length. It can be the same as the input frame length L out = L, but it can also be different from the input if the decoder sampling rate is different from the encoder sampling rate. Note that the factor 1 / L_out will be used in the calculation for E Δ (m) is canceled out in the expression and can therefore be omitted. E Δ (m)=|E log (m)-E log (m-1)|=|E′ log (m)-E′ log (m-1)|
[0054] In block 513, the shape stability factor stab_fac_lt(m) is determined. Since the LP filter is not used in the first mode, the shape stability factor stab_fac(m) cannot be calculated based on the LP filter. However, an estimate of the shape stability factor can be determined by the estimator 330 as stab_fac_est(m)=Q1+Q2D LP (m)+Q3E Δ,LP (m) Where Q1, Q2, and Q3 are constants.
[0055] These constants may be set experimentally, for example, based on D for a test database running the second mode (ie, mode B) or the third mode (ie, mode C). LP (m) and E Δ,LP (m), using a least squares approximation to match stab_fac_est(m) with the true stab_fac(m) (where stab_fac(m) is available). The estimation can also be done by machine learning methods, such as training a linear regression model using a representative database with cross-validation. Suitable values for these constants can be Q1 = 1.093, Q2 = -5.84·10 -5 and Q3 = 0.125. Note that the result of stab_fac_est(m) can be stored in the same memory location as stab_fac(m) because the memory is not updated in the first mode. In other words, stab_fac_est(m) = stab_fac(m) in the first mode.
[0056] The shape stability is determined by low-pass filtering the estimated shape stability factor. stab_fac_lt(m)=γstab_fac_est(m)+(1-γ)stab_fac_lt(m-1) where γ is the low-pass filter coefficient, where a suitable value may be γ=0.1 or in the range γ∈[0.01,0.5]. For clarity, shape stability is defined as the shape stability factor of the low-pass filter across frames.
[0057] Once the multi-mode decoder has completed decoding of frame m, in block 515 the synthesized frame is output for playback by the audio player 114 or for storage in a decoded format such as pulse codec modulation (PCM).
[0058] Second mode (ie, mode B) or third mode (ie, mode C)
[0059] If the current mode is identified as the second mode in block 503, then in block 517, the multi-mode decoder 310 decodes the second mode. In block 511, the energy stability E is determined. Δ,LP Since the second mode is an ACELP-based mode, a shape stability factor stab_fac(m) is calculated based on the LP filter, and the shape stability is determined in block 513 according to the following formula: stab_fac_lt(m)=γstab_fac(m)+(1-γ)stab_fac_lt(m-1) where γ is the low-pass filter coefficient.
[0060] The determined energy stability and shape stability are stored in the memory 320 along with other memories of the multi-mode decoder 310 .
[0061] If the current mode is identified as the third mode in block 503, the multi-mode decoder 310 decodes the third mode in block 519 and determines the energy stability E in block 511. Δ,LP (m). The third mode is an MDCT based mode, but it still uses an LP filter and calculates a shape stability factor stab_fac(m). The shape stability is determined in the same way as was done for mode B in block 513 as described above.
[0062] Once the multi-mode decoder has completed decoding of frame m, the synthesized frame is output in block 515 for playback by the audio player 114 or for storage in a decoded format, such as pulse codec modulation (PCM).
[0063] Variants of energy calculations
[0064] The most computationally complex part of this method is the energy calculation, which is the energy stability E Δ,LP (m) is the basis of the energy. It may be beneficial to estimate this energy based on parameters that are already calculated or available in the decoder. For example, the pitch codebook gain and innovation codebook gain of the ACELP decoder can be used to estimate the frame energy. Moreover, the evolution of these parameters for several frames may be useful. In addition, the energy of the ACELP synthesized frame can be found by using the existing calculation of the residual energy together with the estimation of the prediction gain of the LP filter. If the ACELP coding mode uses a bandwidth extension (BWE) scheme, the energy of the BWE area is usually expressed as a ratio relative to the low-band energy of the ACELP coding band. The synthesized frame energy can also be calculated in the packet loss compensation (PLC) module, which can be reused for this purpose.
[0065] Figure 6 Some other embodiments using decoder 112 to perform multi-mode decoding are shown. Figure 6In block 601, the decoder receives a packet from a bitstream representing an encoded audio frame, including information indicating the selected codec mode CURRENT_MODE and the encoded parameters required for the multi-mode decoder to perform reconstruction of the encoded audio frame in block 607. The codec mode of the current frame is required to select an appropriate decoding method for the frame. When processing of the current frame is completed, CURRENT_MODE is stored in a variable PREVIOUS_MODE for use in subsequent frames. In response to the current mode CURRENT_MODE=FIRST and the previous mode is different from the first mode PREVIOUS_MODE≠FIRST, D LP (m-1) and I(m-1,b) are outdated or do not exist because they have not been updated for one or more frames. In this case, in block 605, the estimator 330 calculates the energy stability E of the previous frame based on the energy stability E of the previous frame as described above. Δ,LP (m-1) and shape stability stab_fac_lt(m-1) are estimated and proceed to block 607 to decode the current frame based on the current mode.
[0066] If it is determined that CURRENT_MODE=FIRST and PREVIOUS_MODE≠FIRST is false, then in block 607 , the decoder 112 continues decoding the current frame based on the current mode.
[0067] For example, if the current mode is the first mode, env_stab(m) is calculated based on the logarithmic energy index I(m,b) from D LP (m) is calculated and used for decoding. In block 609, the energy stability E is determined. Δ,LP (m). Here, it is defined as the long-term estimate of the absolute logarithmic energy difference according to the following formula: E Δ,LP (m) = βE Δ (m)+(1-β)E Δ,LP (m-1) E Δ (m)=|E log (m)-E log (m-1)| in is the output composite frame m and L out Indicates the output composite frame length. It can be the same as the input frame length L out = L, but it can also be different from the input if the decoder sampling rate is different from the encoder sampling rate. Note that the factor 1 / L_out will be used in the calculation for E Δ (m) is canceled out in the expression and can therefore be omitted. E Δ (m)=|E log (m)-E log (m-1)|=|E′ log (m)-E′ log (m-1)|
[0068] In block 611, the shape stability stab_fac_lt(m) is determined. Since the LP filter is not used in the first mode, the stability factor stab_fac(m) cannot be calculated based on the LP filter. However, an estimate of the stability factor can be estimated by the estimator 330 as stab_fac_est(m)=Q1+Q2D LP (m)+Q3E Δ,LP (m) Where Q1, Q2, and Q3 are constants.
[0069] These can be set experimentally, for example, based on D against a test database running the second or third mode. LP (m) and E Δ,LP (m), using a least squares approximation to match stab_fac_est(m) with the true stab_fac(m) (where stab_fac(m) is available). The estimation can also be done by machine learning methods, such as training a linear regression model using a representative database with, for example, 5-fold cross validation. Suitable values for these constants can be Q1 = 1.093, Q2 = -5.84·10 -5 = 0.125. Note that the result of stab_fac_est(m) can be stored in the same memory location as stab_fac(m) because the memory is not updated in the first mode. In other words, stab_fac_est(m) = stab_fac(m) in the first mode.
[0070] The shape stability is determined by low-pass filtering the estimated stability factor. stab_fac_lt(m)=γstab_fac_est(m)+(1-γ)stab_fac_lt(m-1) where γ is the low-pass filter coefficient, where a suitable value may be γ=0.1 or in the range γ∈[0.01,0.5]. In other words, shape stability is defined as the stability factor of the low-pass filtering across frames.
[0071] If the current mode is the second mode or the third mode, the shape stability is determined in block 611 according to the following equation: stab_fac_lt(m)=γstab_fac(m)+(1-γ)stab_fac_lt(m-1) where γ is the low pass filter coefficient. In block 613, the decoded frame is output.
[0072] Figure 7 An audio decoder 112 (e.g., decoder) according to some embodiments is shown, wherein the audio decoder 112 is implemented as a standalone device. As used herein, an audio decoder refers to a device capable of, configured, arranged, and / or operable to decode an encoded object and communicate with a network node, encoder, and / or decoder. Examples of audio decoders include, but are not limited to, smartphones, mobile phones, cellular phones, voice over IP (VoIP) phones, wireless local loop phones, desktop computers, personal digital assistants (PDAs), wireless cameras, game consoles or devices, storage devices, playback devices, wearable terminal devices, wireless endpoints, mobile stations, tablet computers, laptops, laptop embedded devices (LEEs), laptop devices (LMEs), smart devices, wireless customer premises equipment (CPEs), vehicle-mounted or vehicle embedded / integrated wireless devices, and the like.
[0073] The audio decoder may support device-to-device (D2D) communication, such as by implementing 3GPP standards for sidelink communication, dedicated short-range communication (DSRC), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), or vehicle-to-everything (V2X). In other examples, the decoder may not necessarily have a user in the sense of a human user owning and / or operating an associated device.
[0074] The audio decoder 112 includes a processing circuit 702 operatively coupled to an input / output interface 706, a power supply 708, a memory 710, a communication interface 712, and / or any other components or any combination thereof via a bus 704. Some decoders may utilize Figure 7 All or a subset of the components shown in . The level of integration between components can vary from one decoder to another. In addition, some decoders may contain multiple instances of components, such as multiple processors, memories, transceivers, transmitters, receivers, etc.
[0075] The processing circuit 702 is configured to process instructions and data and may be configured to implement any sequential state machine operable to execute instructions stored in the memory 710 as a machine-readable computer program. The processing circuit 702 may be implemented as one or more hardware-implemented state machines (e.g., in discrete logic, a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.); programmable logic along with appropriate firmware; one or more stored computer programs, a general-purpose processor (such as a microprocessor or a digital signal processor (DSP)) along with appropriate software; or any combination of the above. For example, the processing circuit 702 may include multiple central processing units (CPUs).
[0076] In this example, the input / output interface 706 can be configured to provide one or more interfaces to an input device, an output device, or one or more input and / or output devices. Examples of output devices include speakers, sound cards, video cards, displays, monitors, actuators, transmitters, smart cards, another output device, or any combination thereof. An input device can allow a user to collect information into the audio decoder 112. Examples of input devices include touch-sensitive or presence-sensitive displays, cameras (e.g., digital cameras, digital video cameras, webcams, etc.), microphones, sensors, direction pads, trackpads, scroll wheels, smart cards, etc. A presence-sensitive display can include a capacitive or resistive touch sensor to sense input from the user. The sensor can be, for example, an accelerometer, a gyroscope, a tilt sensor, a force sensor, a magnetometer, an optical sensor, a proximity sensor, a biometric sensor, etc., or any combination thereof. The output device can use the same type of interface port as the input device. For example, a universal serial bus (USB) port can be used to provide input devices and output devices.
[0077] In some embodiments, the power supply 708 is configured as a battery or battery pack. Other types of power supplies may be used, such as an external power source (e.g., a power outlet), a photovoltaic device, or a battery. The power supply 708 may also include a power circuit for delivering power from the power supply 708 itself and / or an external power source to various parts of the audio decoder 112 via an input circuit or an interface such as a power cable. The delivered power may be, for example, for charging the power supply 708. The power circuit may perform any formatting, conversion, or other modification on the power from the power supply 708 to make the power suitable for the corresponding components of the audio decoder 112 to which it is supplied.
[0078] The memory 710 may be or be configured to include a memory such as a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic disk, an optical disk, a hard disk, a removable cartridge, a flash drive, etc. In one example, the memory 710 includes one or more application programs 714 (such as an operating system, a web browser application, a widget, a widget engine, or other application) and corresponding data 716. The memory 710 may store any of a variety of operating systems or combinations of operating systems for use by the audio decoder 112.
[0079] The memory 710 may be configured to include multiple physical drive units, such as a redundant array of independent disks (RAID), flash memory, a USB flash drive, an external hard drive, a thumb drive, a pen drive, a key drive, a high-density digital versatile disk (HD-DVD) optical drive, an internal hard drive, a Blu-ray disc drive, a holographic digital data storage (HDDS) optical drive, an external miniature dual in-line memory module (DIMM), synchronous dynamic random access memory (SDRAM), an external miniature DIMM SDRAM, a smart card memory (such as a tamper-resistant module in the form of a universal integrated circuit card (UICC) including one or more subscriber identity modules (SIMs) such as a USIM and / or an ISIM), other memories, or any combination thereof. The UICC may be, for example, an embedded UICC (eUICC), an integrated UICC (iUICC), or a removable UICC commonly referred to as a "SIM card." The memory 710 may allow the audio decoder 112 to access instructions, applications, and the like stored on a transient or non-transitory memory medium to offload data or upload data. An article of manufacture, such as one utilizing a communication system, may be tangibly embodied as or in memory 710 , which may be or include a device-readable storage medium.
[0080] The processing circuit 702 can be configured to communicate with an access network or other network using a communication interface 712. The communication interface 712 may include one or more communication subsystems and may include an antenna 722 or be communicatively coupled to the antenna 722. The communication interface 712 may include one or more transmitters for communicating, such as by communicating with one or more remote transmitters of another device capable of wireless communication (e.g., another UE or a network node in an access network). Each transceiver may include a transmitter 718 and / or a receiver 720 suitable for providing network communications (e.g., optical, electrical, frequency allocation, etc.). In addition, the transmitter 718 and the receiver 720 may be coupled to one or more antennas (e.g., antenna 722) and may share circuit components, software, or firmware, or alternatively be implemented separately.
[0081] In the illustrated embodiment, the communication functionality of the communication interface 712 may include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communication such as Bluetooth, near field communication, location-based communication such as using a global positioning system (GPS) to determine location, another similar communication functionality, or any combination thereof. Communication may be implemented according to one or more communication protocols and / or standards, such as IEEE 802.11, code division multiple access (CDMA), wideband code division multiple access (WCDMA), GSM, LTE, new radio (NR), UMTS, WiMax, Ethernet, transmission control protocol / internet protocol (TCP / IP), synchronous optical networking (SONET), asynchronous transfer mode (ATM), QUIC, hypertext transfer protocol (HTTP), etc.
[0082] Regardless of the type of sensor, the audio decoder may provide an output of decoded data through its communication interface 712 via a wireless connection to a network node.
[0083] The audio decoder, when in the form of an Internet of Things (IoT) device, may be a device for use in one or more application domains including, but not limited to, urban wearable technology, extended industrial applications, and healthcare. Non-limiting examples of such IoT devices are devices embedded in: a connected refrigerator or freezer, a TV, connected lighting, an electric meter, a robotic vacuum cleaner, a voice-controlled smart speaker, a home security camera, a thermostat, an electric door lock, a connected doorbell, an autonomous vehicle, a surveillance system, a weather monitoring device, a vehicle parking monitoring device, an electric vehicle charging station, a smartwatch, a fitness tracker, a head-mounted display for augmented reality (AR) or virtual reality (VR), a wearable device for tactile enhancement or sensory enhancement. The decoder in the form of an IoT device has, in addition to the Figure 7 The audio decoder 112 is shown as depicted, among other components, including circuitry and / or software depending on the intended application of the IoT device.
[0084] Figure 8 800 is a block diagram of a host computer 800 according to various aspects described herein. As used herein, the host computer 800 can be or include various combinations of hardware and / or software, including processing resources in a standalone server, a blade server, a cloud-enabled server, a distributed server, a virtual machine, a container, or a server farm. The host computer 800 can provide one or more services to one or more UEs.
[0085] Host 800 includes processing circuitry 802 operatively coupled to input / output interface 806, network interface 808, power supply 810, and memory 812 via bus 804. Other components may be included in other embodiments. The features of these components may be substantially similar to those described with respect to previous figures (such as Figure 7 ) features of the device description so that its description is generally applicable to the corresponding components of the host 800.
[0086] The memory 812 may include one or more computer programs, including one or more host applications 814 and data 816, which may include user data, such as data generated by a UE for the host 800 or data generated by the host 800 for the UE. An embodiment of the host 800 may utilize only a subset or all of the components shown. The host application 814 may be implemented in a container-based architecture and may provide support for video codecs (e.g., Versatile Video Codec (VVC), High Efficiency Video Codec (HEVC), Advanced Video Codec (AVC), MPEG, VP9) and audio codecs (e.g., FLAC, Advanced Audio Codec (AAC), MPEG, G.711, EVS, IVAS), including transcoding for multiple different categories, types, or implementations of UEs (e.g., mobile phones, desktop computers, wearable display systems, head-up display systems). The host application 814 may also provide user authentication and permission checks and may periodically report health, routing, and content availability to a central node (such as a device in or on the edge of a core network). Thus, the host 800 can select and / or instruct a different host for over-the-top services for the UE. The host application 814 can support various protocols such as HTTP Live Streaming (HLS), Real-Time Messaging Protocol (RTMP), Real-Time Streaming Protocol (RTSP), HTTP Dynamic Adaptive Streaming (MPEG-DASH), etc.
[0087] Figure 91 is a block diagram illustrating a virtualized environment 900 in which the functions implemented by some embodiments of the audio decoder 112 or the components of the audio decoder 112 can be virtualized. In this context, virtualization means creating a virtual version of an apparatus or device, which may include a virtualized hardware platform, storage device, and network resources. As used herein, virtualization can be applied to any device described herein or its components, and relates to an implementation in which at least a portion of a function is implemented as one or more virtual components. Some or all of the functions described herein can be implemented as virtual components performed by one or more virtual machines (VMs), which are implemented in one or more virtual environments 900 hosted by one or more hardware nodes, such as hardware computing devices operating as decoders, encoders, network nodes, UEs, core network nodes, or hosts. In addition, in embodiments in which the virtual nodes do not require a radio connection (e.g., a core network node or host), the nodes can be fully virtualized.
[0088] Application 902 (which may alternatively be referred to as a software instance, a virtual appliance, a network function, a virtual node, a virtual network function, etc.) runs in virtualized environment 900 to implement some features, functions and / or benefits of some embodiments disclosed herein.
[0089] The hardware 904 includes processing circuitry, memory storing software and / or instructions executable by the hardware processing circuitry, and / or other hardware devices as described herein, such as network interfaces, input / output interfaces, and the like. The software may be executed by the processing circuitry to instantiate one or more virtualization layers 906 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs 908A and 908B (one or more of which may be generally referred to as VMs 908), and / or perform any of the functions, features, and / or benefits described with respect to some embodiments described herein. The virtualization layer 906 may present a virtual operating platform that appears to be networked hardware to the VMs 908.
[0090] The VMs 908 include virtual processing, virtual memory, virtual networking or interfaces, and virtual storage, and may be run by a corresponding virtualization layer 906. Different embodiments of instances of virtual devices 902 may be implemented on one or more VMs 908 and may be implemented in different ways. Virtualization of hardware is referred to in some contexts as network function virtualization (NFV). NFV can be used to consolidate many network device types onto industry-standard, high-capacity server hardware, physical switches, and physical storage that can be located in data centers and customer premises equipment.
[0091] In the context of NFV, VMs 908 can be software implementations of physical machines that run programs as if they were executed on physical, non-virtualized machines. Each VM 908 and the portion of hardware 904 on which it executes (whether dedicated to that VM and / or shared with other VMs in the VM stack) form a separate virtual network element. Still in the context of NFV, a virtual network function is responsible for handling specific network functions running in one or more VMs 908 on hardware 904 and corresponding to an application 902.
[0092] The hardware 904 may be implemented in a standalone network node with general or specific components. The hardware 904 may implement some functionality via virtualization. Alternatively, the hardware 904 may be part of a larger hardware cluster (e.g., such as in a data center or CPE) where many hardware nodes work together and are managed via management and orchestration 910, which oversees, among other things, the lifecycle management of the applications 902. In some embodiments, the hardware 904 is coupled to one or more radio units, each of which includes one or more transmitters and one or more receivers that may be coupled to one or more antennas. The radio units may communicate directly with other hardware nodes via one or more appropriate network interfaces, and may be used in combination with virtual components to provide a virtual node with radio capabilities, such as a radio access node or base station. In some embodiments, a control system 912 may be used to provide some signaling, which may alternatively be used for communication between the hardware nodes and the radio units.
[0093] Although the computing devices (e.g., decoders, encoders, hosts) described herein may include the illustrated combinations of hardware components, other embodiments may include computing devices having different combinations of components. It should be understood that these computing devices may include any suitable combination of hardware and / or software required to perform the tasks, features, functions, and methods disclosed herein. The determination, calculation, acquisition, or similar operations described herein may be performed by processing circuitry that may process information by, for example, converting the acquired information into other information, comparing the acquired information or the converted information with information stored in a network node, and / or performing one or more operations based on the acquired information or the converted information, and making a determination as a result of the processing. In addition, although the components are depicted as being located within a larger box or as a single box nested within multiple boxes, in practice, a computing device may include multiple different physical components that make up a single illustrated component, and functionality may be divided between separate components. For example, a communication interface may be configured to include any of the components described herein, and / or functionality of a component may be divided between the processing circuitry and the communication interface. In another example, the non-computationally intensive functions of any such component may be implemented in software or firmware, and the computationally intensive functions may be implemented in hardware.
[0094] In certain embodiments, some or all of the functionality described herein may be provided by a processing circuit that executes instructions stored in a memory, which in certain embodiments may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functionality may be provided by a processing circuit without executing instructions stored on a separate or discrete device-readable storage medium, such as in a hardwired manner. In any of those specific embodiments, the processing circuit may be configured to perform the described functionality regardless of whether instructions stored on a non-transitory computer-readable storage medium are executed. The benefits provided by such functionality are not limited to separate processing circuits or other components of a computing device, but are enjoyed by the computing device as a whole and / or generally by end users and wireless networks.
[0095] Example Embodiments 1. A method of decoding coded audio in a decoder (112, 902), the coded audio being coded using at least two modes, the method comprising: receiving (501) a packet from a bitstream, the packet representing a coded audio frame, the packet comprising information indicating a selected codec mode and coded parameters required to perform reconstruction of the coded audio frame; determining (503) whether the selected codec mode is the first mode; In response to determining that the selected codec mode is the first mode, determining (505) whether the previous codec mode was the first mode; In response to the selected codec mode being the first mode and the previous codec mode not being the first mode, estimating (509) an envelope stability metric using energy stability and shape stability and using the estimated envelope stability metric as the envelope stability metric; decoding the encoded audio frame using the envelope stability metric (507); Determine (511) energy stability; determining (513) shape stability; and The decoded audio frames are output (515) to one of a storage device and an audio playback device. 2. The method according to embodiment 1, further comprising: In response to the selected codec mode being the first mode and the previous codec mode being the first mode, an envelope stability metric is determined as part of decoding the first mode of the encoded audio frame. 3. The method according to any one of embodiments 1 to 2 further includes: in response to determining that the selected codec mode is a second mode based on algebraic code excited linear prediction (ACELP), decoding the encoded audio frame using the ACELP-based codec mode (517). 4. The method according to any one of embodiments 1 to 3 further includes: in response to determining that the selected codec mode is a third mode based on modified discrete cosine transform (MDCT) codec, decoding the encoded audio frame using MDCT-based codec (519). 5. The method of any one of embodiments 1-2, wherein determining the envelope stability env_stab(m) comprises: Determine the long-term estimate of the logarithmic energy change D LP (m); env_stab(m) is derived by mapping the long-term estimate of the log energy change to the range [0,1]. 6. The method of embodiment 5, wherein estimating the long-term estimate of the log energy change comprises determining D according to LP (m): D LP (m) = αD(m) + (1-α)D LP (m-1) where N bands is the number of energy bands, I(m,b) and I(m-1,b) are logarithmic energy indices, and α is the low-pass filter coefficient. 7. The method of any one of embodiments 5-6, wherein deriving env_stab(m) comprises deriving env_stab(m) according to where b, c, and d are constants. 8. The method of embodiment 5, wherein estimating the long-term estimate of log energy change comprises determining D according to LP (m): D LP (m) = D LP,est (m)=P1+P2stab_fac_lt(m-1)+P3E Δ,LP (m-1) where P1, P2, and P3 are constants, stab_fac_lt(m-1) is the shape stability from frame m-1, and E Δ,LP (m-1) is the energy stability from frame m-1. 9. The method of embodiment 8, wherein deriving env_stab(m) comprises deriving env_stab(m) according to Where -a / b is the midpoint of the transition, where env_stab(m) = 0.5. 10. The method according to any one of embodiments 1 to 9, wherein the energy stability E is determined Δ,LP (m) is a long-term estimate of the absolute logarithmic energy difference between the synthesized frames derived from: E Δ,LP (m) = βE Δ (m)+(1-β)E Δ,LP (m-1) E Δ (m)=|E log (m)-E log (m-1)| in is the output synthetic frame m, and L out Indicates the output composite frame length. 11. The method according to any one of embodiments 1 to 9, wherein the energy stability E is determined Δ,LP (m) including determining E based on Δ,LP (m): E Δ,LP (m) = βE Δ (m)+(1-β)E Δ,LP (m-1) E Δ (m)=|E′ log(m)-E′ log (m-1)| 12. The method of embodiment 1, wherein determining the shape stability stab_fac_lt(m) comprises determining stab_fac_lt(m) according to stab_fac_lt(m)=γstab_fac_est(m)+(1-γ)stab_fac_lt(m-1) stab_fac_est(m)=Q1+Q2D LP (m)+Q3E Δ,LP (m) Where γ is the low-pass filter coefficient, Q1, Q2 and Q3 are constants, and D LP (m) is a long-term estimate of the logarithmic energy change, and E Δ,LP (m) is the long-term estimate of the absolute log energy difference between the synthesized frames. 13. The method of any one of embodiments 1 to 12, wherein the first mode is a modified discrete cosine transform (MDCT) based coding mode. 14. The method of any one of embodiments 3 to 4, wherein determining the shape stability stab_fac_lt(m) comprises determining stab_fac_lt(m) according to stab_fac_lt(m)=γstab_fac(m)+(1-γ)stab_fac_lt(m-1) where γ is the low-pass filter coefficient and stab_fac(m) is a shape stability factor based on the Euclidean distance between the line spectral frequency (LSF) representation of the linear predictor (LP) filters of the current frame and the previous frame. 15. A method of decoding coded audio in a decoder (112, 902), the coded audio being coded using a plurality of modes, the method comprising: receiving (601) a packet from a bitstream, the packet representing a coded audio frame, the packet comprising information indicating a selected codec mode and coded parameters required to perform reconstruction of the coded audio frame; determining (603) whether the current mode is the first mode and the previous mode was not the first mode; In response to determining yes, estimating (605) an envelope stability metric using energy stability and shape stability; Decoding the encoded audio frame based on the current mode (607); Determine (609) energy stability; determining (611) shape stability; and The decoded audio frame is output (613). 16. The method of embodiment 15, wherein determining (603) whether the current mode is the first mode and the previous mode was not the first mode; In response to a negative determination, an envelope stability metric is determined as part of the decoding of the first mode 607 . 17. A method according to embodiment 15, wherein decoding the encoded audio frame based on the current mode includes decoding the encoded audio frame using ACELP decoding in response to determining that the selected codec mode is a second mode based on algebraic code excited linear prediction (ACELP). 18. A method according to any one of embodiments 15 to 17, wherein decoding the encoded audio frame based on the current mode includes decoding the encoded audio frame using an MDCT-based codec in response to determining that the selected codec mode is a third MDCT-based mode (519). 19. The method of embodiment 15, wherein determining env_stab(m) comprises: Determine the long-term estimate of the logarithmic energy change D LP (m); env_stab(m) is derived by mapping the long-term estimate of the log energy change to the range [0,1]. 20. The method of embodiment 19, wherein determining the long-term estimate of log energy change comprises determining D according to LP (m): D LP (m) = αD(m) + (1-α)D LP (m-1) where N bands is the number of energy bands, I(m,b) and I(m-1,b) are logarithmic energy indices, and α is the low-pass filter coefficient. 21. The method of any one of embodiments 19-20, wherein deriving env_stab(m) comprises deriving env_stab(m) according to where b, c, and d are constants. 22. The method of any one of embodiments 19 to 21, wherein estimating the long-term estimate of the logarithmic energy variation comprises estimating D according to LP (m): D LP (m) = D LP,est (m)=P1+P2stab_fac_lt(m-1)+P3EΔ,LP (m-1) where P1, P2, and P3 are constants, stab_fac_lt(m-1) is the shape stability from frame m-1, and E Δ,LP (m-1) is a long-term estimate of the absolute log energy difference between the synthesized frames. 23. The method of embodiments 19-20, wherein deriving env_stab(m) comprises deriving env_stab(m) according to Where -a / b is the midpoint of the transition, where env_stab(m) = 0.5. 24. The method according to any one of embodiments 15 to 23, wherein the energy stability E is determined Δ,LP (m) including the determination of E based on Δ,LP (m): E Δ,LP (m) = βE Δ (m)+(1-β)E Δ,LP (m-1) E Δ (m)=|E log (m)-E log (m-1)| in is the output synthetic frame m, and L out Indicates the output composite frame length. 25. The method according to any one of embodiments 15 to 23, wherein the energy stability E is determined Δ,LP (m) including the determination of E based on Δ,LP (m): E Δ,LP (m) = βE Δ (m)+(1-β)E Δ,LP (m-1) E Δ (m)=|E′ log (m)-E′ log (m-1)| 26. The method of embodiment 15, wherein determining the shape stability stab_fac_lt(m) comprises determining stab_fac_lt(m) for the first mode according to stab_fac_lt(m)=γstab_fac_est(m)+(1-γ)stab_fac_lt(m-1) stab_fac_est(m)=Q1+Q2DLP (m)+Q3E Δ,LP (m) Where γ is the low-pass filter coefficient, and Q1, Q2, and Q3 are constants, D LP (m) is a long-term estimate of the logarithmic energy change, and E Δ,LP (m) is the long-term estimate of the absolute log energy difference between the synthesized frames. 27. The method of any one of embodiments 15 to 26, wherein the first mode is a modified discrete cosine transform (MDCT) based coding mode. 28. The method of any one of embodiments 15 to 27, wherein determining the shape stability stab_fac_lt(m) comprises determining stab_fac_lt(m) for the second mode and the third mode according to stab_fac_lt(m)=γstab_fac(m)+(1-γ)stab_fac_lt(m-1) where γ is the low-pass filter coefficient. 29. A decoder (112, 902) for decoding encoded audio, the encoded audio being encoded using a plurality of modes, the decoder being adapted to perform according to any one of embodiments 1 to 28. 30. A decoder (112, 902) for decoding encoded audio, the encoded audio being encoded using a plurality of modes, the decoder (112) comprising: Processing circuit (702); A memory (710) coupled to the processing circuit, wherein the memory includes instructions that, when executed by the processing circuit, cause the decoder to perform according to any one of embodiments 1 to 28. 31. A computer program comprising program code to be executed by a processing circuit (702) of a decoder (112, 902), wherein execution of the program code causes the decoder (112, 902) to perform the operations according to any one of embodiments 1 to 28. 32. A computer program product comprising a non-transitory storage medium comprising program code to be executed by a processing circuit (702) of a decoder (112, 902), wherein execution of the program code causes the decoder (112, 902) to perform operations according to any one of embodiments 1 to 28.
Claims
1. A method for controlling noise filling in a decoder (112, 902) for decoding coded audio, the audio being coded using a plurality of modes, the method comprising: Receiving (601) information indicating a selected codec mode; determining (603) whether the current mode is the first mode and the previous mode was not the first mode; In response to the current mode being the first mode and the previous mode not being the first mode, estimating (605) an envelope stability metric using energy stability and shape stability of a previous frame; decoding (607) an encoded audio frame based on the current mode, wherein the decoding includes applying adaptive attenuation to a noise-filled frequency band based on the estimated envelope stability metric; determining (609) the energy stability of the current frame; as well as The shape stability of the current frame is determined (611).
2. The method of claim 1 , wherein determining (603) whether the current mode is the first mode and the previous mode is not the first mode comprises: Responsive to the current mode being the first mode and the previous mode being the first mode, determining an envelope stability metric as part of the decoding of the first mode.
3. The method of claim 1 , wherein decoding the encoded audio frame based on the current mode comprises decoding the encoded audio frame using ACELP decoding in response to determining that the selected codec mode is a second mode based on Algebraic Code Excited Linear Prediction (ACELP).
4. The method according to any one of claims 1 to 3, wherein decoding the encoded audio frame based on the current mode includes decoding the encoded audio frame using MDCT-based decoding in response to determining that the selected codec mode is a third MDCT-based mode (519).
5. The method of claim 1 , wherein determining the envelope stability metric env_stab(m) comprises: Determine the long-term estimate of the logarithmic energy change D LP (m); The env_stab(m) is determined by mapping the long-term estimate of the log energy change to the range [0, 1].
6. The method of claim 5, wherein determining the long-term estimate of the logarithmic energy variation comprises determining D according to LP (m): D LP (m)=αD(m)+(1-α)D LP (m-1) where N bands is the number of energy bands, I(m,b) and I(m-1,b) are the logarithmic energy indices of the current frame and the previous frame, and α is a low-pass filter coefficient.
7. The method according to any one of claims 5 to 6, wherein determining env_stab(m) comprises determining env_stab(m) according to where b, c, and d are constants.
8. The method of claim 5, wherein determining the long-term estimate of the logarithmic energy variation comprises determining D according to LP (m): D LP (m)=D LP,est (m)=P1+P2stab_fac_lt(m-1)+P3E Δ,LP (m-1) where P1, P2, and P3 are constants, stab_fac_lt(m-1) is the shape stability from frame m-1, and E Δ,LP (m-1) is the energy stability from frame m-1.
9. The method of claim 8, wherein determining env_stab(m) comprises determining env_stab(m) according to Where -a / b is the midpoint of the transition, where env_stab(m) = 0.
5.
10. The method according to claim 1, wherein the energy stability E is determined Δ,LP (m) including the determination of E based on Δ,LP (m): It is Δ,LP (m)=βE Δ (m)+(1-β)E Δ,LP (m-1) HAVE BEEN Δ (m)=|E log (along with log (m-1)| in is the output composite frame m, L consisting of samples n=0,...L-1 out represents the output synthesis frame length, and β is the low-pass filter coefficient.
11. The method according to claim 1 , wherein the energy stability E is determined Δ,LP (m) including the determination of E based on Δ,LP (m): It is Δ,LP (m)=βE Δ (m)+(1-β)E Δ,LP (m-1) HAVE BEEN Δ (m)=|E′ log (along with' log (m-1)| in is the output synthesized frame m comprising samples n=0, ...L-1, and β is the low-pass filter coefficient.
12. The method of claim 1 , wherein determining the shape stability stab_fac_lt(m) comprises determining stab_fac_lt(m) for the first mode according to stab_fac_lt(m)=γstab_fac_est(m)+(1-γ)stab_fac_lt(m-1) stab_fac_est(m)=Q1+Q2D LP (m)+Q3E Δ,LP (m) where γ is the low-pass filter coefficient, stab_fac_est(m) is the estimate of the shape stability, and Q1, Q2 and Q3 are constants, D LP (m) is a long-term estimate of the logarithmic energy change, and E Δ,LP (m) is the energy stability.
13. The method according to any one of claims 1 to 12, wherein the first mode is a Modified Discrete Cosine Transform (MDCT) based coding mode.
14. The method according to any one of claims 1 to 13, wherein determining the shape stability stab_fac_lt(m) comprises determining stab_fac_lt(m) for the second mode and the third mode according to stab_fac_lt(m)=γstab_fac(m)+(1-γ)stab_fac_lt(m-1) where γ is a low-pass filter coefficient and stab_fac(m) is a shape stability factor based on the Euclidean distance between the line spectral frequency (LSF) representations of the linear predictor (LP) filters of the current and previous frames.
15. A decoder (112, 902) for decoding coded audio, the coded audio being coded using a plurality of modes, the decoder being adapted to perform the method according to at least one of claims 1 to 14.