Packet loss concealment based on adaptive cross-band filtering
Adaptive cross-band filtering for packet loss concealment in frame-based audio predicts missing frames using previous frames and random sign blending, addressing the limitations of existing methods by maintaining audio quality and reducing computational complexity.
Patent Information
- Application Number
- PCT/EP2025/051633
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-31
AI Technical Summary
Existing packet loss concealment methods for frame-based audio, especially in low latency codecs, are inadequate for general audio and often result in unrealistic predictions that can alter the intended message, particularly for longer bursts of lost packets.
A method for packet loss concealment using adaptive cross-band filtering, where time-frequency coefficients are predicted based on previous frames and updated continuously during successful transmission, incorporating cross-band prediction and random sign blending to enhance prediction accuracy and reduce computational burden.
The method effectively predicts missing frames in frame-based audio, maintaining audio quality by adapting to previous frames and reducing computational complexity, especially during longer packet losses.
Smart Images

Figure EP2025051633_31072025_PF_FP_ABST
Abstract
Description
[0001] PACKET LOSS CONCEALMENT BASED ON ADAPTIVE CROSS-BAND FILTERING Technical Field [1] The present disclosure relates to techniques for packet loss concealment (frame loss concealment) for frame-based audio. In particular, the present disclosure describes methods and apparatus for packet loss concealment for frame-based audio based on frequency-domain cross- band prediction. Cross-Reference to Related Applications [2] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 624,066 filed on 23 January 2024 and European Patent Application 24164281.8, each of which are incorporated by reference herein in their entirety. Background [3] For audio coding with low latency constraints, lost packets (e.g., lost frames) cannot always be replaced with coded representations of the original audio. Hence, there is a need to be able to extrapolate the audio at the receiver end. For the special case of speech signals, there are many existing solutions based on time and frequency domain prediction, including emerging machine learning methods. For longer bursts of lost packets it is important that the method is not too realistic and inventive, as this would risk changing the intended message. For general audio, the speech-specific solutions tend to not work well, and a commonly used method in transform- based coding is to repeat the previous frames’ transform bins with random flip of sign. A speech- specific improvement for this method is described in US Patent Application 16 / 126,940 “Packet Loss Concealment for Critically-Sampled Filter Bank-Based Codecs Using Multi-Sinusoidal Detection”. Here, the cross-band filters were derived from a sinusoidal signal model, but applying the method for general audio has shown less promise. Summary [4] Thus, there is a need for improved techniques for packet loss concealment for frame- based audio, especially for general audio. There is particular need for such techniques that are applicable to low latency codecs. Low latency usually depends on various factors such as, for example, frame size, sampling rate, hardware and / or software computational resources, etc. For example, in some instances, low latency would normally be less than 40 ms, 20 ms, or 10 ms. [5] In view of this need, the present disclosure provides methods of packet loss concealment for frame-based audio as well as corresponding apparatus, programs, and computer-readable storage media, having the features of the respective independent claims. [6] One aspect of the present disclosure relates to a method of packet loss concealment for frame-based audio. The frame-based audio may relate to or include any of mono signals, stereo signals, multi-channel signals, immersive audio, etc., for example. The method may include, in a first mode (e.g., normal mode), when time-frequency coefficients of a current frame of audio are available (e.g., have been validly received), storing the time-frequency coefficients of the current frame in a buffer. Without intended limitation, the time-frequency coefficients may be coefficients of a Modified Discrete Cosine Transform (MDCT), or of a complex modulated filterbank transform, such as a Quadrature Mirror Filter (QMF) transform or Discrete Fourier Transform (DFT), for example. The method may further include, in a second mode (e.g., conceal mode), when time-frequency coefficients of the current frame are not available (e.g., when the current frame has not been validly received), predicting the time-frequency coefficients of the current frame based on a set of prediction parameters and time-frequency coefficients of one or more previous frames stored in the buffer, and storing the predicted time-frequency coefficients in the buffer. The prediction may involve cross-band prediction, for example. The method may further include outputting the time-frequency coefficients of the current frame and / or the predicted time-frequency coefficients of the current frame, for example for audio rendering or further processing. [7] Thereby, the proposed method can extrapolate time-frequency coefficients (e.g., MDCT coefficients) of previous buffered frames to conceal lost frames. This method is applicable in particular to general audio including music. Here, general audio is understood to encompass any audio that is not strictly speech, such as music, audio for movies or TV segments, etc. [8] In some embodiments, the set of prediction parameters may include a plurality of prediction filter coefficients. The prediction filter coefficients may include, for each of a plurality of frequency bands, prediction filter coefficients associated with respective time-frequency coefficients of one or more previous frames in the frequency band and in one or more neighboring frequency bands. The plurality of frequency bands may relate to, for example, a full frequency range of the time-frequency transform or to a low-frequency range of the time- frequency transform (e.g., up to a frequency threshold chosen in the range between 4kHz and 6kHz). In one example, the given frequency band and the one or more neighboring frequency bands may relate to a range of frequency bands centered on the given frequency band, e.g., three frequency bands centered on the given frequency band. [9] Accordingly, cross-band prediction can be applied to realistically predict any missing frames from buffered previous frames.
[0010] In some embodiments, prediction of a time-frequency coefficient of the current frame in a given frequency band may be based on time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands, and on the prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands.
[0011] In some embodiments, a predicted time-frequency coefficient ^^(^^, ^^) for current frame^^ and frequency band ^^ may be given by ^^(^^, ^^) = ∑^^^^^^^^^^=1 ∑^^=−^^^^ ^^^^(^^, ^^)^^(^^ − ^^, ^^ − ^^).Here, ^^^^(^^, ^^) are prediction filter coefficients for ^^ in relation to time-frequency coefficients ^^(^^ − ^^, ^^ − ^^) of frame (^^ − ^^) in frequency band (^^ − ^^), ^^^^ ≥ 1indicates a number frames to be used for the prediction, and ^^^^ ≥ 1 indicates anumber (or range) of neighboring frequency bands to be used for the prediction. The neighboring frequency bands may correspond to a given number (e.g., one or more) of adjacent frequency bands towards higher frequency and the given number of adjacent frequency bands towards lower frequency, relative to the frequency band in question. It is understood that the total number offrequency bands used for the prediction is given by 2^^^^ + 1. For frequency bands with only oneneighboring frequency band in the plurality of frequency bands, zeros may be assumed for the time-frequency coefficients in the respective non-existing neighboring frequency band. Alternatively, the summation range for ^^ in the above equation may be shifted to cover existing frequency bands (in the plurality of frequency bands) only.
[0012] In some embodiments, the method may further include, in the first mode (e.g., normal mode), updating the set of prediction parameters based on the time-frequency coefficients of the current frame. Accordingly, the prediction parameters (e.g., prediction filter coefficients and / or accumulated per-band values) may be continuously updated while there is no packet loss (e.g., frame loss).
[0013] In some embodiments, the prediction parameters may be updated based on time- frequency coefficients of a predetermined number of frames including the current frame and preceding frames. The time-frequency coefficients of the preceding frames may be obtained from the buffer.
[0014] In some embodiments, in the first mode (e.g., normal mode), updated prediction filter coefficients may be determined, for a given frequency band, based on current prediction filter coefficients for the given frequency band and respective adjustment amounts for the current prediction filter coefficients. For each prediction filter coefficient, there may be a corresponding adjustment amount. The updated prediction filter coefficients may be determined by adding respective adjustment amounts to respective current prediction filter coefficients. Thereby, adaptive filtering for purposes of the frame prediction can be implemented.
[0015] In some embodiments, the adjustment amounts for the given frequency band (one for each prediction filter coefficient for the given frequency band) may be determined based on the current prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band and based on the time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band, using a least squares method.
[0016] In some embodiments, determining the adjustment amounts for the given frequency band may include predicting the time-frequency coefficient of the current frame in the given frequency band in dependence (e.g., functional dependence) on the adjustment amounts. Determining the adjustment amounts for the given frequency band may further include minimizing a squared prediction error between the predicted time-frequency coefficient of the current frame in the given frequency band and the time-frequency coefficient of the current frame in the given frequency band for a given change of the adjustment amounts for the given frequency band. This may involve determining a gradient of the squared prediction error with respect to the adjustment amounts for the given frequency band. The given change of the adjustment amounts for the given frequency band may be a change in a sum or squared sum of the adjustment amounts for the given frequency bands, for example. The maximum reduction of the squared prediction error may be found in the opposite direction of the gradient.
[0017] Thereby, the prediction filter coefficients can be continuously updated in normal mode, while frames are validly received. Specifically, the prediction filter coefficients adapt in accordance with the received audio to provide the best prediction for any missing frames, based on the previous frames.
[0018] In some embodiments, determining the adjustment amounts for the given frequency band may include determining a probabilistic model for the time-frequency coefficient of the current frame in the given frequency band based on the current prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band and based on the time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band. Determining the adjustment amounts for the given frequency band may further include determining the adjustment amounts for the given frequency band based on the probabilistic model and based on the time-frequency coefficient of the current frame in the given frequency band. The probabilistic model may comprise a mean and a variance of a normally distributed random variable, for example.
[0019] This provides an alternative to continuously updating the prediction filter coefficients using the least squares method. Also here, the prediction filter coefficients can be continuously updated in normal mode, while frames are validly received, wherein the prediction filter coefficients adapt in accordance with the received audio to provide the best prediction for any missing frames, based on the previous frames.
[0020] In some embodiments, an adjustment amount for the given frequency band and for a given current prediction filter coefficient may be determined based on the probabilistic model, the time-frequency coefficient of the current frame in the given frequency band, and the time frequency coefficient that is associated with the given current prediction filter coefficient.
[0021] In some embodiments, predicted time-frequency coefficients in the buffer may not be used for updating the prediction parameters. Thereby, it can be ensured that the prediction is based on actual frames that have been received in the past. For longer stretches of missing frames (e.g., tens of missing frames or more), it may be advantageous to gracefully fade out instead of indefinitely attempting to predict frames.
[0022] In some embodiments, the method may further include, in the second mode (e.g., conceal mode), determining the prediction filter coefficients based on time-frequency coefficients of a predetermined number of previous frames, using a least squares method.
[0023] Thereby, average computational burden (computational complexity) can be reduced (e.g., in normal mode), while peak computation burden is increased (e.g., upon frame loss). A choice of whether to determine the prediction filter coefficients upon frame loss or continuously in normal mode may be made depending on a desired balance between average and peak computational burden. Further balancing may be achieved by using one of these techniques for a first frequency range and using the other for a second, different frequency range.
[0024] In some embodiments, determining the prediction filter coefficients for a given frequency band among the plurality of frequency bands using the least squares method may be based on the time-frequency coefficients of the predetermined number of previous frames in the given frequency band and in the one or more neighboring frequency bands.
[0025] In some embodiments, a least squares error to be minimized for the given frequency band may be determined based on a sum of prediction errors for the time-frequency coefficients of the predetermined number of previous frames, each predicted time frequency coefficient being based on time-frequency coefficients of one or more preceding frames to the respective frame among the predetermined number of previous frames.
[0026] In some embodiments, predicting the time-frequency coefficients of the current frame may further include limiting a predicted time-frequency coefficient of a given frequency band to a range that is based on an accumulated per-band value of a signal power in the given frequency band. This may reduce audible artifacts such as “plops” that might otherwise result from exponentially growing predicted samples in some cases.
[0027] In some embodiments, the method may further include, in the second mode (e.g., conceal mode), generating, for each frequency band in the plurality of frequency bands, a weighted sum of the predicted time-frequency coefficient of the current frame in that frequency band and a most recently available time-frequency coefficient multiplied by a random sign. Thereby, the conventional random sign repeat method can be “blended” with the proposed techniques, depending on perceptual considerations.
[0028] In some embodiments, a weight for the weighted sum in a given frequency band may be determined based on one or more of a per-band value of a signal power in the given frequency band, a per-band value of a predicted signal power in the given frequency band, a per-band value of a prediction error of the predicted signal power in the given frequency band, and / or a per-band value of a logarithmic relative noise level in the given frequency band. The aforementioned per- band values may be accumulated or propagated values, for example.
[0029] In some embodiments, weights for the weighted sums may be successively reduced for instances of consecutive predicted frames, to decrease a relative importance of the most recent available time-frequency coefficient multiplied by the random sign. Thereby, a gradual fade-out of the predicted signal can be achieved if a long stretch of frames (e.g., tens of frames or longer) is missing.
[0030] In some embodiments, the method may further include, for a pair of channels, adding a mixture of the most recent available time-frequency coefficient of one of the channels multiplied by a random sign to the weighted sum of the other one of the channels. Thereby, the correlation between channels that one would expect for, for example, stereo signals can be recovered.
[0031] In some embodiments, the mixture may be determined based on a per-band value of a signal cross-product with respect to the pair of channels. The per-band value of the signal cross- product may be an accumulated or propagated value, for example. In some embodiments, the plurality of frequency bands may relate to (e.g., correspond to) a lower-frequency portion of the frequency range of the time-frequency coefficients. Then, the method may further include, for at least one frequency band not in the lower-frequency portion, using a most recent available time-frequency coefficient in the at least one frequency band multiplied by a random sign as the time-frequency coefficient for the current frame in the at least one frequency band. Thereby, load balancing based on, for example perceptual considerations can be achieved.
[0032] According to another aspect, an apparatus for packet loss concealment for frame-based audio is provided. The apparatus may include a processor and a memory coupled to the processor and storing instructions for the processor. The processor may be configured to perform all steps of the methods according to the preceding aspect and its embodiments. According to another aspect, a computer program is described. The computer program may comprise executable instructions for performing the methods or method steps outlined throughout the present disclosure when executed by a computing device.
[0033] According to yet another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted for execution on a processor and for performing the methods or method steps outlined throughout the present disclosure when carried out on the processor.
[0034] It should be noted that the methods and apparatus including their preferred embodiments as outlined in the present disclosure may be used stand-alone or in combination with the other methods and apparatus disclosed in this document. Furthermore, all aspects of the methods and apparatus outlined in the present disclosure may be arbitrarily combined. In particular, the features of the claims may be combined with one another in an arbitrary manner.
[0035] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus, and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) (and, e.g., their steps) are understood to likewise apply to the corresponding apparatus (and, e.g., their blocks, stages, units, etc.), and vice versa. Brief Description of the Drawings
[0036] The invention is explained below in an exemplary manner with reference to the accompanying drawings, wherein:
[0037] Fig.1 is a block diagram schematically illustrating an example of a frequency domain decoder according to embodiments of the disclosure;
[0038] Fig.2 is a flowchart illustrating an example method of packet loss concealment for frame- based audio according to embodiments of the disclosure;
[0039] Fig.3 is a block diagram schematically illustrating an example of an element of the frequency domain decoder of Fig.1, according to embodiments of the disclosure;
[0040] Fig.4 schematically illustrates an example of a frequency-domain prediction geometry according to embodiments of the disclosure;
[0041] Fig.5 is a flowchart illustrating an example of a method of updating prediction filter coefficients according to embodiments of the disclosure;
[0042] Fig.6 is a flowchart illustrating another example of a method of updating prediction filter coefficients according to embodiments of the disclosure;
[0043] Fig.7 to Fig.10 are diagrams showing examples of performance results for techniques according to embodiments of the disclosure; and
[0044] Fig.11 is a block diagram of an example of an apparatus for performing methods according to embodiments of the disclosure. Detailed Description Overview
[0045] Broadly speaking, the present disclosure relates to techniques (e.g., methods and apparatus) for packet loss concealment for frame-based audio. The disclosed techniques are applicable to general audio, for example.
[0046] One main idea of embodiments of the present disclosure is to let the filter coefficients of a cross-band filter for prediction of missing frames be updated continuously during successful transmission according to an extended adaptive filter methodology. Specifically, embodiments of the disclosure may include or relate to one or more of • time-frequency adaptive prediction and filter update rules, • mixing rules for combination of predicted and random contributions, • stabilization by limiting predicted time-frequency coefficients depending on an energy estimate, and • providing means to control peak and average complexity by splitting up filter coefficient computation into two variants operating in different frequency regions.
[0047] Fig.1 is a block diagram schematically illustrating an example of a frequency domain decoder 100 with packet loss (e.g., frame loss) concealment according to embodiments of the disclosure. In the example of Fig.1, the decoder is an MDCT-based audio decoder, but audio decoders based on other forms of time-frequency transforms, such as QMF transforms or DFTs, for example, are likewise understood to be covered by the present disclosure. Accordingly, while the present disclosure may make frequent reference to MDCT coefficients or MDCT bins, it is to be understood that this is without intended limitation and that the present disclosure applies to time-frequency coefficients in general.
[0048] In this decoder 100, parameters to be used for concealment during normal mode operation of the decoder 100 are continuously updated. The decoder 100 continuously receives frames of MDCT coefficients 10 for processing. A buffer (MDCT buffer) 20 of the current frame and a predetermined number of previous frames is also maintained, to which any incoming (i.e., current) frame 10 is stored. In some implementations, the buffer 20 may be a first-in-first-out (FIFO) buffer of predetermined length. Prediction parameters for predicting a (missing) frame from one or more previous frames are stored in a parameter buffer 50. As long as frames are validly received (e.g., frames that can be decoded or otherwise processed to yield the necessary MDCT coefficients), the prediction parameters are continuously updated by parameter update module 40, and updated prediction parameters 45 are stored in the parameter buffer 50.
[0049] Upon packet loss, i.e., if a current frame of MDCT coefficients is not available (e.g., is not validly received), the conceal mode is invoked, relying on the current state of the prediction parameters in the parameter buffer 50 and on the time-frequency coefficients in the MDCT buffer 20. Accordingly, conceal module 60 creates a predicted MDCT frame (e.g., output MDCT frame) 65, which is then fed to the output 70, based on the aforementioned parameters and coefficients. The created MDCT frame 65 is also used to update the MDCT buffer 20. The process is repeated for each missing frame until MDCT data is received again, at which point the normal mode is re-entered. A mode selection, depending on whether the current frame of MDCT coefficients is available, may be made by mode selection module (e.g., switch) 30.
[0050] Configured as described above, the decoder 100 can output an MDCT frame 80 for further processing (e.g., in the context of audio rendering) regardless of whether the current frame of MDCT coefficients is validly received or missing. The output MDCT frame 80 may be either the current frame, if available, or a predicted version thereof, which are respectively output by output module (e.g., adder) 70. The aforementioned conceal operations and parameter updates will be described in more detail below.
[0051] Fig.2 is a flowchart illustrating an example method 200 of packet loss concealment (e.g., frame loss concealment) for frame-based audio according to embodiments of the disclosure. Method 200 comprises steps S210 and S220, as well as optional step S230.
[0052] At step S210, in a first mode (e.g., normal mode), when time-frequency coefficients (e.g., MDCT coefficients) of a current frame of audio are available, the time-frequency coefficients of the current frame are stored in a buffer (e.g., MDCT buffer 20 of decoder 100).
[0053] As will be described in more detail below, in some implementations step S210 may further comprise updating a set of prediction parameters (in the first mode) based on the time- frequency coefficients of the (available) current frame. In other words, the prediction parameters may be continuously updated while there is no packet loss (e.g., frame loss). The updating of the prediction parameters may be performed based on time-frequency coefficients of a predetermined number of frames including the current frame and preceding frames, where the latter may be obtained from the buffer. In some implementations, time-frequency coefficients of preceding (e.g., buffered) frames that are itself predicted time-frequency coefficients may not be used for updating the prediction parameters.
[0054] At step S220, in a second mode (e.g., the conceal mode), when time-frequency coefficients of the current frame are not available, the time-frequency coefficients of the current frame are predicted based on a set of prediction parameters and time-frequency coefficients of one or more previous frames stored in the buffer. Step S220 further comprises storing the predicted time-frequency coefficients in the buffer. The prediction may involve cross-band prediction, for example.
[0055] As will be described in more detail below, in some implementations step S220 may further comprise determining at least part of the prediction parameters (e.g., prediction filter coefficients) based on time-frequency coefficients of a predetermined number of previous frames. This may be done for example using a least squares method. Also here, it is understood that in some implementations time-frequency coefficients of preceding (e.g., buffered) frames that are itself predicted time-frequency coefficients may not be used (e.g., may be skipped) for the determination the prediction parameters. Specifically, determination of at least part of the prediction parameters at step S220 may be performed in those implementations that do not foresee updating the prediction parameters at step S210 in the normal mode, and vice versa.
[0056] At step S230, in the second mode, for each frequency band in the plurality of frequency bands, a weighted sum of the predicted time-frequency coefficient of the current frame in that frequency band and a most recent available time-frequency coefficient multiplied by a random sign is generated. This weighted sum may be used as output frame. Step S230 may be an optional step, as indicated above.
[0057] Although not shown in Fig.2, method 200 may further include a step of outputting the time-frequency coefficients of the current frame (if available) and / or the predicted time- frequency coefficients of the current frame (if the current frame is not available), for example for audio rendering or further processing. In the latter case, output may relate to the result of step S220 or step S230, for example. Conceal Operation
[0058] Next, details of the conceal operation, i.e., the prediction of time-frequency coefficients if the current frame is missing, will be described.
[0059] Fig.3 is a block diagram schematically illustrating an example of a conceal mode operation 300 that may be performed for example in conceal module 60 of Fig.1 and / or in step S220 of Fig.2. The conceal mode operation 300 in this example comprises cross-band prediction 310, random sign blend 320, which may be optional, and limiting 330, which may be optional as well.
[0060] The parameter buffer 50 is unaltered during the second mode (e.g., conceal mode) in the sense that the prediction parameters are not updated in this mode. The prediction parameters (e.g., the parameter buffer 50) may include per-band values of • prediction filter coefficients ^^
[0061] They may further include one or more of • a signal power ^^ • guidance to the random sign blend operation in one of the forms 1. a predicted power ^^ 2. a prediction error power ^^ 3. a logarithmic relative noise level ^^
[0062] For example, the per-band signal power ^^ of a given frame ^^ of time-frequencycoefficients (e.g., MDCT coefficients) ^^(^^, ^^), where ^^ indicates a frequency band or bin (e.g.,MDCT line), may be calculated as ^^(^^) = ^^(^^, ^^)2(1)
[0063] Further, the per-band predicted power ^^ of a given frame ^^ of time-frequency coefficients may be calculated as ^^(^^) = ^^(^^, ^^)2(2)where ^^(^^, ^^) are predicted time-frequency coefficients for the current frame.The per-band prediction error power ^^ of a given frame ^^ of time-frequency coefficients may be calculated as ^^(^^) = (^^(^^, ^^) − ^^(^^, ^^))2(3)
[0064] As will be described below, rather than the above per-band quantities ^^, ^^, ^^, and ^^, accumulated values thereof (e.g., running averages, first-order accumulations) may be used that depend on both the per-band quantity of the current frame and the accumulated value (e.g., running average, first-order accumulation) for the preceding frame. Cross-Band Prediction
[0065] Next, cross-band prediction as for example performed in step S220 of Fig.2 will be described in more detail.
[0066] In general, the set of prediction parameters used for the prediction may include a plurality of prediction filter coefficients ^^. For each of a plurality of frequency bands, there may be prediction filter coefficients ^^ associated with respective time-frequency coefficients ^^ of one or more previous frames in the frequency band and in one or more neighboring frequency bands. The plurality of frequency bands may relate to, for example, a full frequency range of the time- frequency transform or to a low-frequency range of the time-frequency transform.
[0067] Then, prediction of a time-frequency coefficient of the current (and missing) frame in a given frequency band at step S220 may be based on time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands, and on the prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands.
[0068] In one example, the given frequency band and the one or more neighboring frequency bands may relate to a range of frequency bands centered on the given frequency band, e.g., three frequency bands centered on the given frequency band.
[0069] Fig.4Error! Reference source not found. illustrates an example frequency-domain prediction geometry or time-frequency receptive field, i.e., previous time-frequency coefficients that are relied on for the prediction of a current time-frequency coefficient). Squares in this figure indicate transform bins, which in turn correspond to time-frequency coefficients (e.g., MDCT coefficients). The horizontal direction represents time and the vertical direction represents frequency. As noted above, each target transform bin (i.e., transform bin of the missing frame) may be predicted by a linear combination of one or more (e.g., multiple) previous bins in time and their neighboring bins in frequency, according to a predetermined time-frequency receptivefield. In the example of Fig. 4, the light-gray bin at (time, frequency) index (^^, ^^) is predictedbased on (e.g., by) a linear combination of bins (^^ − ^^, ^^ − ^^) with ^^ ∈ {1, ⋯ , ^^^^} and ^^ ∈{−^^^^ , ⋯ , ^^^^}. In the present example, the values of the time and frequency ranges ^^^^ and ^^^^ aregiven by ^^^^ = 2 and ^^^^ = 1.
[0070] In other words, a predicted time-frequency coefficient ^^(^^, ^^) for current frame ^^ andfrequency band ^^ may be given by ^^^^^^^^^^ ^^ − ^^ −where ^^^^(^^, ^^) are time-frequency coefficients ^^(^^ − ^^, ^^ − ^^) of frame (^^ − ^^) in frequency band (^^ − ^^),^^^^ ≥ 1 indicates the number of preceding frames to be used for the prediction, and ^^^^ ≥ 1indicates the number of neighboring frequency bands to be used for the prediction.
[0071] For frequency bands with only one neighboring frequency band in the plurality of frequency bands (e.g., frequency bands at the edges of the range spanned by the plurality of frequency bands), zeros may be assumed for the time-frequency coefficients in the respective non-existing neighboring frequency band, or in other words, the available data may be extended by zeros when required for the summation over ^^. Alternatively, the summation range for ^^ in Eq. (4) above may be shifted to cover existing frequency bands. Random Sign Blend
[0072] The prior art random sign repeat method consists of using the last non-concealed bin times a random sign as the concealment for the current bin in each band. Based thereon, the method 200 of Fig.2 may comprise, with step S230, a step of performing arandom sign blend. Denoting the result of the sign repeat method by ^^(^^, ^^), it may be blendedwith the predicted bin ^^(^^, ^^) according to a blending factor (or weight) 0 ≤ ^^(^^) ≤ 1,^^(^^, ^^) = ^^(^^, ^^) + ^^(^^) ^^(^^, ^^)(5)
[0073] Thus, as indicated above, step S230 may amount to, in the second mode, generating, for each frequency band ^^ in the plurality of frequency bands, a weighted sum of the predicted time-frequency coefficient ^^(^^, ^^) of the current frame ^^ in that frequency band ^^ and a randomsign repeat time-frequency coefficient ^^(^^, ^^), which is a most recent validly received availabletime-frequency coefficient multiplied by a random sign. Herein, a weight for the weighted sum in a given frequency band may correspond to the aforementioned blending factor ^^(^^).
[0074] If the prediction filter ^^ is all zeros and ^^ (i.e., the weight) is set to all ones, the random sign repeat method is recovered. On the other hand, if ^^ is all zeros, there is no random contribution to the output.
[0075] The blending factor (or weight) ^^ may be determined (e.g., computed) based on one or more of a per-band value of the signal power ^^ in the given frequency band, a per-band value of the predicted signal power ^^ in the given frequency band, a per-band value of the prediction error ^^ of the predicted signal power in the given frequency band, and / or a per-band value of the logarithmic relative noise level ^^ in the given frequency band. These per-band values may be accumulated or propagated values (e.g., running averages, first-order accumulations), for example.
[0076] For example, the blending factor ^^ may be computed based on the states of the parameter buffer as one of ^^(^^) = √1 − min{1, ^^(^^) / ^^(^^)}(6) ^^(^^) = √min{1, ^^(^^) / ^^(^^)} (8)
[0077] For consecutively concealed frames, the energy of the random sign blend contribution may be reduced over time by recursively attenuating the blending factor in each processing frame. In other words, weights for the weighted sums may be successively reduced for instances of consecutive predicted frames. This may ensure a graceful fade-out for consecutively missing frames. Random Sign Blend for Stereo Signals
[0078] For two stereo signals (or for a pair of channels in general), the prediction processing may be performed on each of the two stereo signals separately, where the processing for the first channel may be identical to what is described above. For the second channel, a signal covariance ^^(^^) may be computed, which can then be used to derive mixing weights for the random sign blend signals such that their covariance approximately matches the measured (i.e., expected) covariance. This may involve mixing a random sign coefficient of the first channel into the random sign blend of the second channel (or vice versa). For example, the random sign blends for the first and second channels may be performed via ^^(^^, ^^, 0) = ^^(^^, ^^, 0) + ^^(^^, 0) ^^(^^, ^^, 0)(9) ^^(^^, ^^, 1) = ^^(^^, ^^, 1) + ^^(^^, 1) ^^′(^^, ^^, 1)(10) where the last indices indicate the channel index (0 or 1), and where^^′(^^, ^^, 1) = ^^0(^^) ^^(^^, ^^, 0) + ^^1(^^) ^^(^^, ^^, 1)(11) max{0, ^^(^^)}^^ =The covariance C(n) may as ^^(^^) = ^^{^^(^^, ^^, 0)^^(^^, ^^, 1)}(13a)where ^^{.. } denotes the expectation value, which may be computed in practice, similarly to theexpectation value of the signal power S(n), for example by a moving average filter, a first order IIR filter or other suitable ways. For the stereo case, the expectation value of the signal power^^(^^, ^^) for channel ^^, ^^ = 0,1 may be given by, for example,^^(^^, ^^) = ^^{^^(^^, ^^, ^^)2}(13b)
[0079] In other words, the random sign blend at step S230 may comprise, for a pair of channels, adding a mixture of the most recent available time-frequency coefficient of one of the channels multiplied by a random sign into the weighted sum (i.e., the random sign blend) of the other one of the channels. This mixture may be determined based on a per-band value of a signal cross- product with respect to the pair of channels. The per-band value of the signal cross-product may be an accumulated or propagated value (e.g., running average, first-order accumulations), for example. Limiting
[0080] There is no simple criterion for the stability of recursive vector valued linear prediction. For concealment during several time frames, the feedback loop through the MDCT buffer 20 of Fig.1 can therefore lead to instabilities in the outputs. A simple remedy foreseen by the present disclosure is to optionally apply limiting to the output bin (predicted time-frequency coefficient) such that its absolute value is never larger than a per-band limit value ^^(^^). Without intended limitation, this limit value may be computed via the signal power parameter, for instance to avoid more than a 3 dB increase of power, ^^(^^) = √2^^(^^).(14) With that, the final output of the processing for conceal may be, when limiting is used, ^^(^^), ^^(^^, ^^) > ^^(^^);^̂^ ^^ ^^
[0081] In other words, involve, before outputting a final result or before applying the random sing blend, limiting a predicted time-frequency coefficient of a given frequency band to a range that is based on a per- band value (e.g., accumulated or propagated value, such as a running average or first-order accumulation) of a signal power in the given frequency band. Parameter Update
[0082] In some implementations (notably, when the prediction parameters are continuously updated and not determined only upon loss of a frame), the states of the parameter buffer (i.e., the prediction parameters) are updated during normal (non-conceal) operation, i.e., in the first mode. In particular, updated prediction filter coefficients may be determined, for a given frequency band, based on current prediction filter coefficients for the given frequency band and respective adjustment amounts for the current prediction filter coefficients. For each prediction filter coefficient, there may be a corresponding adjustment amount. The updated prediction filter coefficients may be determined by adding respective adjustment amounts to respective current prediction filter coefficients. The adjustment amounts may alternatively be referred to as prediction filter coefficient updates, i.e., an update for each of the plurality of prediction filter coefficients.
[0083] Two different embodiments for determining the adjustment amounts (or updates in general) will be described next, namely a least squares method and a method based on a probabilistic model. Least Squares
[0084] The idea is to incrementally improve the estimate of the current MDCT frame by an additive adjustment (adjustment amount) ℎ to the current state of the predictor coefficients (i.e., the current prediction filter coefficients ^^). These updates are performed independently for each target band ^^.
[0085] Further, the adjustment amounts ℎ for the given frequency band ^^ are determined basedon the current prediction filter coefficients ^^^^(^^, ^^) associated with respective time-frequencycoefficients ^^(^^ − ^^, ^^ − ^^) of the one or more previous frames (^^ = 1, … , ^^^^) in the givenfrequency ^^ and in the one or more neighboring frequency bands (^^ = −^^^^ , … , ^^^^) of thegiven frequency band ^^, and based on the time-frequency coefficients ^^(^^ − ^^, ^^ − ^^) of theone or more previous frames in the given frequency band ^^ and in or more neighboring frequency bands of the given frequency band ^^. Specifically, the adjustment amounts ℎ are determined using a least squares method.
[0086] An example of a method 500 of determining the adjustment amounts ℎ(^^, ^^) for thegiven frequency band ^^ according to this approach is shown in the flowchart of Fig.5. Method 500 comprises steps S510 and S520.
[0087] At step S510, the time-frequency coefficient ^^(^^, ^^) of the current frame ^^ in the givenfrequency band ^^ is predicted in dependence on the adjustment amounts ℎ^^, for example via ^^^^^^^^^^ ^^ − ^^ −
[0088] The adjustment amounts ℎ^^(^^, ^^), at step S520 a squared prediction error ^^ between the predictedtime-frequency coefficient ^^(^^, ^^) of the current frame ^^ in the given frequency band ^^ andthe time-frequency coefficient ^^(^^, ^^) of the current frame ^^ in the given frequency band ^^ isminimized for a given change of the adjustment amounts for the given frequency band ^^. For example, the squared prediction error ^^ may be given by^^ = (^^(^^, ^^) − ^^(^^, ^^))2(17)
[0089] Minimizing the squared prediction error ^^ may involve determining a gradient of thesquared prediction error ^^ with respect to the adjustment amounts ℎ^^(^^, ^^) for the givenfrequency band ^^. The given change of the adjustment amounts for the given frequency band may be a change in a sum or squared sum of the adjustment amounts for the given frequency bands, for example. A maximal reduction of squared prediction error ^^ for a given change in thesquared sum of the coefficient adjustments ℎ^^(^^, ^^)2may be obtained in the opposite direction of the gradient of ^^ with respect to ℎ. This may amount to using prediction filter coefficient updates (adjustment amounts) given by ℎ^^(^^, ^^) = −^^ (^^(^^, ^^) − ^^(^^, ^^)) ^^(^^ − ^^, ^^ − ^^)(18)with a positive gain ^^. As the prediction target is a scalar, a vanishing prediction error ^^ = 0 canbe obtained by adjusting ^^ to ^^full = 1 / ‖^^‖2, where ^^^^^^^^^^ − ^^ − 2
[0090] Trade-offs may be prediction. The method proposed by the current disclosure is, in some embodiments, to •aim at using ^^ = ^^^^ = ‖ ‖2full ^^ / ^^ with a predetermined value of ^^ < 1 depending onthe physical time stride of the time-frequency transform (e.g., MDCT coding system). A good choice for a stride of around 5 ms may be ^^ = 0.5, for example• Further limit ^^ to such that the sum of squares of the coefficient update ℎ is less than ^^2This can be expressed in one step as ^^ =^^ −^^
[0091] The prediction words, the updated prediction filter coefficients are determined by adding respective adjustment amounts to respective current prediction filter coefficients. For updating the power estimates ^^(^^), ^^(^^), and ^^(^^), a simple first order accumulation may be applied, for example via ^^(^^) ← ^^ ^^(^^) + ^^ ^^(^^, ^^)2(21) ^^(^^) ← ^^ ^^(^^) + ^^ ^^(^^, ^^)2(22) ^^(^^) ← ^^ ^^(^^) + ^^ (^^(^^, ^^) − ^^(^^, ^^))2(23)A suitable choice for a 5 ms stride system may be ^^ = 0.9 and ^^ = 1 − ^^ = 0.1, for example.Probabilistic Modeling
[0092] An alternative approach for determining the adjustment amounts (or updates in general) may be to consider a predictor that furnishes a parametrized probabilistic model of the target bin.
[0093] An example of a method 600 of determining the adjustment amounts ℎ(^^, ^^) for thegiven frequency band ^^ according to this approach is shown in the flowchart of Fig.6. Method 600 comprises steps S610 and S620.
[0094] According to this approach, at step S610 a probabilistic model for the time-frequency coefficient of the current frame in the given frequency band is determined based on the current prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band, and based on the time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band. The probabilistic model determined at this step may comprise a mean ^^ and a variance of a normally distributed random variable. For example, one may assume ^^ be a normally distributed random variable with mean value ^^^^^^^^^^ ^^ − ^^ −and variance ^^2^^(^^)^^ relative to the current estimate of signal power ^^(^^). To measure the quality of match of ^^ tothe target bin ^^(^^, ^^), the nonnegative log likelihood (NLL) loss may for example be used,which may be given in this case for example by (^^(^^, ^^) − ^^(^^, ^^))2 ^^ = ^^ + ^^ +where ^^ is a constant. It is model and / or maximize the match ^^ given by the probabilistic model to the target bin ^^(^^, ^^).
[0095] In some embodiments, the parameter ^^(^^) may be used directly to obtain the blendingfactor ^^(^^) = exp ^^(^^) in the conceal mode described above.
[0096] At step S620, the adjustment amounts for the given frequency band are determined based on the probabilistic model and the time-frequency coefficient of the current frame in the given frequency band. For example, an adjustment amount for the given frequency band and for a given current prediction filter coefficient may be determined based on the probabilistic model, the time- frequency coefficient of the current frame in the given frequency band, and the time frequency coefficient that is associated with the given current prediction filter coefficient.
[0097] In one example, additive updates (e.g., adjustment amounts) to the prediction coefficients ^^ and the noise level parameter ^^(^^) may be computed by scaling of the negative gradient of ^^ with respect to ^^ and ^^ respectively, yielding ^^(^^, ^^) − ^^(^^, ^^)ℎ^^(^^, ^^) = −^^^^ ^^(^^ − ^^, ^^ − ^^)and (^^(^^, ^^) − ^^(^^, ^^))2 Δ^^ ^^ = −^^ −
[0098] Considerations be used to adjust the values of the update gains ^^ and ^^, e.g. ^^ may be defined as in Eq. (20) and ^^ may be proportional to ^^. Moreover, it may be advantageous to limit the absolute value of the resulting updated ^^ from below, for instance by -10.
[0099] In the case of probabilistic modeling, only the signal power estimates ^^(^^)need to be updated. For that, one can apply the same method as for the least squares case, for example the update procedure of Eq. (21). Modifications for Achieving Lower Complexity Combining with random repeat in high frequencies
[0100] A complexity reduction of the proposed techniques may be achieved if concealment as described above (e.g., with reference to Fig.3) is applied only to certain frequency ranges, whereas conventional techniques are applied to remaining frequency ranges.
[0101] Accordingly, in some embodiments, the aforementioned plurality of frequency bands may relate to (e.g., correspond to) a lower-frequency portion of the frequency range of the time- frequency coefficients (e.g., of the MDCT). Then, for at least one frequency band not in that range, conventional random sign repeat method may be used. That is, for at least one frequency band not in the lower-frequency portion, a most recent available time-frequency coefficient in the at least one frequency band multiplied by a random sign may be used as the time-frequency coefficient for the current frame in the at least one frequency band.
[0102] For example, it may be of advantage to run the adaptive filtering only in the low to mid frequencies (e.g., up to a threshold chosen in the range of 4-6 kHz), and then use the cheaper repeat random sign method for frequencies above that range.
[0103] Additionally or alternatively, the prediction filter length ^^ at higher frequencies may beset to be smaller (e.g., ^^ = 3 with ^^^^ = ^^^^ = 1) than at lower frequencies (for example ^^ = 6with ^^^^ = 2, ^^^^ = 1). This would be reasonable also perceptually. Block-Based Computation of Prediction Filter Coefficients
[0104] As a means to reduce mean complexity at the expense of peak complexity, the prediction filter coefficients may be computed based on stored previous time-frequency coefficients (e.g., MDCT coefficients) when the frame loss occurs.
[0105] Thus, in some implementations (notably, when the prediction parameters are not continuously updated and determined only upon loss of a frame), the states of the parameter buffer (i.e., the prediction parameters, or prediction filter coefficients) is determined during conceal operation, i.e., in the second mode. Therefore, the prediction filter coefficients may be updated directly before predicting time-frequency coefficients of the lost frame.
[0106] In particular, the prediction filter coefficients may be determined based on time-frequency coefficients of a predetermined number of previous frames. Therein, determining the prediction filter coefficients for a given frequency band among the plurality of frequency bands may be based on the time-frequency coefficients of the predetermined number of previous frames in the given frequency band and in the one or more neighboring frequency bands.
[0107] The minimum number ^^ of stored time-frequency (e.g., MDCT) frames (including the current frame) may depend on the number of prediction filter coefficients ^^ per predicted time- frequency (e.g., MDCT) bin and may be chosen to satisfy, for example ^^ > ^^^^ + ^^^^(28)where ^^ = ^^^^(2^^^^ + 1) and ^^ > 1 to make sure that the optimized prediction coefficientscapture short-term stationary statistics of the signal. Thus, the aforementioned predeterminednumber of frames may be chosen to exceed ^^^^ + ^^^^.
[0108] One approach is to determine the prediction filter coefficients based on time-frequency coefficients of the predetermined number of previous frames using a least squares method.
[0109] For example, the prediction filter coefficients may be computed by running the adaptive least squares procedure described above, on the stored time-frequency coefficients. In another example, prediction filter coefficients may be computed that minimize the total least squares prediction error with respect to all stored time-frequency coefficients (block-based) as shown below for band index ^^, where each stored time-frequency coefficient is predicted based on its own predetermined time-frequency receptive field, 2 0 ^^^^^^^^Here, it is
[0110] In general, a least squares error to be minimized for the given frequency band ^^ may be determined based on a sum of prediction errors for the time-frequency coefficients of the predetermined number of previous frames, where each predicted time frequency coefficient may be based on time-frequency coefficients of one or more preceding frames to the respective frame among the predetermined number of previous frames.
[0111] The minimization of the prediction error of Eq. (29) can be reformulated in matrix notation as ^^^^ ≈ ^^(30)where ^^ is a matrix with ^^ − ^^^^ rows and ^^ columns, ^^ is column vector of ^^ predictionfilter coefficients, and ^^ is a column vector of ^^ target values for bin ^^. The least squares solution can then be found as ^^ = (^^⊺^^)−1(^^⊺^^)(31)
[0112] The block-based method described herein gives advantage in case of multiple lost frames as shown in Fig.7 and Fig.8, which illustrate SNR measurements on concealed MDCT coefficients of size 240 in 20 scale factor bands. The error pattern is 10 lost frames every 100 frames (only SNR in lost and concealed frames shown). Fig.7 shows adaptive prediction coefficient update using 8 stored MDCT coefficients and Fig.8 shows block-based prediction coefficient computation on the same 8 stored MDCT coefficients. Combining block-based and adaptive computation of prediction coefficients
[0113] By splitting up the computation of prediction filter coefficients into the adaptive least squares approach during normal decoding (normal mode) and block-based during frame loss (conceal mode), the peak computational complexity and battery consumption can be balanced. For example, for the lower frequency range (e.g., up to a threshold chosen in the range of 4-6 kHz) prediction coefficients can be computed adaptively during normal operation, while the block-based method may be used for the higher frequencies during frame loss. Thus, for higher frequencies, computational complexity during normal operation is reduced, thereby reducing average computational load, while peak computational load upon packet loss in increased. Preferably, the prediction filter length for the block-based method may be relatively small, which will reduce the buffer size needed to store past MDCT coefficients. The prediction filter length may be chosen based on perceptual considerations, for example. Operation for the First Frame After Concealment
[0114] A possible quality improvement can be achieved by estimating (e.g., predicting) the previous (concealed) time-frequency coefficients (e.g., MDCT coefficients) from the current time-frequency coefficients. Then the right half of the time domain window of the previous frame of time-frequency coefficients can be modified by crossfading with right half of the time domain window of the estimated time-frequency coefficients. This may lead to improved aliasing cancellation for the current frame. Performance Results
[0115] Multi-Stimulus Test with Hidden Reference and Anchor (MUSHRA) listening tests were performed to evaluate the benefit of variants of the proposed method. In Fig.9 and Fig.10, the tested conditions are as follows: • mute: replacing a lost frame by silence. • mdct copy: the conventional random sign repeat method. • mdct pred: the proposed method using 6-tap filters and the least squares estimation. • mdct pred (low comp) the proposed method using 3-tap filters below 6 kHz and the mdct_copy method above.
[0116] Fig.9 illustrates listening test results for random losses with 7.5 % lost frames in a random fashion (random loss). Fig.10 illustrates listening test results for bursts of 10 lost frames in a row with 100 good frames in between (burst loss). As can be seen, the proposed method performs significantly better than conventional methods, both for the cases of random loss and burst loss. Apparatus
[0117] Finally, the present disclosure likewise relates to an apparatus (e.g., computer- implemented apparatus or apparatus having processing capability in general) for performing or implementing methods and techniques described throughout the present disclosure. For example, this apparatus may relate to an audio decoder. The audio decoder may be a low latency audio decoder, for example, that may be applicable to general audio.
[0118] Fig.11 shows an example of such apparatus 1100. In particular, apparatus 1100 comprises a processor 1110 and a memory 1120 coupled to the processor 1110. The memory 1120 may store instructions for the processor 1110. The processor 1110 may also receive, among others, suitable input data 1130 (e.g., input frames of time-frequency coefficients, etc.), depending on use cases and / or implementations. The processor 1110 may be adapted to carry out or implement the methods / techniques described throughout the present disclosure (e.g., method 200 of Fig.2, method 500 of Fig.5, and / or method 600 of Fig.6) and to generate corresponding output data 1140 (e.g., predicted / output frames of time-frequency coefficients, etc.), depending on use cases and / or implementations.
[0119] The present disclosure likewise relates to corresponding computer programs, computer program products, and computer-readable storage media storing such computer programs or computer program products. Interpretation
[0120] Aspects of the methods and apparatus / systems described herein may be implemented in an appropriate computer-based audio processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files. Portions of the audio system may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0121] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer- readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.
[0122] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, the apparatus (e.g., encoders) described above can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0123] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0124] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Enumerated Example Embodiments
[0125] Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.
[0126] EEE1. A method of packet loss concealment for frame-based audio, comprising: in a first mode, when time-frequency coefficients of a current frame of audio are available, storing the time-frequency coefficients of the current frame in a buffer; and in a second mode, when time-frequency coefficients of the current frame are not available, predicting the time-frequency coefficients of the current frame based on a set of prediction parameters and time-frequency coefficients of one or more previous frames stored in the buffer and storing the predicted time-frequency coefficients in the buffer.
[0127] EEE2. The method according to EEE1, wherein the set of prediction parameters comprises a plurality of prediction filter coefficients, including, for each of a plurality of frequency bands, prediction filter coefficients associated with respective time-frequency coefficients of one or more previous frames in the frequency band and in one or more neighboring frequency bands.
[0128] EEE3. The method according to EEE2, wherein prediction of a time-frequency coefficient of the current frame in a given frequency band is based on time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands, and on the prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands.
[0129] EEE4. The method according to EEE2 or EEE3 when depending on EEE2, wherein apredicted time-frequency coefficient ^^(^^, ^^) for current frame ^^ and frequency band ^^ is givenby ^^^^^^^^^^(^^, ^^) = ∑ ∑ ^^^^(^^, ^^)^^(^^ − ^^, ^^ − ^^)^^=1 ^^=−^^^^where ^^^^(^^, ^^) are prediction filter coefficients for frequency band ^^ in relation to time-frequency coefficients ^^(^^ − ^^, ^^ − ^^) of frame (^^ − ^^) in frequency band (^^ − ^^), ^^^^ ≥ 1indicates a number frames to be used for the prediction, and ^^^^ ≥ 1 indicates anumber of neighboring frequency bands to be used for the prediction.
[0130] EEE5. The method according to any one of EEE1 to EEE4, further comprising, in the first mode, updating the set of prediction parameters based on the time-frequency coefficients of the current frame.
[0131] EEE6. The method according to EEE5, wherein the prediction parameters are updated based on time-frequency coefficients of a predetermined number of frames including the current frame and preceding frames.
[0132] EEE7. The method according to EEE5 or EEE6 when depending on EEE2, wherein in the first mode updated prediction filter coefficients are determined, for a given frequency band, based on current prediction filter coefficients for the given frequency band and respective adjustment amounts for the current prediction filter coefficients.
[0133] EEE8. The method according to EEE7, wherein the adjustment amounts for the given frequency band are determined based on the current prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band and based on the time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band, using a least squares method.
[0134] EEE9. The method according to EEE7 or EEE8, wherein determining the adjustment amounts for the given frequency band comprises: predicting the time-frequency coefficient of the current frame in the given frequency band in dependence on the adjustment amounts; and minimizing a squared prediction error between the predicted time-frequency coefficient of the current frame in the given frequency band and the time-frequency coefficient of the current frame in the given frequency band for a given change of the adjustment amounts for the given frequency band.
[0135] EEE10. The method according to EEE7 or EEE8, wherein determining the adjustment amounts for the given frequency band comprises: determining a probabilistic model for the time-frequency coefficient of the current frame in the given frequency band based on the current prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band and based on the time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band; and determining the adjustment amounts for the given frequency band based on the probabilistic model and the time-frequency coefficient of the current frame in the given frequency band.
[0136] EEE11. The method according to EEE10, wherein an adjustment amount for the given frequency band and for a given current prediction filter coefficient is determined based on the probabilistic model, the time-frequency coefficient of the current frame in the given frequency band, and the time frequency coefficient that is associated with the given current prediction filter coefficient.
[0137] EEE12. The method according to EEE6 or any EEE depending on EEE6, wherein predicted time-frequency coefficients in the buffer are not used for updating the prediction parameters.
[0138] EEE13. The method according to EEE2 or any of EEE3 and EEE4 when depending on EEE2, further comprising, in the second mode, determining the prediction filter coefficients based on time-frequency coefficients of a predetermined number of previous frames, using a least squares method.
[0139] EEE14. The method according to EEE13, wherein determining the prediction filter coefficients for a given frequency band among the plurality of frequency bands using the least squares method is based on the time-frequency coefficients of the predetermined number of previous frames in the given frequency band and in the one or more neighboring frequency bands.
[0140] EEE15. The method according to EEE14, wherein a least squares error to be minimized for the given frequency band is determined based on a sum of prediction errors for the time- frequency coefficients of the predetermined number of previous frames, each predicted time frequency coefficient being based on time-frequency coefficients of one or more preceding frames to the respective frame among the predetermined number of previous frames.
[0141] EEE16. The method according to any one of the preceding EEEs, wherein predicting the time-frequency coefficients of the current frame further comprises limiting a predicted time- frequency coefficient of a given frequency band to a range that is based on an accumulated per- band value of a signal power in the given frequency band.
[0142] EEE17. The method according to any one of the preceding EEEs, further comprising, in the second mode, generating, for each frequency band in the plurality of frequency bands, a weighted sum of the predicted time-frequency coefficient of the current frame in that frequency band and a most recent available time-frequency coefficient multiplied by a random sign.
[0143] EEE18. The method according to EEE17, wherein a weight for the weighted sum in a given frequency band is determined based on one or more of a per-band value of a signal power in the given frequency band, a per-band value of a predicted signal power in the given frequency band, a per-band value of a prediction error of the predicted signal power in the given frequency band, and / or a per-band value of a logarithmic relative noise level in the given frequency band.
[0144] EEE19. The method according to EEE18, wherein weights for the weighted sums are successively reduced for instances of consecutive predicted frames.
[0145] EEE20. The method according to any one of EEE17 to EEE19, further comprising, for a pair of channels, adding a mixture of the most recent available time-frequency coefficient of one of the channels multiplied by a random sign to the weighted sum of the other one of the channels.
[0146] EEE21. The method according to EEE20, wherein the mixture is determined based on a per-band value of a signal cross-product with respect to the pair of channels.
[0147] EEE22. The method according to EEE2 or any EEE dependent on EEE2, wherein the plurality of frequency bands correspond to a lower-frequency portion of the frequency range of the time-frequency coefficients; and the method further comprises, for at least one frequency band not in the lower-frequency portion, using a most recent available time-frequency coefficient in the at least one frequency band multiplied by a random sign as the time-frequency coefficient for the current frame in the at least one frequency band.
[0148] EEE23. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEE1 to EEE22.
[0149] EEE24. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEE1 to EEE22.
[0150] EEE25. A computer-readable storage medium storing the program of EEE24.
Claims
CLAIMS 1. A method of packet loss concealment for frame-based audio, comprising: in a first mode, when time-frequency coefficients of a current frame of audio are available, storing the time-frequency coefficients of the current frame in a buffer; in a second mode, when time-frequency coefficients of the current frame are not available, predicting the time-frequency coefficients of the current frame based on a plurality of prediction filter coefficients and time-frequency coefficients of one or more previous frames stored in the buffer and storing the predicted time-frequency coefficients in the buffer; and updating, in the first mode, the plurality of prediction filter coefficients based on the time-frequency coefficients of the one or more previous frames and the time-frequency coefficients of the current frame; or before predicting the time-frequency coefficients of the current frame, updating, in the second mode, the plurality of prediction filter coefficients based on the time-frequency coefficients of the one or more previous frames.
2. The method according to claim 1, wherein the plurality of prediction filter coefficients comprise, for each of a plurality of frequency bands, prediction filter coefficients associated with respective time-frequency coefficients of one or more previous frames in the frequency band and in one or more neighboring frequency bands.
3. The method according to claim 2, wherein prediction of a time-frequency coefficient of the current frame in a given frequency band is based on time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands, and on the prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands.
4. The method according to claim 2 or claim 3, wherein a predicted time-frequencycoefficient ^^(^^, ^^) for current frame ^^ and frequency band ^^ is given by^^^^^^^^^^(^^, ^^) = ∑ ∑ ^^^^(^^, ^^)^^(^^ − ^^, ^^ − ^^)^^=1 ^^=−^^^^where ^^^^(^^, ^^) are prediction filter coefficients for frequency band ^^ in relation to time-frequency coefficients ^^(^^ − ^^, ^^ − ^^) of frame (^^ − ^^) in frequency band (^^ − ^^), ^^^^ ≥ 1indicates a numberframes to be used for the prediction, and ^^^^ ≥ 1 indicates anumber of neighboring frequency bands to be used for the prediction.
5. The method according to any one of the preceding claims, wherein predicted time- frequency coefficients in the buffer are not used for updating the plurality of prediction filter coefficients.
6. The method according to any one of the preceding claims, wherein updating, in the first mode, the plurality of prediction filter coefficients comprises determining prediction filter coefficient updates for a given frequency band based on the current prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band and based on the time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band.
7. The method according to claim 6, wherein determining the prediction filter coefficient updates is based on a least squares method.
8. The method according to claim 7, wherein the least squares method comprises determining the prediction filter coefficient updates based on a gradient of the squared prediction error between the predicted time-frequency coefficient of the current frame in the given frequency band and the time-frequency coefficient of the current frame in the given frequency band.
9. The method according to claim 8 when depending on claim 4, wherein the prediction filter coefficient updates are given byℎ^^(^^, ^^) = −^^ (^^(^^, ^^) − ^^(^^, ^^)) ^^(^^ − ^^, ^^ − ^^)where ^^ is a positive gain and ^^^^^^^^^^ ^^ − ^^ −10. The method according to claim 9, wherein ^^ depends on a physical time stride of a time-frequency transform used to determine the time-frequency coefficients.
11. The method according to claim 9 or 10, wherein ^^ =^^ −^^where ^^ is a12. The method according to claim 6, wherein determining the prediction filter coefficient updates for the given frequency band comprises: determining a probabilistic model for the time-frequency coefficient of the current frame in the given frequency band based on the current prediction filter coefficients associated with respective time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band and based on the time-frequency coefficients of the one or more previous frames in the given frequency band and in the one or more neighboring frequency bands of the given frequency band; and determining the prediction filter coefficient updates for the given frequency band based on the probabilistic model and the time-frequency coefficient of the current frame in the given frequency band.
13. The method according to claim 12, wherein the probabilistic model corresponds to a normal distribution.
14. The method according to claim 13 when depending on claim 4, wherein the mean of the normal distribution is given by^^^^^^^^^^ ^^ − ^^ −and the variance^^2^^(^^)^^(^^)where ^^(^^) is a logarithmic control parameter for a noise level relative to a current estimate of signal power ^^(^^).
15. The method according to claim 13 or 14, wherein the normal distribution is optimized based on a nonnegative log likelihood loss.
16. The method according to claim 14, wherein the prediction filter coefficient updates are given by () ( )ℎ^^(^^, ^^) ^^ ^^, ^^ − ^^ ^^, ^^= −^^ ^^(^^ − ^^, ^^ − ^^)^^ ^^ ^^ where updates(^^(^^, ^^) − ^^(^^, ^^))2 Δ^^ ^^ = −^^ −) with^^ =− ^^where ^^ is a^^ is proportional to ^^.
17. The method according to claim 1, wherein updating, in the second mode, the plurality of prediction filter coefficients based on the time-frequency coefficients of the one or more previous frames comprises determining the prediction filter coefficients by minimizing a sum of least squares prediction errors with respect to time-frequency coefficients of a predetermined number of the one or more previous frames, wherein each predicted time frequency coefficient being based on time-frequency coefficients of one or more preceding frames to the respective frame among the predetermined number of the one or more previous frames.
18. The method according to claim 17 when depending on claim 2, wherein the time- frequency coefficients of the predetermined number of the one or more previous frames are time-frequency coefficients in a given frequency band and in the one or more neighboring frequency bands.
19. The method according to claim 1, wherein performing both updating of the prediction filter coefficients in the first mode and in the second mode, wherein updating the prediction filters coefficients in the first mode is performed for a lower-frequency portion of the frequency range of the time-frequency coefficients according to any one of claims 7 to 11 and updating the prediction filters coefficients in the second mode is performed for a higher-frequency portion of the frequency range of the time-frequency coefficients according to any one of claims 17 to 18.
20. The method according to claim 19, wherein a threshold for determining the lower- frequency portion and the higher-frequency portion is between 4 and 6 kHz.
21. The method according to any one of the preceding claims, wherein predicting the time-frequency coefficients of the current frame further comprises limiting a predicted time- frequency coefficient of a given frequency band to a range that is based on an accumulated per- band value of a signal power in the given frequency band.
22. The method according to any one of the preceding claims, further comprising, in the second mode, generating, for each frequency band in the plurality of frequency bands, a weighted sum of the predicted time-frequency coefficient of the current frame in that frequency band and a most recent available time-frequency coefficient multiplied by a random sign.
23. The method according to claim 22, wherein a weight for the weighted sum in a given frequency band is determined based on one or more of a per-band value of a signal power in the given frequency band, a per-band value of a predicted signal power in the given frequency band, a per-band value of a prediction error of the predicted signal power in the given frequency band, and / or a per-band value of a logarithmic relative noise level in the given frequency band.
24. The method according to claim 22 or 23, wherein weights for the weighted sums are successively reduced for instances of consecutive predicted frames.
25. The method according to any one of claims 22 to 24, further comprising, for a pair of channels, adding a mixture of the most recent available time-frequency coefficient of one of the channels multiplied by a random sign to the weighted sum of the other one of the channels.
26. The method according to claim 25, wherein the mixture is determined based on a per- band value of a signal cross-product with respect to the pair of channels.
27. The method according to claim 2 or any claim dependent on claim 2, wherein the plurality of frequency bands correspond to a lower-frequency portion of the frequency range of the time-frequency coefficients; and the method further comprises, for at least one frequency band not in the lower-frequency portion, using a most recent available time-frequency coefficient in the at least one frequency band multiplied by a random sign as the time-frequency coefficient for the current frame in the at least one frequency band.
28. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of claims 1 to 27.
29. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 to 27.
30. A computer-readable storage medium storing the program of claim 29.
Citation Information
Patent Citations
Packet Loss Concealment for Critically-Sampled Filter Bank-Based Codecs Using Multi-Sinusoidal Detection
US20190081719A1
Linear prediction coefficient generation during frame erasure or packet loss
EP0673018B1
Apparatus and method for error concealment in low-delay unified speech and audio coding
US20130332152A1
Method and apparatus for controlling multichannel audio frame loss concealment
US20220059099A1
General media neural network predictor and a generative model including such a predictor
US20230394287A1