Frame recovery after frame loss using LSF stabilization
The audio decoder addresses frame loss by employing dual decoding units to generate synchronized LPC parameters, ensuring smooth transitions and reduced artifacts through energy-controlled synthesis paths, enhancing audio quality.
Patent Information
- Application Number
- PCT/EP2025/072601
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-08-06
- Publication Date
- 2026-02-12
AI Technical Summary
Existing audio codecs face challenges in managing frame loss during transmission, particularly in synchronizing decoder and encoder memories, leading to noticeable artifacts due to mismatched Linear Prediction Coefficients (LPC) and excitation parameters, and insufficient consideration of energy control during frame recovery.
An audio decoder employs a dual decoding mechanism with a first and second audio parameter decoding unit to generate distinct sets of LPC parameters for frames following a lost frame, allowing selection between two synthesis paths based on energy measurements to ensure smooth transitions and synchronized energy levels.
The solution effectively reduces artifacts and ensures seamless frame recovery by synchronizing decoder and encoder memories, improving audio quality without the need for additional bit transmission.
Smart Images

Figure IMGF000002_0001 
Figure IMGF000002_0002 
Figure IMGF000002_0003
Abstract
Description
[0001] FHIIS24EM08-2024223273. DOCX 1
[0002] Frame Recovery after Frame Loss using LSF stabilization
[0003] The present examples refer to an audio decoder. In particular, some examples refer to frame recovery after frame loss using LSF stabilization. Techniques for energy compensation are also disclosed.
[0004] Introduction
[0005] Frame Recovery is a strategy used in Audio Codecs to offer a smooth transition between a concealed frame and a received frame. If a frame gets lost, packet loss concealment (PLC) is applied, where data from the last good frame is used to render an artificial frame. If a received frame relies on information of the last frame, which is a concealed frame, then there might be a mismatch between concealed and received CELP parameters, like the Linear Prediction Coefficients (LPC) or the adaptive excitation. So, there is need to improve this transition between a lost frame and a received frame. One option is to transfer side information from encoder to decoder like energy or phase information for a guided recovery. Another option is blind recovery. In case of blind recovery, the implantation is only done on decoder side.
[0006] Prior Art
[0007] Modem speech codecs like EVS [1] or the open-source codec iLBC [2] using linear prediction to get a short-term representation of the speech signal. The Linear Prediciton Coefficients with i = 1, ... , M, where M is the filter order, are obtained by performing algorithms like the Levinson- Durbin Recursion. The Linear Prediction filter is given by the formula:
[0008] The filter is used to calculate the residual of the signal, but also to synthesize a signal. The residual or the excitation is obtained by: FHIIS24EM08-2024223273. DOCX 2 where s(ri) is the speech signal, which is usually pre-emphasized, L ist the time length of the current frame and x(n) is the residual. The speech signal can be re-obtained by the synthesis filter:
[0009] E.g. in analysis-by-synthesis codecs usually a weighted speech filter is used to obtain a weighted speech signal, to find the optimum pitch and innovation parameters, by minimizing the squared error between this weighted speech signal and the input signal. Further steps are the openloop pitch search and closed loop pitch search to find the optimum pitch lag and pitch gain, also known as the Long-Term Prediction or adaptive codebook and the innovation excitation search. The optimum parameters on encoder side are quantized and the codebook indices are sent to the decoder.
[0010] If a frame gets lost, the transmitted parameters are not received by the decoder. The PLC on decoder side is enabled to conceal a frame based on the data from the last good frame. If a new good frame is decoded after frame loss, decoder and encoder are not synchronized anymore and noticeable artifacts can occur in the speech.
[0011] In case of a guided concealment and recover speech parameters, for example energy information, phase or periodicity information are transmitted from encoder to decoder to improve the robustness of the frame loss concealment. [3], The phase control is particularly important for the recovery. After one erased frame or a block of erased frames, the decoder memory is desynchronized with the encoder memory. To resynchronize the decoder a rough position of the first glottal pulse is send. The precision the encoder the position of the pulse depends on the closed-loop pitch value for the first subframe pitch lag TO. Then the sample with the maximum amplitude is searched to find the position of the first glottal pulse, where TO is also the length of the search interval. The highest precision for the glottal pulse is achieved with a low pass filtered residual signal. In case of a high bandwidth, a periodicity parameter or voicing parameter can be transmitted. The voicing is estimated based on the normalized correlation of the signal. It can be encoded precisely using 4 bits. The voicing information is needed for highly voiced frames and better voicing resolution. The disadvantage of sending information from encoder to decoder is that bits need to be reserved (this is unwanted, and it will be shown that in the examples proposed here the decoder will not use this information, so that the encoder skips it). The most important part of the recovery is the energy control of the signal due to the strong prediction used in modem speech codecs. The energy needs to FHIIS24EM08-2024223273. DOCX 3 control in that way, that the energy at the beginning of the first good, received frame after frame erasure matches the energy at the end of the concealed frame. Also, the signal is scaled to prevent a strong energy increase in the signal.
[0012] In EVS a scaling gain is applied to the decoded speech signal [3], The scaling is done in the excitation domain to serve the long-term prediction memory for the following frame. The synthesis is done again to achieve a smooth transition from concealed frame to received frame. The excitation signal is scaled as follows: where n is a sample, x(n) is the excitation and xs(n) is the scaled excitation. L is the frame time length (in number of samples) and g(n) is the gain applied to the samples of the excitation (it is anticipated that some optional examples will be shown in which a particular g(n) is generated, different from the prior a). The gain g(n) in the prior art starts from an initial gain go and converts recursively to gi (presenting an exponential-like behavior): with g — 1) = QQ and the attenuation factor fAGC, which was determined experimentally. The gains go and gl are defined as: where E_ is the energy at the end subframe of the previous frame, Eois the energy at the beginning subframe of the current frame and E is the energy at the end subframe of the current frame. Eqis the quantized transmitted energy parameter, which is signalled in the bitstream (while E_ , Eo, and Ei are simply calculated by the decoder). If Eqcannot be transmitted it is set to E±and therefore gl = 1. Cases, where the last good frame before erasure and the first good frame after erasure is classified as a VOICED, VOICED TRANSITION or ONSET [4], Eqis calculated using: FHIIS24EM08-2024223273.DOCX 4 where ELPQis the linear prediction filter gain of the last good frame before erasure and ELP1is the linear prediction filter gain of the current frame after erasure. This is done to compensate a possible energy mismatch between excitation signal energy and the LP filter gain. Further, there are more exceptional cases, where in case of using an artificial onset frame, go is set to 0.5 * gl. If the last good frame is classified as VOICED, VOICED TRANSITION or ONSET and the first good frame after erasure is UNVOICED, go is set to g . Finally, the synthesis is redone with the synthesis filter:
[0013] If Eqis not transmitted, the recovery is done only on decoder side and no side information needs to be transmitted.
[0014] In the internet low bitrate codec (iLBC) an overlap-add procedure is performed to merge the previous excitation smoothly into the current block’s excitation to avoid discontinuity at the frame border. Therefore, a correlation between the excitation of the received frame and the excitation of the concealed frame is done to find the best phase match. First, a closer correlation estimation of the input signal near to the estimated lag is done. If the new correlation is higher than the correlation at the position of the old lag, the new lag is used for the phase match. Then, a portion of the excitation of the previous concealed frame and a portion of excitation of the received frame are copied to a new buffer with a specific length of Ltrans. This buffer is specified by: xotd( ) is the excitation of the concealed frame, Loidis the length of the concealed excitation, xin(ri) is the excitation of the received frame and Tois the estimated lag. Tois limited by Ltrans. FHIIS24EM08-2024223273. DOCX
[0015] Energy limitation is applied to the signal xtrans(ri) in case that the energy of the signalxtrans(n) is higher than xoid(n). The signal xoid(n) and xtrans(n) are than merged doing an overlap-add operation:
[0016] With 11 = 0, ... , Ltra.ns ~ 1
[0017] Drawback of the prior art include the missing or insufficient consideration of a mismatch between LPC coefficients and the excitation in case of a frame loss. In CELP, the excitation is the residual of the LPC analysis filter. The LPC represents short-term characteristic of the signal and the LPC coefficients are interpolated for each subframe through their Linear Spectral Pairs (LSP) from the previous frame and the current frame. To send the LPC to the decoder, they (the LPC coefficients) are converted into Linear Spectral Frequencies (LSFs) before being quantized. The excitation is determined based on the quantized and interpolated set of LPCs and the codebook indexes are determined. On decoder side the LSF are de-quantized.
[0018] Recently, deep neural network models like WaveNet are used as speech synthesizer. The advantage is compared to classic speech coding algorithm is the the improvement of speech quality without increasing the bitrate. In [5] generative adversarial networks (GANs) are used to create a speech signal in combination with the LPC to calculate a glottal excitation from a speech input which is fed to the neural network. The example does not provide concealment of frame losses.
[0019] Summary of the invention
[0020] Independent definitions of the examples are found in the independent claims.
[0021] According to an aspect, there is provided an audio decoder for synthesizing an audio signal from a bitstream which represents the audio signal, the audio decoder including: a bitstream receiver, to receive the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters for each frame of the audio signal, FHIIS24EM08-2024223273. DOCX 6 a first audio parameter decoding unit, to decode, for a current properly received frame, a first set of decoded audio parameters from at least the set of encoded audio parameters, a concealment unit to conceal at least one non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters; a second audio parameter decoding unit, to decode, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters from at least the set of encoded audio parameters, the second set of decoded audio parameters being different from the first set of decoded audio parameters; a synthesizing unit, to output, or derive, an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection is made between: a first version of the synthesized audio signal, synthesized from at least the first set of decoded audio parameters; and a second version of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
[0022] According to an aspect, there is provided a method for synthesizing an audio signal from a bitstream which represents the audio signal, the method including: receiving the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters for each frame of the audio signal, decoding, for a current properly received frame, a first set of decoded audio parameters from at least the set of encoded audio parameters, concealing at least one previously non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters, and to synthesize a concealed frame from the at least one set of concealment audio parameters; decoding, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters from at least the set of encoded audio parameters, the second set of audio parameters being different from the first set of audio parameters, outputting or deriving an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection is made between: a first version of the synthesized audio signal, synthesized from at least the set of encoded audio parameters FHIIS24EM08-2024223273. DOCX 7 a second version of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
[0023] Figures
[0024] Fig. 1 shows an example according to the present disclosure.
[0025] Figs. 2-5 show behaviors in the prior art.
[0026] Figs. 6-12 show examples according to the present disclosure.
[0027] Examples
[0028] Fig. 5 can be taken into account, when reading below, for distinguishing between the different frames and subframes. Fig. 5 distinguishes between different scenarios and shows formulas which are discussed here-below.
[0029] First of all, it is here indicated that the current frame is normally referred to as a q-th frame of the sequence of frames. If the (q-l)-th frame has been a non-properly decoded frame (and therefore has been concealed), then the q-th frame is a recovery frame (provided that the q-th frame is also properly decoded). Otherwise, if the (q-l)-th frame has been a properly decoded frame (and therefore has not been concealed), and the q-th frame is a properly-received frame as well, then the q-th frame is a non-recovery frame. Finally, if the q-th frame is non-properly decoded, then it is a lost frame, and is substituted by a concealed frame synthesized by the concealment unit 40.
[0030] Since many of the pedices (subscripts) of the signs would in principle carry “q” for the current frame and “q-1” for the immediately preceding frame, these pedices will not be used to avoid reading burdens.
[0031] Further, it will be understood that the q-th frame is in general substituted into N subframes. For each frame, the subframes are in general indexed with “k” (with l<k<N). Even in that case, the use of complicated pedices like “k,q” will be preferably avoided.
[0032] When referring to the LPC parameters an , ai2 the index i will refer to the index i of the filter A(z) = 1 +1ai* z~l, while the indices 1 and 2 will refer to the first and second synthesis (see below).
[0033] Further, in some cases the pedex “end” will be used. That pedex “end” is to be understood as globally valid for a whole frame, and not for a particular subframe.
[0034] It will be noted that LSF and LSP parameters are often indicated with pedex “end” (e.g. LSFend, LSPend, LSP’end) despite being often intended as global values for a particular frame. The reason is that it is intended that, at the encoder side, these values are only calculated on a particular, final sub- FHIIS24EM08-2024223273. DOCX 8 frame (i.e. the N-th subframe of the frame) and, rigorously speaking, they should be intended as parameters for that particular N-th, final subframe. Notwithstanding, it is hypothesized that those values are globally valid for the entire frame, at least according to a first approximation. It will be shown that, in some cases (e.g. for calculating the LSP parameter LSPk for each k-th subframe between the 1st(k=l) and the penultimate (k=N) subframe), an interpolation (or another weighting function) will be carried out, e.g. through interpolation weights Wk with Wk increasing (e.g. linearly) for k increasing (e.g. it may be wi being a value between 0 and e.g. 0.3 or 0.2, and WN being a value between 0.7 or 0.8 and 1).
[0035] Fig. 6 shows an example of an audio decoder 100 to synthesize an audio signal from a bitstream which represents the audio signal. The audio decoder 100 may include a bitstream receiver 5, which receives the bitstream (the bitstream receiver 5 may include a bitstream reader, not shown, which e.g. reads the bitstream e.g. from remote or from a storage unit). The bitstream may have, encoded therein, a set of encoded audio parameters (e.g., encoded versions of residual audio parameters, here indicated with ALSFqfor each q-th frame) for each frame of the audio signal. For example, Fig. 6 shows that the bitstream receiver 5 includes a dequantization block 602 (e.g. inputted with the bitstream which has been read by the non-shown bitstream reader). The dequantization block 602 may provide the encoded audio parameters as residual values (ALSFq) of linear spectral frequencies (LSFs). The audio decoder 100 may comprise a first audio parameter decoding unit 10 which can be understood substantially as operating as in the prior art. The first audio parameter decoding unit 10 may decode, for a current properly received frame, a first set of decoded audio parameters (aii) from at least the set of the encoded audio parameters (ALSFq). As shown by Fig. 6, the first audio parameter decoding unit 10 may include a block 604 for computation of first set of parameters LSPs (linear spectral pairs), which provides in output linear spectral pairs (LSPend). The first audio parameter decoding unit 10 may include a block 606 for computation of a first set of LPC parameters which may be LPC coefficients indicated here as an .
[0036] The first audio parameter decoding unit 10 may normally operate for decoding properly-received frames (good frames). The audio decoder 100 may include (not shown in Fig. 6, but shown for example in Fig. 1), a concealment unit 40. The concealment unit 40 may conceal at least one non-properly received frame (e.g. a (q-l)-th frame) based on at least one previously properly received frame (e.g. a (q-2)-th frame). The concealment unit 40 may therefore generate at least one set of concealment audio parameters (here indicated as LSPPLC). The concealment unit 40 may synthesize a concealed frame from the at least one concealment audio parameters.
[0037] Here, the way how the concealment unit 40 operates is left general: we are not really interested in how the concealment unit 40 performs the concealment of the (q-l)-th non-properly received FHIIS24EM08-2024223273. DOCX 9 frame, but, instead, in how the LPC parameters (aii) are obtained for the q-th properly received frame which immediately follows the (q-l)-th concealed frame. The properly received q-th frame which immediately follows the (q-l)-th concealed frame (substituting the non-properly received frame) is here called “recovery frame”. It is mostly intended to discuss the way of how to obtain the synthesis signal (and also the parameters for obtaining the synthesis signal) for the recovery frame.
[0038] In particular, the first audio parameter decoding unit 10 (in case of the current q-th frame being a recovery frame) may obtain the LPC parameters an by taking into account the LPC parameters used by the concealment unit 40 for performing the concealment for the preceding, (q-l)-th non- properly decoded frame. Block 604 may be inputted by a product of a predefined prediction factor m (e.g. a value between 0 and 1 / 3 (e.g. 0.333333) or a value between 1 / 4 and 1 / 2, or another natural number larger than 1 / 10 and less than 1) with concealment audio parameters ALSFq.i (e.g. LSF parameters, which may be differential parameters). Also, block 606 (downstream to block 604) may be also inputted with LSPPLC (concealment parameters, which may be LSP parameters, taken from the concealed frame). In the case that the current q-th frame is a non-recovery frame (i.e. a properly received frame which does not follow a non-properly received frame), then the blocks 604 and 606 operate normally by taking into account parameters from the immediately preceding (properly decoded) (q-l)-th frame (those parameters being also indicated with ALSFq.i). In this case, the LPC parameters an are obtained as usual, by taking into account the parameters of the immediately preceding, properly decoded frame.
[0039] The audio decoder 100 may include a synthesizing unit 50. The synthesizing unit 50 (in particular in block 608 for a first synthesis) may generate a first synthesis signal si(n) from the LPC parameters an. In the prior art, the synthesizing operation would be finished (apart from possible other operations such as post-processing and / or scaling by a gain g(n)) and the synthesized output audio signal si(n) would be provided as the output of the synthesizer unit 50.
[0040] The synthesizing unit 50 (in particular in block 608 for computation of first synthesis, inputted from by the excitation x(n) obtained from the previous frame and the LPC parameters an from the first audio parameter decoding unit 10) may therefore provide a first synthesis signal si(n).
[0041] However, the audio decoder 100 also comprises a second audio parameter decoding unit 20. The second audio parameter decoding unit 20 may decode, for the currently q-th properly received frame which immediately follows the at least one non-properly received frame, a second set of audio parameters (e.g. a second set of LCP parameters) ai2. In particular, the second audio parameter decoding unit 20 may include a block 612, for computation of second set of a LSP parameter (LSP’end) for each frame, block 612 not being inputted with the audio parameters of the immediately preceding audio frame (or, in some examples, which is inputted by a number of concealment FHIIS24EM08-2024223273. DOCX 10 parameters less than the concealment parameters used inputted into the first audio parameter decoding unit 10). Therefore, the linear spectral pair LSP’end may be provided to a block 614 for computation of second set of LPC parameters (which are indicated with ai2). Block 614 is not inputted with concealment parameters LSPPLC (or any other parameter taken from the concealment unit 40) but, rather, with LSPmean (which may be obtained from a table). The second audio parameter decoding unit 20, (in particular block 614) may output or derive the second set of decoded audio parameters (ai?) which is different (and in particular ai2 is different from the first set of decoded audio parameters aii, and in particular is not controlled by the concealment parameters of the immediately preceding concealed frame). The second audio parameter decoding unit 20 is deactivated in the case of the q-th current frame not being a recovery frame: in case of the current q-th frame being a non- properly received frame, then the synthesis is performed by the concealment unit 40, while in the case of the q-th current frame being a non-recovery frame (i.e. a properly decoded frame which immediately follows another, (q-l)-th properly decoded frame), then the synthesis is performed by the first audio parameter decoding unit 10 by keeping into account the parameters of the (q-l)-th immediately preceding audio frame (which is properly decoded), and the second audio parameter decoding unit 20 is deactivated.
[0042] The synthesizing unit 50 (in particular in block 616 for computation of second synthesis, inputted by the excitation x(n), obtained from the previous frame, and the LPC parameters ai2 from block 614) may therefore provide a second synthesis signal S2(n). The synthesis signal S2(n) will compete with the synthesis signal si(n) for being selected as the output or derived signal s(n) for the current frame.
[0043] Block 604 of the first audio parameter decoding unit 10 and block 612 of the second audio parameter decoding unit 20, may each provide, in examples, a parameter which is valid for the whole current q-th frame, while block 606 of the first audio parameter decoding unit 10 and block 614 of the second audio parameter decoding unit 20 may each provide a respective LPC set of parameters (aii, ai2) for each subframe (indeed, we use the wording “set of parameters” both because the parameters vary with the particular subframe, and also because they vary with the index i).
[0044] Summarizing:
[0045] 1) If the current q-th frame is a non-properly received frame, then the synthesis is carried out through the packet concealment unit 40.
[0046] 2) If the current q-th frame is a non-recovery frame (i.e. a properly-received frame which immediately follows a properly-received frame, i.e. both the q-th frame and the (q-l)-th frame are properly-received frames) the synthesis (to derive the output or derived synthesis signal) is carried out through the following path: FHIIS24EM08-2024223273. DOCX 11 a. Block 604, computing the LSPend from parameters ALSFq.i taken from the (q-1)- th immediately preceding frame. b. Block 606, computing the LPC parameters an , from the LSPend and the LSPend-i (which are the non-concealed LSP from the previous properly-received frame) c. Blocks 716 and 718 (blocks 60, 702, 706, 708, 360, 712 being bypassed).
[0047] 3) If the current q-th frame is a recovery frame (i.e. a properly-received frame which immediately follows a concealed frame, i.e. the q-th frame is properly received but the (q-l)-th frame is non-properly received and, therefore, has been concealed) the synthesis is carried out through the following paths: a. A first path, with: i. Block 604, computing LSPend taking into account parameters (e.g. prediction residual parameters) (or more in general concealment parameters) ALSFq.i taken from the immediately preceding frame (concealed frame) and ALSFq from the bitstream for the current q-th frame. ii. Block 606, computing the first set of LPC parameters an from LSPend and concealment parameters LSPPLC. iii. Block 608, processing the first synthesis signal si(n) from LPC parameters aii and the excitation x(n). b. A second path, with i. Block 612, computing LSP’end taking into account parameters (e.g. prediction residual parameters) ALSFqtaken from the bitstream for the current q-th frame, but not from concealment parameters (e.g. residual parameters) of the concealed (q-l)-th frame (or in some alternative examples, by taking into account less concealment parameters than the concealment parameters taken into account by the block 606). ii. Block 614, computing the second set of LPC parameters an, from LSP’end and LSP mean- iii. Block 616, processing the second synthesis signal s?(n) from the second set of LPC parameters ai2 and the excitation x(n). c. A selection (60) including: i. A first energy computation block (or more in general first energy-related measurement block, such as a first envelope measurement block) 610 on the first synthesis signal si(n) (or only an initial subframe or a group of initial subframes of the first synthesis signal si(n)) FHIIS24EM08-2024223273. DOCX 12 ii. A second energy computation block (or more in general second energy- related measurement block, such as a second envelope measurement block) 618 on the second synthesis signal S2(n) (or only an initial subframe or a group of initial subframes of the second synthesis signal S2(n)) iii. A decision (e.g. based on energy-related measurements such as those performed at first / second blocks 610 / 618 and / or on energy related measurements on the concealed frame or on the final frame or final frames of the concealed frame) on whether the first synthesis signal si(n) or the second synthesis signal S2(n) is to become the synthesis signal s(n) for the current frame. iv. Further block 623 with synthesis smoothing / scaling (see also below).
[0048] Fig. 6 also shows an excitation decoding block 622 which outputs an excitation x(n) (or z(n) if expressed as z-transform).
[0049] In order to reduce the computational burden of the two syntheses, it is possible that the following operational steps are performed:
[0050] 1) First operational step: a. Block 608 performs a partial synthesis of the first synthesis signal si(n), by only synthesizing the initial subframe or a group of initial subframes of the first synthesis signal si(n) (hence skipping a last subframe or a group of last subframes of the first synthesis signal si(n)) and b. Block 616 performs a partial synthesis of the second synthesis signal S2(n), by only synthesizing the initial subframe or a group of initial subframes of the second synthesis signal S2(n) (hence skipping a last subframe or a group of last subframes of the second synthesis signal S2(n)).
[0051] 2) Second operational step (which could precede or follow or be simultaneous with the first operational step): a. Block 610 performs an energy-related measurement on the partially synthesized version of the first synthesis signal si(n), e.g. by measuring the energy and / or the envelope (or another energy-related measurement) only in the synthesized initial subframe or synthesized group of initial subframes of the first synthesis signal si(n) (hence skipping measurements on the last subframe or a group of last subframes of the first synthesis signal si(n)) and b. Block 618 performs an energy-related measurement on the partially synthesized version of the second synthesis signal S2(n), e.g. by measuring the energy and / or the envelope FHIIS24EM08-2024223273. DOCX 13
[0052] (or another energy-related measurement) only in the synthesized initial subframe or synthesized group of initial subframes of the second synthesis signal S2(n) (hence skipping measurements on the last subframe or a group of last subframes of the second synthesis signal S2(n)).
[0053] 3) Third operational step (which follows the first operational step and the second operational step): a. Decision block 620 performs the decision based on the energy-related measurements (e.g. energy measurements and / or envelope measurements) obtained from blocks 610 and 618 only on the partially synthesized version of the first synthesis signal si(n) and the partially synthesized version of the second synthesis signal S2(n) (hence without considering the evolution of the synthesis signals si(n) and S2(n) in the last subframe or a group of last subframes of the first synthesis signal).
[0054] 4) Fourth operational step (which follows the third operational step): a. Once the selector 60 has decided, from the partially synthesized version of the first synthesis signal si(n) and the partially synthesized version of the second synthesis signal S2(n), which (between si(n) and S2(n)) is the synthesis signal to be selected as output or derived synthesis signal s(n), then the block 608 or 616 (in accordance to the selected synthesis signal) is reactivated to complete the synthesis of the selected synthesis signal through a second, partial synthesis of the selected signal (the non-selected synthesis signal being therefore disregarded without completing its synthesis).
[0055] It is noted that the first operational step may also interest the first audio parameter decoding unit 10 (e.g. in at least one of blocks 604 and 606) and / or the second audio parameter decoding unit 20 (e.g. in at least one of blocks 612 and 614), because, in order to further save computational power, it is possible to only compute those audio parameters which are in the first subframe(s) of the first and / or second synthesis signals, without computing those audio parameters in the final subframe of the first and / or second synthesis signals (in practice, initially only decoding a first subset, which is a proper subset, of the first set of decoded audio parameters, and only decoding a second subset, which is a proper subset, of the second set of decoded audio parameters): as soon as one of the two synthesis versions of the synthesis signal is selected is selected, then only the first audio parameter decoding unit 10 (in case the first synthesis is selected) or the second audio parameter decoding unit 20 (in case the second synthesis is selected) will perform the decoding of the audio parameters of the remaining, final subframe(s) of the current q-th frame and / or of calculating parameters of the last portions of the selected signal for the current q-th frame (e.g. for calculating the energy compensation gain g(n), see below) (in practice, the non-void subset of the remaining decoded FHIIS24EM08-2024223273. DOCX 14 audio parameters which were initially not decoded is only subsequently decoded, and the synthesis of the selected version of the synthesis signal is processed using the remaining decoded audio parameters of the selected synthesis, disregarding the remaining decoded audio parameters of the nonselected synthesis).
[0056] However, these operational steps are not strictly necessary, and in some (less preferred) examples the complete two synthesis signals si(n) and S2(n) are synthesized before the selection (60). In another example, the complete first and second sets of decoded audio parameters are decoded, but only the first and second partial syntheses are initially performed.
[0057] The audio decoder 100 may include (e.g., within the synthesizing unit 50), a selector 60 which selects between the first synthesis signal si(n) (or the first set of parameters an) and a synthesis signal S2(n) (or the second set of parameters ai2). In particular, the selector 60 may include a block 610 for energy computation and / or envelope evaluation for the first synthesis, which may provide a first energy-related measurement. The selector 60 may include block 618 for energy-related measurement computation (e.g. energy computation and / or envelope computation) for the second synthesis which provides an energy-related measurement (e.g. energy computation and / or envelope computation) on the second synthesis signal S2(n). A block 620 for decision of LPC set and synthesis may decide which synthesis signal, among si(n) and S2(n) to be used as a synthesis signal s(n) (output version or derived version of the synthesis signal). A block 623 synthesis smoothing / scaling may also be used, inputted with the selected signal s(n), to thereby process the output version or derived version of the synthesis signal and render it.
[0058] Fig. 1 also shows elements of the audio decoder 100 in terms of block scheme. In particular, a decision 41 is made between determining whether the current frame is a non-properly decoded frame or a properly decoded frame. If the current frame is a non-properly decoded frame (e.g. by a determination based on cyclical recurrent calculations, such as cyclic redundancy check, CRC, or the like) then the concealment unit 40 may be activated. Otherwise (if the frame is determined as valid), in block 42, it is evaluated whether the previous frame was lost. If the previous frame was lost, then the current frame is a recovery frame. Therefore, at block 43, both the first audio parameter decoding unit 10 and the second audio parameter decoding unit 20 are activated. Subsequently (block 44), the synthesizing unit 50 is invoked. In case at block 42 it is determined that the previous frame was not lost, then the first audio parameter decoding unit 10 only is activated and the synthesizing unit is activated as well in block 44, and there is an updating 45 of the memory for the next frame. This happens both in the case where the present frame is concealed frame (and the concealment unit is therefore activated), or whether the present frame is a properly decoded, recovery frame (and therefore blocks 43 and 44 are activated, therefore using both the first and second audio FHIIS24EM08-2024223273. DOCX 15 parameter decoding units (10, 20), and in the case that the current frame is a properly received frame which follows a properly received frame (i.e., the current frame is a non-recovery frame) and therefore only the decoding block 44 is activated but not block 43. After block 45, a new instance of block 41 is activated.
[0059] It will be shown subsequently that the choice between the first synthesis signal si(n) and the second synthesis signal s?(n) may be made based on the comparison between the energies of the signals, and in particular on the behavior of the envelope of the signals.
[0060] Fig. 7 shows the example 100 with other blocks, which may be optional in some examples. Fig. 7 shows, in particular, the block 620 for decision of LPC set and synthesis. Block 620 may provide both LPC parameters ai (chosen between the LPC parameters an of the first synthesis and a,2 of the second synthesis) and provide those parameters to a zero input response block 702. (This can be a outringing filter. The last e.g. 16 samples (or another amount, e.g. less than 32 samples) of the previous (concealed) frame are inputted and also a zero input excitation. Then an outringing synthesis is performed). The zero input response block 702 may therefore provide a zero input response (ZIR) indicated with szni(n). Block 620 may also output or derive the chosen synthesis signal s(n) between si(n) and s?(n) to the subtractor block 704. The synthesis signal s(n) may therefore be subtracted with the zero impulse response signal SZIR(U) in the subtractor block 704. The energy computation block 706 may compute an energy of the synthesis signal s(n). A gain computation block 708 may be used to obtain a first gain gi and a second gain g2. Values of the first gain gi and of the second gain g2 may be provided to a scaler 360 to provide a gain g(n), which may be defined sample-by- sample, and may, for each sample, take a value between the first gain gi and the second gain g2. It is to be noted that the scaler 360 may be conditioned by the concealment unit 40 and, in particular, by the energy EPLC () °f the concealed frame (it will be shown, in particular, that the scaler 360 may be conditioned by the values of the energy in a final portions (e.g. final half portion) of the concealed frame). The output of the scaler 360 may be added with the zero input response sziR(n) at adder 710. Therefore, a value ss(n) is to be provided to a residual computation block 712. A post-processing block 716 may be used from the output ss(n) of the adder 710. A block 718 of updating excitation for the next frame permits to obtain the residual computation from the output xs(n). As can be seen from Fig. 7, a block 620 for decision is also used to provide a coefficient a value (proportionality coefficient) to the gain computation block 708 and the parameters to the residual computation block 712. The audio signal ss(n) may therefore be post-processed and used and rendered as audio signal.
[0061] 1) LSF in the clean channel (first audio parameter decoding unit 10, blocks 604 and 606) FHIIS24EM08-2024223273. DOCX 16
[0062] First, a description of the behavior of the decoder 100 in the first audio parameter decoding unit 10 (in particular in blocks 604 and 606 of Fig. 6). This section discloses technique which, as such, can be already present in the prior art. Here, it is assumed that both the current frame and the immediately preceding frame are properly decoded (no concealment, no recovery). It is noted that from the bitstream, only parameters like ^LSFq-1and LSFq(residual) are received. No voicing information (e.g., “voiced” vs “unvoiced”) are received from the bitstream. It is noted that an LSF value (and also the residual ^LSFq-1) is common for the whole frame, while LSP values are different for different subframe of a frame.
[0063] First, a prediction of the LSF for the current frame may be determined (at block 604) by: where m is a prediction factor (e.g. a value between 0 and 1 / 3 (e.g. 0.333333) or a value between 1 / 4 and 1 / 2, or another natural number larger than 1 / 10 and less than 1) and ^LSFq-1is the LSF prediction residual decoded for the previous (q-l)-th frame. To get the decoded end LSF, the prediction LSFpredand the received ALSFq(possibly in dequantized form) are added:
[0064] LSFend = LSFpred+ LSFq
[0065] The term m M.SF corresponds to a moving average.
[0066] The so-obtained LSF coefficient LSFendis converted to the LSP coefficient (the following formula is valid for the whole current frame, without any distinction between the subframes): where fsis the sampling frequency. FHIIS24EM08-2024223273. DOCX 17
[0067] To obtain a set of LSP coefficients (precursors of the set of decoded audio parameters) for each subframe of the current frame, the LSP coefficients are interpolated from the LSP coefficients of the previous frame to the LSP coefficients of the current received frame for each k-th subframe out of the N subframes of the current frame:
[0068] LSPk= LSPend-! * (1 - wk) + LSPend* wk, with k = l.. N, wth k = 1.. N, where wkare interpolation weights (e.g., increasing linearly with the increase of k, e.g. it may be wi being a value between 0 and e.g. 0.3 or 0.2, and WN being a value between 0.7 or 0.8 and 1) and N is the number of subframes for the frame. LSPend-1(which a priori should be called LSPq-1 Nsince it is the LSP value of the last, N-th subframe of the immediately preceding, (q-l)-th concealed frame) are the old end LSP coefficients of the end subframe (i.e. of the N-th subframe) of the previous frame, while the LSPendis the LSP coefficient of the end subframe (i.e. of the N-th subframe) of the current frame (as explained above, used globally for the whole current q-th frame). (It is reminded that LSFend(which a priority should be written LSFq N) can be used to approximate the whole frame, even if the encoder had only calculated it on the end of the frame. We assume that the LPC does not change to much about frame, because it is a short-term representation.). The LSPk(which a priority should be written LSPq k) are converted to the LPCk(which are notwithstanding written an). This is a so-called Conversion of LSP parameters to LP coefficients (see ETSI TS 126 445 V16.2.0 (2022-03), currently available at the web page https: / / www.etsi.org / deliver / etsi_ts / 126400_126499 / 126445 / 12.06.00_60 / ts_126445vl20600p.pdf). The signal is than synthesized subframe-wise using the LPC synthesis filter with the excitation as input. See, in particular, Fig. 5.
[0069] 2) LSF while a frame loss (recovery frame, first audio parameter decoding unit 10, blocks 604 and 606)
[0070] This section explains techniques which, as such, are also in the prior art. In the prior art, however, there is not the backup provided by the second audio parameter decoding unit 20 (explained below). It is reminded that, while the ^LSFq-1(residual) is received from the bitstream, no voicing FHIIS24EM08-2024223273. DOCX 18 information (e.g., “voiced” vs “unvoiced”) are received from the bitstream. Further, an LSF value (and also the residual ^LSFq-1) is common for the whole frame, while LSP values are different for different subframes of a frame.
[0071] It is here irrelevant which technique is used for concealment (packet lost concealment, PLC).
[0072] What is important now is how to get the parameters of the recovery frame, i.e. the first properly- decoded frame after a concealed frame (non-properly decoded frame).
[0073] The LSP coefficients of the recovery frame are obtained by:
[0074] LSPk= LSPPLC* (1 - wk) + LSPend* wk, k = 0.. N, where LSPPLC(which could be written LSPend-i) are the LSP coefficients which are used for the concealed frame (i.e. LSPPLChave been inferred by taking into account at least the previously correctly received frame, e.g. the properly-received (q-2)-th frame before the (q-l)-th concealed frame). The LSPendare calculated by: with Jsbeing the sampling frequency.
[0075] The LSFendis calculated by:
[0076] LSPend = LSFmean+ m * &LSFPLC+ &LSFqwhere M,SFPLC, which can also be written as / ^LSFq-1, is a concealed version of the residual in the previous, (q-l)-th, concealed frame. It is not of interest how that delta or residual is concealed, but it is important to know that the residual of the previous concealed frame is used for the recovery frame.
[0077] There can be a great big difference between sl(n) and s2(n) in some cases, and in particular in the cases in which M.SFPLCdiffers from the lost parameters ^LSFq-r, and also in case of LSPPLC(generated by the concealment unit 40) of the concealed frame being different from the parameters LSPena-i of the lost frame. If the &LSFPLCof the concealed frame and the transmitted / ^LSFq-1, in case the frame is not lost and also if the LSPPLCof the concealed frame (i.e., the LPC coefficients FHIIS24EM08-2024223273. DOCX 19 which have been inferred during concealment) and the LSPend-1, in case the frame is not lost, distinguish from each other, here can be a high deviation of the LPC filter response between erroneous signal and clean signal. A big difference in M,SFPLCin a previous concealed frame and ^LSFq-1in a previous received frame can also means a big difference in LSFPLCand LSFqin the recovery frame and the same for LSPPLCin the previous concealed frame and LSPendin the recovery. See for example Figs. 2 and 3. This might result in an unstable synthesized signal in the recovery due to the mismatch of the interpolated LPC and the excitation. A strong oscillation and energy increase is visible. These fast changes of LPC might appear especially in frames, where an onset occurs, i.e. a change from unvoiced to voiced frames. In sharp contrast, Fig. 10 (obtained using the present techniques) shows that it is possible (e.g., by relying on the second audio parameter decoding unit 20) to obtain a more stable output signal (or derived signal).
[0078] 3) First and second set of LSF / LPC (second audio parameter decoding unit 20, blocks 612 and 614)
[0079] This section discloses techniques which are not present in the prior art, to the best knowledge of the inventors. Here, for a recovery frame, audio parameters are generated without taking into account parameters from the (q-l)-th concealed frame (and indeed ^LSFq-1is not an input to bclos 612 and 614). Subsequnelty, a choice (at block 620) will be performed on whether to use the signal si(n) (synthesized by taking into account the parameters from the (q-l)-th concealed frame) and the signal S2(n) (synthesized without taking into account the parameters from the (q-l)-th concealed frame, or taking into account less parameters from the (q-l)-th concealed frame than for si(n)).
[0080] The presented example proposes, inter alia, a technique introducing a second set of LPC coefficients to synthesize a second signal s?(n) in the decoder 100. An idea is to compare (e.g. at selector 620) the energy Eik, E2k) of two synthesis signals si(n) and S2(n) (or of initial portions of si(n) and s2(n)): the first signal si(n) based on a first set of LPC coefficients (an), as described above (e.g. obtained by taking into account parameters of the concealed frame, e.g. through LSPPLC), and the second signal S2(n) based on a second set of LPC (e.g. without taking into account parameters of the concealed frame, but taking into account parameters of the recovery frame) and to choose (in block 620) the signal (si(n) or S2(n)) which, for example, provides a smoother transition in the recovery.
[0081] The second set of LSP parameters may be obtained by using a mean LSP (LSPmean) of the recovery frame instead of LSPPLCof the previous, concealed frame. FHIIS24EM08-2024223273. DOCX 20
[0082] A motivation behind this technique is that the shape of s?(n) might be closer to the LSP in clean channel conditions. Additionally, the parameters (a,?) of s?(n) are independent from the LSP coefficients of the concealed frame. To avoid instability at the end of the recovery frame which could be caused by the previous LSF coefficients, at block 604 (in the first audio parameter decoding unit 10) the prediction of LSPendpreferably doesn’t include the moving average (previously formulated as m * &LSFq-r+ &LSFq= (1 + m z-1) * &LSF) with the previous delta LSF (&LSFq-1) and with m being pre-defined (e.g. a value between 0 and 1 / 3 (e.g. 0.333333) or a value between 1 / 4 and 1 / 2, or another natural number larger than 1 / 10 and less than 1), so that the prediction contains only the mean LSF.
[0083] The second set of LSF parameters, LSF’endis calculated (in block 612, in the second audio parameter decoding unit 20) by:
[0084] J SF' end I S1F mean + kLSFq, where LSFmeanis obtained by from a table with pre-defined values and &LSFqare the LSF coefficients read in (or derived from) the bitstream for the recovery frame (using the formula LSPen£ =
[0085] Summarizing:
[0086] The LSFgndof the second set are converted to the LSPgnd(in block 612) through
[0087] , 2*ir *LSF' end
[0088] LSP end = COS ( - — ^)
[0089] Js FHIIS24EM08-2024223273. DOCX 21
[0090] Independence of the second set of the previous frame
[0091] To avoid more impact of the previous (q-l)-th frame, for the interpolation of the LSP for each subframe, in the second set of end LSF, the LSFq-1are no longer used. Instead, the LSP are obtained by interpolating from the LSPmean, which are obtained by conversion from the LSFmean, to the new LSFend'. where wfcare interpolation weights (e.g., increasing linearly with the increase of k, e.g. it may be wi being a value between 0 and e.g. 0.3 or 0.2, and WN being a value between 0.7 or 0.8 and 1) and N is the number of subframes for the frame.
[0092] So the final LSPkfor each k-th subframe of each frame are:
[0093] First set (at the first decoding unit 10):
[0094] Second set (at the second decoding unit 20):
[0095] LSPkand LSP'kare converted into Linear Prediction Coefficients (LPC), named as and ai2. It is remembered that N is the number of subframes in the recovery frame, and k indicates the k-th subframe in the frame. It is noted that there is a LSP'endfor each k-th subframe and it could therefore be written as LSP’end k.
[0096] In order to get the second synthesis signal Si2(n), in block 614 the second set of LSP parameters LSP’k is converted (in block 614 of the second audio parameter decoding unit 20) to a second set of LPC parameters. The two synthesis signals Sii(n) and Si2(n) are therefore synthesized. The first synthesis signal Sii(n) is synthesized (at block 608 of the synthesizing unit 50) by using the first set of LPC coefficients, and the second synthesis signal Si2(n) is synthesized (at block 616 of the synthesizing unit 50) by using the second set of LPC coefficients: FHIIS24EM08-2024223273. DOCX 22 where x(n) is the excitation (e.g. obtained from block 622), M is the filter order, L is the length of the frame (in samples). The excitation x(n) may be obtained from the concealment signal of the (q- l)-th concealment frame.
[0097] (As explained above, it is actually possible to initially partially synthesize the initial portions Sii_initiaZ(n) and Si2_initiaZ(n) of the first and second signals synthesis signals Sii(n) and with n = 0, ... , L_initial — 1, i = 1, ... , M, (with L initial being the number of samples in the initial frames which are synthesized), and to synthesize the remaining part Si^_final(n) or ^fmaiW (accordingtothe selection)
[0098] In general:
[0099] - in the first audio parameter decoding unit 10 (and in particular in block 604) a first set of linear pairs LSPend are obtained by taking into account the parameters from the concealed frame (which subsequently allow, in block 606, to generate the LPC parameters an by also taking into account the concealment parameters LSPPLC and the first synthesis signal su(n))
[0100] - in the second audio parameter decoding unit 20 (and in particular in block 612) a second set of linear pairs LSP’end are obtained by taking into account the parameters written in the bitstream for the recovery frame but not from the concealed frame (and will subsequently allow, in block 614, to generate the LPC parameters a i 2 by without taking into account the parameters LSPPLC or &LSFq-1from the concealed frame, and the second synthesis signal Si2(n)).
[0101] - a competition between the first synthesis signal su(n) and the second synthesis signal Si2(n) for becoming the output or derived version of the synthesized signal s(n) is then based on comparing the energies of the two synthesis signals or of the first portions of the two synthesis signals (see selector 60, and in particular blocks 610, 618, and 620). The explanation is below. FHIIS24EM08-2024223273. DOCX 23
[0102] Due to the realization that the mismatching LPC coefficients in the recovery might yield in a strong energy increase, the subframe energies of both first and second synthesis signals Si(n) and s2(n) are compared. In general terms, the signal (among signals s1(n) and s2(n)) with the smoother energy envelope compared to the energy in the last subframe of the concealed signal is chosen. The subframe energy of the last subframe in the concealed signal and the energies El kand E2 kof the first three (k=l, 2, 3) subframes of the synthesis signals $i(n) and s2(ii) are determined: with L-L being the number of samples for each subframe (e.g. such that 3 * = L / N (N being the number of subframes, e.g. with N=3 we have L = L / 3 in such a way that Lr+ L2+ L3= IV; the symbol “=” being used instead of “=”, for example, for keeping into account the possibility that some subframes have not exactly the same number of samples, e.g. by virtue of the total number of samples N not being divisible by 3), and k being the index of the subframe. The energies may be scaled or divided by the length of the subframe or the length on which the energy is calculated on to be comparable in case the frame, half frame or subframe length are changing from the concealed frame to the received frame. However, in some examples the formulas above may be substituted by
[0103] At least one condition (e.g. two conditions) can be taken into account. Here below, some conditions may be used alone, but in some examples more than one condition are evaluated. For this reason, it is sometimes written in terms like “if the ... condition is fulfilled, then the first / second synthesis signal is preferentially selected”, where “preferentially” means that the particular synthesis signal is selected either tout-court (e.g. in the examples in which only one condition is present) or that the particular synthesis signal is selected provided that other conditions are fulfilled.
[0104] First condition:
[0105] To check whether the second synthesis signal s2(ii) is to be used, the following first conditions may be considered: FHIIS24EM08-2024223273. DOCX 24 lli 1 ?
[0106] — - > thr&& — — > th2
[0107] E2,I E2I2where “&&” means logical operator “AND”. E1 ±is a measurement of the energy of the first subframe of the first synthesis signal Si(n) and E2 1is a measurement of the energy of the first subframe of the second synthesis signal s2(n) , E1 2is a measurement of the energy of the second subframe of the first synthesis signal Si(n) and E2 2is a measurement of the energy of the second subframe of the second synthesis signal s2(n). th±may have a value of 1.455 (or more in general between 1.4 and 1.5, or even more in general a value which is 1 or larger than 1) and th2may have a value of 1.5 (or more in general between 1.45 and 1.55, or even more in general a value which is 1 or larger than 1), and it may be preferably thr< th2, and it may be thr> 0 and th2> 0. The preferred values were determined empirically. If the first condition is fulfilled, then the second synthesis signal s2(n) (and the second set of LPC) is selected to be used, subjected to the second and / or third conditions, in examples.
[0108] In practice, the first condition may be generalized as: if the energy measurement Ei rof the first subframe of the first synthesis signal Si(n) is larger than the energy measurement E2 1of the first subframe of the second synthesis signal s2(ii) by a predefined amount (e.g. 45% in the case of thr= 1.45), and if the energy measurement E1 2of the second subframe of the first synthesis signal s^ri) is larger than the energy measurement E2 2of the second subframe of the second synthesis signal s2(n) by a predefined amount (e.g. 50% in the case of th2= 1.5) then the second synthesis signal s2(n) is preferentially selected (subjected to the other conditions, if present). Otherwise, if the energy of the first subframe of the first synthesis signal is not larger than the energy of the first subframe of the second synthesis signal by the predefined amount (e.g. 45% in the case of thr= 1.45), or if the energy E1 2of the second subframe of the first synthesis signal $i(n) is not larger than the energy of the second subframe E2 2of the second synthesis signal s2(n) by the predefined amount (e.g. 50% in the case of th2= 1.5) then the first synthesis signal Si(n) is preferentially selected (in some examples, even without checking other conditions, even if present).
[0109] The first condition may be generalized even more as: if, along a number NMAX>2 (with NMAX<N or NMAX<N) of the first consecutive subframes, the energy of the first synthesis signal s^ri) evolves coherently over (e.g. by at least a threshold e.g. of at least 40%) the energy of the first synthesis signal, then the second synthesis signal is preferentially selected (subjected to the other conditions, if present). Otherwise, if in at least one subframe of the first N\[ \x>2 (with NMAX<N or NMAX<N) consecutive subframes the energy of the first synthesis signal does not evolve coherently FHIIS24EM08-2024223273. DOCX 25 over the energy of the first synthesis signal (e.g. if in at least one of the first NMAX consecutive subframes the energy of the first synthesis signal is not larger than the energy of the first synthesis signal in the corresponding subframe) then the first synthesis signal Si(n) is chosen (in some examples, even without checking other conditions, even if present).
[0110] Second condition:
[0111] To avoid an energy decrease by the second set, the energy of the first two subframes (E2 rand E2 2) of the second synthesis signal may be compared to the energy (EPLCsub) of the last subframe in the concealed frame, which is derived in same way like the subframe energy of the received frame: where “&&” means the logical operator “AND”. EPLCsubmight be calculated the same way like
[0112] E-i k and E2jfc, EPLCsub= (but in some examples it could be EPLCsub=
[0113] SPLC [ft]2) on the last subframe, where k is the index of the last subframe of the con¬ cealed frame and SpLC[n] being the concealed signal. thPLCmay have a value of 1.5 (or more in general between 1.45 and 1.55, or even more in general a value which is 1 or larger than 1 ; in some examples, thPLC= th2, and / or thPLC> th2) which was conducted experimentally. Each ratio
[0114] E2I E2 2and - ' - has to be higher than the threshold thPLCto fulfil the second condition. If
[0115] EpLCsub EpLCsub the second condition is fulfilled, the second synthesis signal is preferentially chosen (in some examples, subjected to the fulfilment of other conditions, e.g. the first condition and / or the third condition).
[0116] According to the second condition, if at least one of both and
[0117] ^2 2 1
[0118] - ’■ - > triple is not verified EpLCSUb the first synthesis signal $i(n) is selected. FHIIS24EM08-2024223273. DOCX 26
[0119] Further, the subframe energy (E1;1and E1 2) of the first and second subframe of the first synthesis signal compared to the energy (EPLc) of the last subframe of the concealed signal has to be higher than a certain threshold (e.g., the same thPLC) where “&&” means the logical operator “AND”. This condition also prevents that the second synthesis signal s2(n) is used when the energy difference of the first set compared to the concealed signal (or at least to the last subframe of the concealed signal) is small.
[0120] Therefore, the second condition may be generalized in that: if the energies (Ei,i, EI,2, E2,I, E2,2) of both the first subframe and the second subframe of both the first synthesis signal $i(n) and the second synthesis signal s2(n) are larger than the energy of the last subframe of the concealment signal by a predefined amount (e.g. 50%), then the second synthesis signal s2(ii) is preferentially chosen (in some examples, subjected to the fulfilment of other conditions, e.g. the first condition and / or the third condition). Otherwise, if a least one of the energies (Ei,i, EI,2, E2,I, E2,2) of the first subframe and the second subframe of at least one of the first synthesis signal Si(n) and the second synthesis signal s2(n) is not larger than the energy of the last subframe of the concealment signal by the predefined amount, then the first synthesis signal Si(n) is preferentially selected (in some examples, without examining the fulfilment of other conditions, if present).
[0121] The second condition may be generalized eve more in that: if the energies (e.g. Ei,i, EI,2, E2,I, E2,2) of a number Nmax (with Nmax between 2 and N or between 2 and N-l) of initial consecutive subframes of both the first synthesis signal Si(n) and the second synthesis signal s2(ii) are all larger than the energy of the last subframe of the concealment signal by a predefined amount, then the second synthesis signal s2(ii) is preferentially chosen (in some examples, subjected to the fulfilment of other conditions, e.g. the first condition and / or the third condition). Otherwise, if at least one of the energies (e.g. Ei,i, EI,2, E2,I, E2,2) of the first Nmax consecutive subframes of at least one of the first synthesis signal Si(n) and the second synthesis signal s2(n) is not larger than the energy of the last subframe of the concealment signal by the predefined amount, then the first synthesis signal Si(n) is preferentially chosen (in some examples, without examining the fulfilment of other conditions, if present).
[0122] Third condition (optional) FHIIS24EM08-2024223273. DOCX 27
[0123] Further some heuristic checks are introduced to guaranty that the second synthesis signal is only chosen if the energy envelope of the first synthesis is changing or increasing fast: where “II” means logical operator “OR”. The Threshold th3may have a value of 2 (or more in general between 1.5 and 2.5, or even more in general a value which is 1 or greater than 1; it may be th3= thpic and / or th3= th3and / or th3> t / l^), th4may have a value of 300 (or more in general larger than 100 or between 100 and 500, more in particular between 200 and 400, and even more in particular between 250 and 350; it may be th4> th3and / or t / i4> th2and / or t / i4> th4and / or th4> thPLC) and th3may be 500 (or more in general more than 100 or between 300 and 700, more in particular between 400 and 600, and even more in particular between 450 and 550; it may be th5> th4and / or th3> th3and / or th5> th2and / or th3> th±and / or th$ > thpLc). In examples, it may be that th4 / th3is 150, or more in general between 200 and 300; and / or t / i5 / t / i4=l.666667 (or more in general between 1.5 and 1.8); and / or th5 / th3= 250 (or more in general between 200 and 300). In examples, it may be that th3 / thr>1.36 (in some examples by at least 1.2). At least one of these conditions must fulfilled to allow the second synthesis signal to be used. However, this third condition is optional and can be dropped.
[0124] Summary of the conditions
[0125] Finally, the second synthesis signal is used, when:
[0126] In the case that all the three conditions are used, then the second synthesis is used when: FHIIS24EM08-2024223273. DOCX 28
[0127] Calculation of energy compensating gain
[0128] Since the second set of LPC does not work for every frame especially where the energy differences are small an additional smoothing is done to improve the performance of the recovery.
[0129] Therefore, an energy compensating function g(n) is applied to the selected synthesis signal which is controlled by two gains gi and g2. The first gain gl may be determined by the ratio of the energy of the last half frame (or at least one last portion) of the concealed signal in the concealed frame and the first half frame (or at least one first portion) of the recovery frame.
[0130] For the gains the half frame energies may be used, so the energy calculation is: where L2is the length of a half frame (L2— L / 2) and k are the indices of the half frames (in some examples, it could
[0131] The first gain grmay be obtained by: FHIIS24EM08-2024223273. DOCX 29 which is the start gain of the scaling function g(n). c^may be limited to a ceiling value of 1.2 (or more in general a value D with 1<D<2, e.g. a value between 1.1 and is larger than 1.2, g will be set to 1.2 (or D), so that g±is between 0 and 1.2 (or between 0 and D). EPLChal^ may be calculated in the same way like Ej-rame^ but on the second half frame of the concealed frame be-
[0132] The second gain g2 may be calculated as follows:
[0133] In some cases, however, this formula may have a ceiling in e.g. g2= 1 in caseEPLChalf+a*(Eframe2~ EPLC HALF) .
[0134] - - - — > 1). Therefore, g2may be defined as being always 1 or less than 1, E frame 2 and therefore it may be guaranteed that g(n), at least in its final portion, has an attenuating (or at least a non-amplifying) effect.
[0135] In practice, g(n) permits to scale, sample by sample, the output or derived version s(n) of the synthesized audio signal in the properly received frame immediately following the at least one previously non-properly received frame by an energy compensating gain greater than 0, the energy compensating gain evolving, monotonically (e.g. strictly monotonically) or constantly, from the first value grtowards the second value g2. g may be: comparatively high (e.g. larger than 1) in case of a ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version (s) of the synthesized audio signal in an initial portion of the current frame, is comparatively high; and comparatively low (e.g. closer to 0) in case of the ratio, between the energy (EpLChai ) of the concealed signal in the last portion of the concealed frame and the energy FHIIS24EM08-2024223273. DOCX 30 of the output or derived version (s(n)) of the synthesized audio signal in an initial portion of the current frame, is comparatively low.
[0136] For example, if EPLChalf» Eframe^, then gi is also high (e.g. higher than 1, e.g. reaching the ceiling value), while if EPLChalf« Eframe^ then gi is also low (e.g. closer to O).g2may be conditioned by a conditioning term (a which may be proportional to the difference between the energy of a last portion of the of the current frame and the energy (EPichalf) °f the last portion of the concealed frame. g2may be: .. EPLChalf) comparatively high (e.g. closer to 1) in case ot a ratio [ - - - — ] be-
[0137] Eframe2tween the energy (EPLCfialf) of the concealed signal in the last portion of the concealed frame, added with the conditioning term (a and the energy of the output or derived version (s) of the synthesized audio signal in the last portion of the current frame, is comparatively high; and +- EpLChalf) comparatively low (e.g. closer to 0) in case ot the ratio [ - - - — ] be-
[0138] E frame2tween the energy (EPLCfialf) of the concealed signal in the last portion of the concealed frame, added by the conditioning term (a the energy (Eframe2) of the output or de rived version (s) of the synthesized audio signal in the last portion of the current frame, is comparatively low.
[0139] (Notably, g2 may be the same of the ratio - - - Or, g2 may be de-
[0140] E frame2fined so that, the higher the ratio, the higher g2 (e.g. with a ceiling and / or a floor)).
[0141] E.g., if EPLChalfand Eframe^ are closer to each other, then g2is higher (e.g. closer to 1) than if
[0142] EpLchaifand Eframf>2are distant from each other.
[0143] With reference to Fig. 9 (chart (b)), when Eframe2is smaller than or equal to EPLCfialf, (i.e. if
[0144] ^frwnic
[0145] - - - < 1), then g2may be set to be constantly 1 (ceiling value). For Ejrame2> EPLChalf(i.e.
[0146] ^frwnic ^frwnic
[0147] - - > 1), then g2lies between 0 and 1 (and the higher - the lower g2 apart from a possi-EpLchalfEpLchaifble floor). The term a—EPLchalf) is a conditioning term which ensures that the energy compensation is not too strong if there is a big difference between EPLChalfand Eframf>2. FHIIS24EM08-2024223273. DOCX 31
[0148] The gain g(n) can be between 0 and 1 (and therefore an attenuation is performed) or (at least for some samples) larger than 1 (and in that case being amplifying). In some cases, the gain amplifies in a first part of the frame and attenuates in a last part of the frame.
[0149] In general terms, the gain g(n) causes an energy compensation: the more the energy of the last portion of the concealed signal in the concealed frame is different from the energy of the initial portion of the output or derived synthesis signal in the current q-th frame, the larger the conditioning caused by g(n). In case, for example, In particular:
[0150] 1) If EPLChalf< Eframe^ then the gain g(n) is, at least at the start of the q-th frame, attenuat-
[0151] EpLC alf ing (because 1), meaning that the energy of the first portion of the q-th frame is E frame1attenuated, to avoid an unwanted step from the last portion of the (q-l)-th concealed frame (which has less energy) (and the higher the distance, the higher the attenuation); a. (In particular, if EPLChalf« Erame, then the gain g(n), at least one at the start of the q-th, goes towards 0 or to a floor value of gi)
[0152] 2) If EPLChalf> Eramethen the gain g(n) is, at least at the start of the q-th frame, amplify-
[0153] EpLChalf ing (because 1), meaning that the energy of the first portion of the q-th frame is E frame
[0154] A amplified, to avoid an unwanted step from the last portion of the (q-l)-th concealed frame
[0155] (which has more energy) (and the higher the distance, the higher the amplification); a. (In particular, if EPLChalf» Eframe^ then the gain g(n), at least one at the start of the q-th, goes towards a value greater than 1 or to a ceiling value of gi)
[0156] EPLC half then the gain g(n) is unitary (because 1) at least at the E frame1start of the q-th frame (e.g. at the very first sample of the q-th frame), meaning that the energy of the first portion of the q-th frame does not need to be compensated because is the same of the energy of the last portion of the (q-l)-th concealed frame (which has more energy);
[0157] 4) If EPLChalf+ a <Eframe2- then then the gain g(n) is, at least at the end of the q-th frame, attenuating (because of I < 1) (and the l Eframg2higher the distance, the higher the attenuation, despite the attenuation being to the conditioning term) FHIIS24EM08-2024223273. DOCX 32
[0158] 5) If EPLChalf+ a * Eframe2- EPLChalf) > Eframe2, then then the gain g(n) is, at least at the end of the q-th frame, set to a ceiling value, e.g. 1;
[0159] 6) If EPLChalf= Eframe2, then the gain g(n) is, at least at the end of the q-th frame, unitary (because a. (in particular, if EPLChalfand Eframe 2close with each other, then the gain g(n) goes towards 1 (ceiling value), at least at the end of the q-th frame) b. (if EPLChalf« Eframe2) then the gain g(n) goes towards a value less than 1, at least at the end of the q-th frame)
[0160] It is to be noted that g(n) may evolve constantly (e.g. if gr= g2or monotonically (e.g. strictly monotonically). In some examples, weak monotonicity may be also possible (e.g., following possible quantization of values of g(n) and / or in the case that a ceiling value or floor value is taken for an interval of samples).
[0161] The proportionality coefficient a is a proportionality factor which depends on the energy in the first subframes of the chosen synthesis. The dependency may be described by a linear function a = m * Erei+ c. (See Fig. 12), where Ereimay be the maximum of the ratio between the first subframe and the last subframe of the concealed frame and the second subframe and the last subframe of the concealed frame of the chosen synthesis, m may have a value like -2.14e-05 or -2e-05 < m < -3e-05 (or another value, which may be a negative value) and c a value 0.1999 < c < 0.20001 (or another value, which may be a positive value). For Eretsmaller than 1, which means energy in PLC is bigger in the last subframe, a may have a constant value like 0.2, or 0.1999 < a < 0.20001. More in general, the coefficient a may be obtained as a linear combination of Erei, e.g. with a negative angular coefficient m and / or positive constant term.
[0162] If the first synthesis signal (si(n)) is chosen, Ereimay be obtained by:
[0163] For the case the second synthesis signal (s?(n)) is chosen, Ereimay be obtained by: FHIIS24EM08-2024223273. DOCX 33
[0164] For bigger than a value Emaxlike 7000 or for example 6900 < Eret< 7100, a constant value a is chosen, like a = 0.05. So a may be obtained by:
[0165] An example of the function a = m * Eret+ c is provided by Fig. 12, having in abscissa ordinate the value a = m * Erei+ c (with c=0.2). a may be not negative, and therefore we don't move from the first quadrant.
[0166] In practice, a may be obtained from a linear combination (e.g. with negative coefficient m) of Erei, but it may have a ceiling value (e.g. where 1 < Erei< Emax) and / or a floor value (e.g. a = 0.05 if Erei> Emax).
[0167] The final applied scaling gain g(n) (here below being represented as “g[n]” without any distinction from “g(n)”) is: where L is the length of the frame and is the initial value of g. AGC is the active gain control with a value of 0.98 (or more in general a value between 0.9 and 0.999, e.g. between 0.97 and 0.99). The gain g(n) is computed in block 708. The function can also be written as where AGC is a base and n+1 is an exponent. FHIIS24EM08-2024223273. DOCX 34
[0168] Fig. 9 shows in chart (a) the energy compensating factor g(n) evolving along the L samples of the q-th frame for three different values of alpha (0.2, 0.01, and 0.0005), in the case second synthe-
[0169] E '^I’CLTTLC sis being selected and - - = 1000. In this case, the energy compensation is an attenuation, be-
[0170] EP^half cause the energy compensating gain is between 0 and 1. In this case, the higher the ratio -
[0171] EpLchalf the higher the attenuation (i.e., the closer g(n) is to 0, g(n) being greater than 0). Further, the lower the alpha, the higher the attenuation (i.e., the closer g(n) is to 0, g(n) being greater than 0). The smoothing is done directly in the synthesis domain, so that the impact of the previous filter memory is not considered. This could in principle lead to discontinuities at the frame border. To prevent these discontinuities the zero input response sZIRn) based on the synthesis memory of the erased frame is removed from the signal first (in particular in correspondence of the first subframe). The
[0172] ZIR output is derived (at 702) by using the synthesis filter with a zero-signal input:
[0173] M i=l is the length of a subframe (e.g. with N subframes, e.g. with N>2, such as N=3 or N=4 or N=5, e.g. Li=L / N, eg. L / 3, or L / 4 or L / 5), x0(n) are just zeros (determined at 702), M is the filter order, are the LPC coefficients. The ZIR s'(n) is removed (at 704) from the unsealed synthesis s(n) (e.g. as outputted by block 620):
[0174] Then, the scaling function is applied at 360:
[0175] Additionally, the memory of the synthesis (e.g. in block 360) for the next frame is also scaled:
[0176] Finally, the ZIR is added (at 710) on top of the scaled synthesis: FHIIS24EM08-2024223273. DOCX 35
[0177] The excitation is updated based on the new scaled synthesis by using the analysis filter at 712:
[0178] Fig. 10 shows the result of using the second set of LPC (second audio parameter decoder 20).
[0179] Compared to the signal after a frame loss in Fig. 4, the second synthesis signal S2(n) is clearly stable and similar to the clean signal.
[0180] Here above and below, reference is normally made to energy-related measurements in parti cu-
[0181] , . . . . . . .Flar the form ot average energies in subframes, such ask= and c2k=
[0182] ’ t-i being the length of the subframe for which the energy is calculated. How- ever, in some cases it is also possible to simply measure the integral value of the energy, such as 2T• 1 • 1 • In particular, it may be unnecessary to e ratios -S1 W2tfk-l>LiS2 M calculate th and -2if, for example, the evaluation of the condi-
[0183] Li
[0184] Fa a Fq 1 T tion — - > tht is to be performed, because the ratio — — would notwithstanding cancel the values ^2,1 ^2,1 at the numerator and the denominator. Notwithstanding, it is at least theoretically possible to have that the subframes for which the energy-related measurements are calculated are different. E.g. a first subframe for the first synthesis signal could have length while a second subframe for the second synthesis signal could have length L1 2)- In this case, it could be opportune to average the l t d t energy-related measurements e.g. t thhroug hh F -Z£XV1 [n] 2,P- . Or, it could be possible to modify the thresholds (e.g. thr) to keep into account the different lengths, while using the integral values, such as $i [ft]2and F2= sz [nJ2• isnotwithstanding here supposed that all the subframes have the same length FHIIS24EM08-2024223273. DOCX 36
[0185] (apart for the possibility that some subframes have not exactly the same number of samples, e.g. by virtue of the total number of samples N not being divisible by 3).
[0186] Further, it is noted that the energy-related measurements are not uniquely energy measurements. For example, envelope measurements (which are also energy-related measurements) may be performed. The envelope can be described as the change (progression) of the energy over time (e.g., over the samples). For example, blocks 608 and 610 may obtain an envelope of the first synthesis signal Si(n) and the second synthesis signal s?(n), respectively. As a condition evaluated by block 620 of the selector 60, the envelopes may be evaluated so as: to select the second synthesis signal (s2) in case the envelope of the second synthesis signal (s2) is, at least in the initial subframe or in a sequence of initial subframes, more stable, by at least one predetermined extent (e.g. based on at least a predetermined threshold), than the envelope of the first synthesis signal (si), and to select the first synthesis signal otherwise.
[0187] The stability of the signals may be evaluated, for example, by measuring the fluctuations of the signal over time. Fig. 4 shown, for example, a highly fluctuating signal (“Synthesis signal obtained using the parameters of the concealed frame: set 1”). Measurements of fluctuations of an envelope are known in the art, and can be based, for example, on variance measurements, etc. In general terms, however, where the envelope of the second synthesis signal has a stability which overwhelms the stability of the first synthesis signal by at least the predetermined extent (by the predetermined threshold), then the second synthesis signal is enough stable and can be used instead of the first synthesis signal as selected synthesis signal (and as derived or output synthesis signal).
[0188] Notably, the first and second conditions described above may be used, in some examples, as evaluating the stability of the first and second synthesis signals.
[0189] Further characterization of figures
[0190] Figs. 2 and 3 show the LPC frequency response of the concealed frame (Aq dec) vs. the clean signal (Aq dec clean).
[0191] Fir. 4 shows the clean signal compared to the synthesis signal (e.g. si(n)) obtained using the parameters of the concealed frame (e.g., like in the prior art, or as outputted by the first audio parameter decoding unit 10), and shows evident oscillations.
[0192] Fig. 9 shows:
[0193] • In chart (a), the behavior of g(n) along the samples of the current q-th frame parametrized for different proportionality coefficients a (in this case, the higher the a, the higher the attenuation along the samples, in case of - - = 1000)
[0194] EpLChalf FHIIS24EM08-2024223273. DOCX 37
[0195] • In chart (b), the behavior of g(n) along the samples of the current q-th frame parametrized for different ratios E^rame^ / EPLChalf(notably, the higher the ratio, the higher the attenuation along the samples, in this case).
[0196] Fig. 10 shows the clean signal compared to the synthesis signal (s2(n)) obtained without using the parameters of the concealed frame (as outputted by the second audio parameter decoding unit 10), and shows a stable behavior as compared to that of Fig. 4.
[0197] Some discussion on the present techniques
[0198] It has been explained above that the current q-th frame and the (q-l)-th concealed frame are partitioned both: according to a first partitioning which partitions the current q-th frame and the (q-l)-th concealed frame among a sequence of subframes in a number N of subframes which is 3 or more than 3 (examples of these subframes are Ei r, E21, E1 2, E2 2and EPLCsubetc.), the decoded audio parameters being decoded for each subframe, the selection at 60 being performed only based on an initial subframe or a group (sequence) of initial subframes of the current q-th frame and on the last subframe of the (q-l)-th concealed frame (further, in case of partial synthesis before the selection 60, the synthesis being originally only performed for on single initial subframe or a particular sequence of initial subframes; further and Eretbeing calculated based on the first subframe only); according to a second partitioning which partitions the current frame and the concealed frame among a sequence of portions (e.g. half frames) in a number of portions which is 2 (in the case of half frames) or more than 2, but the number of portions being less than the number N of subframes, and each portion having larger time length than any subframe (this partitioning is preferably used for calculating the energy compensating gain g(n); examples are
[0199] It has been noted that, in this way, diversity is increased.
[0200] It is not strictly necessary to perform the decoding of the audio parameters to obtain LFS parameters, then LSP parameters, and finally LPC parameters. While the examples above have been mainly directed to that technique, other techniques may be implemented.
[0201] Also other representations like ISF (ImmittanceSpectral Frequency) of the LPC or any other could be used with this approach (see [6]).
[0202] While the second audio parameter decoding unit 20 mainly refrains from adopting the parameters from the concealed signal, it is notwithstanding noted that the excitation x(n) can be obtained from the (q-l)-th concealed audio frame. In several examples, however, the encoded parameters are differential parameters and the audio signal is synthesized by taking into account an excitation FHIIS24EM08-2024223273. DOCX 38
[0203] (which can be obtained, or at least inferred, in some examples, from the concealed audio signal in the (q-l)-th concealed frame).
[0204] Some advantages of the new technique
[0205] A major advantage of aspects of the proposed technique is the consideration of the mismatch between LPC parameters of a concealed loss frame and a first well-received, aka recovery, q-th frame. The prior art methods propose smoothing operations of the excitation domain signal but don’t treat the mismatch sets of LPC into account. Further, the proposed technique operates solely on decoder side and requires no extra side information to be transmitted.
[0206] Neural network
[0207] An implementation based on a neural network is illustrated in Fig. 11, where the synthesizing unit 50 of Fig. 1 includes a neural network, NN-based synthesizing unit 90, which implements a NN. In particular, in the example of Fig. 11 block 620 (and the selector 60) is represented as being part of the block “Recovery and Computation / Decision of speech parameters” 91 (which implements blocks 10 and 20, in particular), while block 623 is part of the NN-based synthesizing unit 90.
[0208] Transmitted speech parameters like LPC or any representation of it may be used as input for a neural network NN (in a NN processor 90) with at least one learnable layer (e.g. with a plurality of learnable layers). The learnable layer may be a generative adversarial neural (GAN) layer.
[0209] In case of concealment, concealed parameters (generated by the concealment unit 40, which is not necessary part of the NN-based synthesizing unit) such as LCP parameters an and ai2 (see above) are fed to the NN-based synthesizing unit 90. In case of a properly received frame which is not a recovery frame (i.e. the immediately preceding frame is also a properly received frame), the received speech parameters are used (i.e. the first set of parameters is used, and the first audio parameter decoding unit 10 is activated, while the second audio parameter decoding unit 20 is deactivated). In case of recovery frame (i.e. properly received frame after a concealed frame), the block 620 (which may be a deterministic block) decides which synthesis signal between si(n) and s?(n) is to be used. Speech (audio) parameters are in this case used by the NN-based synthesizing unit 90 (also implementing blocks 612, 614, and 616) to generate the synthesis signal, e.g. based on the excitation (e.g. previously obtained). Here, a NN may be used. If The proposed blocks 604, 608, 610, 614, 616, 618 may be be put in front of a learnable layer of the NN-based synthesizing unit 90 where the prediction is created from the LPC, which are converted from LSP from the first set or FHIIS24EM08-2024223273. DOCX 39 second set in recovery case, LSPpicin case of concealment or LSP from the first set in clean channel case.
[0210] Pseudo - Code
[0211] / * 1. , first audio parameter decoding unit (10), Unquantize LPC * /
[0212] / * goes to block 604 * / de Quantization Isf, Isp, .... );
[0213] / * goes to block 606 * /
[0214] InterpolationOfLSP(lsp ,.... );
[0215] / * 604 Computation of first set of LSP * / dequantization()
[0216] {
[0217] Decode( TransmittedLSF, LSF end, ... );
[0218] Convert( LSF end, LSP end);
[0219] }
[0220] Decode()...
[0221] {
[0222] / * 604: Prediction * /
[0223] Prediction = MeanfLSF + factor * PreviousdeltaLSF;
[0224] AddPredctionToTransmittedLSF( TransmittedLSF, Prediction, ... );
[0225] / * 604 Computation of first set of LSP: Conversion from LSF end to LSP end* /
[0226] }
[0227] Convert()...
[0228] {
[0229] LSPend = cosf ( LSPend * PI / (fs / 2) ))
[0230] }
[0231] / * 606 Computation of first set of LPC * /
[0232] InterpolationOfLSP(lsp ,.... );
[0233] { FHIIS24EM08-2024223273. DOCX 40
[0234] LSPk = LSPendPrev * (wk-1) + LSPend * wk;
[0235] / * conversion from LSP k to LPC* /
[0236] ConversionToLPC(LSPk, al,....) }
[0237] 1*2. second audio parameter unit (20), Calculate second set of LPC* / if ( LastFramewasLost) {
[0238] / * goes into block 612 and 614 * / void decPlcRecoveryinCelp_setSecondLPCSet_bfi() {
[0239] / * block 612 : computation of second set of LSP * /
[0240] AddMeanLSFToTransmittedLSF( TransmittedLSF, SecondSetMeanLsf, SecondSetEndLsf, m
[0241] / *convert to lsp* /
[0242] Convert( SecondSetMeanLsf, SecondSetMeanLsp, ... );
[0243] Convert(SecondSetLsf, SecondSetLsp, ... );
[0244] );
[0245] /
[0246] / * block 614 : computation of second set of LPC * /
[0247] InterpolationOfLSP(SecondSetMeanLsp, SecondSetEndLsp,a2,...);
[0248] / * 3. synthesing unit, e.g block 50 * /
[0249] / *synthesizing unit (50)* / if ( LastFramewasLost) {
[0250] / * selector 60 and and block 608,610,614,618 and 620* /
[0251] LpcSetDecision( al, a2, Synthesisl, Synthesis2 ); } else {
[0252] / * usual decoding if previous and actual frame are received * / GenerationofSynthesis(excitation, Synthesisl, Lsubframe, .... ); } FHIIS24EM08-2024223273. DOCX 41
[0253] LpcSetDecision(
[0254] { if ( UseSynthesisSet2 )
[0255] {
[0256] / * synthesis part of 620: Remaining subframes of the second set are synthesized * /
[0257] GenerationofSynthesis(excitation, Synthesis2, Lsubframe, .... ); } else
[0258] {
[0259] / * 618 and synthesis part of 620: Full snythesis of first set or remaining subframes of the first set are synthesized * /
[0260] GenerationofSynthesis(excitation, Synthesisl, Lsubframe, .... );
[0261] / * decision on specific subframe * / if ( SubframeNumber == 3)
[0262] {
[0263] / *selector 60, all other blocks * /
[0264] UseSynthesisSet2g = EnergyCalculationAndLpcSetDecision( ... )
[0265] }
[0266] }
[0267] 4. Decision of set / * e.g. selector 60 * /
[0268] EnergyCalculationAndLpcSetDecision ( )
[0269] {
[0270] / * 616: Computation of the second sythnesis* /
[0271] GenerationofSynthesis( Excitation ,&Synthesis2, L_subfr,
[0272] ... );
[0273] / * 610: Energy computation for the first synthesis* /
[0274] FirstSetSubframe Energy = EnergyCalculation(Synthesisl, ...)
[0275] / * 618: Energy computation for the second synthesis* /
[0276] SecondSetSubframe Energy = EnergyCalculation(Synthesis2, ...)
[0277] / * Energy Relations to PLC and between the subframes of the two sets* /
[0278] FirstSetEnergyRelationToPLC[ = FirstSetSubframe Energy
[0279] / energy PrevSubframe; FHIIS24EM08-2024223273. DOCX 42
[0280] SecondSetEnergyRelationToPLC = SecondSetSubframeEnergy / energy PrevSubframe;
[0281] FirstSetRelationToSecondSet = FirstSetSubframeEnergy / SecondSetSubframeEnergy;
[0282] / * energy calculation for damping: * /
[0283] Erell = Maximum( FirstSetEnergyRelationToPLC[0], FirstSetEnergyRelationToPLC[l] );
[0284] Erel2 = Maximum( SecondSetEnergyRelationToPLC[0], SecondSetEnergyRelationToPLC[l] );
[0285] / * block 620: Decision of LPC set and synthesis * / if ( ( first condition) && (second condition) && (optional third condition) == true)
[0286] {
[0287] AttenuationFactor = alphaDampingFunction (Erel2 ); return UseSecondSet;
[0288] } else
[0289] {
[0290] AttenuationFactor = alphaDampingFunction ( Erell ); return UseFirstSet;
[0291] }
[0292] / * 5. block 623, Energy Compensation* /
[0293] 5. Energy Compensation void attenuate Synthe sis (. . .
[0294] )
[0295] / * 706 energy computation* /
[0296] / * energy at the beginning of the frame * / energyFirstHalfFrame = EnergyCalculation(Synthesis, FirstHalfframe)
[0297] / * energy at the end of the frame energySecondHalfFrame = EnergyCalculation(Synthesis, SecondHalfframe)
[0298] / * 708 gain computation* /
[0299] / / gains gl =squareroot( energyPrevPLC / energyFirstHalfFrame ); FHIIS24EM08-2024223273. DOCX 43 g2 = squareroot( energyPrevPLC + alpha * ( energyFirstHalfFrame - energyPrevPLC ) / ener- gyFirstHalfFrame)) ;
[0300] / * block 702: compute zero input response and remove it from synthesis* /
[0301] ZeroInputResponseGeneration( ZeroInputResponse );
[0302] / * block 704: Remove ZIR from, signal * /
[0303] RemoveZeroInputResponse( Synthesis, ZeroInputResponse );
[0304] / * scaling of the synthesis and the memory * /
[0305] / * scaler 360 * /
[0306] ScalingofSynthesis(Synthesis, gl, g2)
[0307] / * re-add zero input response and update excitation* /
[0308] / * block 706: Remove ZIR from, signal * /
[0309] AddZeroInputResponse( ( Synthesis, ZeroInputResponse);
[0310] / *block 712: Calculate residual * /
[0311] CalculateTheResidualfromSynthesis( Synthesis, Exciation );
[0312] }
[0313] 6. Updates
[0314] / * block 718: update excitationPrev * /
[0315] UpdateExcitationforNextFrame(Excitation, ExcitationPrevious)
[0316] Aspects
[0317] According to an aspect, there is provided an audio decoder (e.g. 100) for synthesizing an audio signal (e.g. s) from a bitstream which represents the audio signal, the audio decoder (e.g. 100) including: a bitstream receiver (e.g. 5), to receive the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters (e.g. LSFq) for each frame of the audio signal, FHIIS24EM08-2024223273. DOCX 44 a first audio parameter decoding unit (e.g. 10), to decode, for a current properly received frame, a first set of decoded audio parameters (e.g. an) from at least the set of encoded audio parameters, a concealment unit (e.g. 40) to conceal at least one non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters; a second audio parameter decoding unit (e.g. 20), to decode, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters (e.g. a,?) from at least the set of encoded audio parameters, the second set of decoded audio parameters (e.g. a,?) being different from the first set of decoded audio parameters (e.g. aii); a synthesizing unit (e.g. 50), to output, or derive, an output or derived version (e.g. s) of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection (e.g. 60) is made between: a first version (e.g. si) of the synthesized audio signal (e.g. s), synthesized from at least the first set of decoded audio parameters; and a second version (e.g. s2) of the synthesized audio signal (e.g. s), synthesized from the second set of decoded audio parameters.
[0318] According to a further aspect, the first audio parameter decoding unit (e.g. 10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters from the at least one set of concealment parameters of the immediately preceding non-properly received frame and the set of encoded audio parameters of the current frame.
[0319] According to a further aspect, the first audio parameter decoding unit (e.g. 10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters through a first prediction (e.g. 604) of the first set of audio parameters obtained from the version of the set of concealment parameters of the immediately preceding non-properly received frame and a pre-defined value.
[0320] According to a further aspect, the second audio parameter decoding unit (e.g. 20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode (e.g. 612) the second set of decoded audio parameters from the encoded audio parameters of the current frame. FHIIS24EM08-2024223273. DOCX 45
[0321] According to a further aspect, the second audio parameter decoding unit (e.g. 20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the second set of decoded audio parameters from encoded audio parameters of the current frame but not from the set of concealment parameters of the immediately preceding non-properly received frame. according to a further aspect, the first audio parameter decoding unit is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters, from a first number of concealment parameters of the set of concealment parameters of the immediately preceding non-properly received frame and from the encoded audio parameters of the current frame, wherein the second audio parameter decoding unit is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non- properly received frame, to decode the second set of decoded audio parameters from the encoded audio parameters of the current frame and from a second number of concealment parameters of the set of concealment parameters of the immediately preceding non-properly received frame which is smaller than the first number of concealment parameters.
[0322] According to a further aspect, the audio synthesizer may be configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. si) and the second synthesis signal (e.g. s2) based on a comparison between at least energy- related measurements on the first version (e.g. si) of the synthesized audio signal with energy-related measurements on the second version (e.g. s2) of the synthesized audio signal, so as to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the second version (e.g. s2) of the synthesized audio signal in case at least one of the following condition or a combination of at least two of the following conditions is satisfied: the energy of the first version (e.g. si) of the synthesized audio signal is, at least in an initial subframe or sequence of initial subframes, larger than the energy of the second version (e.g. s2) of the synthesized audio signal by at least one threshold ratio (e.g. thl) greater than or equal to 1; the energy of the first version (e.g. si) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial subframes of the first version (e.g. si) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame; the energy of the second version (e.g. s2) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial FHIIS24EM08-2024223273. DOCX 46 subframes of the second version (e.g. s2) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame; and the energy of the first version (e.g. si) of the synthesized audio signal is larger by at least one threshold ratio, in at least one initial subframe or sequence of initial subframes of the first version (e.g. si) of the synthesized audio signal, than the energy of the second version (e.g. s2) of the synthesized audio signal in at least one initial subframe or sequence of initial subframes of the second audio signal (e.g. s2) corresponding to the at least one initial subframe or sequence of initial subframes of the first version (e.g. si), and, in case the condition is not satisfied, to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the first version (e.g. si) of the synthesized audio signal.
[0323] According to a further aspect, the audio synthesizer may be configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. si) and the second synthesis signal (e.g. s2) at least based on a comparison between at least energy-related measurements on the first version (e.g. si) of the synthesized audio signal with energy- related measurements on the second version (e.g. s2) of the synthesized audio signal, so as to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the second version (e.g. s2) of the synthesized audio signal in case of at least one of the following conditions, or a combination of at least one of the following conditions, is satisfied:
[0324] - the energy of at least one initial subframe, or of a sequence of initial subframes (e.g. Ei,i, Ei, 2), of the first version (e.g. si) of the synthesized audio signal is, subframe by subframe, larger than the energy of at least one initial subframe, or of a sequence of initial subframes (e.g. £2,1, £2,2), of the second version (e.g. s2) of the synthesized audio signal according to at least one predetermined ratio threshold (e.g. thi, th2) equal to or greater than 1;
[0325] - the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (e.g. Ei,i, Ei, 2), of the first version (e.g. si) of the synthesized audio signal is larger than the energy (e.g. EpLCsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (e.g. thpLc) equal to or greater than 1; and
[0326] - the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (e.g. E2,I, £2,2), of the second version (e.g. s2) of the synthesized audio signal is larger than the energy (e.g. Eppcsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (e.g. thpLc) equal to or greater than 1; and FHIIS24EM08-2024223273. DOCX 47 in case the at least one condition or combination of conditions is not satisfied, to output or derive, as the output or derived version (e.g. s) of the synthesized audio signal, the first version (e.g. si) of the synthesized audio signal.
[0327] According to a further aspect, the audio synthesizer may be configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. si) and the second synthesis signal (e.g. s2) based on a comparison between an envelope of the first synthesis signal and an envelope of the second synthesis signal, so as to select the second synthesis signal (e.g. s2) in case the envelope of the second synthesis signal (e.g. s2) is, at least in the initial subframe or in a sequence of initial subframes, more stable, by at least one predetermined extent, than the envelope of the first synthesis signal (e.g. si), and to select the first synthesis signal otherwise.
[0328] According to a further aspect, the audio decoder is configured to scale (e.g. 360), sample by sample, the output or derived version (e.g. s) of the synthesized audio signal in the properly received frame immediately following the at least one previously non-properly received frame by an energy compensating gain greater than 0, the energy compensating gain reducing, in at least one portion of the current frame, the energy gap between the concealed signal in a last portion of the concealed frame and the output or derived version of the synthesized audio signal in the at least one portion of the current frame.
[0329] According to a further aspect, the energy compensating gain evolves, monotonically or constantly, from a first value towards a second value, the first value being: comparatively high in case of a ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version (e.g. s) of the synthesized audio signal in an initial portion of the current frame, is comparatively high; and comparatively low in case of the ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version (e.g. s) of the synthesized audio signal in an initial portion of the current frame, is comparatively low, the second value (e.g. g2), conditioned by a conditioning term which is proportional to the difference between the energy of a last portion of the of the current frame and the energy of the last portion of the concealed frame, wherein the second value (e.g. g2) is: comparatively high in case of a ratio between the energy of the concealed signal in the last portion of the concealed frame, added with the conditioning term, and the energy of FHIIS24EM08-2024223273. DOCX 48 the output or derived version (e.g. s) of the synthesized audio signal in the last portion of the current frame, is comparatively high; and comparatively low in case of the ratio between the energy of the concealed signal in the last portion of the concealed frame, added by the conditioning term, and the energy of the output or derived version (e.g. s) of the synthesized audio signal in the last portion of the current frame, is comparatively low.
[0330] According to a further aspect, the current frame and the concealed frame are partitioned both according to a first partitioning which partitions the current frame and the concealed frame among a sequence of subframes in a number of subframes which is 3 or more than 3, and according to a second partitioning which partitions the current frame and the concealed frame among a sequence of portions in a number of portions which is 2 or more than 2, but the number of portions being less than the number of subframes, and each portion having larger time length than any subframe.
[0331] According to a further aspect, the conditioning term has a proportionality coefficient a is linearly dependent on a maximum value between a first ratio and a second ratio, where the first ratio is a ratio between the energy of the initial subframe of the output or derived version (e.g. s) of the synthesized audio signal and the energy of the last subframe of the concealed frame, and the second ratio is a ratio between the energy of the second subframe of the output or derived version (e.g. s) of the synthesized audio signal and the energy of the last subframe of the concealed frame.
[0332] According to a further aspect, the energy compensating gain is comparatively close to 1, in at least one portion of the current frame, in case the distance between the energy of the output or derived version (e.g. s) of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively low, and the energy compensating gain is comparatively distant from 1, in the at least one portion of the current frame, in case the distance between the energy of the output or derived version (e.g. s) of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively high.
[0333] According to a further aspect, the energy compensating gain is defined recursively by crossfading the energy compensating gain for an immediately preceding sample with the second gain value.
[0334] According to a further aspect, the synthesizing unit (e.g. 50) is configured to, initially, partially synthesize only an initial subframe, or a group of initial subframes, of the of the first synthesis signal (e.g. si) and, partially synthesize only an initial subframe, or a group of initial subframes, of FHIIS24EM08-2024223273. DOCX 49 the second synthesis signal (e.g. S2), so that the selection (e.g. 60) is based on energy-related measurements on the partially synthesized version of the first synthesis signal (e.g. si) and the partially synthesized version of the second synthesis signal (e.g. S2), so that, only after the selection (e.g. 60), the remaining subframe or subframe of the selected synthesis signal is or are synthesized, without synthesizing the remaining subframe or subframe of the non-selected synthesis signal.
[0335] According to a further aspect, the first audio parameter decoding unit (e.g. 10) is configured to, initially, decode, respectively, only a first subset of the first set of decoded audio parameters and only a second subset of the second decoded audio parameters, the first subset and second subset corresponding to the initial subframe, or the group of initial subframes, so that the selection (e.g. 60) is based on energy-related measurements of the partially synthesized version of the first synthesis signal (e.g. si) obtained from the first subset and the partially synthesized version of the second synthesis signal (e.g. S2) obtained from the second subset, so that, only after the selection (e.g. 60), the remaining audio parameters of the set of audio parameter associated with the selected synthesis signal are decoded, and the remaining audio parameters of the set of audio parameter associated with the non-selected synthesis signal are not decoded.
[0336] According to a further aspect, further comprising a neural network processor using at least one learnable layer, to be inputted with the concealment parameters as well as the decoded parameters and / or the first and second synthesis signals or the output or derived version of the synthesis signal, so as to process the output or derived version of the synthesis signal through the at least one learnable layer.
[0337] According to a further aspect, the at least one learnable layers is a generative adversarial network, GAN, learnable layer.
[0338] According to a further aspect, the encoded audio parameter include, or provide information on, linear spectral frequencies.
[0339] According to an aspect, there is provided a method for synthesizing an audio signal (e.g. s) from a bitstream which represents the audio signal, the method including: receiving the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters (e.g. LSFq) for each frame of the audio signal, decoding, for a current properly received frame, a first set of decoded audio parameters (e.g. aii, LSPend) from at least the set of encoded audio parameters, concealing at least one previously non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters, and to synthesize a concealed frame from the at least one set of concealment audio parameters; FHIIS24EM08-2024223273. DOCX 50 decoding, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters from at least the set of encoded audio parameters, the second set of audio parameters being different from the first set of audio parameters, outputting or deriving an output or derived version (e.g. s) of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection is made between: a first version (e.g. si) of the synthesized audio signal (e.g. s), synthesized from at least the set of encoded audio parameters a second version (e.g. s2) of the synthesized audio signal (e.g. s), synthesized from the second set of decoded audio parameters.
[0340] According to an aspect, there is provided a non-transitory storage unit storing instructions which, when executed by a processor, cause the processor to perform the method above (or any of the methods above and below).
[0341] Further examples and variants
[0342] Depending on certain implementation requirements, examples may be implemented in hardware. The implementation may be performed using a digital storage medium, for example a floppy disk, a Digital Versatile Disc (DVD), a Blu-Ray Disc, a Compact Disc (CD), a Read-only Memory (ROM), a Programmable Read-only Memory (PROM), an Erasable and Programmable Read-only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM) or a flash memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
[0343] Generally, examples may be implemented as a computer program product with program instructions, the program instructions being operative for performing one of the methods when the computer program product runs on a computer. The program instructions may for example be stored on a machine readable medium.
[0344] Other examples comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier. In other words, an example of method is, therefore, a computer program having program instructions for performing one of the methods described herein, when the computer program runs on a computer.
[0345] A further example of the methods is, therefore, a data carrier medium (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for FHIIS24EM08-2024223273. DOCX 51 performing one of the methods described herein. The data carrier medium, the digital storage medium or the recorded medium are tangible and / or non-transitionary, rather than signals which are intangible and transitory.
[0346] A further example comprises a processing unit, for example a computer, or a programmable logic device performing one of the methods described herein.
[0347] A further example comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0348] A further example comprises an apparatus or a system transferring (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0349] In some examples, a programmable logic device (for example, a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some examples, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods may be performed by any appropriate hardware apparatus.
[0350] While this invention has been described in terms of several examples, there are alterations, permutations, and equivalents which fall within the scope of this invention. It should also be noted that there are many alternative ways of implementing the methods and compositions of the present invention. It is therefore intended that the following appended claims be interpreted as including all such alterations, permutations and equivalents as fall within the true spirit and scope of the present invention.
[0351] References
[0352] [1]https: / / www.etsi.org / dliver / etsi_ts / 126400_126499 / 126445 / 12.06.00_60 / ts_126445vl20600p.pdf
[0353] [2] internet Low Bitrate Codec, WEB RTC, https: / / webrtc.github.io / webrtc-org / license / ilbc-free- ware / ilbc-extra-documentation /
[0354] [3] Method and Decive for efficient frame erasure concealment in linear predictive based speech codecs, Voice Age Company, https: / / www.voiceageevs.com / documents / pa- tents / usa / VAEVS%201000%20-%20US%207%20693%20710%20B2.PDF
[0355] [4] Method and Device for efficient frame erasure concealment in speech codecs https: / / www.voiceageevs.com / documents / patents / usa / VAEVS%202200%20- %20US%208%20255%20207%20B2.PDF FHIIS24EM08-2024223273. DOCX 52
[0356] [5] Analysis by Adversarial Synthesis - A Novel Approach for Speech
[0357] Vocodinghttp: / / arxiv.org / pdf / 1907.00772
[0358] [6] http: / / www.jcomputers.us / vol2 / jcp0207-09.pdf
[0359] FH240802PEP -2025232411. DOCX 53
[0360] According to a 1staspect there is provided an audio decoder (also called audio synthesizer) (e.g. 100) for synthesizing an audio signal from a bitstream which represents the audio signal, the audio decoder (e.g. 100) including: a bitstream receiver (e.g. 5), to receive the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters (e.g. LSFq) for each frame of the audio signal, a first audio parameter decoding unit (e.g. 10), to decode, for a current properly received frame, a first set of decoded audio parameters (e.g. an) from at least the set of encoded audio parameters, a concealment unit (e.g. 40) to conceal at least one non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters; a second audio parameter decoding unit (e.g. 20), to decode, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters (e.g. a,?) from at least the set of encoded audio parameters, the second set of decoded audio parameters (e.g. a,?) being different from the first set of decoded audio parameters (e.g. aii); a synthesizing unit (e.g. 50), to output a synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection (e.g. 60) is made between: a first version (e.g. si) of the synthesized audio signal, synthesized from at least the first set of decoded audio parameters; and a second version (e.g. s2) of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
[0361] According to a 2ndaspect when relating back to the 1staspect, the first audio parameter decoding unit (e.g. 10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters from the at least one set of concealment parameters of the immediately preceding non-properly received frame and the set of encoded audio parameters of the current frame. FH240802PEP -2025232411. DOCX 54
[0362] According to a 3rdaspect when relating back to the 2ndaspect, the first audio parameter decoding unit (e.g. 10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters through a first prediction (e.g. 604) of the first set of audio parameters obtained from the version of the set of concealment parameters of the immediately preceding non- properly received frame and a pre-defined value.
[0363] According to a 4thaspect when relating back to any one of the 1stto 3rdaspects, the second audio parameter decoding unit (e.g. 20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode (e.g. 612) the second set of decoded audio parameters from the encoded audio parameters of the current frame.
[0364] According to a 5thaspect when relating back to any one of the 1stto 4thaspects, the second audio parameter decoding unit (e.g. 20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the second set of decoded audio parameters from encoded audio parameters of the current frame but not from the set of concealment parameters of the immediately preceding non-properly received frame.
[0365] According to a 6thaspect when relating back to any one of the 1stto 5thaspects, the audio decoder (audio synthesizer) is configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. si) and the second synthesis signal (e.g. s2) based on a comparison between at least energy-related measurements on the first version (e.g. si) of the synthesized audio signal with energy-related measurements on the second version (e.g. s2) of the synthesized audio signal, so as to output as the synthesized audio signal, the second version (e.g. s2) of the synthesized audio signal in case at least one of the following condition or a combination of at least two of the following conditions is satisfied: the energy of the first version (e.g. si) of the synthesized audio signal is, at least in an initial subframe or sequence of initial subframes, larger than the energy of the second version (e.g. s2) of the synthesized audio signal by at least one threshold ratio (e.g. thl) greater than or equal to 1 ; FH240802PEP -2025232411. DOCX 55 the energy of the first version (e.g. si) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial subframes of the first version (e.g. si) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame; the energy of the second version (e.g. s2) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial subframes of the second version (e.g. s2) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame; and the energy of the first version (e.g. si) of the synthesized audio signal is larger by at least one threshold ratio, in at least one initial subframe or sequence of initial subframes of the first version (e.g. si) of the synthesized audio signal, than the energy of the second version (e.g. s2) of the synthesized audio signal in at least one initial subframe or sequence of initial subframes of the second audio signal (e.g. s2) corresponding to the at least one initial subframe or sequence of initial subframes of the first version (e.g. si), and, in case the condition is not satisfied, to output as the synthesized audio signal, the first version (e.g. si) of the synthesized audio signal.
[0366] According to a 7thaspect when relating back to any one of the 1stto 6thaspects, the audio decoder (audio synthesizer) is configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. si) and the second synthesis signal (e.g. s2) at least based on a comparison between at least energy-related measurements on the first version (e.g. si) of the synthesized audio signal with energy-related measurements on the second version (e.g. s2) of the synthesized audio signal, so as to output, as the synthesized audio signal, the second version (e.g. s2) of the synthesized audio signal in case of at least one of the following conditions, or a combination of at least one of the following conditions, is satisfied:
[0367] - the energy of at least one initial subframe, or of a sequence of initial subframes (e.g. Ei,i, Ei, 2), of the first version (e.g. si) of the synthesized audio signal is, subframe by subframe, larger than the energy of at least one initial subframe, or of a sequence of initial subframes (e.g. £2,1, £2,2), of the second version (e.g. s2) of the synthesized audio signal according to at least one predetermined ratio threshold (e.g. thi, th2) equal to or greater than 1; FH240802PEP -2025232411. DOCX 56
[0368] - the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (e.g. Ei,i, Ei, 2), of the first version (e.g. si) of the synthesized audio signal is larger than the energy (e.g. EpLCsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (e.g. thpLc) equal to or greater than 1; and
[0369] - the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (e.g. E2,I, £2,2), of the second version (e.g. s2) of the synthesized audio signal is larger than the energy (e.g. Eppcsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (e.g. thpLc) equal to or greater than 1; and in case the at least one condition or combination of conditions is not satisfied, to output or derive, as the synthesized audio signal, the first version (e.g. si) of the synthesized audio signal.
[0370] According to a 8thaspect when relating back to any one of the 1stto 7thaspects, the audio decoder (audio synthesizer) is configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (e.g. 620) between the first synthesis signal (e.g. si) and the second synthesis signal (e.g. s2) based on a comparison between an envelope of the first synthesis signal and an envelope of the second synthesis signal, so as to select the second synthesis signal (e.g. s2) in case the envelope of the second synthesis signal (e.g. s2) is, at least in the initial subframe or in a sequence of initial subframes, more stable, by at least one predetermined extent, than the envelope of the first synthesis signal (e.g. si), and to select the first synthesis signal otherwise.
[0371] According to a 9thaspect when relating back to any one of the 1stto 8thaspects, the audio decoder (audio synthesizer) is configured to scale (e.g. 360), sample by sample, the synthesized audio signal in the properly received frame immediately following the at least one previously non-properly received frame by an energy compensating gain greater than 0, the energy compensating gain reducing, in at least one portion of the current frame, the energy gap between the concealed signal in a last portion of the concealed frame and the synthesized audio signal in the at least one portion of the current frame.
[0372] According to a 10thaspect when relating back to the 9thaspect, the energy compensating gain evolves, monotonically or constantly, from a first value towards a second value, the first value being: FH240802PEP -2025232411. DOCX 57 comparatively high in case of a ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the synthesized audio signal in an initial portion of the current frame, is comparatively high; and comparatively low in case of the ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the synthesized audio signal in an initial portion of the current frame, is comparatively low, the second value, conditioned by a conditioning term which is proportional to the difference between the energy of a last portion of the of the current frame and the energy of the last portion of the concealed frame, wherein the second value is: comparatively high in case of a ratio between the energy of the concealed signal in the last portion of the concealed frame, added with the conditioning term, and the energy of the synthesized audio signal in the last portion of the current frame, is comparatively high; and comparatively low in case of the ratio between the energy of the concealed signal in the last portion of the concealed frame, added by the conditioning term, and the energy of the synthesized audio signal in the last portion of the current frame, is comparatively low.
[0373] According to a 11thaspect when relating back to the 10thaspect, the current frame and the concealed frame are partitioned both according to a first partitioning which partitions the current frame and the concealed frame among a sequence of subframes in a number of subframes which is 3 or more than 3, and according to a second partitioning which partitions the current frame and the concealed frame among a sequence of portions in a number of portions which is 2 or more than 2, but the number of portions being less than the number of subframes, and each portion having larger time length than any subframe.
[0374] According to a 12thaspect when relating back to any one of the 9thto 11thaspects, the conditioning term has a proportionality coefficient a is linearly dependent on a maximum value between a first ratio and a second ratio, where the first ratio is a ratio between the energy of the initial subframe of the synthesized audio signal and the energy of the last subframe of the concealed frame, and the second ratio is a ratio between the energy of the second subframe of the synthesized audio signal and the energy of the last subframe of the concealed frame.
[0375] According to a 13thaspect when relating back to any one of the 9thto 12thaspects, the energy compensating gain is comparatively close to 1, in at least one portion of the current frame, in case the FH240802PEP -2025232411. DOCX 58 distance between the energy of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively low, and the energy compensating gain is comparatively distant from 1, in the at least one portion of the current frame, in case the distance between the energy of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively high.
[0376] According to a 14thaspect when relating back to any one of the 9thto 13thaspects, the energy compensating gain is defined recursively by cross-fading the energy compensating gain for an immediately preceding sample with the second gain value.
[0377] According to a 15thaspect when relating back to any one of the 1stto 14thaspects, the synthesizing unit (e.g. 50) is configured to, initially, partially synthesize only an initial subframe, or a group of initial subframes, of the of the first synthesis signal (e.g. si) and, partially synthesize only an initial subframe, or a group of initial subframes, of the second synthesis signal (e.g. S2), so that the selection (e.g. 60) is based on energy-related measurements on the partially synthesized version of the first synthesis signal (e.g. si) and the partially synthesized version of the second synthesis signal (e.g. S2), so that, only after the selection (e.g. 60), the remaining subframe or subframe of the selected synthesis signal is or are synthesized, without synthesizing the remaining subframe or subframe of the non-selected synthesis signal.
[0378] According to a 16thaspect when relating back to the 14thor 15thaspects, the first audio parameter decoding unit (e.g. 10) is configured to, initially, decode, respectively, only a first subset of the first set of decoded audio parameters and only a second subset of the second decoded audio parameters, the first subset and second subset corresponding to the initial subframe, or the group of initial subframes, so that the selection (e.g. 60) is based on energy-related measurements of the partially synthesized version of the first synthesis signal (e.g. si) obtained from the first subset and the partially synthesized version of the second synthesis signal (e.g. S2) obtained from the second subset, so that, only after the selection (e.g. 60), the remaining audio parameters of the set of audio parameter associated with the selected synthesis signal are decoded, and the remaining audio parameters of the set of audio parameter associated with the non-selected synthesis signal are not decoded. FH240802PEP -2025232411. DOCX 59
[0379] According to a 17thaspect when relating back to any one of the 1stto 16thaspects, the audio decoder (audio synthesizer) further comprises a neural network processor using at least one learnable layer, to be inputted with the concealment parameters as well as the decoded parameters and / or the first and second synthesis signals or the synthesized audio signal, so as to process the synthesized audio signal through the at least one learnable layer.
[0380] According to an 18thaspect when relating back to the 17thaspect, the at least one learnable layers is a generative adversarial network, GAN, learnable layer.
[0381] According to a 19thaspect when relating back to any one of the 1stto 18thaspects, the encoded audio parameter include, or provide information on, linear spectral frequencies.
[0382] A 20thaspect relates to a method for synthesizing an audio signal from a bitstream which represents the audio signal, the method including: receiving the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters for each frame of the audio signal, decoding, for a current properly received frame, a first set of decoded audio parameters (e.g. aii, LSPend) from at least the set of encoded audio parameters, concealing at least one previously non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters, and to synthesize a concealed frame from the at least one set of concealment audio parameters; decoding, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters from at least the set of encoded audio parameters, the second set of audio parameters being different from the first set of audio parameters, outputting a synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection is made between: a first version (e.g. si) of the synthesized audio signal, synthesized from at least the set of encoded audio parameters a second version (e.g. s2) of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
Claims
FH240802PEP -2025232411. DOCX 60Claims1. Audio decoder (100) for synthesizing an audio signal from a bitstream which represents the audio signal, the audio decoder (100) including: a bitstream receiver (5), to receive the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters ( LSFq) for each frame of the audio signal, a first audio parameter decoding unit (10), to decode, for a current properly received frame, a first set of decoded audio parameters (an) from at least the set of encoded audio parameters, a concealment unit (40) to conceal at least one non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters; a second audio parameter decoding unit (20), to decode, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters (a,?) from at least the set of encoded audio parameters, the second set of decoded audio parameters (a,?) being different from the first set of decoded audio parameters (ail); a synthesizing unit (50), to output, or derive, an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection (60) is made between: a first version (si) of the synthesized audio signal, synthesized from at least the first set of decoded audio parameters; and a second version (s2) of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
2. The audio decoder of claim 1, wherein the first audio parameter decoding unit (10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters from the at least one set of concealment parameters of the immediately preceding non-properly received frame and the set of encoded audio parameters of the current frame.
3. The audio decoder of claim 2, wherein the first audio parameter decoding unit (10) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the first set of decoded audio parameters through a first prediction (604) of the first set of audio parameters obtained from the version ofFH240802PEP -2025232411. DOCX 61 the set of concealment parameters of the immediately preceding non-properly received frame and a pre-defined value.
4. The audio decoder of any of the preceding claims, wherein the second audio parameter decoding unit (20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode (612) the second set of decoded audio parameters from the encoded audio parameters of the current frame.
5. The audio decoder of any of the preceding claims, wherein the second audio parameter decoding unit (20) is configured, in the case the current frame is the properly received frame which immediately follows the at least one previously non-properly received frame, to decode the second set of decoded audio parameters from encoded audio parameters of the current frame but not from the set of concealment parameters of the immediately preceding non-properly received frame.
6. The audio decoder of any of the preceding claims, configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (620) between the first synthesis signal (si) and the second synthesis signal (s2) based on a comparison between at least energy-related measurements on the first version (si) of the synthesized audio signal with energy-related measurements on the second version (s2) of the synthesized audio signal, so as to output or derive, as the output or derived version of the synthesized audio signal, the second version (s2) of the synthesized audio signal in case at least one of the following condition or a combination of at least two of the following conditions is satisfied: the energy of the first version (si) of the synthesized audio signal is, at least in an initial subframe or sequence of initial subframes, larger than the energy of the second version (s2) of the synthesized audio signal by at least one threshold ratio (thl) greater than or equal to 1 ; the energy of the first version (si) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial subframes of the first version (si) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame;FH240802PEP -2025232411. DOCX 62 the energy of the second version (s2) of the synthesized audio signal is larger, by at least one threshold ratio equal to or greater than 1, in at least one initial subframe or sequence of initial subframes of the second version (s2) of the synthesized audio signal, than the energy of the concealed signal in at least a final subframe or sequence of final subframes of the concealed frame; and the energy of the first version (si) of the synthesized audio signal is larger by at least one threshold ratio, in at least one initial subframe or sequence of initial subframes of the first version (si) of the synthesized audio signal, than the energy of the second version (s2) of the synthesized audio signal in at least one initial subframe or sequence of initial subframes of the second audio signal (s2) corresponding to the at least one initial subframe or sequence of initial subframes of the first version (si), and, in case the condition is not satisfied, to output or derive, as the output or derived version of the synthesized audio signal, the first version (si) of the synthesized audio signal.
7. The audio decoder of any of the preceding claims, configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (620) between the first synthesis signal (si) and the second synthesis signal (s2) at least based on a comparison between at least energy-related measurements on the first version (si) of the synthesized audio signal with energy-related measurements on the second version (s2) of the synthesized audio signal, so as to output or derive, as the output or derived version of the synthesized audio signal, the second version (s2) of the synthesized audio signal in case of at least one of the following conditions, or a combination of at least one of the following conditions, is satisfied:- the energy of at least one initial subframe, or of a sequence of initial subframes (Ei,i, EI,2), of the first version (si) of the synthesized audio signal is, subframe by subframe, larger than the energy of at least one initial subframe, or of a sequence of initial subframes (E2,I, £2,2), of the second version (s2) of the synthesized audio signal according to at least one predetermined ratio threshold (thi, th2) equal to or greater than 1;- the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (Ei,i, Ei, 2), of the first version (si) of the synthesized audio signal is larger than the energy (EpLCsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (thpLc) equal to or greater than 1; andFH240802PEP -2025232411. DOCX 63- the energy of at least one initial subframe, or of each subframe of a sequence of initial subframes (E2,I, £2,2), of the second version (s2) of the synthesized audio signal is larger than the energy (EpLCsub) of a final subframe of the audio signal in the concealed frame according to at least one predetermined ratio threshold (thpLc) equal to or greater than 1; and in case the at least one condition or combination of conditions is not satisfied, to output or derive, as the output or derived version of the synthesized audio signal, the first version (si) of the synthesized audio signal.
8. The audio decoder of any of the preceding claims, configured, in the case the current frame is the properly decoded frame which immediately follows the at least one previously non-properly received frame, to perform the selection (620) between the first synthesis signal (si) and the second synthesis signal (s2) based on a comparison between an envelope of the first synthesis signal and an envelope of the second synthesis signal, so as to select the second synthesis signal (s2) in case the envelope of the second synthesis signal (s2) is, at least in the initial subframe or in a sequence of initial subframes, more stable, by at least one predetermined extent, than the envelope of the first synthesis signal (si), and to select the first synthesis signal otherwise.
9. The audio decoder of any of the preceding claims, wherein the audio decoder is configured to scale (360), sample by sample, the output or derived version of the synthesized audio signal in the properly received frame immediately following the at least one previously non-properly received frame by an energy compensating gain greater than 0, the energy compensating gain reducing, in at least one portion of the current frame, the energy gap between the concealed signal in a last portion of the concealed frame and the output or derived version of the synthesized audio signal in the at least one portion of the current frame.
10. The audio decoder of claim 9, wherein the energy compensating gain evolves, monotonically or constantly, from a first value towards a second value, the first value being: comparatively high in case of a ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version of the synthesized audio signal in an initial portion of the current frame, is comparatively high; and comparatively low in case of the ratio, between the energy of the concealed signal in a last portion of the concealed frame and the energy of the output or derived version of the synthesized audio signal in an initial portion of the current frame, is comparatively low,FH240802PEP -2025232411. DOCX 64 the second value, conditioned by a conditioning term which is proportional to the difference between the energy of a last portion of the of the current frame and the energy of the last portion of the concealed frame, wherein the second value is: comparatively high in case of a ratio between the energy of the concealed signal in the last portion of the concealed frame, added with the conditioning term, and the energy of the output or derived version of the synthesized audio signal in the last portion of the current frame, is comparatively high; and comparatively low in case of the ratio between the energy of the concealed signal in the last portion of the concealed frame, added by the conditioning term, and the energy of the output or derived version of the synthesized audio signal in the last portion of the current frame, is comparatively low.
11. The audio decoder of claim 10, wherein the current frame and the concealed frame are partitioned both according to a first partitioning which partitions the current frame and the concealed frame among a sequence of subframes in a number of subframes which is 3 or more than 3, and according to a second partitioning which partitions the current frame and the concealed frame among a sequence of portions in a number of portions which is 2 or more than 2, but the number of portions being less than the number of subframes, and each portion having larger time length than any subframe.
12. The audio decoder of any of claims 9-11, wherein the conditioning term has a proportionality coefficient a is linearly dependent on a maximum value between a first ratio and a second ratio, where the first ratio is a ratio between the energy of the initial subframe of the output or derived version of the synthesized audio signal and the energy of the last subframe of the concealed frame, and the second ratio is a ratio between the energy of the second subframe of the output or derived version of the synthesized audio signal and the energy of the last subframe of the concealed frame.
13. The audio decoder of any of claims 9-12, wherein the energy compensating gain is comparatively close to 1, in at least one portion of the current frame, in case the distance between the energy of the output or derived version of the synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively low, and the energy compensating gain is comparatively distant from 1, in the at least one portion of the current frame, in case the distance between the energy of the output or derived version of theFH240802PEP -2025232411. DOCX 65 synthesized audio signal in the at least one portion of the current frame and the energy of concealed signal in the end portion of the concealed frame is comparatively high.
14. The audio decoder of any of claims 9-13, wherein the energy compensating gain is defined recursively by cross-fading the energy compensating gain for an immediately preceding sample with the second gain value.
15. The audio decoder of any of the preceding claims, wherein the synthesizing unit (50) is configured to, initially, partially synthesize only an initial subframe, or a group of initial subframes, of the of the first synthesis signal (si) and, partially synthesize only an initial subframe, or a group of initial subframes, of the second synthesis signal (S2), so that the selection (60) is based on energy- related measurements on the partially synthesized version of the first synthesis signal (si) and the partially synthesized version of the second synthesis signal (S2), so that, only after the selection (60), the remaining subframe or subframe of the selected synthesis signal is or are synthesized, without synthesizing the remaining subframe or subframe of the non-selected synthesis signal.
16. The audio decoder of claim 14 or 15, wherein the first audio parameter decoding unit (10) is configured to, initially, decode, respectively, only a first subset of the first set of decoded audio parameters and only a second subset of the second decoded audio parameters, the first subset and second subset corresponding to the initial subframe, or the group of initial subframes, so that the selection (60) is based on energy-related measurements of the partially synthesized version of the first synthesis signal (si) obtained from the first subset and the partially synthesized version of the second synthesis signal (S2) obtained from the second subset, so that, only after the selection (60), the remaining audio parameters of the set of audio parameter associated with the selected synthesis signal are decoded, and the remaining audio parameters of the set of audio parameter associated with the non-selected synthesis signal are not decoded.
17. The audio decoder of any of the preceding claims, further comprising a neural network processor using at least one learnable layer, to be inputted with the concealment parameters as well as the decoded parameters and / or the first and second synthesis signals or the output or derived version of the synthesis signal, so as to process the output or derived version of the synthesis signal through the at least one learnable layer.FH240802PEP -2025232411. DOCX 6618. The audio decoder of claim 17, wherein the at least one learnable layers is a generative adversarial network, GAN, learnable layer.
19. The audio decoder of any of the preceding claims, wherein the encoded audio parameter include, or provide information on, linear spectral frequencies.
20. A method for synthesizing an audio signal from a bitstream which represents the audio signal, the method including: receiving the bitstream representative of the audio signal, the bitstream having, encoded therein, a set of encoded audio parameters ( LSFq) for each frame of the audio signal, decoding, for a current properly received frame, a first set of decoded audio parameters (an, LSPend) from at least the set of encoded audio parameters, concealing at least one previously non-properly received frame based on at least one previously properly received frame, so as to generate at least one set of concealment audio parameters, and to synthesize a concealed frame from the at least one set of concealment audio parameters; decoding, for a current properly received frame which immediately follows the at least one previously non-properly received frame, a second set of decoded audio parameters from at least the set of encoded audio parameters, the second set of audio parameters being different from the first set of audio parameters, outputting or deriving an output or derived version of the synthesized audio signal in such way that, if the current properly received frame immediately follows the at least one previously non-properly received frame, a selection is made between: a first version (si) of the synthesized audio signal, synthesized from at least the set of encoded audio parameters a second version (s2) of the synthesized audio signal, synthesized from the second set of decoded audio parameters.
Citation Information
Patent Citations
Systems and methods for mitigating potential frame instability
WO2014130087A1