Method for performing packet loss concealment in complex filter bank domain

By adaptively selecting the time-frequency block value of the replacement frame through a sinusoidal spread or linear prediction process in the complex orthogonal mirror filter domain, the problems of low quality and high computational complexity in audio signal transmission in the prior art are solved, and low latency and high quality audio recovery are achieved.

CN121444166APending Publication Date: 2026-01-30DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480031776.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-05
Filing Date
2024-03-22
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing packet loss concealment techniques suffer from lower-than-expected quality in audio signal transmission, often introducing perceptible audio artifacts, and are computationally complex, making it difficult to achieve low latency and high-quality concealment on power-constrained terminal devices.

Method used

By adaptively selecting either sinusoidal expansion or linear prediction processes in the complex orthogonal mirror filter domain to generate time-frequency block values ​​for replacement frames, and choosing an appropriate process based on pitch to generate more realistic and reliable replacement frames, computational complexity is reduced.

Benefits of technology

It improves the quality of packet loss concealment, reduces computational complexity, is suitable for power-constrained terminal devices, and achieves low-latency and high-quality audio recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121444166A_ABST
    Figure CN121444166A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and system for hiding lost or corrupted frames of audio content in a sequence of audio frames. The method includes obtaining at least one first frame associated with a sequence of frames and identifying whether a second frame after the at least one first frame is valid or invalid. When the second frame is identified as invalid, the method includes generating a set of time-frequency (TF) block values for replacing a replacement frame of the second frame by applying a sinusoidal extension process to generate a set of TF block values for the replacement frame, or applying a linear prediction process to generate a set of TF block values for the replacement frame, based on a tone degree of at least one previous frame.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to a system, device and method for processing audio. BACKGROUND

[0002] The methods described in this section are not prior art to the claims in this application and are not admitted to be prior art by inclusion in this section.

[0003] There are many types of codecs for encoding audio content into a format suitable for inclusion in a bitstream. For example, there are several codecs that operate in the frequency domain, whereby a segment of a time-domain audio signal is transformed into vector-valued samples, each sample being represented in the encoded bitstream by a set of time-frequency, TF, tiles in conjunction with some side information. Each TF tile is associated with a particular frequency band, and at the decoding end, the samples carrying the TF tiles are reconstructed by decoding the bitstream, after which the time-domain samples are regenerated by applying an inverse transform to the TF tiles.

[0004] There are a variety of time-frequency transforms (and corresponding inverse transforms). Examples include the discrete Fourier transform (DFT), the discrete cosine transform, DCT, the modified discrete cosine transform (MDCT), or subband techniques that represent the original time-domain signal through a sample representation in a subband domain. An example of a subband domain transform is the use of a quadrature mirror filter (QMF) bank, which has the property that each bandpass (or subband) signal can be critically sampled so as to still allow for an inverse operation to perfectly reconstruct the original time signal.

[0005] One class of QMF-based transforms operates in the complex-valued domain, in which case the time-domain signal is said to be represented in the CQMF domain after being transformed to the complex-valued domain. In some cases, the CQMF bank is designed to introduce minimal delay in the forward direction (CQMF analysis) as well as the backward direction (CQMF synthesis), which makes the CQMF bank particularly useful in conversational speech applications and applications that benefit from low latency, such as telephony or teleconference applications.

[0006] A particular type of CQMF bank is used in the codec being standardized for the 3GPP immersive voice and audio services (IVAS) codec. This CQMF bank is referred to as the complex low-delay filter bank (CLDFB), and according to the CLDFB design, when the CLDFB analysis is applied to a 48 kHz-sampled audio signal, one TF tile (represented with one TF tile value) represents a frequency band of 400 Hz bandwidth, and the CLDFB samples cover a time slot of 1.25 ms.

[0007] Services like 3GPP IVAS are typically deployed over error-prone radio channels between a transmitting device and a receiving unit. In practice, TF blocks of multiple CQMF samples are referred to as frames, which are encoded into data blocks, where each frame contains CQMF samples representing audio content for a predetermined duration (e.g., 20 ms). The encoded frames are transmitted over the radio channel in the form of packet batches containing one or more encoded frames. However, there is a risk that one or more frames are corrupted or not delivered at all to the receiving unit, resulting in so-called ‘packet loss’.

[0008] Some error detection mechanisms like cyclic redundancy check CRC can be applied in the receiving unit to detect whether the packets (carrying one or more frames) have been received without errors or whether the packets are corrupted. In some cases, the packets can also arrive too late to be used by the receiving unit. In any case, only the audio frames contained in the packets that are received in time and without errors can be used for audio decoding and extraction of the time audio signal.

[0009] For frames that are not available for decoding, the receiving unit can use techniques to generate substitute frames. One technique for forming substitute frames is to use zero frames (i.e., frames where each spectral coefficient is equal to zero) as substitute signals. Another technique is to repeat the latest frame that was received without errors.

[0010] Techniques for generating substitute frames are referred to as frame loss concealment or packet loss concealment (PLC). The overall goal of these techniques is that frame loss should become as inaudibly perceptible as possible, thereby minimizing the impact of frame loss on the audio quality at the receiver. In case of relying on waveform coding, packet loss concealment can be based on frame repetition techniques, and in case of relying on parametric coding, the encoding parameters of the latest good frame are repeated.

[0011] In addition, various PLC techniques are known from, for example, the 3GPP codec for enhanced voice services (EVS). This codec contains multiple PLC techniques, which can be selected depending on whether the signal is encoded using a speech coding technique (ACELP) or using an audio coding technique (transform coding).

[0012] With these and other considerations in mind, the disclosure presented herein was made. SUMMARY

[0013] Techniques for processing audio signals are described. Various embodiments described herein provide devices, systems, and methods for mitigating the impact of lost frames that can occur in the transmission of audio signals. Various examples are described that improve concealment of lost frames by processing the audio in the complex quadrature mirror filter (CQMF) domain.

[0014] Disadvantages of existing concealment techniques (e.g., PLCs) are that they provide a lower quality than desired, often introducing perceptible audio artifacts. In addition, many techniques are computationally complex (e.g., in terms of computation time, memory requirements, and number of computation operations), requiring powerful processing hardware and / or increasing the latency of the audio encoding-decoding chain. At the same time, audio decoding is often performed on terminal devices with very limited power, such as earbuds or AR glasses, which means that many concealment techniques are not directly suitable for achieving low latency and high-quality concealment in some applications. To this end, the present disclosure proposes an improved concealment technique that can recreate high-quality replacement frames through reduced-complexity operations.

[0015] According to a first aspect of the present invention, there is provided a method for concealing lost or corrupted frames having audio content in a sequence of frames. The method comprises obtaining at least one first frame associated with the sequence of frames, the at least one first frame being a frame of the sequence of frames or a replacement frame generated; and identifying whether a second frame following the at least one first frame is valid or invalid. When the second frame is identified as invalid, the method comprises generating a set of time-frequency (TF) tile values for a replacement frame to replace the second frame by: analyzing the at least one first frame to determine a degree of tonality and comparing the degree of tonality to a threshold to determine whether the degree of tonality exceeds the threshold. Where the degree of tonality exceeds the threshold, the method comprises applying a sinusoidal spreading process to generate the set of TF tile values for the replacement frame, and where the degree of tonality does not exceed the threshold, the method comprises applying a linear prediction process to generate the set of TF tile values for the replacement frame. The method further comprises outputting the replacement frame to replace the second frame.

[0016] Frames that have been completely successfully received at the time of decoding and are thus determined to be uncorrupted are identified as valid frames. Frames that are determined to be corrupted or not received or otherwise unavailable (e.g., lost) at the time of decoding are identified as invalid frames. For at least the invalid frames, measures are taken to generate a replacement frame in order to conceal the invalid frame(s).

[0017] To generate the set of TF tile values for the replacement frame, the previous first frame can be a valid frame, or a replacement frame of a previous frame that was previously determined to be invalid.

[0018] By the method of the first aspect, the TF tile values of the replacement frame are adaptively generated by adaptively selecting one of the two processes for generating the TF tile values of the replacement frame. Which of the two processes, sinusoidal expansion and linear prediction, to use is based on the degree of tonality of the previous (first) frame. Since sinusoidal expansion is expected to generate more accurate TF tile values when the degree of tonality is high, sinusoidal expansion is used when the degree of tonality exceeds a threshold, and since linear prediction is expected to generate more accurate TF tile values when the degree of tonality is low, linear prediction is used when the degree of tonality is below the threshold. In this way, TF tile values can be generated that make the replacement frame (used as a continuation of the previous first frame) more realistic and believable, which makes the quality of the PLC technique higher.

[0019] There are various methods for quantifying the quality of a PLC technique. For example, subjective listening tests can be performed in which a listener scores the quality of an audio signal generated by simulating an error-prone transmission system that would cause frame loss, and in which the PLC technique to be tested is used to conceal the perceptual effects of the frame loss. The test subject can be asked to score the perceived intelligibility or quality of the audio. Using these tests, different PLC techniques can be compared based on the scores of the different PLC techniques. As an example, signal discontinuities that occur when an invalid frame is replaced with a generated replacement frame are known to cause perceptual artifacts, whereby the quality of a PLC technique can be roughly quantified by analyzing the smoothness of the audio signal at the transition between a valid frame and a generated replacement frame for an invalid frame.

[0020] According to a second aspect of the application, there is provided a method for concealing a lost or corrupted frame in a complex quadrature mirror filter, CQMF, domain. The method comprises receiving at least one first frame of CQMF samples, wherein each CQMF sample spans one or more time-frequency, TF, tiles, and wherein each TF tile is respectively associated with a frequency bin and a complex TF tile value. The method further comprises identifying whether a second frame of CQMF samples, subsequent to the at least one first frame, is valid or invalid, and when the second frame is identified as invalid, the method comprises, for at least one frequency bin of the respective frequency bins of one or more TF tiles of the CQMF samples of the at least one first frame, determining a complex parameter based on at least two of the TF tile values in the at least one first frame, and generating a replacement TF tile value by modifying at least one TF tile value of the first frame with the complex parameter. The method further comprises forming a replacement frame for the second frame based on the replacement TF tile value, and outputting the replacement frame.

[0021] In the CQMF domain, the computational efficiency of the PLC technology disclosed herein can be extremely high. Complex parameters (based on at least two TF block values ​​from at least one previous frame) can be used to generate a set of TF block values ​​for the replacement frame. In some embodiments, the same complex parameters are used iteratively to generate two or more (or all) TF block values ​​for the replacement frame.

[0022] According to a third aspect of the invention, an apparatus is provided, comprising a processor and a memory, the apparatus being configured to perform the method of the first aspect or the second aspect.

[0023] According to a fourth aspect of the invention, a non-transitory medium having software stored thereon is provided, the software including instructions for controlling one or more devices to perform the methods of either the first or second aspect.

[0024] According to a fifth aspect of the invention, a computer program product is provided, the computer program product comprising instructions that, when executed by a computer, cause the computer to perform a method according to either the first aspect or the second aspect.

[0025] According to a sixth aspect of the invention, a receiving unit is provided, the receiving unit comprising a decoder configured to receive a bitstream. The bitstream comprises data packets, wherein each data packet carries a frame comprising a plurality of time-frequency TF block values, wherein the decoder is configured to decode the received data packets of the bitstream to obtain the TF block values ​​of the current frame of the current data packet. The receiving unit further comprises a frame replacement module configured to identify whether the current frame is invalid, and in response to identifying the current frame as invalid, generate a set of TF block values ​​for a replacement frame to replace the currently invalid frame. The frame replacement module comprises a tone extractor configured to analyze at least one previous frame (the at least one previous frame preceding the current frame) to determine the tone level of the previous frame; and an adaptive sine spread and linear prediction module configured to determine whether the tone level exceeds a threshold, and when the tone level exceeds the threshold, generate a set of TF block values ​​for the replacement frame through a sine spread process, and when the tone level does not exceed the threshold, generate a set of TF block values ​​for the replacement frame through a linear prediction process. The receiving unit further includes a synthesis filter bank configured to receive the TF block values ​​of the replacement frame and convert the TF block values ​​of the replacement frame into time-domain audio segments.

[0026] The embodiments described herein can generally be described as technology, wherein the term “technology” can refer to (multiple) systems, (multiple) devices, (multiple) methods, (multiple) computer-readable instructions, (multiple) modules, (multiple) components, hardware logic and / or (multiple) operations as suggested by the context to which this is applied.

[0027] Features and technical benefits beyond those explicitly described above will become apparent upon reading the following detailed description and consulting the accompanying drawings. This summary is provided to illustrate the choice of techniques in a simplified form and is not intended to identify key or essential features of the claimed subject matter as defined by the appended claims.

[0028] The invention according to the third, fourth, fifth, and sixth aspects is characterized by the same or equivalent benefits as the invention according to the first or second aspect. Any function described with respect to the method may have corresponding features in an apparatus or computer program product. Attached Figure Description

[0029] The invention will be described in more detail with reference to the accompanying drawings.

[0030] Figure 1 a It is a block diagram illustrating the encoder, decoder, and frame replacement module.

[0031] Figure 1 b This is a block diagram illustrating an adaptive selection frame replacement module with sinusoidal spread and linear prediction.

[0032] Figure 2 It is a diagram illustrating a frame that includes multiple samples.

[0033] Figure 3 It is a diagram illustrating a frame sequence, in which one frame is invalid and is replaced by a generated replacement frame.

[0034] Figure 4 This is a flowchart illustrating the processing in a frame replacement module according to some implementations.

[0035] Figure 5 It is a diagram illustrating the sine wave extension process.

[0036] Figures 6a to 6d It is a diagram illustrating the linear prediction process.

[0037] Figure 7 This is a diagram illustrating the linear prediction process using a weighted combination of alternative frames.

[0038] Figure 8 It is a diagram illustrating the crossfade in and out from the generated replacement frame to the subsequent valid frames.

[0039] Figure 9 The illustration shows a schematic block diagram of an example device or architecture that can be used to implement embodiments of the present invention. Detailed Implementation

[0040] In the following description, numerous details such as system, device configuration, timing, and operation are set forth to provide an understanding of one or more aspects of this disclosure. It will be apparent to those skilled in the art that these specific details are merely examples and are not intended to limit the scope of this application.

[0041] In this document, the terms “and,” “or,” and “and / or” are used. These terms should be understood to have inclusive meanings. For example, “A and B” can at least mean: “both A and B,” or “at least both A and B.” As another example, “A or B” can at least mean: “at least A,” “at least B,” “both A and B,” or “at least both A and B.” As yet another example, “A and / or B” can at least mean: “A and B,” or “A or B.” When XOR is intended to be used, it will be specifically indicated (e.g., “either A or B,” or “at most one of A and B”).

[0042] The term "includes" and its variations should be interpreted as an open-ended term meaning "including but not limited to". The terms "one example implementation" and "an example implementation" should be understood as "at least one example implementation". The term "another implementation" should be understood as "at least one other implementation". The term "determine" should be understood as obtaining, receiving, calculating, estimating, predicting, or obtaining. Furthermore, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0043] This document describes the various processing functions associated with structures such as blocks, components, parts, and circuits. Typically, these structures can be implemented by one or more processors controlled by one or more computer programs.

[0044] Various acronyms that may appear throughout this disclosure and its associated claims and / or drawings are listed below. For the sake of brevity, other commonly used acronyms and technical terms may be excluded from this list. Therefore, a brief list of acronyms is provided below for the reader's quick reference.

[0045] IVAS – Immersive Voice and Audio Services PLC – Group Loss Hiding CRC – Cyclic Redundancy Check QMF - Quadrature Mirror Filter CQMF - Complex Quadrature Mirror Filter CLDFB – Complex Low Delay Filter Bank ACELP – Algebraically Activated Linear Prediction 3GPP – Third Generation Partnership Project EVS – Enhanced Voice Service IP – Internet Protocol UDP – User Datagram Protocol RTP – Real-time Transport Protocol BFI – Bad Frame Indicator Figure 1 a The transmitting unit 11 and the receiving unit 12 are schematically shown. The transmitting unit 11 includes an analysis filter bank 13 and an encoder 18. The receiving unit 12 includes a decoder 19, a frame replacement module 15, and a synthesis filter bank 14.

[0046] In the transmitting unit 11, the analysis filter bank 13 is configured to acquire time-domain audio segments of the audio signal and responsively convert the time-domain audio segments into time-frequency (TF) representations. The time-domain audio segments include time-domain samples of the audio signal, and the analysis filter bank 13 converts the time-domain audio samples into TF samples. Each TF sample can represent multiple time-domain samples. Each TF sample can further include multiple TF blocks, where each TF block represents a corresponding frequency band. The encoder 18 is configured to receive the TF-represented TF samples and responsively encode the TF samples into a bitstream, which can be output for use by the receiving unit 12. The encoder 18 can form a bitstream by collecting a predetermined number of successive TF samples to form frames of TF samples, thereby encoding these frames into data packets constituting the bitstream.

[0047] In the receiving unit 12, the decoder 19 is configured to receive data packets of the bitstream and responsively decode the data packets of the bitstream to preserve frames of TF samples. The frame replacement module 15 is configured to receive frames of TF samples from the decoder 19, and, if a frame of a TF sample is invalid, the frame replacement module 15 generates a replacement frame for the TF sample. The synthesis filter bank 14 is configured to receive frames of TF samples (or their replacement frames) and generate temporal samples of temporal audio segments for each TF sample. The synthesis filter bank 14 in the receiving unit 12 is configured to perform an inverse filtering operation relative to the analysis filter bank 13 of the transmitting unit 11.

[0048] As an overview Figure 1 aThe encoding and decoding process employed by the system can be described as follows: The time-domain audio segment received by the transmitting unit 11 is processed by the analysis filter bank 13 to obtain TF samples in the TF domain. The TF samples are grouped into frames, which are encoded into a bitstream by the encoder 18 and subsequently output (e.g., stored, transmitted, or sent). The bitstream is received by the receiving unit 12 and processed by the decoder 19 to obtain frames of TF samples. When necessary, the frame replacement module 15 processes the frames of TF samples to generate replacement frames with TF samples. The replacement frames with TF samples are processed by the synthesis filter bank 14 to generate time-domain samples of the recovered time-domain audio segment. As will be further described below, the function of the frame replacement module 15 is to ensure the mitigation of the adverse effects of frames with TF samples that are not received, damaged, or otherwise unusable by generating replacement frames to replace unreceived, damaged, or otherwise unusable frames.

[0049] In some implementations, the analysis filter bank 13 is a quadrature mirror filter (QMF) bank, whereby the TF samples are QMF samples in the QMF domain.

[0050] In some other implementations, the analysis filter bank 13 is a complex quadrature mirror filter (CQMF) bank, such as the complex low-latency filter bank (CLDFB) specified in the 3GPP Immersive Speech and Audio Services (IVAS) codec. In such implementations, the TF samples are CQMF samples in the CQMF domain (e.g., the CLDFB domain).

[0051] The analysis filter bank 13 processes the received audio segment in the time domain. The time-domain audio segment consists of time-domain audio samples. For example, the audio samples in the time domain are converted into TF representations in the time-frequency domain using the Fourier transform method as described above. The TF representations are then used for output or transmitted as a bitstream including TF representation samples. For example, each TF representation sample in the bitstream may correspond to a QMF or CQMF sample, as will be described in more detail below.

[0052] Each analysis filter bank 13 consists of multiple filters, each having a passband defined within a specific frequency range. For example, a simple analysis filter bank with three filters may include filters with passbands defined for frequencies below a first frequency (…). A low-pass filter with a passband at a frequency of ) having a specific frequency for the first frequency ( ) and the second frequency ( A bandpass filter with a passband between frequencies between ) and , and a filter with a passband for frequencies above the second frequency ( A high-pass filter with a passband for the frequency range of a given audio segment. Each TF sample comprises one or more TF blocks, wherein each TF block may correspond to a specific frequency range, such that each TF block can be mapped to the corresponding frequency band of a specific filter in the analysis filter bank 13. When the filters in the analysis filter bank 13 are applied to a time-domain audio segment, the time-domain sample of the audio segment is multiplied by the filter characteristics (e.g., passband or stopband filter coefficients, or their interpolation) to produce the output value for each time-domain sample. The output values ​​can then be decimated. Optionally, the decimated output values ​​form TF block values, each TF block value occupying a TF block of the TF sample.

[0053] When the QMF group is used as analysis filter group 13, the output from analysis filter group 13 is in the form of QMF samples, where each QMF sample represents a portion of the audio content, and where each QMF sample may include multiple TF blocks, each TF block representing a corresponding frequency band. Therefore, a QMF sample may include multiple TF blocks, where each TF block is mapped to a specific frequency range as described above. The number of TF blocks in a QMF sample corresponds to the number of filters in analysis filter group 13. Therefore, when a time-domain audio segment is processed with analysis filter group 12, each filter in analysis filter group 12 outputs a corresponding stream of TF block values ​​associated with the frequency band of the specific filter.

[0054] In other words, each QMF sample corresponds to a time slot of a time-domain audio segment, and each TF block value in one or more TF block values ​​of the same QMF sample (e.g., the same time slot) represents the frequency band of that time slot in the time-domain audio segment.

[0055] The specific passbands of the filters in filter bank 13 may not overlap or may partially overlap. For example, each QMF filter may be defined as a bandpass filter with a passband centered on a single center frequency. A first example bandpass filter may include two frequencies (BW1) that define a first bandwidth of the first passband. , The adjacent second bandpass filter may include two frequencies that define the second bandwidth BW2 of the second passband. , When the example passband frequency is limited to such that < < < When the frequency is such that the first and second passbands do not overlap, however, when the example passband frequency is limited to such that... < < < When the first passband and the second passband partially overlap, then the first passband and the second passband partially overlap.

[0056] Alternatively, the analysis filter bank 13 may include a CQMF group. Each filter in the CQMF group outputs a stream of complex-valued TF block values, wherein the complex-valued TF block values ​​are grouped together in time to form a CQMF sample. As a further example, if the analysis filter bank 13 is a CLDFB of a 3GPP IVAS codec, each filter has a 400 Hz passband, is substantially non-overlapping, and outputs complex-valued TF block values.

[0057] Because each filter in the analysis filter bank 13 is relatively narrow-band (e.g., when compared to a full-band audio representation sampled at, for example, 48 kHz), it is generally understood that the analysis filter bank 13 includes a large number of filters, such as ten or more, or one hundred or more. For example, when CLDFB is used as the analysis filter bank 13, there could be 60 filters, each with a 400 Hz non-overlapping passband for transforming a time-domain segment into a CLDFB sample, and each CLDFB sample containing 60 complex TF block values.

[0058] The QMF (e.g., CQMF) samples output by the analysis filter bank 13 can be encoded by the encoder 18 and subsequently transmitted (e.g., wirelessly), stored, or otherwise delivered to the receiving unit 12. As described above, the encoder 18 can group multiple QMF (e.g., CQMF) samples into frames and encode each frame into a coded data packet.

[0059] Decoder 19 is configured to decode the bitstream by decoding encoded data packets to obtain frames of QMF samples. The QMF samples of the decoded data packet frames are provided to synthesis filter bank 14, which reconstructs samples of the time-domain audio segment. Synthesis filter bank 14 in receiver 12 is configured to perform inverse filtering relative to analysis filter bank 13 in transmitter 11.

[0060] The sequence of encoded data packets (each packet carrying a frame of QMF samples) is thus transmitted via a bitstream to the receiving unit 12, where the decoder 19 performs a decoding operation to obtain QMF samples of the encoded data packets. Typically, the encoding of QMF samples may be a lossy operation (e.g., due to quantization errors), where the QMF samples obtained after decoding may be an approximation of the encoded QMF samples.

[0061] If receiving unit 12 successfully receives all data packets associated with a time-domain audio segment (i.e., all frames), all the information in synthesis filter bank 14 for recreating the time-domain audio segment (or at least reconstructing an approximation of the time-domain audio segment) is available, allowing receiving unit 12 to regenerate the time-domain audio segment (or at least regenerate an approximation of the time-domain audio segment). However, when receiving unit 12 does not receive one or more data packets associated with the time-domain audio signal, or when one or more received data packets are determined to be corrupted, receiving unit 12 will not be able to access all frames to fully reconstruct the corresponding time-domain audio segment. For this purpose, frame replacement module 15 is used to generate a set of TF block values ​​for replacement frames, which replace one or more frames of the unreceived or corrupted data packets(s).

[0062] exist Figure 1 a In this illustration, frame replacement module 15 is depicted as a separate module from decoder 19. However, this is merely an example, and frame replacement module 15 can be a component of decoder 19. Typically, decoder 19 and frame replacement module 15 will be implemented as hardware, firmware, or software in a client device or server that acts as receiving unit 12.

[0063] Figure 2 A frame 20 of a QMF (e.g., CQMF) sample is schematically shown. Frame 20 includes multiple TF blocks 21, which are grouped into multiple QMF samples 22, with time indices of... Each QMF sample 22 includes a quantity of The TF block 21, wherein each TF block 21 of the QMF sample corresponds to a frequency band. Related, where the index is .

[0064] exist Figure 2 In the diagram, a single frequency band As indicated by row 23, all TF blocks 21 in the same row contain TF block values ​​associated with the same frequency band and the same corresponding QMF in the QMF group used to extract TF block values.

[0065] The frame replacement module is configured to use, for example Figure 2 The frame 20, as depicted, operates. For example, the frame replacement module is configured to generate TF block values ​​for all TF blocks 21 used to fill a replacement frame (given one or more previous frames) that is used to replace frames that have been lost or otherwise unavailable.

[0066] Figure 2 The illustrative frame 20 includes five frequency bands (i.e., ) and five QMF samples 22 (i.e., This results in a total of 25 TF blocks 21. Figure 2 Frame 20 shown is merely an example, and it should be understood that the frame size can be determined almost arbitrarily depending on the implementation. For example, the Complex Low Delay Filter Bank (CLDFB) used in the 3GPP IVAS codec uses 60 filters, each with a bandwidth of 400 Hz, and each sample 22 represents 1.25 ms. A commonly used frame size represents 10 ms to 20 ms of audio content, or approximately 8 to 16 CLDFB samples, which in frame 20 total [number missing]. Each TF block 21 to 21 TF blocks. CLDFB can be applied to a 48 kHz sampled audio signal, whereby 60 filters with a 400 Hz bandwidth capture the lower 24 kHz of the 48 kHz sampled audio signal. However, since the spectrum of the 48 kHz audio signal is symmetrical about half the sampling frequency (24 kHz, also known as the folding frequency), the spectrum above 24 kHz can be reconstructed after decoding by mirroring the spectrum of the lower frequencies about 24 kHz.

[0067] Figure 1 b A frame replacement module 15 arranged according to some embodiments is schematically shown. The frame replacement module 15 includes a pitch extractor 17 and an adaptive sine spread and linear prediction module 16.

[0068] Frame replacement module 15 is configured to generate TF block values ​​for a replacement frame 20B' given at least one previous frame 20A, which may be referred to as the first frame 20A. Frame replacement module 15 includes an adaptive sine spread and linear prediction module 16 that generates TF block values ​​for the replacement frame 20B' using a sine spread process or a linear prediction process. Frame replacement module 15 further includes a tone extractor 17 configured to extract tone levels for each frequency band of the previous first frame 20A. Tone levels are provided to the adaptive sine spread and linear prediction module 16, which generates TF block values ​​for the replacement frame 20B' using a sine spread process or a linear prediction process, wherein the type of process employed is based on the tone level of each frequency band.

[0069] exist Figure 1 b In this process, the previous first frame 20A of the QMF sample is input to the frame replacement module 15. The first frame 20A is provided to the tone extractor 17, which extracts the tone for each frequency band. Extracting pitch For example, tone extractor 17 is a frequency band. Each of the values ​​in the text is extracted individually to indicate its pitch level. Each frequency band in the first frame 20A Pitch level Along with the first frame 20A, it is provided to the adaptive sine spread and linear prediction module 16. The adaptive sine spread and linear prediction module 16 is based on the frequency band-specific pitch level. Determining is for each frequency band Whether to use a sinusoidal spreading process or a linear prediction process. Using either a sinusoidal spreading process or a linear prediction process, the adaptive sinusoidal spreading and linear prediction module 16 generates the TF block values ​​for the replacement frame 20B' and outputs the frequency band values ​​according to the selected processing type. The replacement frame 20B' of the generated TF block value.

[0070] As an overview, according to some embodiments, the frame replacement module 15 is configured to analyze at least one previous first frame 20A and generate a set of TF block values ​​that can be used to form a replacement frame 20B' that replaces the subsequent second frame. The frame replacement module 15 may include an adaptive sine spread and linear prediction module 16 that uses a sine spread process or a linear prediction process to generate the set of TF block values. The type of process employed can be selected individually for each frequency band based on the pitch level of each frequency band of the first frame, as determined by the pitch extractor 17 of the frame replacement module 15. Thus, based on the pitch level of each individual frequency band, sine spread can be used to generate TF block values ​​for one or more frequency bands in the replacement frame 20B', and linear prediction can be used to generate TF block values ​​for one or more other frequency bands.

[0071] The frame replacement module 15 may be further configured to store one or more of the previous first frames 20A so that the TF block values ​​of these previous first frames can be accessed when the replacement frame 20B of the current second frame is to be generated.

[0072] Figure 3 A frame sequence containing three frames, 20A, 20B, and 20C, is schematically illustrated. The frame sequence includes a first frame 20A, a second frame 20B that follows the first frame 20A in time, and a third frame 20C that follows the second frame 20B in time. Each frame includes multiple QMF samples (columns), wherein each QMF sample includes multiple TF blocks, and each TF block corresponds to a specific frequency band. (Line) Related.

[0073] Each of frames 20A, 20B, and 20C in the frame sequence can be encoded by the encoder in the transmitting unit into a corresponding number of encoded data packets (each data packet carrying one frame), which are then transmitted as a bit stream to the receiving unit. However, one or more data packets may be lost or corrupted en route to the receiving unit, thus rendering the frames included in these lost or corrupted data packets unusable by the receiving unit.Figure 3 In the frame sequence, the second frame 20B may be a frame that is unavailable at the receiving unit (e.g., lost or damaged). In this case, the purpose of the frame replacement module of the receiving unit is to generate a replacement frame 20B' based on information from the previous frame. This replacement frame can be used to replace the unavailable second frame 20B.

[0074] refer to Figure 4 Flowcharts and Figure 3 The frame sequence will now be described, along with a method for forming replacement frames. This method can be executed by the frame replacement module 15 described above.

[0075] At step S1, “Obtain First Frame”, a first frame 20A is obtained. For example, the first frame 20A is associated with a bitstream received by the frame replacement module and / or the decoder. The first frame 20A may be a replacement frame (e.g., a frame previously generated by the frame replacement module in response to the loss or corruption of the original first frame) or the original first frame passed to the frame replacement module and / or the decoder.

[0076] At step S2a, “Check for a valid second frame”, the decoder and / or frame replacement module checks whether the second frame 20B following the first frame 20A is available and undamaged. Frames that are available for decoding and are undamaged are identified as valid frames. If a frame is unavailable, damaged, or otherwise unavailable, it is identified as invalid.

[0077] Frames (and / or data packets carrying frames) are intended to arrive at the receiving unit as a continuous sequence. However, if a frame is lost, delayed, or arrives corrupted (i.e., invalid) when the receiving unit should process the frame (or data packet) (e.g., to maintain uninterrupted audio playback), this situation should be identified to allow the generation of a replacement frame. Therefore, at step S2a, the receiving unit checks whether a valid second frame is available to allow, for example, sufficient time to generate a replacement frame.

[0078] At step S2b, "Is the second frame valid?", the result from step S2a is used to identify whether the second frame 20B is available or whether it has been lost, corrupted, or otherwise unavailable. This identification is called identifying whether the second frame 20B is a valid or invalid frame. Therefore, if the second frame 20B is unavailable at step S2a, or if the second frame 20B is determined to be corrupted at step S2a, the method determines that the second frame is invalid at step S2b. On the other hand, if the second frame 20B is available at step S2a and the second frame is not corrupted, the method determines that the second frame 20B is valid at step S2b.

[0079] The decoder (and / or similar frame replacement module) can operate in a frame-synchronized manner, meaning that the decoder... Decoding one frame per second, where, TIt is a predetermined time interval. For example, This is equal to the duration of a frame, such as 20 ms. In some implementations, frames 20A and 20B are transmitted as data packets to the receiving unit using one or more data transmission protocols. Example data transmission protocols include IP, UDP, and RTP. The receiver's decoder collects the data packets and can store them in a buffer. Data packets may arrive asynchronously, and some data packets may not arrive at all. For example, one or more data packets may fail to arrive due to buffer overflow caused by congestion along the transmission path. Additionally, some packets may arrive carrying errors, which are detected by the CRC mechanism. A packet that arrives too late to be available when the decoder needs the packet to operate in a frame-synchronized manner is an example of an invalid data packet, and the frame associated with that data packet is an example of an invalid frame. A data packet that does not arrive at the receiving unit at all (e.g., a lost packet) is another example of an invalid data packet, and the frame associated with that data packet is an example of an invalid frame. A data packet that arrives carrying an error detected by the CRC mechanism is yet another example of an invalid data packet, and the frame associated with that data packet is an example of an invalid frame.

[0080] For each instance where an invalid frame (e.g., each invalid packet) is identified, the Bad Frame Indicator (BFI) flag is triggered. When the decoder is about to decode a frame associated with the triggered BFI flag, it ignores the portion of the buffer that normally contains that frame and activates the packet loss concealment process. Therefore, identifying whether a frame is valid or invalid may include identifying whether the BFI flag has been triggered for that frame; thus, if the BFI flag is triggered, the frame is identified as invalid, and if no triggered BFI flag is present, the frame is identified as invalid.

[0081] Therefore, even if no frame is received, the decoder and / or frame replacement module can identify the frame as invalid. In frame synchronization operation, it can be assumed that frames should arrive at the receiver unit's decoder and / or frame replacement module within a predetermined time interval. The predetermined time interval can be measured from the last successfully received frame or data packet. Therefore, when a frame or data packet is determined to be unavailable after the predetermined time interval following the (successful) reception of a previous frame or data packet has expired, the frame can be marked (identified) as invalid.

[0082] When it is determined at step S2b that the second frame 20B is invalid, a set of TF block values ​​for the replacement frame 20B' will be generated.

[0083] In some implementations, the processing type for generating TF block values ​​is selected individually for each frequency band based on the tonal level of the corresponding frequency band in at least one previous frame. For this purpose, the generated replacement frame may include a mixture of frequency bands containing TF block values ​​generated by sinusoidal expansion and frequency bands containing TF block values ​​generated by linear prediction. Of course, the same general processing can be applied to frames with a single frequency band.

[0084] To generate a set of TF block values ​​for the replacement frame, the method may proceed to step S3, "obtaining the pitch of the first frame," which includes analyzing the first frame 20A to determine the pitch level of the first frame 20A. In some implementations, the pitch level is determined separately for each frequency band, so that a separate pitch level is generated for each of the multiple frequency bands.

[0085] Determining the pitch level of a frame (e.g., the first frame 20A) may involve analyzing the TF block values ​​of the frame (e.g., analyzing them individually in each frequency band). For example, the first frame 20A may be temporarily stored and accessed, where determining the pitch level includes analyzing the TF block values ​​of the first frame 20A.

[0086] In some implementations, determining the pitch level of a frame (e.g., a first frame 20A) includes evaluating at least one TF block value of the first frame 20A to identify (e.g., complex) phase differences between different TF block values ​​of the frame, and calculating a phase standard deviation based on the identified phase differences between different TF block values ​​of the frame. The phase standard deviation can be used as a pitch level, wherein a lower phase standard deviation can be used as an indication of a higher pitch level.

[0087] Alternatively, determining the pitch level of a frame (e.g., the first frame 20A) includes evaluating at least one TF block value of the first frame to identify amplitude ratios between different TF block values, and calculating an amplitude standard deviation based on the identified amplitude ratios between different TF block values. The amplitude standard can be used as a pitch level, wherein a lower amplitude standard deviation can be used as an indication of a higher pitch level.

[0088] As mentioned above, the pitch level can be determined individually for each frequency band. It is conceivable that the same method for determining pitch could be used in every frequency band, or that different methods could be used in different frequency bands. Non-limiting examples of the different methods considered are further described below.

[0089] Further reference Figure 1 b The image depicts a replacement module 15 according to some embodiments. The frame replacement module 15 includes a tone extractor 17 that analyzes the first frame 20A to individually determine the tone level of each frequency band. The tone level of each frequency band is represented as... And in Figure 1 bIn the example shown, the pitch levels of the four frequency bands The pitch extractor 17 provides the adaptive sine extension and linear prediction module 16.

[0090] The method can then proceed to step S4, “Is the pitch above the threshold?”, which involves determining the pitch level for each frequency band. Whether a predetermined threshold is exceeded. This step can be performed by the adaptive sine spread and linear prediction module 16. The predetermined threshold can be the same for each frequency band, or different predetermined thresholds can be used.

[0091] In one example, the predetermined threshold is a phase threshold. Whether a specific frequency band is tonal can be determined by comparing its standard deviation from the phase threshold. When the standard deviation is below the phase threshold, the specific frequency band is identified as tonal. When the standard deviation is equal to or greater than the phase threshold, the specific frequency band is identified as noise-like. The phase threshold of the standard deviation can be determined through calculation, experimentation, measurement, and other such techniques. In one example, the phase threshold of the standard deviation is approximately... .

[0092] As another example, the predetermined threshold is the amplitude threshold. By comparing the standard deviation of the amplitude to the amplitude threshold, it can be determined whether a specific frequency band is tonal. When the standard deviation of the amplitude is below the amplitude threshold, the specific frequency band is identified as tonal. When the standard deviation of the amplitude of a specific frequency band is equal to or greater than the amplitude threshold, the specific frequency band is identified as noise-like. The amplitude threshold for the standard deviation can be determined through calculation, experimentation, measurement, and other such techniques. In one example, the amplitude threshold for the standard deviation is approximately 0.1.

[0093] For the first frame 20A, the pitch level indicator TF block value is for each frequency band of the pitch, the method can proceed to step S5, "Generate TF block values ​​by sinusoidal expansion," which involves generating a set of TF block values ​​for the frequency bands of the replacement frame using a sinusoidal expansion process. (As in...) Figure 1 b As seen in the example, sinusoidal spread (abbreviated as SE) has been used to reconstruct the individual TF block values ​​of the intermediate frequency band number B2 in the replacement frame 20B'.

[0094] The sinusoidal spreading process may further include computing spread TF blocks. The spread TF block values ​​for a frame (e.g., the first frame 20A) are QMF samples that are temporally later than that frame (e.g., the first frame 20A). The TF block values ​​will be discussed below.

[0095] On the other hand, for frequency bands where the pitch level indicator TF block value is a noise-like signal block value, the method proceeds to step S6a, "Generate TF block values ​​via a linear predictor," which includes generating TF block values ​​for replacement frames using a linear prediction process. (As in...) Figure 1 b As seen in the example, linear prediction (abbreviated as LP) has been used to reconstruct the individual TF block values ​​of frequency bands B1, B3 and B4 in replacement frame 20B'.

[0096] In this way, the adaptive sine spread and linear prediction module 16 generates TF block values ​​for frequency bands with relatively high pitch through sine spread and generates TF block values ​​for frequency bands with relatively low pitch through linear prediction.

[0097] The linear prediction process may include calculating the extended TF block value based on at least two TF block values ​​in the first frame 20A through a linear prediction process.

[0098] Optionally, the method then proceeds to step S6b, “forming a weighted superposition,” which includes, for each frequency band, forming a weighted superposition between the TF block values ​​generated by the linear prediction process and the replacement frame. The replacement frame may be a copy of the previous first frame 20A. In some implementations, the weighted superposition is performed on a predetermined number of time-earliest TF block values ​​in the replacement frame, as will be described in further detail below.

[0099] It should be understood that the frame structure of the replacement frame can be equal to the frame structure of the invalid second frame 20B. The sinusoidal expansion process and the linear prediction process can be used to generate a subset (e.g., a proper subset or a strict subset) of the TF block values ​​of the replacement frame or to generate all TF block values ​​of the replacement frame. As described above, it is further understood that the steps of generating TF block values ​​by sinusoidal expansion or linear prediction are performed separately for each frequency band at steps S5 and S6a. This enables the generation of TF block values ​​for one or more frequency bands by sinusoidal expansion based on the tone level calculated for each frequency band at step S3, and the generation of TF block values ​​for one or more other frequency bands of the same frame by linear prediction.

[0100] It should be understood that when a frame includes TF block values ​​spanning more than one frequency band, a set of TF block values ​​for generating a replacement frame can be performed separately for each of the multiple frequency bands. As described, a sinusoidal spreading process or a linear prediction process is used separately for each frequency band to generate TF block values.

[0101] After generating a set of TF block values ​​for the replacement frame through sinusoidal spreading and / or linear prediction at steps S5 and / or S6a, S6b, the method may proceed to the optional step S7 "pre-compute extended TF block values," which includes determining the extended TF block values ​​using the selected TF block value generation process. The extended block values ​​are samples that are temporally later than the replacement frame. The TF block value is a block value of the TF block values. As will be described below, the TF block value generation process can, in principle, operate continuously, thereby generating any number of extended samples after generating a set (or optionally all) of TF block values ​​for the replacement frame. By generating extended TF block values ​​at step S7, the extended TF block values ​​can be stored or otherwise made available for use in the third frame 20C following the second frame 20B. Then, when the third frame 20C is a valid frame, the extended TF block values ​​are used to form a perceptually high-quality transition to the third frame 20C. If the third frame 20C is another invalid frame, the extended TF block values ​​can be retained or regenerated when generating the replacement frame for the third frame 20C.

[0102] The method can then proceed to step S8a, "output replacement frame," which involves outputting the replacement frame as a replacement for the second frame. The replacement frame output at step S8a is still in the QMF (e.g., CQMF) domain. To convert the QMF (e.g., CQMF) frame into a temporal representation, the method can proceed to step S8b, "perform synthesis filtering and output temporal frame," which may involve processing the replacement frame with a synthesis filter to form the output temporal audio frame.

[0103] Returning to step S2, when it is determined that the second frame 20B is a valid frame, the method proceeds to step S9, "Is the first frame valid?", which includes determining whether the first frame 20A is a valid or invalid frame. Determining whether the first frame 20A is valid may include identifying whether the BFI flag associated with the first frame 20A (or the data packet carrying the first frame 20A) has been triggered. For example, for each frame 20A, 20B, 20C processed by the replacement module and / or decoder, the BFI flag value indicates whether the frame is valid or invalid. The BFI flag may be temporarily stored to allow identification of whether previous frames are valid or invalid.

[0104] If it is determined at step S9 that the previous frame is valid, the method may proceed to step S10a, "outputting an unmodified second frame," which may include outputting an unmodified second frame. That is, if both the first frame 20A and the second frame 20B are identified as valid frames, no frame replacement needs to be performed, and all information is available to the decoder, thereby outputting an unmodified second frame 20B. For example, if the frame replacement module and / or the decoder obtain a sequence containing only valid frames, frame replacement is not required, and the frame replacement module outputs an unmodified frame. Optionally, to convert the QMF frame to a temporal representation, the method may proceed to step S10b, "performing synthesis filtering and outputting a temporal frame," which may include processing the replaced frame with a synthesis filter to form the output temporal audio frame.

[0105] If it is determined at step S9 that the previous frame is an invalid frame, the method proceeds to step S11, "Generate crossfade TF block values," which includes obtaining a set of extended TF block values ​​for the first frame 20A and generating one or more crossfade TF block values ​​for the second frame 20B. Generating crossfade TF block values ​​may include crossfading the extended TF block values ​​with the TF block values ​​of the second frame in the crossfade region. Note that at step S11, the first frame 20A is invalid (and has been replaced by the replacement frame), while the second frame 20B is valid. Thus, a perceptibly high-quality transition from the replacement frame to the second frame 20B is achieved. Similar to step S7, the set of extended TF block values ​​for the first frame 20A can be generated through a sinusoidal extension process or a linear predictor process; these values ​​may already be available because the extended TF block values ​​were generated when the first frame 20A was replaced by the replacement frame.

[0106] Step S11 can also be repeated individually for each frequency band. It is envisioned that the extended TF block value of one frequency band can be generated by sinusoidal extension, while the extended TF block value of another frequency band can be generated by linear prediction, depending on the tonal level of the frequency band determined in the previous frame.

[0107] The cross-fade-in / fade-out area may include a predetermined number of frames 20B of the second frame. The earliest sample in time.

[0108] In some implementations, the crossfade-in / fade-out (TF) block values ​​are formed as a crossfade-in / fade-out weighted sum of the TF block values ​​of the second frame 20B and the TF block values ​​from a set of extended TF block values ​​from the first frame 20A. The two terms of the crossfade-in / fade-out weighted sum are weighted using corresponding crossfade-in / fade-out weighting factors. The crossfade-in / fade-out weighting factor of the TF block values ​​of the second frame 20B increases over time between successive QMF samples. The crossfade-in / fade-out weighting factor of the TF block values ​​from the set of extended TF block values ​​decreases over time in a complementary manner between successive QMF samples.

[0109] In other words, the cross-fade-in / fade-out weighting factor can be related to the sample index. Inversely proportional. This results in each cross-fade-in / fade-out TF block value including values ​​that vary with the sample index. Increase and decrease the number of extended TF block values ​​and increase the number of TF block values ​​in the second frame 20B. In some examples, the crossfade weighting factor may vary linearly or non-linearly (e.g., exponentially) with the sample index to control the transition from the generated TF block values ​​to the TF block values ​​in the second frame 20B.

[0110] The method can then proceed to step S12a, "outputting a replacement frame with crossfade-in / fade-out TF block values," which includes outputting a replacement frame for replacing the second frame 20B, wherein the replacement frame includes crossfade-in / fade-out TF block values. The replacement frame may include crossfade-in / fade-out TF block values ​​in the crossfade-in / fade-out region, while in the tail region that is later in time than the crossfade-in / fade-out region, the replacement frame includes the unmodified TF block values ​​of the second frame 20B.

[0111] Optionally, in order to convert the replacement QMF frame into a temporal representation, the method may proceed to step S12b, "perform synthesis filtering and output temporal frame", which includes processing the replacement frame with a synthesis filter to form the output temporal audio frame.

[0112] The sinusoidal expansion and linear prediction processes used to generate a set of TF block values ​​are applicable to any type of frame carrying at least two temporally consecutive QMF samples. That is, each frame includes at least two temporally separated TF block values. In some implementations, each frame includes complex TF block values. For example, each frame is generated using a CQMF group (such as CLDFB).

[0113] Now we will assume Figure 3 Frames 20A, 20B, and 20C are multiframes generated using CQMF groups. Frames 20A, 20B, and 20C carry TF blocks 21, where each TF block includes a complex TF block value representing the frequency band and duration of the time-domain audio signal.

[0114] During normal operation, when frames 20A, 20B, and 20C are not identified as invalid (e.g., lost or corrupted), the stream containing frames 20A, 20B, and 20C can be decoded in the expected order. When a given frame is identified as invalid (e.g., lost or corrupted), there will typically be at least one frame that has been received without errors preceding that given frame. However, the sinusoidal spreading and linear prediction processes can operate recursively and generate TF block values ​​to replace any number of invalid frames, as described below.

[0115] To describe the sinusoidal expansion process and the linear prediction process, without loss of generality, it will be assumed that the second frame 20B is an invalid frame (for which a replacement frame will be generated), and that the previous first frame 20A was available or at least a replacement frame of the first frame 20A was available.

[0116] The third frame 20C following the second frame 20B can be identified as a valid frame or an invalid frame, as will be discussed below.

[0117] The two TF block value generation processes (i.e., the sinusoidal expansion process and the linear prediction process) can be used independently. That is, implementations utilizing only sinusoidal expansion or linear prediction are envisioned. Alternatively, these two processes can be used sequentially and / or in parallel for different frequency bands. In some implementations, an adaptive frame replacement technique is used, wherein one of the processes in the sinusoidal expansion process or the linear prediction process is used for each invalid frame or for each frequency band of each invalid frame, based on the tonal level of the previous frame (or its replacement frame) or for each frequency band of the previous frame.

[0118] exist Figure 3 In the frame sequence, the second frame 20B is an invalid frame, and a set of TF block values ​​(e.g., all TF block values ​​of the replacement frame 20B' with the same frame structure as the second frame 20B) are generated to form the replacement frame 20B'.

[0119] When operating in the CQMF domain, both the sinusoidal spread process and the linear prediction process determine complex parameters by using at least two TF block values ​​based on the first frame 20A. This is executed. The TF block value in the first frame 20A is used to determine the complex parameters. TF block 21 can belong to the same frequency band as the set of TF block values ​​to be generated. For example, when generating the second frame 20B belonging to the frequency band... When calculating the complex TF block value of TF block 21, the complex parameter... It is based on the frequency band in the first frame 20A. It is determined by the TF block value.

[0120] Complex parameters This can describe the ratio between the two TF block values ​​of the first frame (20A). Complex parameters This can describe the difference between two parameters obtained from two TF blocks. For example, complex parameters. Describes the phase difference between two TF block values. Complex parameters. Allowed by using complex parameters The TF block values ​​in the second frame 20B are generated by extrapolating from the TF block values ​​in the first frame 11. The extrapolated TF block values ​​can be used to determine complex parameters. One of the TF block values ​​or different TF block values. In some implementations, the TF block value used for extrapolation is the latest TF block value in time at frame 20A of the first frame.

[0121] Complex parameters It can be a phase parameter (e.g., or Amplitude parameters and one or more complex predictor coefficients One or more of them, as will be described below.

[0122] It is also possible to perform sinusoidal extension processes and linear prediction processes in the real-valued QMF domain (with some limitations), where complex parameters... It is replaced with a real-valued parameter that describes, for example, the ratio between two TF block values ​​in the first frame 20A.

[0123] The sinusoidal spread and linear prediction processes will now be described in detail. In this specification, it is assumed that each technique operates within a specific frequency band. These processes are carried out in the frequency bands, but it should be understood that they can operate in a similar manner across multiple frequency bands.

[0124] exist Figure 5 The diagram schematically illustrates a sinusoidal spreading process for generating TF block values ​​of a replacement frame 20B' based on at least one first frame 20A. The sinusoidal spreading process is performed by a frame replacement module to generate TF block values ​​for one or more frequency bands of the replacement frame.

[0125] The first frame 20A is either a valid frame or an invalid replacement frame for the previous first frame. The subsequent second frame is an invalid frame, in which the TF block value of the replacement frame 20B' should be generated for the second frame.

[0126] The first frame 20A uses multiple frequency bands. (including specific frequency bands) ) and range from 0 to CQMF samples The second frame and the replacement frame 20B' have the same overall structure as the first frame 20A. Since the second frame is invalid, a set of TF block values ​​will be generated for the replacement frame 20B'.

[0127] The sinusoidal spread process assumes that for a specific frequency band A sinusoidal signal component exists in the first frame 20A.

[0128] In the real-valued QMF domain, the sinusoidal signal component It can be modeled as (1) in, It is the amplitude parameter. It is the frequency of the sinusoidal signal component. It is the sampling frequency. It is the phase shift at the beginning of the first frame 20A, and This represents the QMF sample index. If the sinusoidal signal component in the QMF domain representation of the first frame (20A) is determined... parameters , , and Therefore, the generation of the QMF domain representation of the replacement frame 20B' can be performed based on the assumption that the same sinusoidal signal component will also continue into the second frame. In other words, it is possible to assume that the sinusoidal signal component... The replacement frame 20B' of the second frame is generated, continuing into the second frame (and therefore existing in the replacement frame 20B'). Sine wave component. The extension in the replacement frame 20B' of the frame is represented as ,in, In this context, it represents the QMF sample index of the replacement frame 20B' of the second frame.

[0129] Therefore, the QMF domain representation of the frame to the sinusoidal spread in the replacement frame 20B' can be determined as follows: (1) in, This represents the last temporal sample of the first frame 20A. The phase of the sinusoidal component in the equation, and This is a phase parameter describing the phase difference between successive time-domain samples in the first frame 20A, i.e., a sample-by-sample phase evolution parameter. The sample-by-sample phase evolution parameter can be expressed as: (2) If sample-by-sample phase evolution If this is known, then the last temporal sample of the first frame 20A can be extracted from the phase offset starting at the beginning of the first frame 20A. sinusoidal phase As shown below (3) However, reliably determining the parameters , and Generating replacement frames is challenging. For example, real-world signals are rarely pure sine waves and often have characteristics such as transients and noise.

[0130] Turning to the CQMF group domain, each TF block 21 includes complex TF block values, rather than real numbers as in the case of a real-valued QMF group domain or time-domain audio signal. To apply the sinusoidal extension technique in the CQMF domain, it is assumed that the complex sinusoidal signal components in the first frame 20A can be described as... (4) in, It is the complex-valued start parameter, whose value is equal to the first CQMF sample in the first frame 20A. The TF block value. Generally, it should be understood that the complex sinusoidal signal component can be identified individually for each frequency band. Based on the further assumption that this complex sinusoidal signal component extends from the first frame 20A to the replacement frame 20B', the extension of the complex sinusoidal signal component into the replacement frame 20B' can be represented as... (6) in, It is a complex-valued terminal parameter, having an amplitude equal to the aforementioned sine wave. The amplitude and the above The corresponding phase.

[0131] The complex sinusoidal signal component described in Equation 6 can be completely derived from a single complex evolution parameter. Describes the sample-by-sample phase evolution. Complex evolution parameters. It can be represented as (7) Furthermore, it should be recognized that, ideally, the TF block value from a CQMF sample To successive CQMF samples The phase evolution follows a pattern , where the function Extract the phase of the parameters. Typically, it is assumed that the phase evolution from one CQMF sample to the next is constant, meaning that complex evolution parameters can be used. To determine the phase of any TF block values ​​in the replacement frame 20B' of the invalid second frame. For example, it can be assumed that the CQMF sample The phase of the TF block value can be determined based on the index. The phase of the TF block value at the location is determined as follows: ,in, It can be any positive or negative integer.

[0132] TF block value of the replacement frame 20B' in the second frame By changing the terminal parameters and of It is generated by multiplying by powers. In other words... It can be obtained as follows (8) It can also be recursively represented as (9) in, As an alternative to the recursion in Equation 9 (which generates successive TF block values ​​based on previous TF block values), the last TF block value can also be calculated, for example, using Equation 8 (in CQMF samples). (place) and then by multiplying Recursively determine the earlier TF block value in the replacement frame 20B' to perform recursion in the opposite direction.

[0133] In other words, using equation 8 or 9, we can base it on... and / or (include Generate a set of TF block values ​​for the replacement frame 20B' of the second frame. If and If both are known, the most accurate sine expansion can be achieved, but it is assumed that either one is sufficient to generate TF block values.

[0134] For example, if only If it is known, then it can be used. The default value. An example is by... It is obtained by setting the center frequency of a specific frequency band. For example, for a critical sampling frequency band with a bandwidth of 400 Hz, A suitable choice would be 200 Hz. While this may not correspond to the sinusoidal component of the frequency band, it might be a sufficient approximation. On the other hand, if only If it is known, then it can be used. The default value. Although this may not correspond to the TF block value of the last CQMF sample in the first frame, it is retained as... The frequency of the described sinusoidal component may be sufficient to form an acceptable replacement frame.

[0135] In some implementations, it is determined that having A vector of elements, each element corresponding to an index of replacement frame 20B'. The index ranges from 0 to Each of the N elements in the vector is... Proportional (e.g., equal), thus by dividing the TF block values ​​of the first frame 20A (e.g., end parameters) The vector is multiplied to generate a set of TF block values ​​for the replacement frame 20B'.

[0136] Phase parameters This can be determined through various methods. Assuming the sample-by-sample phase evolution is approximately constant in the first frame 20A, any two successive TF block values ​​in the first frame 20A can be selected, and the phase parameter... This is determined to be the phase difference between the two TF block values. In practice, individual TF block values... The phase can be obtained by taking the TF block value. The imaginary part and TF block value The ratio between the real parts is determined by the arctangent of the phase modulus. Since the implementation of the arctangent function (e.g., atan2) returns the phase modulus... Therefore, the phase can be expanded to produce an expanded phase. .

[0137] Then, the phase parameter It can be determined as the difference between the (expanded) phases of two successive TF block values, for example, for any , .

[0138] It should also be noted that the sample-by-sample phase difference can also be determined based on the selection of non-sequential TF block values. For example, if the first TF block value is determined... With the second TF block value The phase difference between them, then for any This can be achieved by dividing the difference by To calculate the phase difference per sample.

[0139] To mitigate the effects of statistical bias, the phase parameter It can be identified as a phase context window The average phase difference between the TF block values ​​of two (e.g., consecutive) CQMF samples.

[0140] Phase Context Window It can span multiple consecutive CQMF samples across the first 20A frame. For example, the phase context window. The phase context window spans all CQMF samples across the first frame 20A, or a subset of the CQMF samples of the first frame 20A (e.g., a proper subset or a strict subset). As another example, the phase context window spans CQMF samples across more than one frame, such as at least two first frames that are consecutive but precede the second frame to be replaced by the replacement frame 20B'. Typically, the TF block value generation process defined in this paper is not limited to analyzing only the first frame that comes immediately before, but can also analyze multiple previous first frames (whether they are valid frames or replacement frames for invalid frames) to reduce statistical fluctuations.

[0141] Because the CQMF samples are closer to the end of the time frame of the first frame 20A (i.e., when...) near The phase context window is closer in time to the second frame, and these CQMF samples are likely more important for generating the TF block values ​​of the replacement frame 20B' compared to the earlier CQMF samples in the first frame 20A. It can contain at least The last CQMF sample and the second-to-last CQMF sample at the location And optionally the penultimate CQMF sample And optionally, additional consecutive CQMF samples in the end portion of the first frame 20A.

[0142] Average phase parameter The phase difference between the (expanded) phase differences of the TF block values ​​of consecutive CQMF samples can be summed and divided by the phase context window. The number of CQMF samples in the window is used for calculation. That is, if, for example, the phase context window... It has For the entire first valid frame 20A of CQMF samples, the average phase parameter It can be determined as (10) in, The range is from 0 to .

[0143] In addition, when When expanding the phase, what can be recognized is the average phase parameter. This can be determined (for a phase context window covering the entire valid frame 20A). (11) Of course, equations 10 and 11 can be modified to fit shorter or longer phase context windows. .

[0144] As a parameter for determining the average phase An alternative is to identify and use the median or mode phase evolution parameter instead of the mean phase parameter. .

[0145] In some implementations, each TF block value can be provided in polar coordinates (i.e., each TF block value is represented by its magnitude and expanded phase values), or each TF block value is converted to polar coordinates. That is, each TF block value can be represented in the following form: (12) in, This represents the magnitude of the TF block value, and This represents the expanded phase of the TF block values. When the TF block values ​​are represented in polar coordinates, the expanded phase... It can be directly extracted from each TF block value of each CQMF sample and used, for example, to determine the average phase difference parameter according to the above. .

[0146] Go to end parameter According to some implementation methods, this parameter can be directly determined as the last CQMF sample point of the first frame 20A (in The TF block value at (location). That is, Alternatively, multiple TF block values ​​in the first frame 20A can be analyzed to determine... The expected value thus mitigates any statistical bias, where, The expected value is used as For example, a linear predictor (as described below) can be used to predict based on one or more previous TF block values ​​in the first frame 20A. The value. Optionally, a linear predictor and the first frame 20A are used. The observed values ​​will The expected value is identified as The predicted value.

[0147] thus, The expected value can be based on It is determined by the observed values. For example, The expected value can be determined through a smoothing operation that removes statistical fluctuations (e.g., noise). The smoothing operation can be based, for example, on a CQMF sample. Regression of one or more previous CQMF samples.

[0148] A useful further improvement to the above sinusoidal expansion process is to consider the case of amplitude decay or increase. Analysis of amplitude decay or increase can improve the perceived quality of the replacement frame compared to simply expanding a sine wave with a fixed amplitude. It should be understood that the values ​​derived from Equation 7 above... The amplitude is one, which means that for All CQMF samples, The amplitude will remain unchanged.

[0149] By recursion parameters Redefining (13) Amplitude parameters can be used For exponentially decreasing ( ) or increasing exponentially ( Modeling is performed using a sine wave. The amplitude parameters of the first frame (20A) are determined (or at least approximated). The trend of amplitude decay or increase can be extended from the first frame 20A to the TF block value of the replacement frame 20B' of the second frame 20B.

[0150] In some implementations, to determine the amplitude parameter κ, the phase parameter in equations 10 and 11 above can be used. A similar approach is to determine the amplitude context window. The average magnitude ratio between successive TF block values. Amplitude context window. Can be used with the phase context window They may have the same or different sizes.

[0151] Since κ is an exponential parameter, it is determined to fall within the magnitude context window. Sample-by-sample amplitude ratio of all CQMF samples within The average value is appropriate.

[0152] It should also be noted that κ can also be calculated as the geometric mean (rather than the arithmetic mean) across the amplitude context window, which produces (14) Similar to phase parameters It should be understood that the average magnitude parameter can be replaced with the median or mode magnitude parameter determined in the magnitude context window.

[0153] In some implementations, the amplitude parameter exceeds one. This may be undesirable because it leads to an exponential increase in amplitude, which could cause problems with the resulting audio quality and / or result in an overload effect. Therefore, the amplitude parameter can be adjusted... Implementing an upper limit of 1 makes the amplitude parameter It is equal to or less than 1. In practice, this limitation allows the amplitude of the sinusoidal extension of the replacement frame 20B of the second frame 20B to remain unchanged or decrease, but not increase.

[0154] The sinusoidal expansion process expands the sinusoidal signal components of the first frame 20A to generate the TF block values ​​of the replacement frame 20B' for the second frame 20B. The spectrum of the pure sinusoidal signal components will be located at the frequencies of the sinusoidal signal components. The extended sinusoidal signal component is a single spectral peak at a single frequency. That is, the extended sinusoidal signal component is a single tone at a single frequency, which is an extremely narrowband signal. Theoretically, a single frequency peak can be described as having zero bandwidth. In some implementations, it is desirable for the signal generated by the narrowband (single frequency) signal, especially when the signal represented by the first frame 20A is also narrowband and has a high degree of pitch. Then, the TF block values ​​generated by the replacement frame 20B' of the second frame 20B will be a reasonable and accurate extension of the tone signal in the first frame 20A.

[0155] On the other hand, in some implementations, the signal of the first frame 20A has a low pitch and is characterized by a non-zero bandwidth. For example, many types of audio signals will be characterized by a spectral maximum at a specific frequency, while being surrounded by additional frequency components that form a non-zero bandwidth. To more accurately generate the TF block values ​​of the replacement frame 20B' of the second frame 20B for this type of non-tonal signal, the sinusoidal spreading process can be modified to implement phase perturbation parameters. and amplitude disturbance parameters One or both of them.

[0156] Phase perturbation parameters and amplitude disturbance parameters The purpose is to modify the phase parameters determined by the sinusoidal extension process. and amplitude parameters In order to generate non-zero bandwidth extension This non-zero bandwidth extension is closer to the TF block value of the first frame (20A). The non-zero bandwidth of the signal represented. Phase perturbation parameter. and amplitude disturbance parameters It can be described as additive perturbation noise, which can be determined individually for each generated TF block value or drawn from a random or pseudo-random sequence.

[0157] For each generated TF block value, the phase perturbation parameter It can be determined and used for phase parameters Perturbation is performed to form perturbation phase parameters This perturbation phase parameter is used to generate the TF block values ​​for the replacement frame 20B', rather than the phase parameter. For example, the phase perturbation parameter... Additive perturbations can be defined, and the perturbation phase parameters can be defined accordingly. Calculated as (6) For each CQMF sample .

[0158] For each CQMF sample of replacement frame 20B' Phase perturbation parameters It can be obtained from a source with a predetermined standard deviation. The predetermined standard deviation is determined by randomly selecting parameter values ​​from the distribution. The phase difference can be based on the complete first frame 20A or on the sample-by-sample phase difference of valid frames 20A within a context window spanning at least a portion of at least one first frame 20A (which may be equal to or different from the phase context window). The standard deviation. For example, the predetermined standard deviation. It can be equal to all CQMF samples of the first frame 20A (or its context window). Sample-by-sample phase difference The standard deviation or proportional to it.

[0159] In some implementations, operating the random number generator for each TF block value to be generated is undesirable because it increases computational complexity and can potentially increase latency. Therefore, in some implementations, the phase perturbation parameter for a given TF block value to be generated is... The sample-by-sample phase difference is obtained from the first frame 20A (or its context window). Compared with the average sample-by-sample phase difference The difference between them. Another phase perturbation parameter is needed to obtain another TF block value to be generated. Calculate the sample-by-sample phase difference of another pair of TF block values ​​in the first frame 20A. Compared with the average sample-by-sample phase difference The difference between them. Then, this process is repeated for multiple pairs of TF block values ​​in the first frame 20A to generate new phase perturbation parameters. , used to generate TF block values ​​in the replacement frame 20B' of the second frame, these generated TF block values ​​will inherently be characterized by the same standard deviation, without operating the random number generator.

[0160] In some implementations, the sample-by-sample phase difference from the valid frame 20A (or its context window) Compared with the average sample-by-sample phase difference The difference between them is scaled before the TF block value used to generate the replacement frame 20B' for the second frame is generated.

[0161] Although the phase perturbation parameter Obtain the sample-by-sample phase difference With average phase difference The difference between them allows for the determination of individual phase perturbation parameters for each TF block value to be generated. However, this may not always be necessary. For example, to reduce computational complexity, a set of phase perturbation parameters can be determined for a subset of the TF block value pairs in the first frame 20A. And then the set of phase perturbation parameters can be repeated (e.g., looped) for the TF block values ​​to be generated for the replacement frame 20B' of the second frame. .

[0162] Similarly, amplitude perturbation parameters can be determined for each TF block value to be generated. And use it for the amplitude parameter Perturbation is performed to form perturbation phase parameters Use the perturbation phase parameter instead of the original amplitude parameter To generate TF block values. For example, amplitude perturbation parameters. Additive perturbations can be defined, and the perturbation phase parameters can be defined accordingly. Obtained as (16) For each CQMF sample in the second frame, the amplitude perturbation parameter It can be obtained from a source with a predetermined standard deviation. The predetermined standard deviation is determined by randomly selecting parameter values ​​from the distribution. It can be based on the first frame 20A or the context window of the first frame 20A (which can be equal to or different from the amplitude context window). The per-sample amplitude ratio The standard deviation. For example, the predetermined standard deviation. It can be equal to the per-sample amplitude ratio of the first frame 20A (or its context window). The standard deviation or proportional to it.

[0163] Alternatively, to avoid using a random number generator, the relative amplitude change and average amplitude parameter of two adjacent CQMF samples in valid frame 20A can be calculated. The difference between them determines the amplitude disturbance parameter. Similarly, this can be achieved by calculating the relative amplitude changes of different pairs. and average amplitude parameter To determine the additional amplitude disturbance parameters In this case, the applied perturbation can be multiplicative according to the following formula. (17) In some implementations, amplitude perturbation parameters The calculation is performed in the logarithmic domain, and in the logarithmic domain, the amplitude perturbation parameter... The phase perturbation parameters can be similar to those described above. To determine.

[0164] Figures 6a to 6d A linear predictor process for generating a set of TF block values ​​for the replacement frame 20B' of the second frame is schematically illustrated. The linear predictor process involves defining a set of complex predictor coefficients. A linear predictor recursively generates TF block values ​​for one or more frequency bands of the replacement frame. The linear predictor process is executed by the frame replacement module to generate TF block values ​​for one or more frequency bands of the replacement frame.

[0165] exist Figures 6a to 6d In this context, the TF block value of the first frame 20A (or the replacement frame thus generated) is used to generate the replacement frame 20B' of the subsequent invalid frames.

[0166] The linear prediction process is a replacement process for the sinusoidal expansion process. Based on the tonal level of the frequency band in the first frame 20A, the TF block replacing frame 20B' can be generated by the frame replacement module using either the sinusoidal expansion process or the sinusoidal expansion process.

[0167] During the linear prediction process, specific frequency bands of the first frame 20A are analyzed. The TF block values ​​are analyzed. For example, the TF block values ​​of at least two CQMF samples in the first frame are analyzed. The TF block values ​​are then analyzed to determine a linear predictor, which is based on one or more previous TF block values. Generate predicted TF block values .

[0168] The first-order linear predictor is based on a previous TF block value (e.g., Generate predicted TF block values The second-order linear predictor is based on two previous TF block values ​​(e.g., and Generate predicted TF block values The third-order linear predictor generates predicted TF block values ​​based on three previous TF block values. And this applies to fourth-order linear predictors and higher-order linear predictors as well.

[0169] In some implementations, the linear predictor is first-, second-, or third-order to maintain low computational complexity. Since the linear predictor according to some implementations operates in a single specific frequency band with limited bandwidth (e.g., about 400 Hz for CLDFB, or at least less than 1 kHz), the order of the linear predictor can be very low, such as only first- or second-order.

[0170] A linear predictor can be defined as minimizing a specific frequency band in the first frame 20A (or its context window). The prediction error for all TF block values. For example, the prediction error is for the first frame 20A (or the predictor context window in the first frame 20A). In or across multiple frames in the predictor context window All TF block values ​​in (the middle), predict the TF block value. TF block value of the first frame 20A The expected value of the mean squared error between the two. For example, the prediction error can be expressed as... .

[0171] In some implementations, the linear predictor is defined as (18) Among them, index The range is from 1 to , It is the order of the predictor, and These are complex-valued parameters. This type of linear predictor generates TF block values ​​based on at least one previous and immediately adjacent TF block value. As understood from Equation 18, the linear predictor includes as many predictor parameters as the predictor order. In other words, a first-order linear predictor only includes parameters. The second-order linear predictor includes predictor parameters. and As an example, a third-order linear predictor will divide the TF into blocks of values. Generate as (19) To determine the parameters of the complex predictor This can be used to formulate an optimization problem to fit the prediction error of all TF block values ​​for a specific frequency band in the first frame 20A (or across the context window of the first frame and optionally earlier frames). Minimize the complex predictor parameters .

[0172] In some implementations, one or more predictor coefficients It is an autocorrelation sequence based on the TF block values ​​of the first frame 20A. It is calculated according to the Levinson-Durbin method. For a first-order linear predictor, the coefficients of a single complex predictor can be used in this case. Simply extract as (20) To generate the TF block values ​​for the second frame 20B, a linear predictor is used to predict the block values ​​for the second frame 20B using one or more TF block values ​​from the first frame 20A as starting conditions. For example, the linear predictor is used to take the TF block values ​​of an invalid frame 20B as input. Generate as (twenty one) Among them, the negative CQMF sample index CQMF samples Taken from the later end of the previous first frame 20A, which means .(twenty two) Figure 6a and Figure 6b The diagram schematically illustrates how to define a third-order linear predictor from Equation 21 to generate accurate predictions of the TF block values ​​in the first frame 20A. Figure 6a In the middle, the third-order predictor uses previous CQMF samples Three complex predictor parameters for operating on TF block values , , Generate CQMF samples Specific frequency bands The prediction of TF block values ​​in the data. Similarly, the third-order predictor can use data from previous CQMF samples. Three complex predictor parameters for operating on TF block values , , Generate CQMF samples Specific frequency bands Prediction of TF block values ​​in, such as Figure 6b As shown.

[0173] However, although Figure 6a and Figure 6b The illustration shows the task that the linear predictor has been configured to perform (i.e., to generate accurate predictions of the TF block values ​​for the first frame 20A), but the same predictor is also used to generate specific frequency bands in the replacement frame 20B' of the second frame. TF block value.

[0174] Figure 6c The illustration shows the source Figure 6a and Figure 6b The predictor (using the same complex predictor parameters) , , The predictor uses CQMF samples according to equations 21 and 22 above. and The last TF block value of the first frame 20A is used to predict the first CQMF sample of the replacement frame 20B' of the second frame 20B. The TF block value. In the CQMF samples that predicted the replacement frame 20B' of the second frame 20B. In the case of the TF block value, the linear predictor can then continue, using the CQMF sample of replacement frame 20B'. The (predicted) TF block values ​​and CQMF samples and The TF block values ​​of the previous first frame 20A are used to predict the CQMF samples of the replacement frame 20B' of the second frame 20B. TF block values ​​in, such as Figure 6d As shown. In a similar manner, the linear predictor can continue to operate using the TF block values ​​of the replacement frame 20B' of the second frame 20B to operate in a specific frequency band. Generate all TF block values ​​for replacement frame 20B'.

[0175] For a first-order linear predictor of the type that operates on the immediately preceding CQMF sample according to Equation 19, it should be noted that the complex predictor parameters... Corresponding to the complex factors from Equation 10 Therefore, the linear predictor parameters The determination of (e.g., using Equation 20) is the determination of the complex evolution factor. Another alternative is to combine the amplitude parameter. and phase parameters .

[0176] It has been found that linear predictors of the type from Equation 19 tend to predict decaying TF block values. That is, linear predictors tend to predict complex predictor parameters. These parameters gradually generate TF block values ​​with decreasing amplitude. This behavior is because the poles of the inference filter defined by the linear predictor are typically inside the unit circle, and only poles precisely located on the unit circle allow for the generation of TF block values ​​with constant amplitude. The replacement frame 20B' with attenuated TF block values ​​may cause the decoded time-domain audio signal to be perceived as unstable.

[0177] Therefore, it is envisioned that, for a first-order predictor, the calculated complex predictor parameters... It can be modified in the following ways: set its amplitude to one, or if If the value is not equal to one or is lower than a predetermined value close to one, then the amplitude is set to at least that predetermined value close to one. Then, when generating the TF block values ​​for the replacement frame of the second frame, the linear predictor can use the modified complex predictor parameters obtained in this way. To replace the calculated complex predictor parameters .

[0178] Furthermore, it should be noted that for a first-order predictor, the calculated complex predictor parameters... This is a multiplicative value that describes how the TF block value evolves from one CQMF sample to the next. In other words, it represents the complex predictor parameters of the first-order predictor. Equal to the complexization parameter, making Therefore, if the complex predictor parameters If modified to have an amplitude of one, then the complex evolution parameter The phase evolution portion is preserved. Simultaneously, this relationship indicates that the complex predictor parameters of the first-order predictor are... It can be used to determine phase evolution parameters. (as parameters of the complex predictor) (phase) and / or amplitude evolution parameters (as parameters of the complex predictor) (the amplitude). Therefore, the complex predictor parameters Phase evolution parameters for determining the sinusoidal extension process are introduced. and / or amplitude evolution parameters Another approach is to first determine the parameters of the complex predictor. Then based on calculate and / or The opposite also applies, meaning that it can be determined according to... (or and )calculate .

[0179] For higher-order predictors (e.g., second- or third-order predictors), the weakened signal can also be counteracted by modifying the predictor parameters according to their calculated values. In some implementations, this is achieved by calculating the roots of the predictor polynomial and subsequently identifying the modified predictor parameters. This allows the root to maintain its phase value while moving closer to or onto the unit circle. Using these modified predictor parameters, the magnitude of the generated TF block values ​​typically decreases more slowly or not at all, which can help improve the perceptual quality of replacement frames.

[0180] Figure 7 The diagram illustrates the first frame 20A, the generated replacement frame 20B', and the substitute frame 20A'.

[0181] The TF block values ​​of the first frame 20A are used to generate the TF block values ​​of the replacement frame 20B' using a sinusoidal spread process or a linear prediction process in each frequency band. Additionally, the TF block values ​​of the replacement frame 20B' are further modified by weighted summing of the corresponding TF block values ​​of the replacement frame 20A'.

[0182] exist Figure 7 In this process, the linear prediction process described above is used to generate the frequency bands in the replacement frame 20B'. The TF block value is then added to the corresponding TF block value in the replacement frame 20A' to further modify the generated TF block value in the replacement frame 20B'. For example, the replacement frame 20A' is a copy of the first frame 20A.

[0183] In other words, the TF block value of the replacement frame 20B' of the second frame 20B can be formed as a weighted sum of the generated TF block value (using sine spread or linear prediction) and the replacement frame 20A'. In one example, the replacement frame 20A' is equal to the first frame 20A or its cyclically shifted version.

[0184] Amplitude parameter of sinusoidal extension (It is equal to the first-order predictor parameters of the first-order predictor) The exponential decay caused by the amplitude can be mitigated by adding a weighting factor. (or The weighted replacement frame 20A' is used to compensate, so that the generated TF block value of the replacement frame 20B' of the second frame 20B is formed as a weighted sum of the TF block value of the replacement frame 20A' and the TF block value generated from the sinusoidal expansion process or the linear prediction process.

[0185] For example, when using linear prediction, each CQMF sample in the replacement frame 20B' of the second frame 20B It can be generated as (twenty three) in, It is the corresponding CQMF sample that replaces frame 20A' (i.e., has the same frequency band and index). As mentioned above, the appropriate choice for alternative frame 20A' is to use the first frame 20A, such that... .(twenty four) In some implementations, the alternative frame 20A' is a cyclically shifted version of the first frame 20A. For example, the cyclically shifted version applies to all... It can be obtained as And for all It can be obtained as ,in, It is an integer.

[0186] A high-quality replacement frame is provided by reintroducing the replacement frame 20A' with weights that compensate for the exponential decay of the generated TF block values. This creates a crossfade effect from the TF block values ​​generated through a sinusoidal expansion or linear prediction process to the TF block values ​​of the replacement frame 20A', which has the perceived benefit of avoiding signal discontinuities that would otherwise occur if frame repetition techniques were applied instead of generating TF block values ​​through sinusoidal expansion or linear prediction.

[0187] According to the above technique, the TF block value of the replacement frame 20B' for an invalid frame can be generated. Now, the process for transitioning from the replacement frame 20B' generated for the second frame 20B back to the subsequent valid third frame 20C (which has been received without errors) and for generating the TF block value when two or more consecutive frames are invalid will be described.

[0188] Figure 8 The diagram illustrates a replacement frame 20B' that has been generated to replace the second frame, a subsequent third valid frame 20C, and a replacement frame 20C' that has been generated to replace the third valid frame 20C.

[0189] Replacement frames can be generated for all invalid frames in a frame sequence. Additionally, in some cases, replacement frames can also be generated for valid frames. For example, when a replacement frame 20B' has already been generated to replace the second invalid frame and the subsequent third frame 20C is a valid frame, a replacement frame 20C' for the third frame 20C can be generated and used as the replacement frame 20C' for the third frame 20C. The purpose of replacing a valid frame following an invalid frame with a replacement frame can be to smoothly transition from a replacement frame (which is an approximation of the original frame) to a valid frame.

[0190] Replace the TF block value of frame 20B' The third frame 20C has been generated using the sinusoidal spreading process and / or linear prediction process described above. The third frame 20C is the valid frame following the replacement frame 20B' of the second frame. To provide a high-quality transition from the replacement frame 20B' to the third frame 20C, various techniques can be used to avoid signal discontinuities that would interfere with the listener. Each frequency band of the third frame 20C... CQMF samples were labeled as , 。

[0191] In some implementations, crossfade-in / fade-out techniques are applied to form the replacement frame 20C' of the third frame, transitioning the TF block values ​​from the replacement frame 20B' of the second frame to the third frame 20C. The crossfade-in / fade-out technique involves generating extended TF block values ​​for the replacement frame 20B' of the second frame (these extended TF block values ​​exceed the values ​​of the replacement frame 20B' of the second frame). (number of CQMF samples), and use more than These extended CQMF samples of the three CQMF samples are used to achieve the replacement frame 20C' of the third frame 20C by cross-fading in and out within the cross-fade-in / fade-out region to the TF block value of the third frame 20C. That is, the sinusoidal extension process and / or the linear prediction process are used to generate the replacement frame 20B' of the second frame. A series of continuously extended TF block values, where the CQMF sample index is .These The extended TF block value is temporally related to the effective third frame 20C. The TF block values ​​overlap.

[0192] Generated samples exceeding CQMF Number of extended TF block values This can be adjusted based on the desired length of the cross-fade-in / fade-out region. In one example, samples exceeding the CQMF are generated. The TF block values ​​of two CQMF samples, i.e., CQMF samples and The TF block value. In the CLDFB domain, two CLDFB samples correspond to 2.5 ms of audio content.

[0193] The earliest TF block value in the K time intervals of the third frame 20C is marked as And cross-fade-in and cross-fade-out can be performed according to the following formula. (25) in, It is the replacement frame 20C' of the third frame. The cross-fade-in / fade-out TF block values ​​of the earliest CQMF sample in time, and These are cross-fade-in / fade-out interpolation parameters. The replacement frame 20C' for the third frame 20C. The earliest TF block value in time was thus obtained as the CQMF sample of the replacement frame 20B' beyond the second frame. The generated extended TF block value TF block value of the effective third frame 20C Cross-fade in and out.

[0194] Crossfade-in / fade-out interpolation parameters It can be adjusted to follow The value decreases as it increases. This produces replacement frame 20C', where the preceding third frame 20C... The TF block values ​​gradually increase, and the generated extended TF block values... Accordingly, it gradually fades. In one example, the cross-fade-in / fade-out interpolation parameters... Linear interpolation is defined and can be set to However, it should be understood that the cross-fade-in / fade-out interpolation parameters Other definitions are possible, for example, to define a non-linear cross-fade-in / fade-out curve.

[0195] For the replacement frame 20C' of the third frame 20C, which is later in time than the crossfade-in / fade-out area... TF block values ​​(i.e., TF block values) The TF block value of the third frame 20C is set to the TF block value of the third frame 20C. The TF block values ​​of the third frame 20C that are later in time than the crossfade-in and fade-out areas can be referred to as the TF block values ​​of the tail area. These TF block values ​​are equal to the corresponding TF block values ​​of the third frame 20C.

[0196] Although it is possible Adjust the crossfade-in / fade-out interpolation parameters on each extended TF block value. To achieve the desired cross-fade-in / fade-out curves, but further utilization of attenuation-type linear prediction parameters is possible. Or using amplitude parameters A sinusoidal spread less than 1. According to this technique, the crossfade-in / fade-out duration will not be affected by a predetermined number of CQMF samples. Controlled, but subject to linear prediction parameters or amplitude evolution parameters Attenuation rate control.

[0197] In one implementation, a linear predictor process or a sinusoidal expansion process is used to generate CQMF samples beyond the second frame 20B. The extended TF block values ​​are then cross-faded in and out with the TF block values ​​of the third frame 20C using a cross-fade-in / fade-out weighted sum. The weight of the extended TF block values ​​in this cross-fade-in / fade-out weighted sum decreases progressively, and the cross-fade-in / fade-out weighting factor corresponds to the prediction parameters. or amplitude parameter The amplitude attenuation. For example, for a linear prediction process, the crossfade-in and fade-out regions of the replacement frame 20C' of the third frame 20C. The TF blocks were determined as follows: (26) in, If a sinusoidal expansion process is used instead of a linear predictor process, then the magnitude evolution parameter is used. replace Therefore, the length of the cross-fade-in / fade-out region will depend on the prediction parameters. or amplitude parameter The amplitude, of which, close to one or This will result in a longer cross-fade-in / fade-out region, and closer to zero. or This will result in a shorter cross-fade-in and fade-out area.

[0198] The idea is to target The first CQMF sample, or until or If the value is lower than the predetermined value, a replacement frame 20B' for the third frame 20C is generated according to Equation 26, and thereafter the TF block value of the unmodified third frame 13 is used.

[0199] In some scenarios, there will be two or more invalid frames in a row, and in order to generate TF block values ​​for replacement frames of two or more invalid frames, two techniques can be used: update extension and continuation extension.

[0200] In the update extension, the first invalid frame will be replaced with a replacement frame using the previous valid frame, as described above. To generate a replacement frame for the second invalid frame following the first invalid frame, the replacement frame of the first invalid frame is treated as a valid frame, and the above process is repeated by generating a replacement frame for the second invalid frame based on the (generated) TF block value of the replacement frame of the first invalid frame.

[0201] It is worth noting that the phase evolution parameters and / or amplitude evolution parameters It is calculated based on the replacement frame of the first invalid frame, and / or the complex prediction coefficients. It is calculated based on the replacement frame of the first invalid frame. This means that, for sinusoidal expansion, the phase evolution parameters of the TF block values ​​used to perform sinusoidal expansion to generate the replacement frame of the second invalid frame are... and / or amplitude evolution parameters This can be used with the TF block value for performing sinusoidal spreading to generate a replacement frame for the first invalid frame. and / or The difference is also true. This applies equally to any phase or amplitude perturbation parameter that can be updated accordingly. Similarly, for linear prediction, one or more complex prediction coefficients are used to generate the TF block values ​​of the replacement frame for the second invalid frame. This will be combined with one or more complex prediction coefficients of the TF block values ​​used to generate the replacement frame for the first invalid frame. different.

[0202] In the continuation spread, the linear prediction or sinusoidal spread process continues from the replacement frame of the first invalid frame to the replacement frame of the second invalid frame without updating the parameters. , and The sinusoidal expansion process or linear prediction process can, in principle, continue indefinitely and generate CQMF samples beyond the frame. More TF block values. Therefore, in the continuation spread, the linear predictor process or the sinusoidal spread process continues to generate more than the replacement frames that constitute the first invalid frame. The TF block values ​​of each CQMF sample, from which the following... One CQMF sample is used as a replacement frame for the second invalid frame. The process can also continue to generate TF block values ​​of CQMF samples that form the third, fourth, and fifth replacement frames corresponding to the invalid frames, where each batch is generated through a sinusoidal spreading process or a linear prediction process. Each TF block value constitutes a new replacement frame. Compared to the update spread, the complex prediction coefficients... or phase evolution parameters and / or amplitude evolution parameters Therefore, the TF block values ​​will be preserved throughout the entire process of generating all replacement frames. Similarly, the same amplitude perturbation parameters and / or phase perturbation parameters can be preserved.

[0203] Figure 9 A schematic block diagram is shown of an example electronic device or architecture 200 (e.g., device 200) suitable for implementing example embodiments of the present disclosure. Architecture 200 includes, but is not limited to, as referenced... Figure 1 a , Figure 1 b and Figure 3 The described server and client devices, systems, modules, and methods are illustrated. Architecture 200 includes a central processing unit (CPU) 201 capable of executing various processes based on a program stored, for example, in read-only memory (ROM) 202 or loaded from, for example, storage unit 208 into random access memory (RAM) 203. CPU 201 may be, for example, an electronic processor 201, which may include one or more processor cores, and in some examples, processor 201 may be multiple processors. RAM 203 also stores data required by CPU 201 when executing various processes. CPU 201, ROM 202, and RAM 203 are interconnected via bus 204. Input / output (I / O) interface 205 is also connected to bus 204.

[0204] The following components are connected to I / O interface 205: input unit 206, which may include a keyboard, mouse, etc.; output unit 207, which may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 208, which may include a hard disk or another suitable storage device; and communication unit 209, which may include a network interface card such as a network card (e.g., wired or wireless).

[0205] In some implementations, the input unit 206 includes one or more microphones located at different locations (depending on the host device), which enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0206] In some implementations, output unit 207 includes a system with a variety of numbers of speakers. Output unit 207 (depending on the capabilities of the host device) can render audio signals in various formats, such as mono, stereo, immersive, binaural, and other suitable formats.

[0207] In some embodiments, communication unit 209 is configured to communicate with other devices (e.g., via a network). Drive 210 is also connected to I / O interface 205 as needed. Removable media 211 (such as a disk, optical disk, magneto-optical disk, flash memory drive, or another suitable removable media) is mounted on drive 210 such that computer programs read therefrom can be installed into storage unit 208 as needed. Those skilled in the art will understand that while apparatus 200 is described as including the components described above, in practice, some of these components may be added, removed, and / or replaced, and all such modifications or alterations fall within the scope of this disclosure.

[0208] According to exemplary embodiments of this disclosure, the processes described above can be implemented as computer software programs or on computer-readable storage media. For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing methods. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 209, and / or installed from removable media 211, such as... Figure 9 As shown.

[0209] Typically, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the unit discussed above can be implemented by control circuitry (e.g., CPU 201 and...). Figure 9The control circuitry can perform the actions described in this disclosure, as executed by other components. Some aspects may be implemented in hardware, while others may be implemented in firmware or software (which may include the control circuitry) that can be executed by a controller, processor, or other computing device(s). Although various aspects of exemplary embodiments of this disclosure are illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that the blocks, apparatuses, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers, other computing devices, or some combination thereof, as non-limiting examples.

[0210] Additionally, the various blocks shown in the flowchart can be viewed as method steps, and / or operations resulting from the operation of computer program code, and / or multiple coupled logic circuit elements configured to perform associated functions(s). For example, embodiments of this disclosure include a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program containing program code configured to perform the methods described above.

[0211] In the context of this disclosure, a machine-readable medium can be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transitory and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0212] Computer program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. This computer program code may be provided to one or more processors of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having control circuitry, such that when executed by one or more processors of the computer or other programmable data processing apparatus, the program code implements the functions / operations specified in the flowcharts and / or block diagrams. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0213] The embodiments described herein perform TF block values ​​to generate one or more replacement frames for replacing one or more frames identified as invalid and / or for replacing one or more valid frames when a previous frame is valid. The described methods and processes can be implemented iteratively, wherein, based on the validity of each frame in the frame sequence, these methods and processes are repeated for each frame in the frame sequence to generate replacement frames, or replacement frames are avoided.

Claims

1. A method for concealment of lost or corrupted frames having audio content in a sequence of frames, the method comprising: obtaining at least one first frame associated with the sequence of frames, the at least one first frame being a frame of the sequence of frames or a generated replacement frame; identifying whether a second frame following the at least one first frame is valid or invalid; when the second frame is identified as invalid, generating a set of time- frequency (TF) tile values for a replacement frame to replace the second frame by: analyzing the at least one first frame to determine a degree of tonality; and comparing the degree of tonality to a threshold to determine whether the degree of tonality exceeds the threshold; when the degree of tonality exceeds the threshold, applying a sinusoidal extension process to generate the set of TF tile values for the replacement frame; and when the degree of tonality does not exceed the threshold, applying a linear prediction process to generate the set of TF tile values for the replacement frame; and outputting the replacement frame for replacement of the second frame.

2. The method of claim 1, further comprising: when the second frame is identified as valid: identifying whether the at least one first frame is valid or invalid; when the at least one first frame is identified as invalid, generating a set of TF tile values for a replacement frame to replace the second frame by: obtaining a set of extended TF tile values for the at least one first frame; generating one or more cross-fade TF tile values by cross-fading the set of TF tile values of the second frame with the set of extended TF tile values of the at least one first frame in a cross-fade region; and outputting the second frame with the one or more cross-fade TF tile values as the replacement frame for replacement of the second frame. obtaining the set of extended TF tile values for the at least one first frame comprises: generating the set of extended TF tile values based on TF tile values of the at least one first frame based on the sinusoidal extension process or the linear prediction process. each frame comprises N consecutive TF tile values associated with N consecutive samples, and wherein the set of extended TF tile values comprises K consecutive TF tile values that span K consecutive samples later in time than the N samples of the first frame. the cross-fade region comprises a predetermined number of earliest-in-time TF tile values in the set of TF tile values of the second frame.

3. The method of claim 2, wherein, generating one or more cross-fade TF tile values comprises, for each TF tile value in the cross-fade region, computing a cross-fade weighting sum of the TF tile value of the second frame and a TF tile value in the set of extended TF tile values of the at least one first frame. ​ 4. The method of claim 3 or claim 4, wherein, ​ 5. The method of any one of claims 2 to 5, wherein, ​ 6. The method of any one of claims 2 to 5, wherein, ​ 7. The method of claim 6, wherein, The cross-fade weighting factor of the cross-fade weighting sum indicates a portion of the cross-fade weighting sum to be made up by a TF bin value of the set of extended TF bin values, and wherein the weighting factors of the TF bin values of successive cross-fades in the cross-fade region decrease with TF bin values of cross-fades later in time.

8. The method of any one of claims 2 to 8, wherein, The replacement frame comprises cross-fade TF bin values in the cross-fade region, and the replacement frame further comprises a tail region comprising TF bin values of the second frame outside the cross-fade region.

9. The method of any one of claims 2 to 8, further comprising: when the degree of tonality of the at least one first frame exceeds the threshold, calculating an extended TF bin value by a sinusoidal extension based on at least two TF bin values of the at least one first frame; and when the degree of tonality of the first frame does not exceed the threshold, calculating an extended TF bin value by the linear prediction process based on at least two TF bin values of the at least one first frame.

10. The method of any one of the preceding claims, wherein, The at least one first frame and the second frame are CQMF frames each comprising a set of CQMF samples, wherein each of the CQMF samples comprises at least one TF bin value, wherein generating the set of TF bin values of the replacement frame further comprises: determining a complex parameter based on at least two TF bin values of different CQMF samples in the at least one first frame; and generating the set of TF bin values of the replacement frame by modifying TF bin values of the at least one first frame with the complex parameter.

11. The method of claim 10, wherein, Each CQMF sample is a sample of a complex low-delay filter bank (CLDFB) representation as specified in 3GPP Immersive Voice and Audio Services (IVAS) codec.

12. The method of claim 10 or claim 11, wherein, Generating the set of TF bin values of the replacement frame comprises: generating a first replacement TF bin value in a first CQMF sample of the replacement frame by modifying a TF bin value of the at least one first frame with the complex parameter.

13. The method of claim 12, wherein, Generating the set of TF bin values of the replacement frame further comprises: generating a second replacement TF bin value in a second CQMF sample of the replacement frame by modifying the generated first replacement TF bin value of the first CQMF sample in the replacement frame with the complex parameter, the second CQMF sample being after or before the first CQMF sample of the replacement frame.

14. The method of claim 10 or claim 11, further comprising: determining a vector, the vector comprising N elements with indices ranging from 0 to N, wherein each of the N elements is proportional to the power of the complex parameter; and wherein generating the set of TF bin values of the replacement frame comprises: multiplying a TF bin value of the first frame with the vector.

15. The method of any one of claims 10 to 14, wherein, The complex parameter comprises one or more of a real part, an imaginary part, or both a real part and an imaginary part.

16. The method of any one of claims 10 to 15, wherein, Applying the sinusoidal extension process to generate the set of TF bin values of the replacement frame further comprises: calculating the complex parameter as a complex evolution parameter parameterizing a complex sinusoid based on the at least two TF bin values of different CQMF samples in the at least one first frame, and generating the set of TF bin values of the replacement frame by modifying TF bin values of the at least one first frame with the complex parameter. wherein generating the set of TF bin values for the replacement frame comprises performing a sinusoidal expansion based on the complex evolution parameter.

17. The method of claim 16, wherein, computing the complex evolution parameter comprises: computing a phase parameter of the complex evolution parameter as an average phase difference between TF bin values of a phase context window of TF bin values of the at least one first frame.

18. The method of claim 17, further comprising: generating the set of TF bin values for the replacement frame with the phase parameter.

19. The method of claim 17, further comprising: computing at least one phase perturbation parameter based on a deviation of a phase of at least one TF bin value in the first frame from the phase parameter; and adjusting the phase parameter using the phase perturbation parameter prior to generating the set of TF bin values for the replacement frame with the phase parameter.

20. The method of any one of claims 16 to 19, wherein, computing the complex evolution parameter comprises: computing a magnitude parameter of the complex evolution parameter as an average magnitude ratio between successive TF bin values in a magnitude context window of TF bin values of the at least one first frame.

21. The method of claim 20, further comprising: generating the set of TF bin values for the replacement frame with the magnitude parameter.

22. The method of claim 20, further comprising: determining a magnitude perturbation parameter based on a deviation of a magnitude of at least one TF bin value in the at least one first frame from the magnitude parameter; and adjusting the magnitude parameter using the magnitude perturbation parameter prior to generating the set of TF bin values for the replacement frame with the magnitude parameter.

23. The method of any one of claims 10 to 22, wherein, applying the linear prediction expansion process to generate the set of TF bin values for the replacement frame further comprises: computing the complex parameter as at least one complex predictor coefficient of a linear predictor configured to predict a TF bin value in a successive CQMF sample in the at least one first frame based on TF bin values in at least one previous CQMF sample of the at least one first frame; and wherein generating the set of TF bin values for the replacement frame comprises performing linear prediction on the set of TF bin values based on the at least one complex predictor coefficient.

24. The method of claim 23, computing the complex parameter as at least one complex predictor coefficient comprises: identifying a first order complex coefficient to multiply with a TF bin value of a previous CQMF sample to predict a TF bin value in a successive CQMF sample in a predictor context window of the at least one first frame, wherein the first order complex coefficient is identified by minimizing a prediction error in the predictor context window.

25. The method of claim 24, computing the complex parameter as at least one complex predictor coefficient further comprises: determining a second order complex predictor coefficient to multiply with a second previous TF bin value of a second previous CQMF sample value to predict the TF bin value in the successive CQMF sample in the predictor context window, wherein the first order predictor coefficient and the second order predictor coefficient are identified by minimizing the prediction error in the predictor context window.

26. The method of any one of claims 23 to 25, wherein, applying the linear prediction process to generate the set of TF bin values for the replacement frame includes: generating a set of predicted TF bin values with the at least one complex predictor coefficient; forming a set of modified TF bin values as a weighted sum of the set of predicted TF bin values and a set of TF bin values of the at least one first frame; and forming the replacement frame based on the set of modified TF bin values.

27. The method of claim 26, wherein, the weighting factors of the weighted sum are based on a magnitude of the at least one complex predictor coefficient.

28. The method of claim 27, wherein, a first weighting factor of a first modified TF bin value is different than a second weighting factor of a second modified TF bin value, such that a ratio between the first weighting factor and the second weighting factor is equal to a magnitude of the at least one complex predictor coefficient.

29. The method of any one of claims 26 to 28, wherein, the replacement frame includes the set of modified TF bin values in a decay region, and the replacement frame further includes a repetition region that includes unmodified TF bin values of the first frame, wherein the decay region is earlier in time than the repetition region.

30. The method of any one of the preceding claims, wherein, analyzing the at least one first frame to determine the degree of tonality includes: evaluating TF bin values of the at least one first frame to identify phase differences between different TF bin values; calculating a phase standard deviation based on the identified phase differences between different TF bin values; and comparing the phase standard deviation to a phase threshold to determine that the at least one first frame is tonal when the phase standard deviation is below the phase threshold, and to determine that the at least one first frame is noise-like when the phase standard deviation is equal to or above the phase threshold.

31. The method of any of the preceding claims, wherein, analyzing the at least one first frame to determine the degree of tonality includes: evaluating TF bin values of the at least one first frame to identify amplitude ratios between different TF bin values; calculating an amplitude standard deviation based on the identified amplitude ratios between different TF bin values; and comparing the amplitude standard deviation to an amplitude threshold to determine that the at least one first frame is tonal when the amplitude standard deviation is below the amplitude threshold, and to determine that the at least one first frame is noise-like when the amplitude standard deviation is equal to or above the amplitude threshold.

32. The method of any of the preceding claims, wherein, each frame includes a plurality of samples, wherein each sample includes a plurality of TF bin values across a plurality of frequency bands, the method further including: when the second frame is identified as invalid, generating a set of TF bin values for the replacement frame for each frequency band of the plurality of frequency bands by: analyzing the at least one first frame to determine a degree of tonality for each frequency; and for each frequency band, comparing the degree of tonality to a threshold to determine whether the degree of tonality exceeds the threshold; in frequency bands where the degree of tonality exceeds the threshold, applying a sinusoidal expansion process to generate the set of TF bin values for the replacement frame; and in frequency bands where the degree of tonality does not exceed the threshold, applying a linear prediction process to generate the set of TF bin values for the replacement frame.

33. The method of any of the preceding claims, wherein, the second frame is identified as invalid if the second frame is unavailable for decoding or corrupted when decoded.

34. An apparatus comprising a processor and a memory, the apparatus configured to perform the method of any of the preceding claims.

35. A non-transitory medium having stored thereon software comprising instructions for controlling one or more devices to perform the method of any of claims 1-33.

36. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of claims 1-33.

37. A method for concealing lost or corrupted frames in a complex quadrature mirror filter (CQMF) domain, the method comprising: receiving at least one first frame of CQMF samples, wherein each CQMF sample spans one or more time-frequency (TF) tiles, and wherein each TF tile is respectively associated with a frequency bin and a complex TF tile value; identifying whether a second frame of CQMF samples, subsequent to the at least one first frame, is valid or invalid; when the second frame is identified as invalid, for at least one frequency bin in a respective frequency bin of the one or more TF tiles of the CQMF samples of the at least one first frame: determining a complex parameter based on at least two of the TF tile values in the at least one first frame; generating a replacement TF tile value by modifying at least one TF tile value of the first frame with the complex parameter; and forming a replacement frame for the second frame based on the replacement TF tile value; and outputting the replacement frame.

38. The method of claim 37, wherein, Each CQMF sample is a sample of a complex low-delay filter bank (CLDFB) representation as specified in the 3GPP Immersive Voice and Audio Services (IVAS) codec.

39. A receiving unit comprising: a decoder configured to receive data packets in a bitstream, wherein each data packet is associated with a current frame comprising a plurality of time-frequency (TF) tile values, wherein the decoder is further configured to decode the received data packets to obtain the TF tile values of the current frame; a frame replacement module configured to identify whether the current frame is invalid, and when the current frame is identified as invalid, generate a set of TF tile values for a replacement frame, wherein the frame replacement module comprises: a tonality extractor configured to analyze at least one previous frame to determine a degree of tonality of the at least one previous frame, wherein the at least one previous frame is earlier in time than the current frame; and an adaptive sinusoidal expansion and linear prediction module configured to generate the set of TF tile values for the replacement frame based on the determined degree of tonality of the at least one previous frame; and a synthesis filter bank configured to receive the TF tile values of the replacement frame and convert the TF tile values of the replacement frame to a time-domain audio segment.

40. The receiving unit of claim 39, wherein, Each frame comprises a plurality of frequency bins, and each frame comprises a plurality of TF tile values associated with each frequency bin, for each frequency: the tonality extractor is further configured to: analyze the at least one previous frame to determine a degree of tonality; determining whether the degree of tonality exceeds the threshold; and The adaptive sinusoidal and linear prediction module is further configured to: generate the set of TF bin values of the replacement frame by a sinusoidal expansion process when the degree of tonality is determined to exceed the threshold; and generate the set of TF bin values of the replacement frame by a linear prediction process when the degree of tonality is determined to not exceed the threshold.

41. The receiving unit of claim 39 or claim 40, wherein, The frame replacement module is further configured to: identify whether the at least one previous frame is valid or invalid when the current frame is identified as valid; and when the at least one first frame is identified as invalid: obtain a set of expanded TF bin values of the at least one previous frame, and generate one or more cross-fade TF bin values of a replacement frame, the one or more cross-fade TF bin values being a cross-fade of the set of TF bin values of the current frame and the set of expanded TF bin values of the at least one previous frame in a cross-fade region of the replacement frame.

42. The receiving unit of any one of claims 39 to 41, wherein, The synthesis filter bank is a CQMF bank, and each frame comprises a plurality of CQMF samples, wherein each sample comprises at least one TF bin value associated with a frequency bin.

43. The receiving unit of claim 42, wherein, The adaptive sinusoidal expansion and linear prediction module is configured to: determine a complex parameter based on at least two TF bin values of different CQMF samples in the at least one previous frame; and modify a TF bin value of the at least one previous frame with the complex parameter to generate the set of TF bin values of the replacement frame by sinusoidal expansion or linear prediction.

44. A system comprising a transmitting unit and a receiving unit according to any of claims 39 to 43, wherein, The sending unit comprises: an analysis filter bank configured to obtain a time-domain audio segment and output time-frequency, TF, bin values representing the time-domain audio segment, and an encoder configured to receive the TF bin values and encode the TF samples into data packets, wherein each data packet comprises a frame containing a plurality of TF bin values; and send the data packets to the receiving unit in the form of a bitstream.