Method and related decoder for frequency-domain packet loss concealment
By filling the MDCT window with time domain signals in the audio codec and overlapping addition, the problem of endpoint unreliability of the IFFT signal is solved, improving the loss hidden quality of the tone signal and the overall performance of the reconstructed signal.
Patent Information
- Application Number
- CN202080015563.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-21
- Filing Date
- 2020-02-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-02-20
AI Technical Summary
In audio codecs, when using shorter Fast Inverse Fourier Transform (IFFT) and low-latency corrected discrete cosine transform (MDCT) windows, the reconstruction signal is less reliable at the endpoints of the IFFT signal, resulting in a lower quality of the reconstruction signal.
Reduce the impact of unreliable endpoints of the IFFT signal by filling the active part of the MDCT window with time domain signals in the PLC prototype buffer and performing overlapping additions (OLA) between the endpoints of the IFFT signal and the PLC prototype buffer.
The loss hidden quality of tone signals is improved, essentially eliminating the possible synthetic noise floor caused by the lack of sufficiently reliable samples of MDCT analysis windowing steps, and is able to provide high-quality reconstructed signals while maintaining efficient low-latency MDCTs.
Smart Images

Figure CN113439302B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to a method for filling an analysis window length to hide lost audio frames associated with a received audio signal. The present disclosure also relates to a decoder configured to fill an analysis window length to hide lost audio frames associated with a received audio signal. Background Art
[0002] Transmission of voice / audio over modern communication channels / networks is mainly carried out in the digital domain using voice / audio codecs. This may involve obtaining an analog signal and digitizing it using a sampler and an analog-to-digital converter (ADC) to obtain digital samples. These digital samples can be further grouped into frames, which, depending on the application, contain samples for 10–40 ms consecutive periods. These frames can then be processed using a compression algorithm, which reduces the number of bits to be transmitted and still enables the highest possible quality to be achieved. The encoded bitstream is then transmitted as data packets over a digital network to a receiver. In the receiver, the process is reversed. The data packets can first be decoded to recreate frames with digital samples, and then the frames can be input into a digital-to-analog converter (DAC) to recreate an approximation of the input analog signal at the receiver. Figure 1 An example block diagram of audio transmission using an audio encoder and a decoder over a network such as a digital network using the above method is provided.
[0003] When data packets are transmitted over a network, there may be cases where data packets are dropped by the network due to traffic load or due to bit errors that render the digital data undecodable. When these events occur, the decoder needs to replace the output signal during the period when actual decoding cannot be performed. This replacement process is generally referred to as frame / packet loss concealment (PLC). Figure 2 A block diagram of a decoder 200 including packet loss concealment is shown. When a bad frame indicator (BFI) indicates a lost or corrupted frame, the PLC 202 can create a signal to replace the lost / corrupted frame. Otherwise, i.e., when the BFI does not indicate a lost or corrupted frame, the received signal is decoded by the stream decoder 204. A frame erasure signal can be sent to the decoder by setting the bad frame indicator variable of the current frame to active (i.e., BFI = 1). Then, the decoded or concealed frame is input into the DAC 206 to output an analog signal. Frame / packet loss concealment can also be referred to as an error concealment unit (ECU).
[0004] There are many ways to perform packet loss concealment in a decoder. Some examples are replacing the lost frames with silence and repeating the last frame (or decoding the parameters of the last frame). Other methods attempt to replace the frames with the continuation of the most likely audio signal. For noise-like signals, one method may produce noise with a similar spectral structure. For tonal signals, the characteristics of the current tone (frequency, amplitude, and phase) can be estimated first and these parameters can be used to generate the continuation of the tone at the corresponding time positions of the lost frames.
[0005] Another method for the ECU is the Phase ECU, which was initially described in International Patent Application No. WO2014123470 and later in Section 5.4.3.5 of 3GPP TS26.447 V15.0.0, where the decoder can continuously save the prototype of the decoded signal during normal decoding. This prototype can be used in case of lost frames. A spectral analysis of the prototype is performed and the noise and tonal ECU functions are combined in the spectral domain. The Phase ECU identifies the tones in the spectrum and calculates the spectral time replacement of the relevant spectral bins. The other bins (non-tonal) can be treated as noise and scrambled to avoid tonal artifacts in these spectral regions. The generated reconstructed spectrum is transformed to the time domain by an inverse FFT (Fast Fourier Transform) (IFFT) and the signal is processed to create a replacement for the lost frame. When the audio codec is based on the modified discrete cosine transform (MDCT), the creation of the replacement includes windowing, TDA (time domain aliasing), and ITDA (inverse TDA) related to the overlapping MDCT to create an integrated continuation of the decoded signal. This method can ensure the continued use of the MDCT memory and create the MDCT memory that will be used for the first good frame. Summary of the Invention
[0006] For a PLC using sinusoidal modeling in the frequency domain, the reconstructed signal may be less reliable at the endpoints of the IFFT signal. This part may be masked by the window shape of the MDCT analysis window, especially when it is symmetric about leading and trailing zeros. Leading zeros are located on future samples and thus reduce the algorithmic delay of the encoder and can therefore be used frequently. Trailing zeros can be mainly used to make the window simpler, but this can reduce the transform efficiency as they increase complexity without adding any information to the input signal. Therefore, fewer trailing zeros are typically used. Compared to a larger-delay symmetric MDCT window with a similar frequency resolution, a low-delay asymmetric MDCT window can have a sharper (rapidly increasing / rapidly decreasing) window shape. The combination of the sharper window shape and the lower reliability of the IFFT signal can include more of these unreliable parts into the MDCT analysis windowing and subsequent TDA and ITDA steps, which are used to create the final reconstructed signal including the updated MDCT memory (also known as the MDCT OLA (overlap-add) memory buffer). This may result in a lower quality of the reconstructed signal. While one approach could be to increase the length of the prototype, this may be undesirable as it may significantly increase the complexity of the PLC. Another approach could be to use a shorter MDCT window in the actual audio codec, but this may result in a poorer frequency resolution (and performance) of the audio codec.
[0007] The use of a relatively short inverse fast Fourier transform (IFFT) (e.g., having a length corresponding to 16 ms, providing approximately 12 ms of reliable evolving samples in the time domain) in combination with low-delay MDCT analysis and synthesis steps compared to a “windowing -> TDA -> ITDA” input that is much larger (e.g., 18 ms) than the 12 ms of time-domain samples provided is a challenging problem to solve. Using a larger IFFT increases complexity (e.g., the complexity of the PLC), while using a smaller low-delay MDCT (LD-MDCT) window reduces the spectral resolution of the codec, which in turn reduces the compression efficiency.
[0008] According to some embodiments of the inventive concept, there is provided a method of operating a decoder to fill an analysis window length with a time-domain signal to hide lost audio frames associated with a received audio signal. In this method, a first segment of a previously received portion of the received audio signal is copied from a prototype buffer. A second segment of the previously received portion of the received audio signal is overlap-added from the prototype buffer to an initial portion of a reconstructed portion of the received audio signal, followed by the remaining portion of the reconstructed portion of the received audio signal.
[0009] According to other embodiments of the inventive concept, there is provided a decoder that fills an analysis window length with a time-domain signal to hide lost audio frames associated with a received audio signal. The decoder includes a processor and a memory coupled to the processor, wherein the memory includes instructions that, when executed by the processor, cause the decoder to perform operations. The decoder performs operations including copying a first segment of a previously received portion of the received audio signal from a prototype buffer. The decoder performs further operations including overlap-adding a second segment of the previously received portion of the received audio signal from the prototype buffer to an initial portion of a reconstructed portion of the received audio signal, followed by a remaining portion of the reconstructed portion of the received audio signal.
[0010] According to other embodiments of the inventive concept, there is provided a decoder that fills an analysis window length with a time-domain signal to hide lost audio frames associated with a received audio signal, wherein operations suitable for execution by the decoder include: copying a first segment of a previously received portion of the received audio signal from a prototype buffer and overlap-adding a second segment of the previously received portion of the received audio signal from the prototype buffer to an initial portion of a reconstructed portion of the received audio signal, followed by a remaining portion of the reconstructed portion of the received audio signal.
[0011] According to additional embodiments, there is provided a computer program including program code to be executed by at least one processor of a decoder, the program code for filling an analysis window length with a time-domain signal to hide lost audio frames associated with a received audio signal. Execution of the program code causes the decoder to perform operations including: copying a first segment of a previously received portion of the received audio signal from a prototype buffer and overlap-adding a second segment of the previously received portion of the received audio signal from the prototype buffer to an initial portion of a reconstructed portion of the received audio signal, followed by a remaining portion of the reconstructed portion of the received audio signal.
[0012] According to other embodiments, there is provided a computer program product including a non-transitory storage medium including program code to be executed by at least one processor (1306) of a decoder (1300), the program code for filling an analysis window length with a time-domain signal to hide lost audio frames associated with a received audio signal. Execution of the program code causes the decoder to perform operations including: copying a first segment of a previously received portion of the received audio signal from a prototype buffer and overlap-adding a second segment of the previously received portion of the received audio signal from the prototype buffer to an initial portion of a reconstructed portion of the received audio signal, followed by a remaining portion of the reconstructed portion of the received audio signal.
[0013] According to some embodiments, filling the active part of the MDCT window with signals already available in the PLC prototype buffer and the OLA between the end points of the IFFT and the PLC prototype buffer can reduce the impact of the potentially unreliable end points of the IFFT signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings illustrate certain non - limiting embodiments of the inventive concept, which are included to provide a further understanding of the present disclosure and are incorporated into and form a part of this application. In the drawings:
[0015] Figure 1 A block diagram showing the use of an audio encoder and an audio decoder over a network is shown;
[0016] Figure 2 A block diagram showing a decoder including packet loss concealment is shown;
[0017] Figure 3 A time alignment and signal diagram of a phase ECU and the reconstruction of a signal from a PLC prototype are shown;
[0018] Figure 4 is a flowchart of a PLC signal reconstruction operation;
[0019] Figure 5 is a flowchart for generating a time PLC signal according to some embodiments of the inventive concept;
[0020] Figure 6 is a signal diagram of PLC processing using a PLC prototype according to some embodiments of the inventive concept;
[0021] Figure 7 Two windows for OLA summation according to some embodiments of the inventive concept are shown;
[0022] Figure 8 A phase ECU buffer completion copy length table according to some embodiments of the inventive concept is shown;
[0023] Figure 9 A phase ECU buffer completion OLA length table according to some embodiments of the inventive concept is shown;
[0024] Figures 10 - 12 is a flowchart showing the operation of a decoder according to some embodiments of the inventive concept; and
[0025] Figures 13 - 15 is a block diagram of a decoder according to some embodiments of the inventive concept. DETAILED DESCRIPTION
[0026] In the following, the inventive concept will be described more fully with reference to the accompanying drawings, in which examples of embodiments of the inventive concept are shown. However, the inventive concept may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the inventive concept to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Components from one embodiment may be assumed to be present in / used in another embodiment by default.
[0027] The following description presents various embodiments of the disclosed subject matter. These embodiments are presented as teaching examples and are not to be construed as limiting the scope of the disclosed subject matter. For example, certain details of the embodiments may be modified, omitted, or extended without departing from the scope of the subject matter.
[0028] With a PLC using sine modeling in the frequency domain, the reconstructed signal may not be very reliable at the endpoints of the IFFT signal. This part may be masked by the window shape of the MDCT analysis window, especially when it is symmetric about leading zeros and trailing zeros. Leading zeros are located on future samples and thus reduce the algorithmic delay of the encoder and can therefore be used frequently. Trailing zeros can be mainly used to make the window simpler, but this can reduce the transform efficiency because they increase complexity without adding any information to the input signal. Therefore, fewer trailing zeros are usually used.
[0029] Compared with a larger-delay symmetric MDCT window with a similar frequency resolution, a low-delay asymmetric MDCT window may have a sharper (rapidly increasing / rapidly decreasing) window shape. The combination of the sharper window shape and the lower reliability of the IFFT signal may include more of these unreliable parts into the MDCT analysis windowing and subsequent TDA and ITDA steps, which are used to create the final reconstructed signal including an updated MDCT memory (also known as the MDCT OLA memory buffer). This may result in a lower quality of the reconstructed signal. While one approach could be to increase the length of the prototype, this may be undesirable as it may significantly increase the complexity of the PLC. Another approach could be to use a shorter MDCT window in an actual audio codec, but this may lead to a poorer frequency resolution (and poorer performance) of the audio codec.
[0030] The use of a combination of a relatively short inverse fast Fourier transform (IFFT) (e.g., having a length corresponding to 16 ms, providing approximately 12 ms of reliable evolution samples in the time domain) and low-latency MDCT analysis and synthesis steps with an input of "windowing -> TDA -> ITDA" that is much larger (e.g., 18 ms) than the 12 ms time-domain samples provided is a challenging problem to solve. Using a larger IFFT increases complexity (e.g., the complexity of PLC), while using a smaller LD-MDCT window reduces the spectral resolution of the codec, which in turn reduces the compression efficiency, as Figure 3 shown. Figure 3 Shows the time alignment and signal diagram of the phase ECU and the reconstruction of the signal from the PLC prototype using the phase evolution and new time positions (phases) of the prototype. Final reconstruction is performed using MDCT analysis windowing, TDA, and ITDA steps, as well as MDCT synthesis windowing, where the last synthesis windowing step also uses the previously stored MDCT transform memory (MDCT-OLA buffer) and creates a new one (MDCT-OLA signal buffer) for use during the next frame (good or bad). Figure 3 An asymmetric MDCT window with look-ahead zeros (3 / 8 of the LA_ZEROS-frame length) is used in
[0031] Some embodiments disclosed by the inventive concept can fill the active part of the MDCT window with signals already available in the PLC prototype buffer and can use OLA between the end points of the IFFT and the PLC prototype buffer to reduce the impact of the possibly unreliable end point part of the IFFT signal.
[0032] Some embodiments disclosed by the inventive concept can use the prototype time-domain signal. With PLC using sinusoidal modeling, the decoder holds the prototype time-domain signal in the PLC prototype buffer in the form of the last decoded signal. That is, an alternative frame for the lost frame is calculated by applying a sinusoidal model to the frames of the previously synthesized good frame signals, where the frame is used as the prototype frame. A good frame refers to a correctly received frame, while a bad frame refers to an erased frame, i.e., a lost or damaged frame.
[0033] Turning to Figure 4, lost or damaged packets can be recognized by the transport layer handling the connection and sent to the decoder as "bad frames" via a Bad Frame Indicator (BFI). When the Bad Frame Indicator (BFI) indicates that a lost or damaged frame has occurred, the PLC using sinusoidal modeling in block 401 performs a spectral analysis of the last synthesized good frame signal (i.e., the prototype frame) and identifies the peaks of the amplitude signal. Then, the fractional frequency of the peak is estimated using the peak frequency bin. In block 403, a phase shift is performed on the peak frequency bin corresponding to the peak and the neighbors of the peak frequency bin. For the remaining frequency bins of the frame, the previously synthesized amplitude is retained and the phase can be randomized. In block 405, an IFFT is used to create a time-domain signal. This is followed by MDCT windowing, TDA, and ITDA, as shown in block 407.
[0034] Some embodiments disclosed by the inventive concept can use a PLC prototype buffer to create a high-quality approximation signal to fill the full MDCT analysis window length. This may involve two operations; one operation can copy a segment of the PLC prototype buffer into the MDCT processing buffer. Another operation can overlap and add the remaining last part of the prototype buffer to the MDCT processing buffer with the IFFT signal of the initial time evolution, as Figure 4 shown. This is followed by MDCT windowing, TDA, and ITDA, as Figure 4 shown in block 407 of
[0035] In the operations discussed in section 6.2.4.1 of 3GPP TS 26.447 V15.1.0, the time-domain prototype signal from the previous frame is only used during the loss of the first frame. For continuous sinusoidal modeling, the spectral representation of the prototype is saved and used for consecutive lost frames. However, in order to be able to perform the copy and OLA operations of some embodiments of the present disclosure, the relevant part of the prototype buffer (or a separate time continuity buffer) may need to be continuously updated during consecutive frame errors (i.e., during bad frames). Since the spectral representation is saved, it is not necessary to update the full prototype buffer during the processing of consecutive lost frames.
[0036] In some embodiments, the quality can be improved when the method for packet hiding is sinusoidal modeling in the frequency domain. By ensuring that the reconstructed frame is well integrated with the current decoded signal, even when the frame reconstruction step after the IFFT of the phase ECU only provides a limited frame (i.e., a limited number of appropriately time-evolved samples) and when the MDCT window is asymmetric with respect to leading and trailing zeros.
[0037] In various embodiments of the inventive concept disclosed herein, unreliable endpoints of the IFFT signal can be partially replaced, and an asymmetric MDCT window can be used for windowing, TDA, and ITDA steps. This can reduce frame repetition discontinuities that would otherwise be introduced. The inventive concept disclosed herein improves the quality of PLC of tonal signals and can substantially eliminate the synthetic background noise that might otherwise be generated due to the lack of an MDCT analysis windowing step with sufficiently reliable samples.
[0038] In various embodiments of the inventive concept disclosed herein, even though the MDCT analysis / synthesis window requires more than 12 ms of reliable signal to provide high-quality analysis and synthesis reconstruction, as well as an interface to the core audio codec (via the MDCT OLA memory buffer), a rather short IFFT (e.g., 16 ms, yielding approximately 12 ms of reliable time-evolving samples) can be used in combination with an efficient low-latency MDCT.
[0039] In some embodiments of the inventive concept, the MDCT can be performed 10 ms in advance over a 20 ms window. The length of the PLC prototype frame saved after a good frame can be 16 ms. The associated transient detector can use two short FFTs of length 4 ms - i.e., a quarter of the PLC prototype frame. The actual lengths of these items depend on the sampling frequency used and can be from 8 kHz to 48 kHz. These lengths affect the number of spectral bins in each transform.
[0040] In some embodiments disclosed herein, an evolved and time-corrected signal x ph (n) can be created according to the core phase ECU method, see, for example, the core phase ECU method discussed in section 5.4.3.5 of 3GPP TS 26.477 V15.0.0.
[0041] In some embodiments disclosed herein, then, the signal x ph (n) can be extended in both directions to x ph_ext (n) to achieve the same length as a normal MDCT window, thereby creating a smooth transition from the last decoded frame to the newly reconstructed signal before windowing. The left (oldest) part may require two steps, a segment copy and a segment overlap and add. In the first step, a portion (time-domain samples) can be copied from the prototype buffer, corresponding to the synthetic part before the evolved and reconstructed signal. In the second step, a segment can be overlap-added between the final part from the prototype buffer and the initial part of the reconstructed signal x ph (k). The part after the reconstructed signal can be zero-extended. Then the signal x ph_ext(n) Windowing (MDCT analysis window) and time-domain aliasing are performed. The leading zero samples in the asymmetric MDCT window are referred to as LA_ZEROS. Then, the resulting windowed and time-domain aliased signal can be processed using the memory / state of the MDCT from the previous frame described in a conventional MDCT decoder for overlap-and-add (OLA).
[0042] Some embodiments of the inventive concept are presented on the window for OLA used in a PLC prototype and a non-limiting example of the trailing reconstruction PLC signal as a Hann window, while the embodiments of the present disclosure can also be applied to other types of windows, such as, for example, a Hamming window, a Kaiser window, etc.
[0043] In some embodiments, OLA can first apply a per-sample window to achieve the fade-out of a portion from the PLC prototype buffer. The window scales each sample in the buffer according to the following function:
[0044]
[0045] where L ola is the length of the OLA segment. A second window can be used to achieve the fade-in of the initial (oldest) portion of the reconstructed IFFT time-domain signal, where the scaling of each sample is defined as:
[0046]
[0047] The windowed and scaled samples from the PLC prototype and the IFFT tail can be added together and form a new estimate for the OLA period. Two windows (w old and w new ) are preferably constructed such that the sum at any point n is 1.0.
[0048] Figure 8 Shows a phase ECU buffer completion copy length table according to some embodiments of the inventive concept. Figure 8 Shows that the length of the copy portion (Lcopy) depends on the sampling frequency. Figure 9 Shows the length of the Hamming portion depending on the sampling frequency f s according to some embodiments of the inventive concept.
[0049] Some embodiments of the present disclosure can provide a smooth and nearly noise-free synthesized signal for a steady-state sine wave.
[0050] In some embodiments of the present disclosure, the lengths of the COPY segment and the OLA segment can be dynamically adjusted based on the analysis of past synthesized signals and somewhat unreliable phase evolution IFFT signals.
[0051] For example, corresponding to the oldest samples from an unreliable IFFT signal, in cases where the default lengths (2.0 ms, 1.75 ms) result in strong transients in the 0.75 ms region, the copy portion can be extended to 2.75 ms while the OLA portion is reduced to 1 ms. The transient detector can compare the RMS (root mean square) value of the first 2 ms with the RMS of the 0.75 ms adjustment region (with two length sets (COPY, OLA)) and can use the length set that provides the lowest RMS energy difference.
[0052] In some embodiments of the present disclosure, complexity can be reduced by preprocessing the copy segment with an MDCT analysis window in a GOOD non-PLC frame. The same operation can be performed on the copy segment and the OLA segment using a combined OLA_old window and MDCT window. This reallocation of windowing complexity can allow for more complexity to be used in BAD frames.
[0053] Figure 5 is a flowchart for generating a time PLC signal according to some embodiments of the inventive concept. Blocks 401, 403, and 405 are described above in the Figure 4 description. At block 501, a complete time PLC signal of full MDCT length can be generated using copy and OLA from a prototype signal. At block 503, operations including performing an MDCT analysis window, TDA, ITDA, and MDCT synthesis and MDCT-OLA can be performed to reconstruct the PLC signal. Figure 5 The generation of the time signal of
[0054] Figure 6 is a signal diagram of PLC processing according to some embodiments of the inventive concept. Window 601 shows using a PLC prototype 605 to fill an MDCT frame and performing overlap-add (OLA) 607 between the PLC prototype 605 and the possibly unreliable initial endpoints of the IFFT reconstruction signal 609. The other endpoint of 609 overlaps with the look-ahead zeros of the analysis MDCT window 601.
[0055] Figure 7 shows how two windows 701 (old) and 702 (new) can be used for OLA to facilitate summing two windowed signal portions 703.
[0056] Now, with reference to the Figures 10 - 12 flowchart, the operation of the decoder 1300 (see Figures 13 - 15 ) will be discussed according to some embodiments of the inventive concept. For example, modules can be stored in Figures 13 - 15in the memory 1308, and these modules can provide instructions such that when the instructions of the modules are executed by the processor 1306, the processor 1306 performs the corresponding operations of the corresponding flowchart.
[0057] Figure 10 illustrates the operation of a decoder that fills an analysis window length with a time-domain signal to hide lost audio frames associated with a received audio signal. At Figure 10 block 1000, the processor 1306 of the decoder 1300 copies the first segment of the previously received portion of the received audio signal from the prototype buffer 1318 to the processing buffer 1320. At block 1002, the processor 1306 of the decoder 1300 overlaps and adds the second segment of the previously received portion of the received audio signal from the prototype buffer 1318 to the initial portion of the reconstructed portion of the received audio signal into the processing buffer 1320, followed by the remaining portion of the reconstructed portion of the received audio signal.
[0058] Figure 11 illustrates further operations of a decoder that fills an analysis window length with a time-domain signal to hide lost audio frames associated with a received audio signal when there are consecutive lost frames.
[0059] At Figure 11 block 1100, the processor 1306 of the decoder 1300 copies the first segment of the previously received portion of the received audio signal from the time continuity buffer 1316. At block 1102, the processor 1306 of the decoder 1300 overlaps and adds the second segment of the previously received portion of the received audio signal from the time continuity buffer 1316 to the initial portion of the reconstructed portion of the received audio signal into the processing buffer 1320, followed by the remaining portion of the reconstructed portion of the received audio signal.
[0060] The prototype buffer 1318 can be updated with newly decoded signals, and the time continuity buffer 1316 can be updated with newly reconstructed signals after MDCT OLA.
[0061] Figure 12 illustrates further operations of a decoder that fills an analysis window length with a dynamically adjusted signal segment length. At Figure 12 block 1200, the processor 1306 of the decoder 1300 can dynamically adjust the lengths of the first and second segments based on an analysis of the previously synthesized time-domain signal from the filled analysis window length.
[0062] Referring to some embodiments of a decoder and related methods, from the process Figure 11 and 12The various operations can be optional. For example, the blocks 1100, 1102, and 1200 can be optional.
[0063] The various embodiments described above are applied to a controller in a decoder, such as Figures 13 - 15 shown. Figure 13 FIG. is a schematic block diagram of a decoder according to some embodiments. The decoder 1300 includes an input unit 1302 configured to receive an encoded audio signal. Figure 13 FIG. shows frame loss concealment performed by a logical frame loss concealment unit 1304 according to various embodiments described herein, the logical frame loss concealment unit 1304 indicating that the decoder is configured to implement concealment of lost audio frames. Further, the decoder 1300 includes a controller 1306 (also referred to herein as a processor or a processor circuit) for implementing the various embodiments described herein. The controller 1306 is coupled to an input terminal (IN) and a memory 1308 (also referred to herein as a memory circuit), and the memory 1308 is coupled to the processor 1306. The decoded and reconstructed audio signal obtained from the processor 1306 is output from an output terminal (OUT). The memory 1308 may include computer-readable program code 1310 which, when executed by the processor 1306, causes the processor to perform operations according to the embodiments disclosed herein. According to other embodiments, the processor 1306 may be defined to include a memory such that a separate memory is not required.
[0064] The controller 1306 is configured to fill an analysis window length with a time-domain signal to conceal lost audio frames associated with the received audio signal. The controller 1306 may copy a first segment of a previously received portion of the received audio signal from a prototype buffer to a processing buffer. The controller 1306 may overlap and add (OLA) a second segment of the previously received portion of the received audio signal from the prototype buffer to an initial portion of the reconstructed portion of the received audio signal into the processing buffer, followed by the remaining portion of the reconstructed portion of the received audio signal. As Figure 15 shown, the copying may be performed by a copy unit 1312, and the OLA may be performed by an OLA unit 1314. The signals to be processed by the processor 1306 (including the copy unit 1312 and the OLA unit 1314) may be provided from the memory 1308, including from a temporal continuity buffer 1316, a prototype buffer 1318, and a processing buffer 1320, as Figure 15 shown.
[0065] The decoder with its copy unit and OLA unit can be implemented in hardware. The functions of the units of the decoder can be implemented using and combining various variants of circuit elements. Such variants are included in various embodiments. A specific example of the hardware implementation of the decoder is implemented in digital signal processor (DSP) hardware and integrated circuit technology, including general-purpose electronic circuits and application-specific circuits.
[0066] Abbreviations
[0067] At least some of the following abbreviations can be used in this disclosure. If there is an inconsistency between abbreviations, precedence should be given to how it is used above. If listed multiple times below, the first listing should take precedence over any subsequent listings.
[0068] Abbreviation Explanation
[0069] ADC Analog-to-Digital Converter
[0070] BFI Bad Frame Indicator
[0071] DAC Digital-to-Analog Converter
[0072] FFT Fast Fourier Transform
[0073] IFFT Inverse Fast Fourier Transform
[0074] ITDA Inverse Time-Domain Aliasing
[0075] LA_ZEROS Look-Ahead Zeros
[0076] MDCT Modified Discrete Cosine Transform
[0077] OLA Overlap and Add
[0078] TDA Time-Domain Aliasing
[0079] References:
[0080] 3GPP TS 26.447 V15.0.0, Section 5.4.3.5
[0081] 3GPP TS 26.445 V15.1.0, Section 5.3.2.2
[0082] 3GPP TS 26.445 V15.1.0, Section 6.2.4.1
[0083] References 3GPP TS 26.447 V15.0.0, clause 5.4.3.5, 3GPP TS 26.445 V15.1.0, clause 5.3.2.2, and 3GPP TS 26.445 V15.1.0, clause 6.2.4.1 are hereby incorporated herein by reference as if their contents were fully set forth herein.
[0084] List of exemplary embodiments of the inventive concept:
[0085] Exemplary embodiments are discussed below. Reference numerals / letters are provided in parentheses by way of example / illustration without limiting the exemplary embodiments to the specific elements indicated by the reference numerals / letters.
[0086] 1. A method of filling an analysis window length with a time-domain signal to hide lost audio frames of a received audio signal, the method comprising:
[0087] copying (1000, 1312) a first segment of a previously received portion of the received audio signal from a prototype buffer (1318) to a processing buffer (1320); and
[0088] overlap-adding (1002, 1314) a second segment of a previously received portion of the received audio signal from the prototype buffer (1318) to an initial portion of a reconstructed portion of the received audio signal, followed by a remaining portion of the reconstructed portion of the received audio signal, to the processing buffer (1320).
[0089] 2. The method according to embodiment 1, wherein the previously received portion of the received audio signal is a time-domain signal.
[0090] 3. The method according to any one of embodiments 1 to 2, wherein the reconstructed portion of the received audio signal includes a time-evolution transform signal.
[0091] 4. The method according to any one of embodiments 1 to 3, wherein the processing buffer is a folded-transform modified discrete cosine transform (MDCT) buffer.
[0092] 5. The method according to any one of embodiments 1 to 4, wherein the MDCT analysis window is asymmetric.
[0093] 6. The method according to any one of embodiments 1 to 5, further comprising copy and overlap-add for consecutive lost frames, comprising:
[0094] copying (1100, 1312) a first segment of a previously received portion of the received audio signal from a time-continuity buffer (1316); and
[0095] Overlap-add (1002, 1314) a second segment of a previously received portion of the received audio signal from the time continuity buffer (1316) to an initial portion of the reconstructed portion of the received audio signal to the processing buffer (1320), followed by the remaining portion of the reconstructed portion of the received audio signal.
[0096] 7. The method according to any one of embodiments 1 to 6, wherein the time continuity buffer is updated with the newly reconstructed signal after the MDCT overlap-addition.
[0097] 8. The method according to any one of embodiments 1 to 7, wherein the overlapping addition portion of the analysis window comprises:
[0098] Applying a first window (701, 1314) to obtain first scaled samples of a previously received portion of the received audio signal from the prototype buffer;
[0099] Applying a second window (702, 1314) to obtain second scaled samples of the reconstructed portion of the received audio signal;
[0100] Summing (703, 1314) the first scaled samples and the second scaled samples to form the overlapping addition portion of the analysis window.
[0101] 9. The method according to any one of embodiments 1 to 8, wherein the length of the first segment depends on the sampling frequency.
[0102] 10. The method according to any one of embodiments 1 to 9, wherein the length of the overlapping addition portion of the analysis window depends on the sampling frequency.
[0103] 11. The method according to any one of embodiments 1 to 10, further comprising:
[0104] Dynamically adjusting (1200, 1306) the lengths of the first and second segments based on an analysis of a previously synthesized time-domain signal from a filled analysis window length.
[0105] 12. A decoder (1300) for filling an analysis window length with a time-domain signal to hide lost audio frames of a received audio signal, the decoder comprising:
[0106] A processor (1306); and
[0107] A memory (1308) coupled to the processor, wherein the memory comprises instructions that, when executed by the processor, cause the decoder to perform the operations according to any one of embodiments 1 - 11.
[0108] 13. A decoder (1300) that fills an analysis window length with a time-domain signal to hide lost audio frames of a received audio signal, wherein the decoder is adapted to perform the operations according to any one of Embodiments 1-11.
[0109] 14. A computer program comprising program code to be executed by at least one processor (1306) of a decoder (1300), the program code for filling an analysis window length with a time-domain signal to hide lost audio frames of a received audio signal, whereby execution of the program code causes the decoder (1300) to perform the operations according to any one of Embodiments 1-11.
[0110] 15. A computer program product comprising a non-transitory storage medium including program code to be executed by at least one processor (1306) of a decoder (1300), the program code for filling an analysis window length with a time-domain signal to hide lost audio frames of a received audio signal, whereby execution of the program code causes the decoder (1300) to perform the operations according to any one of Embodiments 1-11.
[0111] Additional Notes
[0112] In general, unless explicitly given and / or implied from the context to have a different meaning, all terms used herein will be interpreted according to their ordinary meaning in the relevant technical field. All references to "a / an / element, device, component, apparatus, step, etc." shall be construed openly to refer to at least one instance of the element, device, component, apparatus, step, etc., unless otherwise explicitly stated. The steps of any method disclosed herein need not be performed in the exact order disclosed, unless a step must explicitly be described as after or before another step and / or implicitly a step must be after or before another step. Where appropriate, any feature of any embodiment disclosed herein may be applied to any other embodiment. Similarly, any advantage of any embodiment may apply to any other embodiment, and vice versa. Other objects, features, and advantages of the appended embodiments will be apparent from the following description.
[0113] Any suitable steps, methods, features, functions or benefits disclosed herein may be performed by one or more functional units or modules of one or more virtual devices. Each virtual device may include a plurality of such functional units. These functional units may be implemented by processing circuitry, which may include one or more microprocessors or microcontrollers and other digital hardware (which may include a digital signal processor (DSP), dedicated digital logic, etc.). The processing circuitry may be configured to execute program code stored in a memory, which may include one or several types of memory, such as read-only memory (ROM), random access memory (RAM), cache memory, flash memory devices, optical storage devices, etc. The program code stored in the memory includes program instructions for executing one or more telecommunication and / or data communication protocols, as well as instructions for executing one or more of the techniques described herein. In some implementations, the processing circuitry may be used to cause the corresponding functional unit to perform the corresponding function in accordance with one or one embodiment of the present disclosure.
[0114] The term unit may have its conventional meaning in the field of electronics, electrical equipment and / or electronic devices, and may include, for example, electrical and / or electronic circuits, devices, modules, processors, memories, logical solid-state and / or discrete devices, computer programs or instructions for performing corresponding tasks, processes, calculations, outputs and / or display functions, etc., such as those described herein.
[0115] In the above description of the various embodiments of the inventive concept, it is to be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the inventive concept. Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the inventive concept pertains. It should also be understood that terms such as those defined in a general dictionary should be interpreted as having a meaning consistent with their meaning in the context of this specification and the relevant art, and not as having an ideal or overly literal meaning, unless expressly so defined herein.
[0116] When an element is referred to as being "connected", "coupled", "responsive" or a variant thereof with respect to another element, it can be directly connected, coupled to or responsive to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected", "directly coupled", "directly responsive" or a variant thereof with respect to another element, no intervening element is present. Throughout the specification, like reference numerals refer to like elements. Further, as used herein, "coupled", "connected", "responsive" or a variant thereof may include wireless coupling, connection or response. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. For brevity and / or clarity, well-known functions or constructions may not be described in detail. The term "and / or" includes any and all combinations of one or more of the associated listed items.
[0117] It will be understood that although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments without departing from the teachings of the inventive concept. Throughout the specification, the same reference numerals or the same reference signs denote the same or similar elements.
[0118] The terms "comprising", "including", "containing", "covering", "consisting of", "counting", "having", "owning", "possessing" or variations thereof as used herein are open-ended and include one or more of the recited features, integers, elements, steps, components, or functions, but do not preclude the presence or addition of one or more other features, integers, elements, steps, components, functions or combinations thereof. Further, as used herein, the common abbreviation "e.g." is derived from the Latin phrase "exempli gratia" and may be used to introduce or specify one or more general examples of the item(s) previously mentioned, and is not intended to be limiting of that item. The common abbreviation "i.e." is derived from the Latin phrase "id est" and may be used to specify a more specific item of a more general recitation.
[0119] This document describes example embodiments with reference to block diagrams and / or flowchart illustrations of computer-implemented methods, apparatuses (systems and / or devices), and / or computer program products. It should be understood that the blocks of the block diagrams and / or flowchart illustrations, and combinations of the blocks in the block diagrams and / or flowchart illustrations, can be implemented by computer program instructions executed by one or more computer circuits. These computer program instructions can be provided to a processor circuit of a general-purpose computer circuit, a special-purpose computer circuit, and / or other programmable data processing circuits to produce a machine, such that the instructions executed by the processor of the computer and / or other programmable data processing devices transform and control transistors, values stored in memory locations, and other hardware components within such circuits to implement the functions / actions specified in the block diagrams and / or flowchart blocks, and thereby create means (functional entities) and / or structures for implementing the functions / actions specified in the block diagrams and / or flowchart blocks.
[0120] These computer program instructions can also be stored in a tangible computer-readable medium, which can direct a computer or other programmable data processing device to act in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instructions for implementing the functions / actions specified in the blocks of the block diagrams and / or flowchart. Thus, embodiments of the inventive concept can be implemented in hardware and / or in software (including firmware, stored software, microcode, etc.) running on a processor such as a digital signal processor, which can be collectively referred to as "circuits", "modules", or variants thereof.
[0121] It should also be noted that in some alternative implementations, the functions / actions marked in the blocks may not occur in the order marked in the flowchart. For example, depending on the functions / actions involved, two consecutive blocks shown may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order. In addition, the functions of a given block of the flowchart and / or block diagram can be divided into multiple blocks, and / or the functions of two or more blocks of the flowchart and / or block diagram can be at least partially integrated. Finally, other blocks can be added / inserted between the blocks shown, and / or blocks / operations can be omitted without departing from the scope of the inventive concept. In addition, although some blocks include arrows regarding communication paths indicating the main direction of communication, it should be understood that communication can occur in a direction opposite to that of the represented arrows.
[0122] Many changes and modifications can be made to the embodiments without materially departing from the principles of the inventive concept. All such changes and modifications are intended to be included within the scope of the inventive concept herein. Accordingly, the foregoing subject matter should be understood as illustrative and not restrictive, and the examples of embodiments are intended to cover all such modifications, improvements, and other embodiments falling within the spirit and scope of the inventive concept. Thus, to the fullest extent permitted by law, the scope of the inventive concept shall be determined by the broadest permissible interpretation of this disclosure, which includes the examples of embodiments and their equivalents, and shall not be limited or restricted to the previous specific embodiments.
Claims
1. A method for filling a modified discrete cosine transform (MDCT) analysis window length with a time-domain signal to conceal lost or corrupted audio frames associated with a received audio signal, the method comprising: Copying (1000) a first segment of a previously received and decoded audio signal from a prototype buffer (1318), wherein the last decoded signal remains in the prototype buffer; Overlap-adding (1002) a second segment of the previously received and decoded audio signal from the prototype buffer (1318) to an initial portion of a time-domain substitute signal; Generating a time-domain signal by filling the MDCT analysis window with the copied segment, followed by the overlap-added segment, followed by the remaining portion of the time-domain substitute signal; and Reconstructing a concealed signal using the generated time-domain signal.
2. The method according to claim 1, wherein Copying the first segment of the previously received and decoded audio signal from the prototype buffer (1318) comprises: copying the first segment of the previously received and decoded audio signal to a processing buffer (1320), and overlap-adding the second segment of the previously received and decoded audio signal from the prototype buffer to the initial portion of the time-domain substitute signal comprises: overlap-adding the second segment of the previously received and decoded audio signal from the prototype buffer to the initial portion of the time-domain substitute signal to the processing buffer (1320).
3. The method according to claim 1, wherein, The previously received and decoded audio signal is a time-domain signal.
4. The method according to any one of claims 1 to 3, wherein, The time-domain substitute signal comprises an inverse fast Fourier transform signal of time evolution generated based on the previously received and decoded audio signal.
5. The method according to claim 2, wherein, The processing buffer is a folded modified discrete cosine transform (MDCT) buffer.
6. The method according to any one of claims 1 to 3, wherein The MDCT analysis window is asymmetric.
7. The method according to any one of claims 1 to 3, for continuously lost or corrupted frames, comprising: Copying (1100) a first segment of a reconstructed audio signal from a time continuity buffer (1316); And Overlap-adding (1102) a second segment of the reconstructed audio signal from the time continuity buffer (1316) to an initial portion of a time-evolution reconstructed audio signal to the processing buffer (1320), followed by the remaining portion of the time-evolution reconstructed audio signal.
8. The method according to claim 7, wherein Updating the time continuity buffer with the newly reconstructed signal.
9. The method according to any one of claims 1 to 3, wherein The overlap-added portion of the analysis window comprises: Applying a first window (701) to obtain a first scaled sample of the previously received and decoded audio signal from the prototype buffer; Applying a second window (702) to obtain a second scaled sample of the time-domain substitute signal; Summing (703) the first scaled sample and the second scaled sample to form the overlap-added portion of the analysis window.
10. The method according to any one of claims 1 to 3, wherein, The length of the first segment depends on the sampling frequency.
11. The method according to any one of claims 1 to 3, wherein, The length of the overlap-added portion of the analysis window depends on the sampling frequency.
12. The method according to any one of claims 1 to 3, further comprising: Dynamically adjusting (1200, 1306) the lengths of the first segment and the second segment based on an analysis of a previously synthesized time-domain signal from a filled analysis window length.
13. A decoder (1300) that fills a Modified Discrete Cosine Transform (MDCT) analysis window length with a time-domain signal to conceal a lost or damaged audio frame associated with a received audio signal, the decoder comprising: a processor (1306); and a memory (1308) coupled to the processor, wherein the memory includes instructions that, when executed by the processor, cause the decoder to perform operations that include: copying (1000) a first segment of a previously received and decoded audio signal from a prototype buffer (1318), wherein the last decoded signal remains in the prototype buffer; and overlap-adding (1002) a second segment of the previously received and decoded audio signal from the prototype buffer (1318) to an initial portion of a time-domain replacement signal; generating a time-domain signal by filling the MDCT analysis window with the copied segment, followed by the overlap-added segment, followed by the remaining portion of the time-domain replacement signal; and reconstructing a concealed signal using the generated time-domain signal.
14. The decoder (1300) according to claim 13, wherein, The decoder is adapted to perform the method according to any one of claims 2-12.
15. A computer-readable storage medium comprising program code to be executed by at least one processor (1306) of a decoder (1300), the program code for filling an analysis window length with a time-domain signal to conceal a lost or damaged audio frame associated with a received audio signal, whereby execution of the program code causes the decoder (1300) to perform the method according to any one of claims 1-12.
16. A computer program product comprising a non-transitory storage medium including program code to be executed by at least one processor (1306) of a decoder (1300), the program code for filling an analysis window length with a time-domain signal to conceal a lost or damaged audio frame associated with a received audio signal, whereby execution of the program code causes the decoder (1300) to perform the method according to any one of claims 1-12.
Citation Information
Patent Citations
Audio frame loss concealment
WO2014123470A1
Method and apparatus for packet loss concealment, and decoding method and apparatus employing same
EP3176781A2
Packet loss concealment apparatus and method, and audio processing system
WO2015003027A1