Method and apparatus for time-reversed audio subframe error concealment
By using time inversion technology to generate hidden audio subframes in the decoding device, the spectrum characteristic differences and computational complexity problems caused by frame loss are solved, and efficient audio signal processing and storage are achieved.
Patent Information
- Application Number
- CN202510545723.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2020-05-25
- Publication Date
- 2025-06-13
AI Technical Summary
When channel conditions change, frame loss may occur during the decoding process of the audio signal, resulting in increased spectrum characteristics differences and storage requirements, and high computational complexity.
By generating hidden audio subframes of audio signals in the decoding device, using time inversion technology, an error hiding unit (ECU) of multiple subframes is generated from a single subframe spectrum, ensuring that the subframes have similar spectral characteristics and reducing memory occupancy and computational complexity.
It is implemented to efficiently generate hidden audio subframes when the subframe windows are time-inverted versions of each other, ensuring the consistency of spectrum characteristics between frames and reducing storage and computing requirements.
Smart Images

Figure CN120148527A_ABST
Abstract
Description
This application is a divisional application of Chinese Patent Application No. 202080042683.0. The international filing date of the original application is May 25, 2020, and the invention title is "Time-Reversed Audio Sub-Frame Error Concealment". Technical Field The present disclosure generally relates to communications, and more particularly to methods and apparatus for controlling packet loss concealment for mono, stereo, or multi-channel audio encoding and decoding. Background Art Modern telecommunication services typically provide reliable connections between end users. However, such services still need to handle varying channel conditions under which occasional data packets may be lost due to, for example, network congestion or poor cell coverage. To overcome the problems of transmission errors and lost packets, telecommunication services can use packet loss concealment techniques (PLC). In the case where data packets are lost due to poor connections, network congestion, etc., the missing information of the lost packets on the receiver side can be replaced by a synthesized signal in the decoder. PLC techniques are typically closely associated with the decoder, where internal states can be used to generate signal continuation or extrapolation to compensate for packet loss. For multi-mode codecs with multiple operating modes for different signal types, there are typically multiple PLC techniques to handle concealment. Many different terms are used for packet loss concealment techniques, including frame error concealment (FEC), frame loss concealment (FLC), and error concealment unit (ECU). For linear prediction (LP)-based speech coding modes, PLC can adjust the glottal pulse position based on using the estimated frame tail pitch information and a copy of the pitch period of the previous frame [1]. The gain of the long-term predictor (LTP) converges to zero, and the convergence speed depends on the number of consecutive lost frames and the stability of the last good (i.e., error-free) frame [2]. Frequency domain (FD)-based coding modes are designed to handle general or complex signals, such as music. Different techniques can be used depending on the characteristics of the last received frame. Such analysis can include the number of detected tonal components and the period of the signal. If a frame loss occurs during a highly periodic signal (such as valid speech or a single musical instrument), time-domain PLC (similar to LP-based PLC) may be appropriate. In this case, FD PLC can mimic the LP decoder by estimating LP parameters and the excitation signal based on the last received frame [2]. If a lost frame occurs during a non-periodic or noise-like signal, the last received frame can be repeated in the spectral domain, where the coefficients are multiplied by a random sign signal to reduce the metallic sound of the repeated signal. For stationary tonal signals, it has been found that using methods based on prediction and extrapolation of the detected tonal components is advantageous. More details about the above techniques can be found in [1][2][3]. A general error concealment method that operates in the frequency domain is the Phase ECU (Error Concealment Unit) [4]. The Phase ECU is an independent tool that operates on a buffer of previously decoded and reconstructed time-domain signals. The framework of the Phase ECU is based on the sinusoidal analysis and synthesis paradigm. In this method, the sinusoidal components of the last good frame can be extracted and phase-shifted. When a frame is lost, the sinusoidal frequencies are obtained from past decoded synthesis in the DFT (Discrete Fourier Transform) domain. First, the corresponding frequency bins are identified by finding the peaks in the magnitude spectrum plane. Then, the fractional frequencies of the peaks are estimated using the peak frequency bins. The frequency bins corresponding to the peak and the adjacent peaks are phase-shifted using the fractional frequencies. For the remainder of the frame, the magnitudes of the past synthesis are retained while the phases are randomized. Burst errors are also processed so that the estimated signal is smoothly muted by converging the estimated signal to zero. More details about the Phase ECU can be found in [4]. The concept of the Phase ECU can be used in decoders that operate in the frequency domain. This concept includes encoding and decoding systems that perform decoding in the frequency domain, as Figure 1 shown, and also includes decoders that perform time-domain decoding and additional frequency-domain processing, as Figure 2 shown. In Figure 1 , the time-domain input audio signal (sub) frames are windowed 100 and transformed to the frequency domain by the DFT 101. The encoder 102 performs encoding in the frequency domain and provides the encoding parameters for transmission 103. The decoder 104 decodes the received frames or applies PLC 109 in case of frame loss. In the construction of the concealed frame, the PLC can use the memory 108 of the previously decoded frames. The decoded or concealed frames are transformed to the time domain by the inverse DFT 110 and then the output audio signal is reconstructed by the overlap-add operation 111. Figure 2 shows an encoder and decoder pair where the decoder applies the DFT transform to facilitate frequency-domain processing. The received and decoded time-domain signal is first windowed 105 by (sub) frames and then transformed to the frequency domain by the DFT 106 for frequency-domain processing 107, which can be done either before or after the PLC 109 (in case of frame loss). Since the frequency-domain spectra have been generated for each frame, the raw materials for the Phase ECU can be easily obtained by simply storing the last decoded spectrum in the memory. However, if the decoded spectrum corresponds to a frame of a time-domain signal with a different windowing function (see Figure 1) Then the efficiency of the algorithm may be reduced. This may occur when the decoder divides the synthesized frame into shorter sub - frames, for example, to process transient sounds that require a higher time resolution. To obtain good results, the ECU should generate the desired window shape for each frame; otherwise, there may be transition artefacts at each frame boundary. One solution is to store the spectrum of each frame corresponding to a specific window and apply the ECU to them individually. Another solution could be to store a single spectrum for the ECU and correct the windowing in the time domain. This can be achieved by applying an inverted window and then re - applying a window with the desired shape. These solutions have some drawbacks discussed below. One drawback of applying the frequency - domain ECU to individual sub - frames is that there may be differences between the sub - frames that will be replicated for each sub - frame during lost frames. For consecutive frame losses, this may lead to replication artefacts because each sub - frame may have slightly different spectral characteristics. Another problem is the increased memory requirement because the spectrum of each sub - frame needs to be stored. The window readjustment solution (where the windowing is inverted and re - applied) overcomes the problem of different spectral characteristics because the ECU can be based on a single sub - frame. However, applying the inverted window and applying the new window involve division and multiplication for each sample, where division is a computationally complex operation and has a high computational cost. This solution can be improved by storing the pre - calculated readjusted window in memory, but this will increase the required table memory. If the ECU is applied to a sub - part of the spectrum, it may also be necessary to readjust the full spectrum because the full spectrum needs to have the same window shape. Summary of the Invention According to a first aspect, there is provided a method of generating a hidden audio sub - frame of an audio signal in a decoding device. The method includes: generating a spectrum on the basis of a sub - frame, wherein consecutive sub - frames of the audio signal have the following characteristic: the applied window shape of a first sub - frame in the consecutive sub - frames is a mirror image version or a time - reversed version of a second sub - frame in the consecutive sub - frames. The method further includes: detecting peaks of the signal spectrum of a previously received audio signal on a fractional frequency scale; estimating the phase of each of the peaks; and based on the estimated phase, deriving a time - reversed phase adjustment to be applied to the peaks of the signal spectrum to form time - reversed phase - adjusted peaks. The method further includes: applying time - reversal to the hidden audio sub - frame. The potential advantage provided is to generate multi - sub - frame ECU from a single sub - frame spectrum by applying inverted time synthesis. This generation can be suitable for the case where the sub - frame windows are time - reversed versions of each other. Generating all ECU frames from a single stored decoded frame ensures that the sub - frames have similar spectral characteristics while keeping the memory footprint and computational complexity to a minimum. According to a second aspect, a decoder device configured to generate a hidden audio subframe of an audio signal is provided. The decoder device is configured to: generate a spectrum on a basis of subframes, wherein consecutive subframes of the audio signal have the following property: an applied window shape of a first subframe among the consecutive subframes is a mirror version or a time-reversed version of a second subframe among the consecutive subframes. The decoder device is further configured to: detect peaks of a signal spectrum of a previously received audio signal on a fractional frequency scale; and estimate a phase of each of the peaks. The decoder device is further configured to: derive, based on the estimated phase, a time-reversed phase adjustment to be applied to the peaks of the signal spectrum; and form time-reversed phase-adjusted peaks by applying the time-reversed phase adjustment to the peaks of the signal spectrum. The decoder device is further configured to: apply time reversal to the hidden audio subframe. According to a third aspect, a computer program is provided. The computer program includes program code to be executed by a processing circuit of a decoder device configured to operate in a communication network, whereby execution of the program code causes the decoder device to perform the operations according to the first aspect. According to a fourth aspect, a computer program product is provided. The computer program product includes a non-transitory storage medium storing program code to be executed by a processing circuit of a decoder device configured to operate in a communication network, whereby execution of the program code causes the decoder device to perform the operations according to the first aspect. According to a fifth aspect, a method of generating a hidden audio subframe of an audio signal in a decoding device is provided. The method includes: generating a spectrum on a basis of subframes, wherein consecutive subframes of the audio signal have the following property: an applied window shape of a first subframe among the consecutive subframes is a mirror version or a time-reversed version of a second subframe among the consecutive subframes. Storing a signal spectrum corresponding to the second subframe among the first two consecutive subframes. The method further includes: receiving a bad frame indicator for a second two consecutive subframes. The method further includes: obtaining the signal spectrum; detecting peaks of the signal spectrum on a fractional frequency scale; estimating a phase of each of the peaks; and deriving, based on the estimated phase, a time-reversed phase adjustment to be applied to the peaks of the stored spectrum for the first subframe among the second two consecutive subframes. The method further includes: applying the time-reversed phase adjustment to the peaks of the signal spectrum to form time-reversed phase-adjusted peaks. The method further includes: applying time reversal to the hidden audio subframe; combining the time-reversed phase-adjusted peaks with a noise spectrum of the signal spectrum to form a combined spectrum for the first subframe among the second two consecutive subframes; and generating a synthesized hidden audio subframe based on the combined spectrum. According to a sixth aspect, there is provided a decoder device configured to generate a hidden audio subframe of an audio signal. The decoder device includes processing circuitry and a memory operatively coupled to the processing circuitry, wherein the memory includes instructions that, when executed by the processing circuitry, cause the decoder device to perform the operations according to the first or fifth aspect. According to a seventh aspect, there is provided a decoder device. The decoder device is configured to generate a hidden audio subframe of an audio signal, wherein the decoder device is adapted to perform the method according to the fifth aspect. According to an eighth aspect, there is provided a computer program. The computer program includes program code to be executed by the processing circuitry of a decoder device configured to operate in a communication network, whereby execution of the program code causes the decoder device to perform the operations according to the fifth aspect. According to a ninth aspect, there is provided a computer program product. The computer program product includes a non-transitory storage medium storing program code to be executed by the processing circuitry of a decoder device configured to operate in a communication network, whereby execution of the program code causes the decoder device to perform the operations according to the fifth aspect. According to another aspect of the present disclosure, there is provided an audio decoding method. The method includes: a decoder generating a spectrum on a subframe basis. Successive subframes of an audio signal have the following characteristic: the window shape applied to a first subframe in the successive subframes is a mirror image version or a time-reversed version of the window shape applied to a second subframe in the successive subframes. The method includes: storing a signal spectrum corresponding to the second subframe. The audio decoding method further includes: in response to a frame loss, obtaining a previously generated signal spectrum corresponding to the second subframe. The audio decoding method further includes: detecting peaks of the signal spectrum and estimating the phase of each of the peaks. The audio decoding method further includes: calculating a phase adjustment for each detected peak based on the estimated phase. The audio decoding method further includes: adjusting the detected peaks by applying the phase adjustment to a peak interval of each detected peak to form a phase-adjusted peak interval and taking a complex conjugate of the phase-adjusted peak interval to form a time-reversed phase-adjusted peak. The audio decoding method further includes: combining the time-reversed phase-adjusted peak interval with a noise component of the spectrum derived from a non-peak interval of the signal spectrum to form a combined spectrum for a first hidden subframe of a hidden audio frame. In one embodiment, the synthesized hidden audio frame includes two successive hidden subframes. The method further includes: combining the phase-adjusted peak interval with the non-peak interval of the signal spectrum to form a combined spectrum for a second hidden subframe of the hidden audio frame. In one embodiment, the method further includes: associating each of the detected peaks with a plurality of peak frequency intervals representing the peaks. In one embodiment, the phase adjustment for the peaks of the hidden audio subframe is calculated according to the following formula: Δφ = -2φ 0 -2πf(N step21 +N lost ·N) / N, where φ 0 is the estimated phase of the peak and f is the frequency of the peak, N lost represents the number of consecutively lost frames, N represents the length of the full frame and N step21 is the sample distance between the start of the second subframe of the last received frame and the start of the first subframe of the hidden audio frame. In one embodiment, the detected peaks are adjusted by applying the phase adjustment to the peak intervals of each detected peak to form phase-adjusted peak intervals, and taking the complex conjugate of the phase-adjusted peak intervals to form time-reversed phase-adjusted peaks, according to the following formula: where * is the complex conjugate, is the signal spectrum, and Δφ i is the phase adjustment. In one embodiment, estimating the phase of each of the peaks includes: calculating a phase estimate for the peaks in the time-reversed phase-adjusted peaks according to the following formula: f frac = f i - k i where φ i is the estimated phase at frequency f i , is the spectrum of the previously received audio signal at the angle in the frequency interval k i , f frac is the rounding error, and φ C is the tuning constant. According to another aspect of the present disclosure, an audio decoder is provided that is configured to generate a spectrum on a subframe basis. Successive subframes of an audio signal have the following characteristic: the window shape applied to the first subframe in the successive subframes is a mirror version or a time-reversed version of the window shape applied to the second subframe in the successive subframes. The audio decoder is configured to store the signal spectrum corresponding to the second subframe. The audio decoder is further configured to, in response to a frame loss, obtain a previously generated signal spectrum corresponding to the second subframe. The audio decoder is further configured to detect peaks of the signal spectrum and estimate the phase of each of the peaks. The audio decoder is further configured to calculate a phase adjustment for each detected peak based on the estimated phase. The audio decoder is further configured to adjust the detected peaks by applying the phase adjustment to the peak interval of each detected peak to form a phase-adjusted peak interval and taking the complex conjugate of the phase-adjusted peak interval to form a time-reversed phase-adjusted peak. The audio decoder is further configured to combine the time-reversed phase-adjusted peak interval with a noise component of the spectrum derived from a non-peak interval of the signal spectrum to form a combined spectrum for a first hidden subframe for hiding an audio frame. In one embodiment, the synthesized hidden audio frame includes two successive hidden subframes. The audio decoder is further configured to combine the phase-adjusted peak interval with the non-peak interval of the signal spectrum to form a combined spectrum for a second hidden subframe of the hidden audio frame. In one embodiment, the audio decoder is further configured to associate each of the detected peaks with a plurality of peak frequency intervals representing the peak. In one embodiment, the audio decoder is configured to calculate the phase adjustment for the peaks of the hidden audio subframe according to the following formula: Δφ = -2φ 0 -2πf(N step21 +N lost ·N) / N, where φ 0 is the estimated phase of the peak and f is the frequency of the peak, N lost represents the number of successive lost frames, N represents the length of a full frame, and N step21 is the sample distance between the start of the second subframe of the last received frame and the start of the first subframe of the hidden audio frame. In one embodiment, adjusting the detected peaks by applying the phase adjustment to the peak interval of each detected peak to form a phase-adjusted peak interval and taking the complex conjugate of the phase-adjusted peak interval to form a time-reversed phase-adjusted peak is according to the following formula: where * is the complex conjugate, is the signal spectrum, and Δφ i is the phase adjustment. In one embodiment, estimating the phase of each of the peaks includes: calculating a phase estimate for the peak among the peaks of the time-reversed phase-adjusted signal according to the following formula: f frac = f i - k where φ i is the estimated phase at frequency f i is the spectrum of the previously received audio signal is the angle at frequency interval k is the rounding error, and φ i is the tuning constant. frac is the rounding error, and φ C is the tuning constant. According to another aspect of the present disclosure, there is provided a user equipment including the audio decoder as described in the foregoing aspect. BRIEF DESCRIPTION OF THE DRAWINGS The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application, illustrate specific non-limiting embodiments. In the drawings: Figure 1 is a block diagram showing an encoder and decoder pair, where encoding is performed in the DFT domain; Figure 2 is a block diagram showing an encoder and decoder pair, where the decoder applies a DFT transform to facilitate frequency-domain processing; Figure 3 is an illustration of two sub-frame windows of a decoder, where the window applied to the second sub-frame is the time-reversed version or mirror version of the window applied to the first sub-frame; Figure 4 is a block diagram showing an encoder and decoder system including a PLC method according to some embodiments, the PLC method performing phase estimation and applying ECU synthesis in reverse time using a time-reversed phase calculator; Figure 5 is a flowchart showing the operation of a decoder device performing time-reversed ECU synthesis according to some embodiments; Figure 6 is an illustration of a time-reversed window on a sine wave according to some embodiments; Figure 7 is an illustration of how the window in reverse time affects the DFT coefficients in the complex plane according to some embodiments; Figure 8 is a diagram of φ ε - frequency f; Figure 9 is a block diagram showing a decoder device according to some embodiments; Figure 10 is a flowchart showing the operation of a decoder device according to some embodiments; Figure 11 is a flowchart showing the operation of a decoder device according to some embodiments. Detailed Description Aspects of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which examples of embodiments are shown. However, the embodiments can be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the embodiments to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Components from one embodiment can be assumed to be present / used in another embodiment. The following description provides various embodiments of the disclosed subject matter. These embodiments are provided as teaching examples and are not to be construed as limiting the scope of the disclosed subject matter. For example, specific details of the described embodiments can be modified, omitted, or extended without departing from the scope of the described subject matter. Figure 9 is a block diagram showing the units of a decoder device 900 configured to provide wireless communication. The decoder device 900 can be part of a mobile terminal, a mobile communication terminal, a wireless communication device, a wireless terminal, a wireless communication terminal, a user equipment UE, a user equipment node / terminal / device, etc. As shown, the decoder 900 can include a network interface circuit 906 (also referred to as a network interface) configured to provide communication with other devices / entities / functions, etc. The decoder 900 can also include a processor circuit 902 (also referred to as a processor) operatively coupled to the network interface circuit 906 and a memory circuit 904 (also referred to as a memory) operatively coupled to the processor circuit. The memory circuit 904 can include computer-readable program code that, when executed by the processor circuit 902, causes the processor circuit to perform operations according to the embodiments disclosed herein. According to other embodiments, the processor circuitry 902 can be defined to include memory such that a separate memory circuitry is not required. As discussed herein, the operations of the decoder 900 can be performed by the processor 902 and / or the network interface 906. For example, the processor 902 can control the network interface 906 to send communications to a multi-channel audio player and / or receive communications from one or more other network nodes / entities / servers (such as an encoder node, a repository server, etc.) via the network interface 906. Additionally, modules can be stored in the memory 904, and these modules can provide instructions such that when the instructions of the modules are executed by the processor 902, the processor 902 performs corresponding operations. In the following description, sub-frame notation will be used to describe the embodiments. Here, a sub-frame represents a part of a larger frame, where the larger frame includes a set of sub-frames. The described embodiments can also be used with frame notation. In other words, sub-frames can form a set of frames having the same window shape as the window shapes described herein, and sub-frames do not need to be part of a larger frame. Consider the decoder in an encoder-decoder pair, where the decoding method generates a spectrum on a sub-frame basis. Consecutive sub-frames can have the following property: the applied window shapes are mirror versions or time-reversed versions of each other, as Figure 3 shown, where sub-frame 2 is a mirror version or time-reversed version of sub-frame 1. The decoder obtains the reconstructed sub-frame spectrum for each frame m. In one embodiment, the sub-frame spectrum can be obtained from the reconstructed time domain synthesis where n is the sample index. Figure 2 The dashed box in indicates that frequency domain processing can be performed before or after the memory and the PLC module. The spectrum can be obtained by multiplying 1 with the sub-frame windowing functions w 2 (n) and w where N represents the length of the sub-frame window, and N step12 is the sample distance between the start points of the first sub-frame and the second sub-frame. The sub-frame windowing functions w 1 (n) and w 2 (n) are mirror versions or time-reversed versions of each other. Here, the sub-frame spectrum is obtained from the decoder time domain synthesis, similar to Figure 2 the system outlined in Figure 1 It should be noted that the embodiments are equally applicable to the system outlined in is stored in the memory. For correctly received frames, the decoder device 900 can continue to perform frequency-domain processing steps, perform the inverse DFT transform, and use an overlap-add strategy to reconstruct the output audio. Missing or corrupted frames can be identified by the processing-connected transport layer and signaled to the decoder as "bad frames" via a bad frame indicator (BFI), which can take the form of a flag. When the decoder device 900 detects a bad frame via the bad frame indicator (BFI), the PLC algorithm is activated. The PLC follows the principle of the phase ECU [4]. The stored spectrum is input into a peak detector algorithm that detects peaks on a fractional frequency scale. A set of peaks F = {f i}, i = 1, 2, … N peaks These peaks are represented by their estimated fractional frequencies f i and among them N peaks is the number of detected peaks. Similar to the sinusoidal coding paradigm, the peaks of the spectrum are modeled using sine waves with specific amplitudes, frequencies, and phases. The fractional frequency can be represented as a fraction of the DFT bins such that, for example, the Nyquist frequency is found at f = N / 2 + 1. Each peak can be associated with multiple frequency bins representing that peak. These bins are found by rounding the fractional frequency to the nearest integer and including adjacent bins (e.g., N near peaks on each side): where [·] represents the rounding operation and G i is the group of bins representing the peak at frequency f i . The quantity N near is a tuning constant that can be determined when designing the system. A larger N near provides higher accuracy in each peak representation but also introduces a larger distance between the peaks that can be modeled. A suitable value for N near can be 1 or 2. The peaks of the hidden spectrum can be formed by using these groups of bins, where a phase adjustment has been applied to each group. The phase adjustment takes into account the phase change in the underlying sine wave, assuming that the frequency remains the same between the last correctly received and decoded frame and the hidden frame. The phase adjustment is based on the fractional frequency and the number of samples between the analysis frame of the previous frame and the start position of the current frame. As Figure 3 shown, between the start of the second subframe of the last received frame and the start of the first subframe of the first ECU frame, this number of samples is N step21and between the first sub-frame of the last received frame and the first sub-frame of the first ECU frame, the number of samples is N full . Note that N full also gives the distance between the second sub-frame of the last received frame and the second sub-frame of the first ECU frame. Figure 4 An encoder and decoder system according to an embodiment described below is shown, where the PLC block 109 uses a phase estimator 112 to perform phase estimation and applies ECU synthesis in reverse time using a time-reversed phase calculator. Figure 5 is a flowchart showing the steps of the time-reversed ECU synthesis described below. For the hiding of the first sub-frame, ECU synthesis can be performed in reverse time to obtain a desired window shape. For the first sub-frame, the phase adjustment or phase correction or phase progression (these terms can be used interchangeably throughout the specification) of peak i can be expressed as Δφ i = -2φ i -2πf i (N + N step21 +(N lost -1)N full ) / N, where N lost represents the number of consecutive lost frames, and φ i represents the phase of the sine wave at frequency f i . The term (N lost -1)N full handles the phase progression of burst errors, where the step size increases with the frame length N full of the entire frame. For the first lost frame, N lost = 1. For a frequency centered on the frequency interval of the spectrum , the phase φ i can be easily obtained by simply extracting the angle: where k i = [f i . Generally speaking, the frequency f i is a fraction and the phase needs to be estimated in operation 501. One estimation method is to use linear interpolation of the phase spectrum. where and respectively represent operators for rounding down and rounding up. However, it has been found that this estimation method is unstable. This estimation method also requires two phase extractions, which, if using complex numbers in the standard form a + bi to represent the spectrum, requires computationally complex arctangent (arctan) functions. Another phase estimation that has been found to be reliable at relatively low computational complexity is: f frac = f i - k i where f frac is the rounding error, and φ C is the tuning constant, which depends on the window shape applied. For the window shape of this embodiment, a suitable value has been found to be φ C = 0.33. For another window shape, a suitable value has been found to be φ C = 0.48. Generally, it is expected that suitable values can be found within the range [0.1, 0.7]. In operation 502, the time-reversed phase adjustment Δφ i is derived as described above. The peak of the hidden spectrum can be formed by applying the phase adjustment to the stored spectrum in operation 503. The asterisk "*" represents the complex conjugate, which gives the time reversal of the signal in operation 504. This results in the time reversal of the first ECU subframe. It should be noted that the reversal can also be performed in the time domain after the inverse DFT. However, if only represents a part of the complete spectrum, this requires preprocessing of the remaining spectrum, for example, by time reversal before DFT analysis. The remaining interval not occupied by the peak interval G i of can be referred to as the noise spectrum or noise component of the spectrum. These remaining intervals can be filled with the coefficients of the stored spectrum to which a random phase has been applied: where φ rand represents the random phase value. The remaining intervals can also be filled with spectrum coefficients that preserve the desired characteristics of the signal (e.g., the correlation with the second channel in a multi-channel decoder system). In operation 505, the peak spectrum (where k ∈ G i ) is combined with the noise spectrum (where ) to form a combined spectrum. In an embodiment where noise is generated in the time domain and windowed and transformed, the time reversal of the noise (to match the windowing of the peak component) and the combination with the peak spectrum should be performed before applying the above-mentioned time reversal. For the generation of the second sub-frame synthesized in normal (non-reversed) time, conventional phase adjustment can be used. Δφ i = 2πf i N full N lost / N The ECU synthesis for the second sub-frame can be formed similarly to the first sub-frame, but omitting the complex conjugate on the peak coefficient. Once the combined hidden spectrum is generated in operation 505, the combined hidden spectrum can be fed in operation 506 to subsequent processing steps, including the inverse DFT and the overlap-add operation resulting in the output audio signal. The output audio signal can be sent to one or more speakers, such as the speakers for playback. The speakers can be part of the decoding device, a separate device, or part of another device. Derivation of the Phase Correction Formula for ECU Synthesis for Time Reversal Assume the starting phase of the sine component is φ 0 , and the frequency of the sine wave is f. After advancing N step samples, the desired phase φ 1 of the sine wave is then: φ 1 = φ 0 + 2πfN step / N For the continuation of the time reversal of the sine wave, the phase needs to be mirrored in the real axis by applying the complex conjugate or by simply taking the negative phase -φ 1 Since this phase angle now represents the end point of the ECU synthesis frame, the phase needs to wrap around the length of the analysis frame to reach the desired starting phase φ 2 . φ 2 = -φ 1 - 2πf(N To obtain the phase correction Δφ, the starting phase needs to be subtracted, i.e., Substituting φ 2 gives Δφ = -2φ 0 - 2πf(N step + N - 1) / N To add the progression of consecutive frame losses (burst losses), a factor N corresponding to the number of samples between the start point of the full frame can be added. offset =(N lost -1)N full . This provides the final phase correction: Δφ = -2φ 0 -2πf(N + N step -1+(N lost -1)N full ) / N, By using complex conjugation and single-sample circular shifting, the desired time reversal can be achieved in the DFT domain. This circular shifting can be achieved by a phase correction of 2πk / N, which can be included in the final phase correction. Δφ = -2φ 0 -2πf(N + N step -1+(N lost -1)N full ) / N + 2πk / N For the coefficient representing a single peak, the frequency interval k of the circular shift can be approximated as the fractional frequency k ≈ f, and the phase correction can be simplified to: Δφ = -2φ 0 -2πf(N + N step -1+(N lost -1)N full ) / N + 2πf / N = -2φ 0 -2πf(N + N step +(N lost -1)N full ) / N The window can be designed such that N = N full , in which case the expression can be further simplified to: Δφ = -2φ 0 -2πf(N step +N lost ·N) / N Alternative Embodiment of Inverted Time ECU Synthesis In another embodiment, the phase correction is performed in two steps. In the first step, the phase is advanced, ignoring the window mismatch. Δφ = 2πf(N step +(N lost -1)N full ) In the second step, the phase can be made to advance by -φ mReturn, apply the complex conjugate, and use φ m Restore the phase to implement windowed time reversal: By studying the effect of the time-reversal window on a sine wave as Figure 6 shown, the motivation for this operation can be found. In Figure 6 , the upper figure shows the window applied in the first direction, and the lower figure shows the window applied in the reverse direction. In Figure 7 , three coefficients representing the sine wave are shown, and this figure shows how the window with reversed time affects the DFT coefficients in the complex plane. Approximately Figure 6 the three DFT coefficients of the sine wave in the upper figure of Figure 6 are marked with circles, while the corresponding coefficients in the lower figure of m are marked with asterisks. The diamond represents the position of the original phase of the sine wave, and the dashed line shows the observed mirror plane through which the coefficients of the time-reversal window are projected. The time-reversal window gives the mirror image of the coefficients in the mirror plane and the angle is φ φ m = φ 0 + φ frac Through experiments, it is found that φ frac can be expressed as: φ frac = πf fraqc f fraqc = f i - k i k i = [f i where [·] represents the rounding operation. It is also found that φ ε (expressed as a positive angle) can be approximated as a linear relationship with f frac . In Figure 8 , the angle φ ε is expressed as a function of the frequency f. Studying Figure 8 the sawtooth shape, it is found that a good approximation of φ ε is: φ ε = - f frac φ C where φ C is a constant. In one embodiment, φ C can be set to φ C = 0.33, which gives a close approximation. Since φ 0 is not explicitly known, an alternative approximation of φ m can be denoted as: Among them, is the phase of the maximum peak coefficient found at the rounded frequency interval k after the first phase adjustment step i at the place, The operations of aligning the mirror plane with the real axis, applying the complex conjugate, and reversing the phase again can be understood as adjusting the phase of the shaped sine wave to a phase position (0 or π) that is neutral to the complex conjugate, thereby only reversing the time shape of the signal. The two-step method is computationally more complex than the previously described embodiments. However, the observation can also lead to an approximation of φ 0 From Figure 7 it can be seen that φ 0 can be expressed as: φ 0 = φ ki + φ ε - φ ftac = φ ki - f frac (φ C + π) This is the phase approximation used above. Now, the operation of the decoder device 900 (implemented using the structure of the block diagram of Figure 10 ) will be discussed with reference to the flowchart of some embodiments. For example, the modules can be stored in Figure 9 the memory 904 of Figure 9 , and these modules can provide instructions such that when the instructions of the modules are executed by the corresponding decoder device processing circuit 902, the processing circuit 902 performs the corresponding operations of the flowchart. In operation 1000, the processing circuit 902 generates a spectrum on a subframe basis, where the consecutive subframes of the audio signal have the following characteristics: the window shape applied to the first subframe in the consecutive subframes is a mirror version or a time-reversed version of the second subframe in the consecutive subframes. For example, generating a spectrum for each of the first two consecutive subframes includes determining: Among them, N represents the length of the subframe window, and the subframe windowing function w 1 (n) is the subframe windowing function of the first subframe in the consecutive subframes , w 2 (n) is the subframe windowing function of the second subframe in the consecutive subframes , and N step12 is the number of samples between the first subframe and the second subframe in the first two consecutive subframes. In operation 1002, processing circuit 902 determines whether a bad frame indicator (BFI) has been received. The bad frame indicator provides an indication that an audio frame has been lost or corrupted. In operation 1004, for each correctly decoded audio frame, processing circuit 902 stores the spectrum corresponding to the second subframe in a memory. For example, for correctly decoded frame m, the spectrum corresponding to the second subframe is stored in the memory, e.g., For correctly received frames, decoder device 900 may continue to perform frequency domain processing steps, perform an inverse DFT transform, and use an overlap - add strategy to reconstruct the output audio, as described above and Figure 4 shown. Note that the principle of overlap - add is the same for both subframes and frames. Creation of a frame requires application of overlap - add to subframes, and the final output frame is the result of an overlap - add operation between frames. When processing circuit 902 detects a bad frame via the bad frame indicator (BFI) in operation 1002, PLC operations 1006 to 1030 are performed. In operation 1006, processing circuit 902 obtains the signal spectrum corresponding to the second subframe in the first two consecutive subframes that were previously correctly decoded and processed. For example, processing circuit 902 may obtain the signal spectrum from memory 904 of the decoding device. In operation 1008, processing circuit 902 detects peaks in the signal spectrum of a previously received audio frame of the audio signal on a fractional frequency scale, where the previously received audio frame was received before the bad frame indicator was received. In operation 1010, processing circuit 902 determines whether a hidden frame is used for the first subframe in two consecutive subframes. If a hidden frame is used for the first subframe, then in operation 1012, processing circuit 902 estimates the phase of each peak. In one embodiment, the phase estimate is calculated for the peak in the time - reversed phase - corrected peak according to the following formula: f frac = f i - k i where φ i is the estimated phase at frequency f i , is the angle of the spectrum i at frequency interval k , f frac is the rounding error, φ C is the tuning constant, and k i is [f i . The tuning constant φ CIt can be a value within the range between 0.1 and 0.7. In operation 1014, the processing circuit 902 derives a phase correction for time reversal to be applied to the peak of the signal spectrum based on the estimated phase. In operation 1016, the processing circuit 902 applies the phase correction for time reversal to the peak of the signal spectrum to form a peak with time-reversed phase correction. In operation 1018, the processing circuit 902 applies time reversal to the hidden audio subframe. In one embodiment, time reversal can be applied by applying the complex conjugate to the hidden audio subframe. In operation 1020, the processing circuit 902 combines the peak with time-reversed phase correction and the noise spectrum of the signal spectrum to form a combined spectrum of the hidden audio subframe. Go to Figure 11 , in one embodiment, the processing circuit 902 can perform 1016 and 1018 by associating each peak with a plurality of peak frequency intervals in operation 1100. The processing circuit 902 can apply the phase correction for time reversal by applying the phase correction for time reversal to each of the plurality of frequency intervals in operation 1102. In operation 1104, the remaining intervals are filled with the coefficients of the signal spectrum to which a random phase has been applied. Return to Figure 10 , in operation 1022, the processing circuit 902 generates a synthesized hidden audio subframe based on the combined spectrum. If it is determined in operation 1010 that the hidden frame is not used for the first subframe, the processing circuit 902 derives a non-time-reversed phase correction to be applied to the peak of the signal spectrum for the second hidden subframe among at least two consecutive hidden subframes in operation 1024. In operation 1026, the processing circuit 902 applies the non-time-reversed phase correction to the peak of the signal spectrum for the second subframe to form a peak with non-time-reversed phase correction. In operation 1028, the processing circuit 902 combines the peak with non-time-reversed phase correction and the noise spectrum of the signal spectrum to form a combined spectrum for the second hidden subframe. In operation 1030, the processing circuit 902 generates a second synthesized hidden audio subframe based on the combined spectrum. Go to Figure 11, in one embodiment, processing circuit 902 may perform 1026 and 1028 by associating each peak with a plurality of peak frequency intervals in operation 1100. The association by processing circuit 902 may apply a non-time-reversed phase correction by applying a non-time-reversed phase correction to each of the plurality of frequency intervals in operation 1102. In operation 1104, the remaining intervals are filled with coefficients of the signal spectrum to which a random phase has been applied. For some embodiments of decoder devices and related methods, various operations of the Figure 10 flowchart may be optional. For example, with respect to the method of Example Embodiment 1 (set forth below), Figure 10 the operations of blocks 1004 and 1022-1030 may be optional. For example, with respect to the method of Example Embodiment 19 (set forth below), Figure 10 the operations of blocks 1010 and 1022-1030 may be optional. Example embodiments are discussed below. 1. A method of generating a hidden audio subframe of an audio signal in a decoding device, the method comprising: generating (1000) a spectrum on a subframe basis, wherein consecutive subframes of the audio signal have the following property: the applied window shape of the first subframe in the consecutive subframes is a mirror image version or a time-reversed version of the second subframe in the consecutive subframes; receiving (1002) a bad frame indicator; detecting (1008) peaks of a signal spectrum of a previously received audio frame of the audio signal on a fractional frequency scale, the previously received audio frame being received before the bad frame indicator; estimating (1012) the phase of each peak; deriving (1014) a time-reversed phase correction to be applied to the peaks of the signal spectrum based on the estimated phase; applying (1016) the time-reversed phase correction to the peaks of the signal spectrum to form time-reversed phase-corrected peaks; applying (1018) time-reversal to the hidden audio subframe; combining (1020) the time-reversed phase-corrected peaks with a noise spectrum of the signal spectrum to form a combined spectrum for the hidden audio subframe; and generating (1022) a synthesized hidden audio subframe based on the combined spectrum. 2. The method according to embodiment 1, wherein synthesizing the hidden audio frame includes at least two consecutive hidden sub - frames, and wherein deriving the time - reversed phase correction, applying the time - reversed phase correction, applying time reversal, and combining the peaks of the time - reversed phase correction are performed for the first hidden sub - frame of at least two consecutive hidden sub - frames. The method further includes: For the second hidden sub - frame of at least two consecutive hidden sub - frames, deriving (1024) a non - time - reversed phase correction to be applied to the peak of the signal spectrum; For the second sub - frame, applying (1026) the non - time - reversed phase correction to the peak of the signal spectrum to form a peak with non - time - reversed phase correction; Combining (1028) the peak with non - time - reversed phase correction with the noise spectrum of the signal spectrum to form a combined spectrum for the second hidden sub - frame; and Generating (1030) a second synthesized hidden audio sub - frame based on the combined spectrum. 3. The method according to any one of embodiments 1 - 2, wherein the hidden audio sub - frame includes a hidden audio sub - frame for one of a lost audio frame and a corrupted audio frame. 4. The method according to any one of embodiments 1 - 3, wherein the bad frame indicator provides an indication of audio frame loss or corruption. 5. The method according to any one of embodiments 1 - 4, further including: obtaining the signal spectrum of a previously received audio signal frame from the memory of the decoder. 6. The method according to any one of embodiments 1 - 5, wherein applying time reversal includes: applying a complex conjugate to the hidden audio sub - frame. 7. The method according to any one of embodiments 1 - 6, further including: Associating (1100) each peak of the plurality of peaks with a plurality of peak frequency intervals representing the peak. 8. The method according to embodiment 7, wherein for each peak of the plurality of peaks, one of the time - reversed phase correction and the non - time - reversed phase correction is applied (1102) to the peak. 9. The method according to any one of embodiments 8, further including: Filling (1104) the remaining intervals of the signal spectrum with coefficients of the stored signal spectrum with a random phase applied. 10. The method according to any one of embodiments 1 - 9, wherein estimating the phase of each peak includes: Calculating a phase estimate for the peak in the peak with time - reversed phase correction according to the following formula: ffrac = f i - k i where φ i is the estimated phase at frequency f i , is the spectrum at frequency interval k i , is the angle of f frac , f C is the rounding error, φ i is the tuning constant, and k i is [f 11. The method according to embodiment 10, wherein φ C has a value in the range between 0.1 and 0.7. 12. The method according to embodiment 10, wherein the phase estimation for peak calculation for non-time-reversed phase correction is calculated according to the following formula: Δφ i = 2πf i N full N lost / N where Δφ i represents the phase correction of the sine wave at frequency f i , N full represents the number of samples between two frames, N lost represents the number of consecutive lost frames, and N represents the length of the subframe window. 13. The method according to any one of embodiments 1-12, further comprising: applying a random phase to the noise spectrum of the signal spectrum. 14. The method according to embodiment 13, wherein applying a random phase to the noise spectrum includes: applying a random phase to the noise spectrum before combining the non-time-reversed phase-corrected peak with the noise spectrum. 15. A decoder device (900) configured to generate a hidden audio subframe of a received audio signal, wherein the decoding method of the decoding device generates a spectrum on a subframe basis, and wherein consecutive subframes have the following characteristics: the applied window shapes are mirror versions or time-reversed versions of each other, and the decoder device includes: a processing circuit (902); and a memory (904) coupled to the processing circuit, wherein the memory includes instructions that, when executed by the processing circuit, cause the decoder device to perform the operations according to any one of embodiments 1-14. 16. A decoder device (900) configured to generate hidden audio sub - frames of a received audio signal, wherein a decoding method of the decoding device generates a spectrum on a sub - frame basis, and wherein consecutive sub - frames have the following property: the applied window shapes are mirror versions or time - reversed versions of each other, and wherein the decoder device is adapted to perform according to any one of embodiments 1 - 14. 17. A computer program comprising program code to be executed by a processing circuit (902) of a decoder device (900) configured to operate in a communication network, whereby execution of these program codes causes the decoder device (900) to perform the operations according to any one of embodiments 1 - 14. 18. A computer program product comprising a non - transitory storage medium storing program code to be executed by a processing circuit (902) of a decoder device (900) configured to operate in a communication network, whereby execution of these program codes causes the decoder device (900) to perform the operations according to any one of embodiments 1 - 14. 19. A method of generating hidden audio sub - frames of an audio signal in a decoding device, the method comprising: Generating (1000) a spectrum on a sub - frame basis, wherein consecutive sub - frames of the audio signal have the following property: the applied window shape of the first sub - frame in consecutive sub - frames is a mirror version or a time - reversed version of the second sub - frame in consecutive sub - frames; Storing (1004) the signal spectrum corresponding to the second sub - frame in the first two consecutive sub - frames; Receiving a bad - frame indicator (1002) for the second two consecutive sub - frames; Obtaining (1006) the signal spectrum; Detecting (1008) peaks of the signal spectrum on a fractional - frequency scale; Estimating (1012) the phase of each peak; Based on the estimated phase, deriving (1014) a time - reversed phase correction to be applied to the peaks of the stored spectrum for the first sub - frame in the second two consecutive sub - frames; Applying (1016) the time - reversed phase correction to the peaks of the signal spectrum to form time - reversed phase - corrected peaks; Applying (1018) time - reversal to the hidden audio sub - frame; Combining (1020) the time - reversed phase - corrected peaks with the noise spectrum of the signal spectrum to form a combined spectrum for the first sub - frame in the second two consecutive sub - frames; and Generating (1022) a synthetic hidden audio sub - frame based on the combined spectrum. 20. The method according to embodiment 19, wherein synthesizing a hidden audio frame includes at least two consecutive hidden sub - frames, and wherein deriving the time - reversed phase correction, applying the time - reversed phase correction, and combining the peaks of the time - reversed phase correction are performed for the first hidden sub - frame of at least two consecutive hidden sub - frames, the method further comprising: For the second sub - frame of the second two consecutive sub - frames, derive (1024) a non - time - reversed phase correction to be applied to the peak of the signal spectrum; For the second sub - frame of the second two consecutive sub - frames, apply (1026) the non - time - reversed phase correction to the peak of the signal spectrum to form a non - time - reversed phase - corrected peak; Combine (1028) the non - time - reversed audio sub - frame with the noise spectrum of the signal spectrum to form a second combined spectrum for the second sub - frame of the second two consecutive sub - frames; and Generate (1030) a second synthesized audio sub - frame based on the second combined spectrum. 21. The method according to any one of embodiments 19 - 20, wherein the hidden audio sub - frame includes a hidden audio sub - frame of one of a lost audio frame and a corrupted audio frame. 22. The method according to any one of embodiments 19 - 21, wherein the bad - frame indicator provides an indication of audio frame loss or corruption. 23. The method according to any one of embodiments 19 - 22, further comprising: obtaining the signal spectrum from the memory of the decoder. 24. The method according to any one of embodiments 19 - 23, wherein applying time - reversal includes: applying a complex conjugate to the hidden audio sub - frame. 25. The method according to any one of embodiments 18 - 24, further comprising: Associate each peak with a plurality of peak frequency intervals representing the peak. 26. The method according to embodiment 25, further comprising: for each peak of the plurality of peaks, applying one of a time - reversed phase correction and a non - time - reversed phase correction to the peak. 27. The method according to embodiment 26, further comprising: Filling the remaining intervals of the signal spectrum with coefficients of the stored spectrum having a random phase applied. 28. The method according to any one of embodiments 19 - 27, wherein estimating the phase includes: Calculating a phase estimate for the peak of the time - reversed phase - corrected peak according to the following formula: f frac =f i -ki where φ i is the estimated phase at frequency f i ; is the angle of the spectrum at frequency interval k i ; f is the rounding error, φ frac is the tuning constant, and k C is [f i . i 29. The method according to embodiment 28, wherein φ C has a value in the range between 0.1 and 0.7. 30. The method according to embodiment 28, further comprising: calculating a phase estimate for a peak of non-time-reversed phase correction according to the following formula: Δφ i = 2πf i N full N lost / N where Δφ i represents the phase correction of a sine wave at frequency f i ; N full represents the number of frame samples between two frames, N lost represents the number of consecutive lost frames, and N represents the length of the subframe window. 31. The method according to any one of embodiments 19-30, wherein generating a spectrum for each of the first two consecutive subframes includes determining the following: where N represents the length of the subframe window, the subframe windowing function w 1 (n) is the subframe windowing function for the first subframe in consecutive subframes ; w 2 (n) is the subframe windowing function for the second subframe in consecutive subframes ; and N step1 is the number of samples between the first subframe and the second subframe in the first two consecutive subframes. 32. The method according to any one of embodiments 19-31, further comprising: applying a random phase to the noise spectrum of the signal spectrum. 33. The method according to embodiment 32, wherein applying a random phase to the noise spectrum includes: applying a random phase to the noise spectrum before combining the non-time-reversed phase-corrected peak with the noise spectrum. 34. A decoder device (900) configured to generate hidden audio sub - frames of a received audio signal, wherein a decoding method of the decoding device generates a spectrum on a sub - frame basis, and wherein consecutive sub - frames have the following property: the applied window shapes are mirror versions or time - reversed versions of each other. The decoder device includes: A processing circuit (902); and A memory (904) coupled to the processing circuit, wherein the memory includes instructions which, when executed by the processing circuit, cause the decoder device to perform the operations according to any one of Embodiments 19 - 33. 35. A decoder device (900) configured to generate hidden audio sub - frames of a received audio signal, wherein a decoding method of the decoding device (900) generates a spectrum on a sub - frame basis, and wherein consecutive sub - frames have the following property: the applied window shapes are mirror versions or time - reversed versions of each other, and wherein the decoder device is adapted to perform according to any one of Embodiments 19 - 33. 36. A computer program comprising program code to be executed by a processing circuit (902) of a decoder device (900) configured to operate in a communication network, whereby execution of these program codes causes the decoder device (900) to perform the operations according to any one of Embodiments 19 - 33. 37. A computer program product comprising a non - transitory storage medium storing program code to be executed by a processing circuit (902) of a decoder device (900) configured to operate in a communication network, whereby execution of these program codes causes the decoder device (900) to perform the operations according to any one of Embodiments 19 - 33. A description of various abbreviations / acronyms used in the present disclosure is provided below. Abbreviation Explanation DFT Discrete Fourier Transform IDFT Inverse Discrete Fourier Transform LP Linear Prediction PLC Packet Loss Concealment ECU Error Concealment Unit FEC Frame Error Correction / Concealment References are provided below.
[0001] "Efficient Frame Erasure Concealment in Predictive Speech Codecs using Glottal Pulse Resynchronisation" by T. Vaillancourt, M. Jelinek, R. Salami, and R. Lefebvre, Proceedings of the 2007 IEEE International Conference on Acoustics, Speech, and Signal Processing - ICASSP'07, Honolulu, Hawaii, 2007, pp. IV-1113-IV-1116.
[0002] "Packet-loss concealment technology advances in EVS" by J. Lecomte et al., Proceedings of the 2015 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Brisbane, Queensland, 2015, pp. 5708-5712.
[0003] 3GPP TS26.447, Codec for Enhanced Voice Services (EVS); Error Concealment of Lost Packets (Release 12)
[0004] "A novel sinusoidal approach to audio signal frame loss concealment and its application in the new evs codec standard" by S. Bruhn, E. Norvell, J. Svedberg, and S. Sverrisson, Proceedings of the 2015 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Brisbane, Queensland, 2015, pp. 5142-5146. In general, unless a different meaning is clearly given and / or is implicit in the context in which the term is used, all terms used herein will be interpreted according to their ordinary meaning in the relevant technical field. Unless otherwise specified, all references to an element, apparatus, component, part, step, etc. shall be construed openly as referring to at least one instance of the element, apparatus, component, part, step, etc. The steps of any method disclosed herein need not be performed in the exact order disclosed, unless a step is explicitly described as after or before another step and / or it is implicit that a step must be after or before another step. In appropriate cases, any feature of any embodiment disclosed herein may be applied to any other embodiment. Similarly, any advantage of any embodiment may apply to any other embodiment, and vice versa. Other objects, features, and advantages of the accompanying embodiments will be apparent from the following description. In the foregoing description of the various embodiments, it will be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the context of this specification and the relevant art, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. When a unit is referred to as being "connected to", "coupled to", "responsive to" (or a variant thereof) another unit, it can be directly connected to, coupled to, or responsive to the other unit, or there may be an intermediate unit. In contrast, when a unit is referred to as being "directly connected to", "directly coupled to", "directly responsive to" (or a variant thereof) another unit, there is no intermediate unit. The same number within this document refers to the same unit. Additionally, as used herein, "coupled", "connected", "responsive", or a variant thereof may include wireless coupling, connection, or responsiveness. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. For brevity and / or clarity, well-known functions or structures may not be described in detail. The term "and / or" includes any and all combinations of one or more of the listed associated items. It will be understood that although the terms first, second, third, etc. may be used herein to describe various units / operations, these units / operations should not be limited by these terms. These terms are only used to distinguish one unit / operation from another. Thus, a first unit / operation in some embodiments may be referred to as a second unit / operation in other embodiments without departing from the teachings of this disclosure. The same reference numeral or the same reference indicator within this specification indicates the same or similar unit. As used herein, the terms "comprising", "including", "having" or variations thereof are open-ended and include one or more of the stated features, integers, units, steps, components or functions, but do not preclude the presence or addition of one or more other features, integers, units, steps, components, functions or combinations thereof. Further, as used herein, the general abbreviation "e.g.", which is derived from the Latin phrase "exempli gratia", may be used to introduce or specify one or more general examples of the previously mentioned items and is not intended to be a limitation of such items. The general abbreviation "i.e.", which is derived from the Latin phrase "id est", may be used to specify a particular item from a more general recitation. Example embodiments are described herein with reference to block diagrams and / or flowcharts of computer-implemented methods, apparatus (systems and / or devices) and / or computer program products. It will be understood that the blocks of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by computer program instructions executed by one or more computer circuits. These computer program instructions can be provided to a processor circuit of a general purpose computer circuit, a special purpose computer circuit and / or other programmable data processing circuit to produce a machine such that the instructions, when executed via the processor of the computer and / or other programmable data processing device, transform and control transistors, values stored in storage units and other hardware components within such circuits to implement the functions / operations specified in one or more of the blocks of the block diagrams and / or flowcharts, thereby producing an apparatus (function) and / or structure that implements the functions / operations specified in the blocks of the block diagrams and / or flowcharts. These computer program instructions can also be stored in a tangible computer-readable medium, which can cause a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instructions for implementing the functions / operations specified in one or more of the blocks of the block diagrams and / or flowcharts. Accordingly, embodiments of the present disclosure can be embodied in hardware and / or software (including firmware, resident software, microcode, etc.) that runs on a processor such as a digital signal processor, and the hardware and / or software can be collectively referred to as "circuitry", "module" or variations thereof. It should also be noted that in some alternative implementations, the functions / operations marked in the boxes may occur in a different order than that marked in the flowchart. For example, two consecutive boxes may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions / operations involved. In addition, the functions of a given box in the flowchart and / or block diagram may be divided into multiple boxes, and / or the functions of two or more boxes in the flowchart and / or block diagram may be at least partially integrated. Finally, other boxes may be added / inserted between the boxes shown, and / or boxes / operations may be omitted without departing from the scope of the embodiments. Further, although some of the figures include arrows on communication paths to indicate the primary direction of communication, it will be understood that communication may occur in a direction opposite to that shown by the arrows. Many changes and modifications can be made to the embodiments without substantially departing from the principles of the present disclosure. All such changes and modifications are intended to be included within the scope of the present disclosure herein. Accordingly, the subject matter disclosed above is considered illustrative and not restrictive, and examples of embodiments are intended to cover all such modifications, enhancements, and other embodiments that fall within the spirit and scope of the present disclosure. Thus, to the maximum extent permitted by law, the scope of the present disclosure will be determined by the broadest permissible interpretation of the present disclosure (including examples of embodiments and their equivalents) and should not be limited or restricted by the foregoing detailed description.
Claims
1. An audio decoding method, the method comprises: A decoder generates (1000) a spectrum on the basis of sub - frames, wherein consecutive sub - frames of an audio signal have the following characteristics: the window shape applied to the first sub - frame in the consecutive sub - frames is a mirror version or a time - reversed version of the window shape applied to the second sub - frame in the consecutive sub - frames, and stores (1004) the signal spectrum corresponding to the second sub - frame. The audio decoding method further comprises: In response to frame loss, obtaining (1006) a previously generated signal spectrum corresponding to the second sub - frame; Detecting (1008) peaks of the signal spectrum and estimating (1012) the phase of each of the peaks; Calculating (1014) a phase adjustment for each detected peak based on the estimated phase; Adjusting the detected peaks by applying (1016) the phase adjustment to the peak intervals of each detected peak to form phase - adjusted peak intervals, and taking the complex conjugate of the phase - adjusted peak intervals to form time - reversed phase - adjusted peaks; and Combining (1020) the time - reversed phase - adjusted peak intervals with the noise component of the spectrum derived from the non - peak intervals of the signal spectrum to form a combined spectrum for a first hidden sub - frame for hiding an audio frame.
2. The method according to claim 1, wherein, The synthesized hidden audio frame comprises two consecutive hidden sub - frames. The method further comprises: Combining (1028) the phase - adjusted peak intervals with the non - peak intervals of the signal spectrum to form a combined spectrum for a second hidden sub - frame of the hidden audio frame.
3. The method according to claim 1 or 2, further comprises: Associating each of the detected peaks with a plurality of peak frequency intervals representing the peak.
4. The method according to any one of claims 1 - 3, wherein, The phase adjustment for the peaks of the hidden audio sub - frame is calculated according to the following formula: Δφ = -2φ 0 -2πf(N step21 +N lost ·N) / N, where φ 0 is the estimated phase of the peak and f is the frequency of the peak, N lost represents the number of consecutively lost frames, N represents the length of the full frame and N step21 is the sample distance between the start of the second subframe of the last received frame and the start of the first subframe of the hidden audio frame.
5. The method according to any one of claims 1 - 4, wherein, Adjusting the detected peaks by applying the phase adjustment to the peak intervals of each detected peak to form phase - adjusted peak intervals, and taking the complex conjugate of the phase - adjusted peak intervals to form time - reversed phase - adjusted peaks is according to the following formula: where * is the complex conjugate, is the signal spectrum, and Δφ i is the phase adjustment.
6. The method according to any one of claims 1 - 5, wherein, Estimating the phase of each of the peaks comprises: Calculating a phase estimate for the peaks in the time - reversed phase - adjusted peaks according to the following formula: where φ i is the estimated phase at frequency f i , is the spectrum of the previously received audio signal at the angular frequency interval k i , f frac is the rounding error, and φ C is the tuning constant.
7. An audio decoder (900), which is configured to: generate a spectrum on the basis of sub - frames, wherein, Consecutive sub - frames of an audio signal have the following characteristics: the window shape applied to the first sub - frame in the consecutive sub - frames is a mirror version or a time - reversed version of the window shape applied to the second sub - frame in the consecutive sub - frames, and stores the signal spectrum corresponding to the second sub - frame. The audio decoder is further configured to: In response to frame loss, obtain a previously generated signal spectrum corresponding to the second sub - frame; Detect the peaks of the signal spectrum and estimate the phase of each of the peaks; Calculate a phase adjustment for each detected peak based on the estimated phase; Adjust the detected peaks by applying the phase adjustment to the peak intervals of each detected peak to form phase-adjusted peak intervals and taking the complex conjugate of the phase-adjusted peak intervals to form time-reversed phase-adjusted peaks; And Combine the time-reversed phase-adjusted peak intervals with the noise components of the spectrum derived from the non-peak intervals of the signal spectrum to form a combined spectrum for a first hidden sub-frame for hiding an audio frame.
8. The audio decoder according to claim 7, wherein, The synthesized hidden audio frame includes two consecutive hidden sub-frames, and the audio decoder is further configured to: Combine the phase-adjusted peak intervals with the non-peak intervals of the signal spectrum to form a combined spectrum for a second hidden sub-frame of the hidden audio frame.
9. The audio decoder according to claim 7 or 8, further configured to: Associate each of the detected peaks with a plurality of peak frequency intervals representing the peak.
10. The audio decoder according to any one of claims 7-9, wherein, The audio decoder is configured to calculate the phase adjustment for the peaks of the hidden audio sub-frame according to the following formula: Δφ = -2φ 0 -2πf(N step21 +N lost ·N) / N, where φ 0 is the estimated phase of the peak and f is the frequency of the peak, N lost represents the number of consecutive lost frames, N represents the length of a full frame and N step21 is the sample distance between the start of the second sub-frame of the last received frame and the start of the first sub-frame of the hidden audio frame.
11. The audio decoder according to any one of claims 7-10, wherein, Adjusting the detected peaks by applying the phase adjustment to the peak intervals of each detected peak to form phase-adjusted peak intervals and taking the complex conjugate of the phase-adjusted peak intervals to form time-reversed phase-adjusted peaks is according to the following formula: where * is the complex conjugate, is the signal spectrum, and Δφ i is the phase adjustment.
12. The audio decoder according to any one of claims 7-11, wherein, Estimating the phase of each of the peaks includes: Calculating a phase estimate for the peaks in the time-reversed phase-adjusted peaks according to the following formula: f frac = f i - k i where φ i is the estimated phase at frequency f i , is the spectrum of the previously received audio signal at the angle in the frequency interval k i , f frac is the rounding error, and φ C is the tuning constant.
13. A user equipment, comprising the audio decoder according to claim 7.