Frame loss concealment for low frequency effects channel
LPC-based frame loss concealment techniques address the rumble issue in LFE channels by generating substitute frames using audio filters and resonators, ensuring high-quality audio reproduction even with multiple frame losses.
Patent Information
- Application Number
- JP2025197871
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-27
- Filing Date
- 2025-11-19
- Publication Date
- 2026-02-18
AI Technical Summary
Existing frame loss concealment techniques for the low-frequency effects (LFE) channel in multi-channel audio, such as 5.1 or 7.1 audio, produce an unpleasant low-frequency rumble when applied to the LFE channel due to their non-optimization for very low frequency content.
A method using linear predictive coding (LPC)-based frame loss concealment that involves generating a substitute frame by determining an audio filter from preceding valid audio frames, applying bandwidth sharpening to the LPC synthesis filter, and using a resonator to extend samples into the lost audio frame.
The method effectively generates a substitute frame that maintains audio quality without the unpleasant rumble, suitable for LFE channels and other audio signals, and can handle multiple consecutive frame losses with appropriate attenuation.
Smart Images

Figure 2026027497000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following priority applications: U.S. Provisional Application No. 63 / 037,673 (Reference No. D20058USP1), filed June 11, 2020, and U.S. Provisional Application No. 63 / 193,974 (Reference No. D20058USP2), filed May 27, 2021, which are incorporated herein by reference.
[0002] technology This disclosure generally relates to a method and apparatus for frame loss concealment for a low frequency effects (LFE) channel. More specifically, this disclosure relates to linear predictive coding (LPC)-based frame loss concealment for the LFE channel of a multi-channel audio signal. The presented techniques may be applied, for example, to 3GPP IVAS coding.
[0003] Although some embodiments are described herein with particular reference to that disclosure, it will be understood that the disclosure is not limited to such fields of use but is applicable in a broader context. [Background technology]
[0004] Any discussion of background art throughout this disclosure should in no way be taken as an admission that such art is widely known or forms part of the common general knowledge in the art.
[0005] LFE (Low Frequency Effects) is the low-frequency effects channel in multi-channel audio, such as 5.1 or 7.1 audio. This channel is intended to drive the subwoofer in the loudspeaker playback system for such multi-channel audio. As the term LFE implies, this channel is expected to provide only bass information, with a typical upper frequency limit of 120 Hz.
[0006] However, this frequency limit may not necessarily be very sharp: in practice it may happen that the LFE channel contains some higher frequency components, for example up to 400 Hz or even 700 Hz.
[0007] Multi-channel audio may be rendered through stereo headphones. Specific rendering techniques are used to create an equivalent sound experience as if the multi-channel audio were listened to through a multi-speaker system. This also applies to the LFE channel. In that case, the appropriate rendering techniques ensure that the sound experience of the LFE channel is close to that of a subwoofer system used for playback. Because the LFE channel typically has very limited frequency content, it can be encoded and transmitted at a relatively low bit rate. One suitable coding technique for LFE is transform-based coding using the modified discrete cosine transform (MDCT). This technique allows for the representation of LFE at bit rates of, for example, approximately 2000-4000 bits per second.
[0008] One particular circumstance in multi-channel audio transmission, especially over wireless channels, is that the transmission is prone to errors. The transmission is typically packet-based, and a transmission error can result in one or more complete coded frames of multi-channel audio being erased. There are so-called packet or frame loss concealment techniques used by multi-channel audio decoding systems that aim to make the effect of lost audio frames as inaudible as possible.
[0009] For normal signal channels of multi-channel audio, there are well-established frame loss concealment techniques. For example, a set of suitable techniques is part of the 3GPP® EVS codec [3GPP® TS26.447].
[0010] For an MDCT-encoded LFE channel, the same techniques are in principle applicable. For example, it is possible to reuse MDCT coefficients from the most recent valid audio frame and use these coefficients after gain scaling (attenuation) and code prediction or randomization. The EVS standard also provides other techniques, such as reconstructing missing audio frames in the time domain according to a sinusoidal approach.
[0011] The main problem with applying these current techniques to the LFE channel is that they are not designed or optimized for very low frequency content: they are very powerful for audio channels with normal frequency content, but when applied to the LFE channel they produce an unpleasant low frequency rumble.
[0012] Thus, the purpose of this disclosure is to describe a novel technique that overcomes the problems and limitations of prior art frame loss concealment techniques applied to the LFE channel, although the scope of application of the novel method is not limited to the LFE channel. Summary of the Invention [Means for solving the problem]
[0013] According to a first aspect of the present disclosure, a method for generating a substitute frame for a lost audio frame of an audio signal is presented. The method may include determining an audio filter based on samples of a valid audio frame preceding the lost audio frame. The method may include generating a substitute frame based on the audio filter and the samples of the valid audio frame preceding the lost audio frame. Generating a substitute frame based on the audio filter and the samples of the valid audio frame may include initializing a filter memory of the audio filter with samples of the valid audio frame. The method may include determining a modified audio filter based on the audio filter. The modified audio filter may replace the audio filter, and generating a substitute frame based on the audio filter may include generating a substitute frame based on the modified audio filter and the samples of the valid audio frame.
[0014] The audio filter may be an all-pole filter. The audio filter may be a linear predictive coding (LPC) synthesis filter. The audio filter may be derived from an all-pass filter operated on at least the samples of the valid frame.
[0015] The method may include determining the audio filter based on a denominator polynomial of a transfer function of an all-pass filter. The step of determining the modified audio filter may include bandwidth sharpening. The bandwidth sharpening may be applied such that the duration of the impulse response of the modified audio filter is extended relative to the duration of the impulse response of the audio filter. The bandwidth sharpening may be applied such that the distance between a pole of the modified audio filter and a unit circle is shortened compared to the distance between the corresponding pole of the audio filter and the unit circle. The bandwidth sharpening may be applied such that the pole of the modified audio filter having the largest magnitude is equal to or at least close to 1. The bandwidth sharpening may be applied such that the frequency of the pole of the modified audio filter having the largest magnitude is equal to the frequency of the pole of the audio filter having the largest magnitude.
[0016] The method may include determining the pole magnitudes and frequencies of the audio filter using a root-finding method. Bandwidth sharpening may be applied such that the pole magnitudes of the modified audio filter are set equal to or at least close to 1, and the pole frequencies of the modified audio filter are identical to the pole frequencies of the audio filter. The pole magnitudes of the modified audio filter may be set equal to or at least close to 1 only if the corresponding pole magnitudes of the audio filter have magnitudes above a threshold.
[0017] The method may include determining filter coefficients of the audio filter. γ Optionally, applying bandwidth sharpening using a bandwidth sharpening factor such that (z)=S(z / γ), where S γwhere σ represents the transfer function of the modified audio filter, S represents the transfer function of the audio filter, and γ represents a bandwidth sharpening factor. The method may include generating a substitute frame based on filter coefficients of the audio filter, samples of a valid audio frame preceding the lost audio frame, and a bandwidth sharpening factor γ. The bandwidth sharpening factor may be determined in an iterative procedure by gradually increasing and / or decreasing the bandwidth sharpening factor. The method may include checking whether the poles of the modified audio filter are within a unit circle by converting the polynomial coefficients of the modified audio filter into reflection coefficients. In this case, converting the polynomial coefficients of the modified audio filter into reflection coefficients may be based on a backward Levinson recursion. The bandwidth sharpening factor may be determined such that the pole of the modified audio filter with the largest magnitude is moved as close to the unit circle as possible, and at the same time, all poles of the modified audio filter are located within the unit circle. The substitute frame is determined using the formula
number
number
number
[0018] The method may include determining filter coefficients of an audio filter and applying bandwidth sharpening by reducing a distance between a pair of line spectral frequencies representing the audio filter coefficients, thereby generating modified line spectral frequencies. The method may include deriving coefficients of the modified audio filter from the modified line spectral frequencies. The method may include generating a substitute frame based on the filter coefficients of the modified audio filter and samples of a valid audio frame preceding the lost audio frame.
[0019] The lost audio packets may be associated with a low frequency effect (LFE) channel of a multi-channel audio signal. In particular, the lost audio packets may have been transmitted over a wireless channel from a transmitter to a receiver. The method may be performed at the receiver.
[0020] The method can include downsampling the samples of the valid audio frame before generating the alternate samples for the alternate frame. The method can include upsampling the alternate samples for the alternate frame after generating the alternate frame.
[0021] Multiple audio frames may be lost, and the method may include determining a first modified audio filter by scaling audio filter coefficients of an audio filter with a first bandwidth sharpening factor. The method may include determining a second modified audio filter by scaling the audio filter coefficients with a second bandwidth sharpening factor. The method may include generating substitute frames based on the first modified audio filter for the first M lost audio frames. The method may include generating substitute frames based on the second modified audio filter for the (M+1)th lost audio frame and all subsequent lost audio frames, such that the audio signal is attenuated for the latter frames.
[0022] The method may include dividing an audio signal into a first subband signal and a second subband signal. The method may include generating a first subband audio filter for the first subband signal. The method may include generating a first subband alternate frame based on the first subband audio filter. The method may include generating a second audio filter for the second subband signal. The method may include generating a second subband alternate frame based on the second subband audio filter. The method may include generating an alternate frame by combining the first subband alternate frame and the second subband alternate frame.
[0023] The audio filter may be configured to operate as a resonator. The resonator may be tuned on samples of a valid audio frame preceding the lost audio frame. The resonator may be initially excited with at least one sample of the valid audio frame preceding the lost audio frame. A substitute frame may be generated by using ringing of the resonator to extend the at least one sample into the lost audio frame.
[0024] According to a second aspect of the present disclosure, there is provided a system, which may include one or more processors and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of the method described above.
[0025] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable medium, the non-transitory computer-readable medium may store instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of the method described above. [Brief explanation of the drawings]
[0026] Exemplary embodiments of the present disclosure will now be described with reference to the accompanying drawings. [Figure 1] 1 is a flowchart of an example process for frame loss concealment. [Figure 2] 1 illustrates an exemplary mobile device architecture for implementing the functions and processes described in this document. DETAILED DESCRIPTION OF THE INVENTION
[0027] One key idea of this disclosure is to extrapolate samples of a lost audio frame from the most recent valid audio samples by implementing a resonator. The resonator is tuned on the most recent valid audio samples and then operated to extend those audio samples into the lost audio frame. As an example, if the most recent valid audio samples are a sine wave in frequency and phase, a suitable resonator would be an oscillator tuned to extend that sine wave into the lost audio frame.
[0028] In this example, the most recent valid signal can be expressed as:
number
number
[0029] One possible implementation of this resonator is the following all-pass filter:
number
number
number
[0030] In other words, the extrapolated sample is constructed as the ringing of a resonator filter that is first excited with the most recent audio sample, which thus determines the initial filter state memory, and then lets the filter ring (oscillate) by itself, i.e. without any further (non-zero) input.
[0031] The sample extrapolation approach described above would be possible if the signal could be well approximated by a sine wave, however this still requires identifying the sine wave frequency and the resonant frequency of the resonator.
[0032] A more general approach that overcomes the limitation to a single sinusoid and also solves the problem of determining the resonant frequency of the resonator is to apply a linear prediction (LPC) approach. Linear prediction synthesis filter ringing has traditionally been used in frame-based analysis-by-synthesis speech coding systems, where the LPC filter excitation for the current frame is calculated taking into account the synthesis filter ringing of the previous frame. LPC synthesis filter ringing has also been used to extrapolate a small number of samples in case of ACELP codec mode switching, where a small number of future samples are unavailable [3GPP TS 26.445].
[0033] Similar to the all-pass filter described above, the filter H(z) is constructed as follows:
number
[0034] The approach to extrapolating the signal samples is similar to that for the oscillator described above:
number
number
[0035] In particular, the analysis filter A(z) can be generated / determined by traditional approaches such as the Levinson-Durbin approach. The all-pass filter H(z) can be constructed from A(z) as described above. In case of frame loss, the synthesis filter part of H(z), i.e., the LPC synthesis filter S(z) = σ / A(z), can be used to construct a replacement frame for the lost frame.
[0036] Furthermore, as will be described later, it is noted that the LPC approach solves the problem of determining the resonant frequency of a resonator: one property of LPC analysis, well known from speech coding, is that the frequency response of the corresponding LPC synthesis filter matches the speech format. In general, this means that the synthesis filter matches, at its resonant frequency, the dominant spectral component (dominant frequency) of the decomposed input signal. Thus, the LPC approach is suitable for determining a resonator with a matching resonant frequency.
[0037] A drawback of the LPC synthesis filter ringing approach is that the impulse response of the LPC synthesis filter typically decays very quickly (almost exponentially). Therefore, this approach is not sufficient to generate a replacement frame for a 20 ms lost audio frame. For several consecutive lost frames, multiple 20 ms of the replacement signal must be generated accordingly. A typical LPC synthesis filter is already faded out and cannot generate a useful replacement signal.
[0038] To overcome this limitation, the LPC synthesis filter does not have to be used as is, as calculated using standard techniques such as the Levinson / Durbin approach. Rather, by bandwidth sharpening, the filter is modified so that its poles are as close to the unit circle as possible while still maintaining stability. According to one such approach, the poles of the LPC synthesis filter are calculated using standard root-finding techniques. The original pole locations z i =r i ·exp(jω i ), the pole size r i is replaced by a magnitude equal to 1, or at least close to 1. The effect of this operation is to i =f s ω i The filter response for / 2π does not fade out, while the pole frequencies are maintained. A slight modification of this method is that only poles with magnitudes above a certain threshold, say 0.75, are moved towards the unit circle.
[0039] A practical drawback of the described method may be the numerical complexity required for root finding in some implementations. One way to avoid that processing step is to take the given LPC synthesis filter and modify it by a bandwidth sharpening factor γ as follows: S γ (z)=S(z / γ) This operation has the effect of moving all of the filter poles towards the unit circle by the factor γ. However, since the pole locations are unknown, it is possible that a given factor γ is too large, and at least the poles with the largest magnitude will be moved outside the unit circle, resulting in an unstable filter. Thus, after applying a given factor γ, it is possible to check whether the filter has become unstable or whether it is still stable. If the filter is unstable, a smaller γ is selected; otherwise, a larger γ is selected. This procedure can be repeated iteratively (using nested interval techniques) until a bandwidth sharpening factor γ is found that brings the filter very close to instability but still stable.
[0040] In particular, other filter bandwidth sharpening techniques can also be used, such as line-spectral frequency-based sharpening. In this approach, LPC filter coefficients are represented as line-spectral frequency pairs. The sharpening effect is achieved by reducing the distance between pairs of line-spectral frequencies. If the distance is reduced to zero, this is equivalent to moving the filter poles onto the unit circle or pushing the filter to its stability limit. The corresponding modified filter, represented by the modified line-spectral frequencies, can then be represented again by LPC coefficients obtained by back-transforming from the modified line-spectral frequencies to modified LPC coefficients.
[0041] The above LPC-based approach can be summarized as follows: In a first step, an audio filter (which may be seen as a resonator) may be tuned to a previously received and / or reconstructed audio signal (e.g., an LFE audio signal). For example, the LPC coefficients a i , i=1...P may be calculated. Tuning to the previously received signal and / or the reconstructed signal may be performed such that the audio filter resulting in this step has characteristics (e.g., resonant frequency) that are based on (e.g., derived from) the previously received signal and / or the reconstructed signal.
[0042] The corresponding bandwidth sharpening of the LPC synthesis filter is performed by the modified synthesis filter S crit (z)=S(z / γ crit ) where γ crit is chosen so that the LPC filter is at its stability limit. Alternatively, line-spectral frequency-based sharpening can be used. The LPC synthesis filter memory stores the most recent samples of the previously received and / or reconstructed audio signal.
number
number
[0043] The filter stability check in the above procedure can be performed by converting the polynomial coefficients of the modified LPC synthesis filter into reflection coefficients. This can be done using a backward Levinson recursion. The reflection coefficients allow a straightforward stability test: if any of the absolute values of the reflection coefficients are greater than or equal to 1, the filter is unstable; otherwise, it is guaranteed to be stable.
[0044] For implementation reasons, it may be advantageous to perform the above operations in the sub-sampled domain, e.g., using f instead of the original sampling frequency of 48000 Hz, under the assumption that the LFE signal has no significant frequency content above 800 Hz. sUsing a sampling frequency of 1 / 30 = 1600 Hz, it is possible to perform the described frame loss concealment operation in the sub-sampled domain. This allows, for example, to reduce the memory required to store the preceding valid samples by a corresponding factor of 1 / 30 = 1600 Hz / 48000 Hz. The complexity of certain numerical operations is also reduced by the same factor. Under the assumption that the LFE signal is sufficiently band-limited, no further filtering is required before sub-sampling. However, during up-sampling to the original sampling frequency, after calculating the substitute samples, the application of corresponding interpolation filtering, typically a linear-phase low-pass filter, is necessary. The delay introduced by the filter may be taken into account, and a corresponding additional number of substitute samples must be calculated.
[0045] The LPC filter order of P=20 is set to the sampling frequency f s It has been found suitable in practical implementations to be operated in the sub-sampling domain using = 1600 Hz.
[0046] Another factor to consider in frame loss concealment for MDCT-based coding is that the frame to be restored may need to be prepared to match the specific realization of its (lapped) MDCT transform. This means that after applying the frame loss concealment techniques described above, the replacement samples can be windowed and then transformed into the time-folded domain. The time-folded domain transform may then be inverted, and the resulting signal frame is then subjected to the time-reversed window. Note that time folding and unfolding can be combined into one step. After these operations, the recovered frame can be combined with the remainder of the previous (valid) frame to generate replacement samples for the erased frame. Depending on the MDCT frame size and window shape, as well as the interpolation filters mentioned above, this may require reconstructing more samples in the described manner than can be expected by the coding system's nominal stride or frame size, which may be, for example, 20 milliseconds.
[0047] A particular case is when several consecutive frames are lost in a row. In principle, the above process remains unchanged even if the frame loss is the second, third, etc., of a series. The previous frame recovered by the described technique can simply be treated as if it were a valid frame received without error. Alternatively, the ringing can simply be extended to the next lost frame, whereby the resonator or (modified) synthesis filter parameters are maintained from the initial calculation for the first frame loss. However, after a very long burst of frame losses (e.g., more than 10 consecutive frames corresponding to 200 ms), it is advantageous for the listener to start muting the substitute signal. Otherwise, the listener may be confused by a seemingly endless substitute signal despite the connection interruption.
[0048] A particular inventive method suitable for muting is to modify the bandwidth sharpening factor found according to the steps described above. The found factor γ ensures that the modified synthesis filter S(z / γ) produces a sustained alternative signal, but for muting, γ is further modified (scaled) to ensure proper attenuation. This has the effect of moving the poles of the modified synthesis filter towards the inside of the unit circle by the scaling factor, thus exponentially decaying the synthesis filter response.
[0049] For example, 3 dB attenuation (att_per_frame) per 20 ms frame (flen=0.02 s) is desired, and the synthesis filter is set at the sampling frequency f s = 1600Hz, the following scaling factors would apply:
number
number
[0050] Note that in general, muting should only be initiated after a very long burst of frame losses, e.g., after 10 consecutive frame losses. That is, only then should γ be greater than γ mute is replaced by
[0051] The above-described embodiment of the present invention is based on the assumption that the signal on which frame loss concealment is performed is the LFE channel of a multi-channel audio signal. However, similar principles can be applied to any audio signal without bandwidth limitations. One obvious possibility is to perform the operation using a full-band approach at the signal's nominal sampling frequency. However, this can face practical difficulties, especially when using an LPC approach. For a sampling frequency of 48 kHz, it can be difficult to find a sufficiently high-order LPC filter that can adequately represent the spectral characteristics of the extended signal. The difficulties can be both numerical (to calculate a sufficiently high-order LPC filter) and conceptual. A conceptual difficulty can be that low frequencies may require a longer LPC analysis window than high frequencies.
[0052] One effective way to address these challenges is to implement the described operation using a subband / split-band approach. To this end, the initial full-band signal is divided into several subband signals by a bank of analysis filters, each representing a partial frequency band. The split-band approach can be combined with the use of specific quadrature mirror filtering and subsampling (QMF approach), which offers advantages in terms of complexity (due to critical sampling) and memory savings. After the analysis filter operation, which yields the subband signals, the frame loss concealment technique described above can be applied to all subband signals in parallel. This approach allows for the use of a wider LPC analysis window, especially for the low-frequency band than for the high-frequency band, thereby making the LPC approach frequency-selective.
[0053] After the frame loss concealment operation on the initial subbands, the subbands are recombined into a full-band substitute signal. In the case of QMF, the QMF synthesis also includes upsampling and QMF interpolation filtering.
[0054] interpretation Unless otherwise indicated, and as will be apparent from the discussion that follows, discussions throughout this disclosure using terms such as "processing," "computing," "calculating," "determining," "analyzing," and the like will be understood to refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulates and / or transforms data represented as physical quantities, e.g., electronic quantities, into other data also represented as physical quantities.
[0055] Similarly, the term "processor" may refer to any device or portion of a device for processing electronic data, e.g., from registers and / or memory, to form other electronic data, which may be stored, e.g., in registers and / or memory. A "computer" or "calculating machine" or "computing platform" may include one or more processors.
[0056] The methods described herein, in an exemplary embodiment, can be performed by one or more processors that accept computer-readable (also called machine-readable) code, including a set of instructions that, when executed by one or more of the processors, perform at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify actions to be performed is included. Thus, one example is a typical processing system including one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem including main RAM and / or static RAM and / or ROM. A bus subsystem for communicating between components may also be included. The processing system may also be a distributed processing system with processors coupled by a network. If the processing system requires a display, such a display may be included, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data entry is required, the processing system may also include one or more input devices, such as an alphanumeric input unit, such as a keyboard, a pointing control device, such as a mouse, etc. The processing system may also include a storage system, such as a disk drive unit. The processing system in some configurations may include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium that carries computer-readable code (e.g., software), which includes a set of instructions that, when executed by one or more processors, cause the processor to perform one or more of the methods described herein. Note that when a method includes several elements, e.g., several steps, no ordering of such elements is implied unless specifically stated.The software may reside on a hard disk, or, during its execution by a computer system, may reside, completely or at least partially, in RAM and / or in the processor. Thus, the memory and processor also constitute a computer-readable carrier medium carrying computer-readable code. Furthermore, the computer-readable carrier medium may form or be included in a computer program product.
[0057] In alternative exemplary embodiments, the one or more processors may operate as standalone devices or, in a networked deployment, may be connected, e.g., networked to other processors, and the one or more processors may operate in the capacity of a server or user machine in a server-user network environment, or as a peer machine in a peer-to-peer or distributed network environment. The one or more processors may form a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a cellular telephone, a web appliance, a network router, switch, or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine.
[0058] It should be noted that the term "machine" is also intended to include any collection of machines that individually or jointly execute a set (or sets) of instructions for performing any one or more of the methodologies discussed herein.
[0059] Thus, one exemplary embodiment of each method described herein is in the form of a computer-readable carrier medium carrying a set of instructions, e.g., a computer program for execution on one or more processors, e.g., one or more processors that are part of a web server configuration. Thus, as will be appreciated by those skilled in the art, exemplary embodiments of the present disclosure may be embodied as a method, an apparatus such as a special purpose device, an apparatus such as a data processing system, or a computer-readable carrier medium, e.g., a computer program product. The computer-readable carrier medium carries computer-readable code, including a set of instructions that, when executed on one or more processors, cause the processor(s) to perform the method. Thus, aspects of the present disclosure may take the form of a method, an entirely hardware exemplary embodiment, an entirely software exemplary embodiment, or an exemplary embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.
[0060] The software may also be transmitted or received over a network via a network interface device. While the carrier medium is a single medium in the exemplary embodiment, the term "carrier medium" should be taken to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term "carrier medium" should also be taken to include any medium that can store, encode, or carry a set of instructions for execution by one or more of the processors, causing the one or more processors to perform any one or more of the methods disclosed herein. Carrier media can take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus subsystem. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications. For example, the term "carrier medium" should be taken to include, but is not limited to, computer products embodied in solid-state memory, optical and magnetic media; media carrying propagated signals detectable by at least one processor or one or more processors and representing a set of instructions that, when executed, implement a method; and transmission media within a network carrying propagated signals detectable by at least one of the one or more processors and representing a set of instructions.
[0061] It will be understood that the steps of the methods discussed are, in an exemplary embodiment, performed by a suitable processor(s) of a processing (e.g., computer) system executing instructions (computer-readable code) stored on a storage device. It will also be understood that the present disclosure is not limited to any particular implementation or programming technique, and that the present disclosure may be implemented using any suitable technique for implementing the functions described herein. The present disclosure is not limited to any particular programming language or operating system.
[0062] Throughout this disclosure, a reference to "one exemplary embodiment," "some exemplary embodiments," or "an exemplary embodiment" means that a particular feature, structure, or characteristic described in connection with that exemplary embodiment is included in at least one exemplary embodiment of the disclosure. Thus, the appearances of the phrases "in one exemplary embodiment," "in some exemplary embodiments," or "in an exemplary embodiment" in various places throughout this disclosure are not necessarily all referring to the same exemplary embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more exemplary embodiments, as would be apparent to one of ordinary skill in the art from this disclosure.
[0063] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe a common object merely indicates that different instances of similar objects are being referred to, and is not intended to imply that the objects so described must be in a given order, temporally, spatially, in ranking, or in any other way.
[0064] In the claims and the description herein, any of the terms include, including, or having are open terms including at least the recited elements / features but not excluding others. Thus, when used in the claims, the terms include / have should not be interpreted as being limited to the recited means, elements, or steps. For example, the scope of the expression "device having A and B" should not be limited to a device consisting only of elements A and B. As used herein, any of the terms include, include, or include are also open terms including at least the recited elements / features but not excluding others. Thus, including is synonymous with having and means having.
[0065] In the foregoing description of exemplary embodiments of the present disclosure, it should be understood that various features of the present disclosure may be grouped together in a single exemplary embodiment, figure, or description thereof for the purpose of improving the flow of the disclosure and aiding in understanding one or more of its various inventive aspects. This method of disclosure, however, should not be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed exemplary embodiment. Accordingly, the claims following this specification are hereby expressly incorporated herein, with each claim standing on its own as a separate exemplary embodiment of the present disclosure.
[0066] Furthermore, although some exemplary embodiments described herein may include some features included in other exemplary embodiments but not others, combinations of features from different exemplary embodiments are intended to be within the scope of the present disclosure and to form different exemplary embodiments, as will be understood by those skilled in the art. For example, in the following claims, any of the exemplary embodiments recited in the claims can be used in any combination.
[0067] In the description provided herein, numerous specific details are set forth. However, it will be understood that exemplary embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this document.
[0068] Thus, while what is believed to be the best mode of the disclosure has been described, those skilled in the art will recognize that other and further modifications may be made without departing from the spirit of the disclosure, and it is intended to claim all such changes and modifications as falling within the scope of the present disclosure. For example, any formulas set forth above merely represent procedures that may be used. Functions may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to methods described within the scope of the present disclosure.
[0069] Finally, FIG. 1 shows a flowchart of an exemplary process for frame loss concealment. This exemplary process may be performed, for example, by a mobile device architecture 800 shown in FIG. 2. Architecture 800 can be implemented in any electronic device, including, but not limited to, a desktop computer, consumer audio / visual (AV) equipment, radio broadcasting equipment, and mobile devices (e.g., smartphones, tablet computers, laptop computers, wearable devices). In the illustrated exemplary embodiment, architecture 800 is for a smartphone and includes processor(s) 801, peripherals interface 802, audio subsystem 803, loudspeaker 804, microphone 805, sensors 806 (e.g., accelerometer, gyro, barometer, magnetometer, camera), position processor 807 (e.g., GNSS receiver), wireless communication subsystem 808 (e.g., Wi-Fi, Bluetooth, cellular), and I / O subsystem 809, which includes touch controller 810 and other input controllers 811, touch surface 812, and other input / control devices 813. Other architectures having more or fewer components may also be used to implement the disclosed embodiments.
[0070] Memory interface 814 is coupled to processor 801, peripherals interface 802, and memory 815 (e.g., Flash, RAM, ROM). Memory 815 stores computer program instructions and data, including, but not limited to, operating system instructions 816, communications instructions 817, GUI instructions 818, sensor processing instructions 819, telephony instructions 820, electronic messaging instructions 821, web browsing instructions 822, audio processing instructions 823, GNSS / navigation instructions 824, and applications / data 825. Audio processing instructions 823 include instructions for performing the audio processing described with reference to FIG. 1.
[0071] Aspects of the systems described herein may be implemented in a suitable computer-based sound processing network environment for processing digital or digitized audio files. Portions of the adaptive audio system may include one or more networks containing any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route data transmitted between computers. Such networks may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0072] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described in terms of their behavior, register transfers, logical components, and / or other characteristics using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media. The computer-readable media on which such formatted data and / or instructions may be embodied include, but are not limited to, various forms of physical (non-transitory) non-volatile storage media, such as optical, magnetic, or semiconductor storage media.
[0073] While one or more implementations have been described by way of example and in terms of specific embodiments, it is to be understood that the one or more implementations are not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0074] Itemized Exemplary Embodiments Various aspects and implementations of the present invention can also be seen from the following non-claimed enumerated example embodiments (EEE). [EEE1] Methods to recover lost audio frames: tuning a resonator to a sample of a valid audio frame preceding the lost audio frame; adapting the resonator to operate as an oscillator according to samples of the valid audio frame; extending the audio signal generated by the oscillator into the lost audio frame. The resonator may correspond to the audio filter H(z) described above, and the oscillator may correspond to the term S(z)=σ / A(z) described above. [EEE2] The method of EEE1, wherein the resonator / oscillator combination is constructed using linear prediction (LPC) techniques and the oscillator is implemented as an LPC synthesis filter. [EEE3] The method according to EEE2, wherein the LPC synthesis filter is modified using bandwidth sharpening. [EEE4] The LPC synthesis filter is modified using a bandwidth sharpening factor γ, resulting in a modified filter S γ (z)=S(z / γ) The method according to EEE3, wherein [EEE5] The method of EEE4, wherein the bandwidth sharpening factor is selected such that the modified LPC synthesis filter is close to unstable but still stable. [EEE6] 8. The method according to any one of EEE1 to EEE5, wherein the method is operated in the sub-sampling domain. [EEE7] A method for recovering frames from a sequence of consecutive audio frame losses, comprising: applying a first modified LPC synthesis filter with a sharpening factor γ for n consecutive frame losses, where n is less than a threshold M; For the kth consecutive frame loss, a further modified sharpening factor γ mute gradually muting other frame losses in the sequence using a second modified LPC synthesis filter using γ mute The sharpening factor γ is converted into a factor α mute and a step scaled by method. [EEE8] the threshold M and the scaling factor α mute is selected such that a muting behavior is achieved with a 3 dB attenuation per 20 ms audio frame, starting from the 10th consecutive frame loss, as described in EEE7. [EEE9] 9. The method of any one of EEE1 to 8, wherein the method is applied to a low frequency effects (LFE) channel of a multi-channel audio signal. [EEE10] one or more processors; a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations recited in any one of EEE1 to EEE8. system. [EEE11] A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations recited in any one of EEE1 to EEE9.
Claims
1. 1. A method for generating substitute frames for a plurality of lost audio frames of an audio signal, the method comprising: determining an audio filter based on samples of valid audio frames preceding the plurality of lost audio frames; determining a first modified audio filter by scaling audio filter coefficients of said audio filter with a first bandwidth sharpening factor; determining a second modified audio filter by scaling the audio filter coefficients with a second bandwidth sharpening factor; generating substitute frames based on the first modified audio filter for a first M lost audio frames of the plurality of lost audio frames; generating substitute frames based on the second modified audio filter for the (M+1)th lost audio frame and all subsequent lost audio frames of the plurality of lost audio frames, such that the audio signal is attenuated for the latter frames; method.
2. one or more processors; a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operations of the method of claim 1; system.
3. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of the method of claim 1.