Wireless communication device, related method and memory.
The decoder uses mid-channel signals and quantized stereo parameters to align and reconstruct audio signals, addressing alignment challenges in multi-microphone setups and improving decoding accuracy under noisy conditions.
Patent Information
- Application Number
- BR112019023204
- Authority / Receiving Office
- BR · BR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-04-25
- Filing Date
- 2018-04-27
- Publication Date
- 2026-07-28
- Estimated Expiration
- 2038-04-27
AI Technical Summary
Existing audio decoding technologies face challenges in accurately aligning audio signals received from multiple microphones due to delays, leading to inefficient encoding and decoding processes, especially under noisy transmission conditions.
A decoder is configured to receive and decode mid-channel signals with stereo parameters, allowing it to generate left and right channels even when frames are unavailable, using quantized values to ensure precise alignment and reconstruction of audio signals.
This approach enhances the accuracy of audio signal reconstruction by compensating for delays and ensuring high-quality decoding even in noisy conditions, maintaining audio fidelity.
Smart Images

Figure 00000114_0000 
Figure 00000115_0000 
Figure 00000116_0000
Abstract
Description
WIRELESS COMMUNICATION DEVICE, RELATED METHOD AND MEMORY I. Priority Claim
[0001] This application claims the benefit of priority from U.S. Provisional Patent Application No. 62 / 505,041, filed May 11, 2017, entitled STEREO PARAMETERS FOR STEREO CODING and U.S. Non-Provisional Patent Application No. 15 / 962,834, filed April 25, 2018, entitled STEREO PARAMETERS FOR STEREO CODING, which are jointly owned, the content of each of the aforementioned applications being expressly incorporated herein by reference in its entirety. II. Field
[0002] The present invention relates generally to the decoding of audio signals. III. Description of Related Technique
[0003] Technological advancements have resulted in smaller and more powerful computing devices. For example, there is now a variety of portable personal computing devices, including cordless phones such as cell phones and smartphones, tablets, and laptops that are small, lightweight, and easy for users to carry. These devices can communicate voice and data packets over wireless networks. In addition, many of these devices incorporate additional functionalities such as a digital still camera, a digital video camera, a digital recorder, and an audio file player. Furthermore, these devices can process executable instructions, including software applications such as a web browser application, which can be Petition 870250038999, dated 05 / 13 / 2025, page 13 / 22 2 / 105 is used to access the Internet. Therefore, these devices can use significant computing resources.
[0004] A computing device may include or be coupled with multiple microphones to receive audio signals. Generally, a sound source is closer to a first microphone than to a second microphone of multiple microphones. Consequently, a second audio signal received from the second microphone may be delayed relative to a first audio signal received from the first microphone due to the respective distances of the microphones from the sound source. In other implementations, the first audio signal may be delayed relative to the second audio signal. In stereo encoding, audio signals from microphones may be encoded to generate a mid-channel signal and one or more side-channel signals. The mid-channel signal may correspond to the sum of the first audio signal and the second audio signal. A side-channel signal may correspond to the difference between the first audio signal and the second audio signal.The first audio signal may not be aligned with the second audio signal due to a delay in receiving the second audio signal relative to the first. This delay can be indicated by a coded offset value (e.g., a stereo parameter) that is transmitted to a decoder. Precise alignment of the first audio signal with the second audio signal allows for efficient encoding for transmission to the decoder. However, the transmission of high-precision data indicating the alignment of the signals is also crucial. Petition 870190112962, dated 05 / 11 / 2019, page 7 / 134 3 / 105 audio signals utilize high transmission capabilities compared to low-precision data transmission. Other stereo parameters indicative of the characteristics between the first and second audio signals can also be encoded and transmitted to the decoder.
[0005] The decoder can reconstruct the first and second audio signals based at least on the mid-channel signal and stereo parameters that are received at the decoder via a bitstream that includes a sequence of frames. The accuracy at the decoder during audio signal reconstruction can be based on the accuracy of the encoder. For example, the high-precision offset value encoded can be received at the decoder and can allow the decoder to reproduce the delay in reconstructed versions of the first audio signal and the second audio signal with high accuracy. If the offset value is unavailable at the decoder, such as when a data frame transmitted via the bitstream is corrupted due to noisy transmission conditions, the offset value can be requested and retransmitted to the decoder to allow accurate reproduction of the delay between the audio signals.For example, the decoder's accuracy in reproducing the delay may exceed a limitation of audible perceptibility for humans to perceive a variation in the delay. IV. Summary
[0006] According to one implementation of the present invention, an apparatus includes a receiver Petition 870190112962, dated 05 / 11 / 2019, p. 8 / 134 The 4 / 105 unit is configured to receive at least a portion of a bitstream. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a mid-channel and a first value of a stereo parameter, and the second frame includes a second portion of the mid-channel and a second value of the stereo parameter. The unit also includes a decoder configured to decode the first portion of the mid-channel to generate a first portion of a decoded mid-channel. The decoder is also configured to generate a first portion of a left channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter, and to generate a first portion of a right channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter.The decoder is further configured to, in response to the second frame being unavailable for decoding operations, generate a second portion of the left channel and a second portion of the right channel based at least on the first value of the stereo parameter. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0007] According to another implementation, a method of decoding a signal involves receiving at least a portion of a bitstream. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a mid-channel and a first value of a stereo parameter, and the second frame Petition 870190112962, dated 05 / 11 / 2019, page 9 / 134 5 / 105 includes a second portion of the mid-channel and a second value of the stereo parameter. The method also includes decoding the first portion of the mid-channel to generate a first portion of a decoded mid-channel. The method further includes generating a first portion of a left channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter, and generating a first portion of a right channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter. The method also includes, in response to the second frame being unavailable for decoding operations, generating a second portion of the left channel and a second portion of the right channel based at least on the first value of the stereo parameter. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0008] According to another implementation, a non-transient computer-readable medium includes instructions that, when executed by a processor within a decoder, cause the processor to perform operations including receiving at least a portion of a bitstream. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a mid-channel and a first value of a stereo parameter, and the second frame includes a second portion of the mid-channel and a second value of the stereo parameter. The operations also include decoding the first portion of the mid-channel to generate a first portion of a Petition 870190112962, dated 05 / 11 / 2019, page 10 / 134 6 / 105 decoded mid-channel. The operations also include generating a first portion of a left channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter, and generating a first portion of a right channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter. The operations also include, in response to the second frame being unavailable for decoding operations, generating a second portion of the left channel and a second portion of the right channel based at least on the first value of the stereo parameter. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0009] According to another implementation, an apparatus includes means for receiving at least a portion of a bitstream. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a mid-channel and a first value of a stereo parameter, and the second frame includes a second portion of the mid-channel and a second value of the stereo parameter. The apparatus also includes means for decoding the first portion of the mid-channel to generate a first portion of a decoded mid-channel. The apparatus further includes means for generating a first portion of a left channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter, and means for generating a first portion of a right channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter. The Petition 870190112962, dated 05 / 11 / 2019, page 11 / 134 The 7 / 105 device also includes means for generating, in response to the second frame being unavailable for decoding operations, a second portion of the left channel and a second portion of the right channel based at least on the first value of the stereo parameter. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0010] According to another implementation, an apparatus includes a receiver configured to receive at least a portion of a bitstream from an encoder. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a mid-channel and a first value of a stereo parameter. The second frame includes a second portion of the mid-channel and a second value of the stereo parameter. The apparatus also includes a decoder configured to decode the first portion of the mid-channel to generate a first portion of a decoded mid-channel. The decoder is also configured to perform a transformation operation on the first portion of the decoded mid-channel to generate a first portion of a decoded frequency-domain mid-channel.The decoder is further configured to upmix the first portion of the decoded mid-frequency domain channel to generate a first portion of a left-frequency domain channel and a first portion of a right-frequency domain channel. The decoder is also configured to generate a first portion of a left channel based at least on the first portion of the... Petition 870190112962, dated 05 / 11 / 2019, page 12 / 134 The decoder is further configured to generate a first portion of a right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter. The decoder is also configured to determine that the second frame is unavailable for decoding operations. The decoder is further configured to generate, based at least on the first value of the stereo parameter, a second portion of the left channel and a second portion of the right channel in response to the determination that the second frame is unavailable. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0011] According to another implementation, a method for decoding a signal involves receiving, in a decoder, at least a portion of a bitstream from an encoder. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a mid-channel and a first value of a stereo parameter. The second frame includes a second portion of the mid-channel and a second value of the stereo parameter. The method also includes decoding the first portion of the mid-channel to generate a first portion of a decoded mid-channel. The method further includes performing a transformation operation on the first portion of the decoded mid-channel to generate a first portion of a decoded frequency-domain mid-channel. The method Petition 870190112962, dated 05 / 11 / 2019, page 13 / 134 Method 9 / 105 also includes upmixing the first portion of the decoded mid-frequency domain channel to generate a first portion of a left frequency domain channel and a first portion of a right frequency domain channel. The method further includes generating a first portion of a left channel based at least on the first portion of the left frequency domain channel and the first value of the stereo parameter. The method further includes generating a first portion of a right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter. The method also includes determining that the second frame is unavailable for decoding operations. The method further includes generating, based at least on the first value of the stereo parameter, a second portion of the left channel and a second portion of the right channel in response to the determination that the second frame is unavailable.The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0012] According to another implementation, a non-transient computer-readable medium includes instructions that, when executed by a processor within a decoder, cause the processor to perform operations including receiving at least a portion of a bitstream from an encoder. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a medium channel and a first value of a stereo parameter. The second frame includes a second Petition 870190112962, dated 05 / 11 / 2019, page 14 / 134 The operations also include decoding the first portion of the mid-channel to generate a first portion of a decoded mid-channel. The operations further include performing a transformation operation on the first portion of the decoded mid-channel to generate a first portion of a decoded frequency-domain mid-channel. The operations also include upmixing the first portion of the decoded frequency-domain mid-channel to generate a first portion of a left frequency-domain channel and a first portion of a right frequency-domain channel. The operations further include generating a first portion of a left channel based at least on the first portion of the left frequency-domain channel and the first value of the stereo parameter.The operations also include generating a first portion of a right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter. The operations also include determining that the second frame is unavailable for decoding operations. The operations further include generating, based at least on the first value of the stereo parameter, a second portion of the left channel and a second portion of the right channel in response to the determination that the second frame is unavailable. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0013] According to another implementation, a device includes means for receiving at least a portion of Petition 870190112962, dated 05 / 11 / 2019, page 15 / 134 11 / 105 a bitstream of an encoder. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a mid-channel and a first value of a stereo parameter. The second frame includes a second portion of the mid-channel and a second value of the stereo parameter. The apparatus also includes means for decoding the first portion of the mid-channel to generate a first portion of a decoded mid-channel. The apparatus also includes means for performing a transformation operation on the first portion of the decoded mid-channel to generate a first portion of a decoded frequency-domain mid-channel. The apparatus also includes means for upmixing the first portion of the decoded frequency-domain mid-channel to generate a first portion of a left frequency-domain channel and a first portion of a right frequency-domain channel.The apparatus also includes means for generating a first portion of a left channel based at least on the first portion of the left frequency domain channel and the first value of the stereo parameter. The apparatus also includes means for generating a first portion of a right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter. The apparatus also includes means for determining that the second frame is unavailable for decoding operations. The apparatus also includes means for generating, based at least on the first value of the stereo parameter, a second portion of the left channel and a second portion of the right channel. Petition 870190112962, dated 05 / 11 / 2019, page 16 / 134 12 / 105 response to a determination that the second frame is unavailable. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0014] According to another implementation, a device includes a receiver and a decoder. The receiver is configured to receive a bit stream that includes an encoded middle channel and a quantized value representing an offset between a reference channel associated with an encoder and a target channel associated with the encoder. The quantized value is based on an offset value. The offset value is associated with the encoder and has a higher precision than the quantized value. The decoder is configured to decode the encoded middle channel to generate a decoded middle channel and to generate a first channel based on the decoded middle channel. The decoder is further configured to generate a second channel based on the decoded middle channel and the quantized value. The first channel corresponds to the reference channel and the second channel corresponds to the target channel.
[0015] According to another implementation, a method of decoding a signal involves receiving, in a decoder, a bit stream including a middle channel and a quantized value representing an offset between a reference channel associated with an encoder and a target channel associated with the encoder. The quantized value is based on a value of the offset. The value is associated with the encoder and has a higher precision than the value Petition 870190112962, dated 05 / 11 / 2019, page 17 / 134 13 / 105 quantized. The method also includes decoding the average channel to generate a decoded average channel. The method further includes generating a first channel based on the decoded average channel and generating a second channel based on the decoded average channel and the quantized value. The first channel corresponds to the reference channel and the second channel corresponds to the target channel.
[0016] According to another implementation, a non-transient computer-readable medium includes instructions that, when executed by a processor within a decoder, cause the processor to perform operations including receiving, in a decoder, a bit stream including a middle channel and a quantized value representing an offset between a reference channel associated with an encoder and a target channel associated with the encoder. The quantized value is based on a value of the offset. The value is associated with the encoder and has a higher precision than the quantized value. The operations also include decoding the middle channel to generate a decoded middle channel. The operations further include generating a first channel based on the decoded middle channel and generating a second channel based on the decoded middle channel and the quantized value. The first channel corresponds to the reference channel and the second channel corresponds to the target channel.
[0017] According to another implementation, a device includes means for receiving, in a decoder, a bit stream including a medium channel and a quantized value representing a shift between a channel of Petition 870190112962, dated 05 / 11 / 2019, page 18 / 134 14 / 105 reference associated with an encoder and a target channel associated with the encoder. The quantized value is based on an offset value. The value is associated with the encoder and has a higher precision than the quantized value. The apparatus also includes means for decoding the average channel to generate a decoded average channel. The apparatus further includes means for generating a first channel based on the decoded average channel and means for generating a second channel based on the decoded average channel and the quantized value. The first channel corresponds to the reference channel and the second channel corresponds to the target channel.
[0018] According to another implementation, an apparatus includes a receiver configured to receive a bitstream from an encoder. The bitstream includes a mid-channel and a quantized value representing an offset between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value is based on a value of the offset that has a higher precision than the quantized value. The apparatus also includes a decoder configured to decode the mid-channel to generate a decoded mid-channel. The decoder is also configured to perform a transformation operation on the decoded mid-channel to generate a decoded frequency-domain mid-channel. The decoder is further configured to upmix the decoded frequency-domain mid-channel to generate a first frequency-domain channel and a second frequency-domain channel. The decoder is also configured to generate Petition 870190112962, dated 05 / 11 / 2019, page 19 / 134 15 / 105 a first channel based on the first frequency domain channel. The first channel corresponds to the reference channel. The decoder is further configured to generate a second channel based on the second frequency domain channel. The second channel corresponds to the target channel. The second frequency domain channel is shifted in the frequency domain by the quantized value if the quantized value corresponds to a frequency domain shift, and a time domain version of the second frequency domain channel is shifted by the quantized value if the quantized value corresponds to a time domain shift.
[0019] According to another implementation, one method involves receiving, in a decoder, a bitstream from an encoder. The bitstream includes a mid-channel and a quantized value representing an offset between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value is based on a value of the offset that has a higher precision than the quantized value. The method also includes decoding the mid-channel to generate a decoded mid-channel. The method further includes performing a transformation operation on the decoded mid-channel to generate a decoded frequency-domain mid-channel. The method also includes upmixing the decoded frequency-domain mid-channel to generate a first frequency-domain channel and a second frequency-domain channel. The method also includes generating a first channel based on the first frequency-domain channel. Petition 870190112962, dated 05 / 11 / 2019, page 20 / 134 16 / 105 frequency. The first channel corresponds to the reference channel. The method also includes generating a second channel based on the second frequency domain channel. The second channel corresponds to the target channel. The second frequency domain channel is shifted in the frequency domain by the quantized value if the quantized value corresponds to a frequency domain shift, and a time domain version of the second frequency domain channel is shifted by the quantized value if the quantized value corresponds to a time domain shift.
[0020] According to another implementation, a non-transient computer-readable medium includes instructions for decoding a signal. The instructions, when executed by a processor within a decoder, lead the processor to perform operations including receiving a bitstream from an encoder. The bitstream includes a medium channel and a quantized value representing an offset between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value is based on a value of the offset that has a higher precision than the quantized value. The operations also include decoding the medium channel to generate a decoded medium channel. The operations further include performing a transformation operation on the decoded medium channel to generate a decoded frequency domain medium channel.The operations also include upmixing the decoded mid-frequency domain channel to generate a first frequency domain channel and a second frequency domain channel. The operations also... Petition 870190112962, dated 05 / 11 / 2019, page 21 / 134 17 / 105 includes generating a first channel based on the first frequency domain channel. The first channel corresponds to the reference channel. The operations also include generating a second channel based on the second frequency domain channel. The second channel corresponds to the target channel. The second frequency domain channel is shifted in the frequency domain by the quantized value if the quantized value corresponds to a frequency domain shift, and a time domain version of the second frequency domain channel is shifted by the quantized value if the quantized value corresponds to a time domain shift.
[0021] According to another implementation, an apparatus includes means for receiving a bitstream from an encoder. The bitstream includes a mid-channel and a quantized value representing an offset between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value is based on a value of the offset that has greater precision than the quantized value. The apparatus also includes means for decoding the mid-channel to generate a decoded mid-channel. The apparatus also includes means for performing a transformation operation on the decoded mid-channel to generate a decoded frequency-domain mid-channel. The apparatus also includes means for upmixing the decoded frequency-domain mid-channel to generate a first frequency-domain channel and a second frequency-domain channel. The apparatus also includes means for generating a first channel Petition 870190112962, dated 05 / 11 / 2019, page 22 / 134 18 / 105 based on the first frequency domain channel. The first channel corresponds to the reference channel. The device also includes means for generating a second channel based on the second frequency domain channel. The second channel corresponds to the target channel. The second frequency domain channel is shifted in the frequency domain by the quantized value if the quantized value corresponds to a frequency domain shift, and a time domain version of the second frequency domain channel is shifted by the quantized value if the quantized value corresponds to a time domain shift.
[0022] Other implementations, advantages and features of the present invention will become apparent after analysis of the entire application, including the following sections: Brief Description of the Drawings, Detailed Description and the Claims. V. Brief Description of the Drawings
[0023] Figure 1 is a block diagram of a particular illustrative example of a system that includes an operable decoder for estimating stereo parameters for missing frames and for decoding audio signals using quantized stereo parameters;
[0024] Figure 2 is a diagram illustrating the decoder of Figure 1;
[0025] Figure 3 is a diagram of an illustrative example of predictive stereo parameters for a missing frame in a decoder;
[0026] Figure 4A is an illustrative example not Petition 870190112962, dated 05 / 11 / 2019, page 23 / 134 19 / 105 limiting factor of a method for decoding an audio signal;
[0027] Figure 4B is a non-limiting illustrative example of a more detailed version of the audio signal decoding method in Figure 4A;
[0028] Figure 5A is another illustrative, non-limiting example of a method for decoding an audio signal;
[0029] Figure 5B is a non-limiting illustrative example of a more detailed version of the audio signal decoding method in Figure 5A;
[0030] Figure 6 is a block diagram of a particular illustrative example of a device that includes a decoder for estimating stereo parameters for missing frames and for decoding audio signals using quantized stereo parameters; and
[0031] Figure 7 is a block diagram of a base station that is operable for estimating stereo parameters for missing frames and for decoding audio signals using quantized stereo parameters. VI. Detailed Description
[0032] Particular aspects of the present invention are described below with reference to the drawings. In the description, common features are referred to as common reference numbers. As used herein, various terminologies are used for the purpose of describing specific implementations only and are not intended to limit implementations. For example, the singular forms a / an and the / a are also intended to include also Petition 870190112962, dated 05 / 11 / 2019, p. 24 / 134 20 / 105 as plural forms, unless the context clearly indicates otherwise. It may still be understood that the terms comprehends and comprehending may be used interchangeably with includes or including. Furthermore, it will be understood that the term in which may be used interchangeably with where. As used here, an ordinal term (e.g., first, second, third, etc.) used to modify an element, such as a structure, a component, an operation, etc., does not by itself indicate any priority or order of the element in relation to another element, but only distinguishes the element from another element with the same name (but for use of the ordinal term). As used here, the term set refers to one or more of a specific element, and the term plurality refers to multiples (e.g., two or more) of a specific element.
[0033] In the present invention, terms such as determine, calculate, move, adjust, etc. may be used to describe how one or more operations are performed. It should be noted that these terms should not be interpreted as limiting, and other techniques may be used to perform similar operations. Additionally, as referred to in this document, generate, calculate, use, select, access, and determine may be used interchangeably. For example, generating, calculating, or determining a parameter (or a signal) may refer to actively generating, calculating, or determining the parameter (or signal), or it may refer to using, selecting, or accessing the parameter (or signal) that has already been generated. Petition 870190112962, dated 05 / 11 / 2019, page 25 / 134 21 / 105 is generated, just like by another component or device.
[0034] Operable systems and devices for encoding multiple audio signals are disclosed. A device may include an encoder configured to encode multiple audio signals. Multiple audio signals may be captured simultaneously in time using multiple recording devices, for example, multiple microphones. In some examples, multiple audio signals (or multichannel audio) may be generated synthetically (e.g., artificially) by multiplexing several audio channels recorded at the same time or at different times. As illustrative examples, simultaneous recording or multiplexing of audio channels may result in a 2-channel configuration (i.e., stereo: left and right), a 5.1-channel configuration (left, right, center, left surround, right surround, and low-frequency emphasis (LEE) channels), a 7.1-channel configuration, a 7.1+4-channel configuration, a 22-channel configuration.2 channels or an N-channel configuration.
[0035] Audio capture devices in teleconferencing rooms (or telepresence rooms) may include multiple microphones that acquire spatial audio. Spatial audio may include speech as well as background audio that is encoded and transmitted. Speech / audio from a given source (e.g., a speaker) may reach the multiple microphones at different times, depending on how the microphones are arranged, as well as where the source (e.g., the speaker) is located relative to the microphones. Petition 870190112962, dated 05 / 11 / 2019, page 26 / 134 22 / 105 microphones and room dimensions. For example, a sound source (e.g., a speaker) may be closer to a first microphone associated with the device than to a second microphone associated with the device. Therefore, a sound emitted from the sound source may reach the first microphone sooner than the second microphone. The device may receive a first audio signal through the first microphone and a second audio signal through the second microphone.
[0036] Mid-side (MS) coding and parametric stereo (PS) coding are stereo coding techniques that can provide greater efficiency compared to dual-mono coding techniques. In dual-mono coding, the left channel (or signal) (L) and the right channel (or signal) (R) are coded independently, without using inter-channel correlation. MS coding reduces redundancy between a pair of correlated L / R channels by transforming the left and right channels into a sum channel and a difference channel (e.g., a side channel) before coding. The sum signal and the difference signal are coded or encoded waveforms based on an MS coding model. Relatively more bits are spent on the sum signal than on the side signal. PS coding reduces redundancy in each sub-band by transforming the L / R signals into a sum signal and a set of side parameters.Lateral parameters can indicate an interchannel intensity difference (IID), an interchannel phase difference (IPD), and an interchannel time difference (ITD). Petition 870190112962, dated 05 / 11 / 2019, page 27 / 134 23 / 105 lateral or residual prediction gains, etc. The summation signal is a waveform encoded and transmitted along with the lateral parameters. In a hybrid system, the lateral channel can be the waveform encoded in the lower bands (e.g., less than 2 kilohertz (kHz)) and PS encoded in the upper bands (e.g., greater than or equal to 2 kHz), where interchannel phase preservation is perceptually less critical. In some implementations, PS encoding can be used in the lower bands as well to reduce interchannel redundancy before waveform encoding.
[0037] MS coding and PS coding can be done in the frequency domain, subband domain, or time domain. In some instances, the left and right channels may not be correlated. For example, the left and right channels may include uncorrelated synthetic signals. When the left and right channels are uncorrelated, the coding efficiency of MS coding, PS coding, or both, can approach the coding efficiency of dual-mono coding.
[0038] Depending on the recording setup, there may be a time shift between a left and right channel, as well as other spatial effects such as echo and room reverberation. If the time shift and phase mismatch between channels are not compensated for, the sum channel and the difference channel may contain comparable energies, reducing the coding gains associated with MS or PS techniques. The reduction Petition 870190112962, dated 05 / 11 / 2019, page 28 / 134 24 / 105 gains in encoding can be based on the amount of time (or phase) shift. The comparable energies of the sum signal and the difference signal may limit the use of MS encoding in certain frames where channels are temporally shifted but are highly correlated. In stereo encoding, a middle channel (e.g., a sum channel) and a side channel (e.g., a difference channel) can be generated based on the following formula: M=(L+R / 2, S=(LR) / 2 (Formula 1)
[0039] where M corresponds to the middle canal, S corresponds to the lateral canal, L corresponds to the left canal, and R corresponds to the right canal.
[0040] In some cases, the mid-canal and lateral canal can be generated based on the following formula: M=c(L+R), S=c(LR) Formula 2
[0041] where c corresponds to a complex value that is frequency dependent. The generation of the mid-channel and side-channel based on Formula 1 or Formula 2 can be referred to as downmixing. A reverse process of generating the left-channel and right-channel from the mid-channel and side-channel based on Formula 1 or Formula 2 can be referred to as upmixing.
[0042] In some cases, the average channel may be based on other formulas, such as: M=(L+gDR) / 2 Formula 3 M = g²L + g²R Formula 4
[0043] where g2+ g2= 1.0, and where gD is a gain parameter. In other examples, the downmix can be Petition 870190112962, dated 05 / 11 / 2019, p. 29 / 134 25 / 105 performed in bands, where mid(b) = ciL(b) + C2R(b), where Ci and C2 are complex numbers, where (b) = C3L(b)C4R(b), and where C3 and C4 are complex numbers.
[0044] An ad-hoc approach used to choose between MS encoding or dual-mono encoding for a specific frame might involve generating an average signal and a side signal, calculating the energies of the average and side signals, and determining whether or not to perform MS encoding based on the energies. For example, MS encoding might be performed in response to the determination that the ratio of side signal and average signal energies is less than a threshold. To illustrate, if a right channel is shifted at least once (e.g., about 0.001 seconds or 48 samples at 48 kHz), a first energy of the average signal (corresponding to a sum of the left and right signals) might be comparable to a second energy of the side signal (corresponding to a difference between the left and right signals) for speech frames with voice.When the first energy is comparable to the second energy, a larger number of bits can be used to encode the side channel, thus reducing the encoding efficiency of MS encoding compared to dual-mono encoding. Dual-mono encoding can therefore be used when the first energy is comparable to the second energy (e.g., when the ratio between first energy and second energy is greater than or equal to the threshold). In an alternative approach, the decision between MS encoding and dual-mono encoding for a specific frame can be made based on a one-to-one comparison. Petition 870190112962, dated 05 / 11 / 2019, page 30 / 134 26 / 105 limit and normalized cross-correlation values of the left and right channels.
[0045] In some examples, the encoder may determine a mismatch value indicative of an amount of temporal misalignment between the first audio signal and the second audio signal. As used herein, a time offset value, an offset value, and a mismatch value may be used interchangeably. For example, the encoder may determine a time offset value indicative of an offset (e.g., temporal mismatch) of the first audio signal relative to the second audio signal. The temporal mismatch value may correspond to an amount of time delay between the reception of the first audio signal at the first microphone and the reception of the second audio signal at the second microphone. Furthermore, the encoder may determine the temporal mismatch value on a frame-by-frame basis, for example, based on each 20-millisecond (ms) speech / audio frame.For example, the time mismatch value could correspond to the amount of time that a second frame of the second audio signal is delayed relative to a first frame of the first audio signal. Alternatively, the time mismatch value could correspond to the amount of time that the first frame of the first audio signal is delayed relative to the second frame of the second audio signal.
[0046] When the sound source is closer to Petition 870190112962, dated 05 / 11 / 2019, page 31 / 134 27 / 105 If the sound source is closer to the first microphone than to the second microphone, the frames of the second audio signal may be delayed relative to the frames of the first audio signal. In this case, the first audio signal may be referred to as the reference audio signal or reference channel, and the delayed second audio signal may be referred to as the target audio signal or target channel. Alternatively, when the sound source is closer to the second microphone than to the first microphone, the frames of the first audio signal may be delayed relative to the frames of the second audio signal. In this case, the second audio signal may be referred to as the reference audio signal or reference channel, and the delayed first audio signal may be referred to as the target audio signal or target channel.
[0047] Depending on where the sound sources (e.g., speakers) are located in a conference room or telepresence system, or how the position of the sound source (e.g., speaker) changes relative to the microphones, the reference channel and the target channel may change from frame to frame; similarly, the time delay value may also change from frame to frame. However, in some implementations, the time mismatch value may always be positive to indicate an amount of delay of the target channel relative to the reference channel. Furthermore, the time mismatch value may correspond to a non-causal offset value, whereby the delayed target channel is pushed back in time such that the target channel is aligned (e.g., maximally aligned) with the reference channel. Petition 870190112962, dated 05 / 11 / 2019, page 32 / 134 28 / 105 reference. The downmix algorithm for determining the mid-channel and side-channel can be run on the reference channel and the non-causal offset target channel.
[0048] The encoder can determine the time mismatch value based on the reference audio channel and a plurality of time mismatch values applied to the target audio channel. For example, a first frame from the reference audio channel, X, can be received at a first time (mx). A specific first frame from the target audio channel, Y, can be received at a second time (ηχ) corresponding to a first time mismatch value, for example, offsetl = ηχ-πίχ. Furthermore, a second frame from the reference audio channel can be received at a third time (np). A specific second frame from the target audio channel can be received at a fourth time (ηχ) corresponding to a second time mismatch value, for example, offset! = ηχ-πρ
[0049] The device may execute a framing or buffering algorithm to generate a frame (e.g., 20 ms samples) at a first sampling rate (e.g., 32 kHz sampling rate (i.e., 640 samples per frame)). The encoder may, in response to the determination that a first frame of the first audio signal and a second frame of the second audio signal arrive at the same time at the device, estimate a temporal mismatch value (e.g., offset) equal to zero samples. A channel Petition 870190112962, dated 05 / 11 / 2019, page 33 / 134 The left 29 / 105 channel (e.g., corresponding to the first audio signal) and the right channel (e.g., corresponding to the second audio signal) may be temporarily aligned. In some cases, the left and right channels, even when aligned, may differ in power due to various reasons (e.g., microphone calibration).
[0050] In some instances, the left and right channels may be temporarily misaligned due to various reasons (for example, a sound source, such as a speaker, may be closer to one microphone than the other, and the two microphones may be further apart than a certain threshold (for example, 1 to 20 centimeters)). The location of the sound source relative to the microphones may introduce different delays in the left and right channels. In addition, there may be a difference in gain, a difference in power, or a difference in level between the left and right channels.
[0051] In some examples, where there are more than two channels, a reference channel is initially selected based on the channel levels or energies and subsequently refined based on the time mismatch values between different pairs of channels, for example, tl(ref, ch2), t2(ref, ch3), t3(ref, ch4), ..., where chi is the initial reference channel, tl(.), t2(.), etc. are the functions to estimate the mismatch values. If all time mismatch values are positive, then chi is treated as the channel. Petition 870190112962, dated 05 / 11 / 2019, pp. 34 / 134 30 / 105 of the reference. If any of the mismatch values is a negative value, then the reference channel is reconfigured to the channel that was associated with a mismatch value that resulted in a negative value, and the above process will continue until the best selection (e.g., based on the maximum number of maximally uncorrelated side channels) of the reference channel is achieved. A hysteresis can be used to overcome any sudden variation in the selection of the reference channel.
[0052] In some examples, the arrival time of audio signals at the microphones of multiple sound sources (e.g., announcers) may vary when multiple announcers are speaking alternately (e.g., without overlapping). In this case, the encoder can dynamically adjust a temporal mismatch value based on the announcer to identify the reference channel. In some other examples, multiple announcers may be speaking at the same time, which may result in variable temporal mismatch values depending on which announcer is speaking loudest, closest to the microphone, etc. In this case, the identification of the target and reference channels may be based on the variable temporal offset values in the current frame and the temporal mismatch values estimated in previous frames, and based on the energy or temporal evolution of the first and second audio signals.
[0053] In some examples, the first audio signal and the second audio signal may be synthesized or Petition 870190112962, dated 05 / 11 / 2019, page 35 / 134 31 / 105 artificially generated when the two signals potentially show less (e.g., no) correlation. It should be understood that the examples described here are illustrative and may be instructive in determining a relationship between the first audio signal and the second audio signal in similar or different situations.
[0054] The encoder can generate comparison values (e.g., difference values or cross-correlation values) based on a comparison of a first frame of the first audio signal and a plurality of frames of the second audio signal. Each frame in the plurality of frames can correspond to a specific temporal mismatch value. The encoder can generate an estimated first temporal mismatch value based on the comparison values. For example, the estimated first temporal mismatch value might correspond to a comparison value indicating a higher temporal similarity (or smaller difference) between the first frame of the first audio signal and a corresponding first frame of the second audio signal.
[0055] The encoder can determine a final temporal mismatch value by refining, in multiple stages, a series of estimated temporal mismatch values. For example, the encoder can first estimate an experimental temporal mismatch value based on comparison values generated from pre-processed and resampled stereo versions of the first audio signal and the second audio signal. The Petition 870190112962, dated 05 / 11 / 2019, page 36 / 134 The 32 / 105 encoder can generate interpolated comparison values associated with temporal mismatch values close to the estimated experimental temporal mismatch value. The encoder can determine a second estimated interpolated temporal mismatch value based on the interpolated comparison values. For example, the second estimated interpolated temporal mismatch value might correspond to a specific interpolated comparison value that indicates greater temporal similarity (or less difference) than the remaining interpolated comparison values and the first estimated experimental temporal mismatch value.If the second estimated interpolated temporal mismatch value of the current frame (e.g., the first frame of the first audio signal) differs from the final temporal mismatch value of a previous frame (e.g., a frame of the first audio signal preceding the first frame), then the interpolated temporal mismatch value of the current frame is still altered to improve the temporal similarity between the first audio signal and the shifted second audio signal. In particular, a third estimated altered temporal mismatch value may correspond to a more accurate measurement of temporal similarity by searching around the second estimated interpolated temporal mismatch value of the current frame and the final estimated temporal mismatch value of the previous frame. The third estimated altered temporal mismatch value is still present. Petition 870190112962, dated 05 / 11 / 2019, page 37 / 134 33 / 105 conditioned to estimate the final temporal mismatch value, limiting any spurious changes in the temporal mismatch value between frames and further controlled to not switch from a negative temporal mismatch value to a positive temporal mismatch value (or vice versa) in two successive (or consecutive) frames, as described here.
[0056] In some instances, the encoder may fail to switch between a positive temporal mismatch value and a negative temporal mismatch value or vice versa in consecutive frames or in adjacent frames. For example, the encoder may set the final temporal mismatch value to a particular value (e.g., 0) indicating no temporal shift based on the estimated altered or interpolated temporal mismatch value of the first frame and a corresponding estimated final or altered or interpolated temporal mismatch value in a particular frame preceding the first frame.To illustrate, the encoder can set the final temporal mismatch value of the current frame (e.g., the first frame) to indicate no temporal shift, i.e., shiftl = 0, in response to the determination that one of the altered or interpolated or experimentally estimated temporal mismatch values of the current frame is positive and the other of the altered or interpolated or experimentally estimated temporal mismatch values of the previous frame (e.g., the frame). Petition 870190112962, dated 05 / 11 / 2019, page 38 / 134 34 / 105 preceding the first frame) is negative. Alternatively, the encoder can also set the final temporal mismatch value of the current frame (e.g., the first frame) to indicate no temporal shift, i.e., shiftl = 0, in response to the determination that one of the altered or interpolated or experimentally estimated temporal mismatch values of the current frame is negative and the other of the altered or interpolated or experimentally estimated temporal mismatch values of the previous frame (e.g., the frame preceding the first frame) is positive.
[0057] The encoder can select a frame from the first audio signal or the second audio signal as a reference or target based on the time mismatch value. For example, in response to the determination that the final time mismatch value is positive, the encoder can generate a reference signal or channel indicator having a first value (e.g., 0) indicating that the first audio signal is a reference signal and that the second audio signal is the target signal. Alternatively, in response to the determination that the final time mismatch value is negative, the encoder can generate the reference signal or channel indicator having a second value (e.g., 1) indicating that the second audio signal is the reference signal and that the first audio signal is the target signal.
[0058] The encoder can estimate a gain Petition 870190112962, dated 05 / 11 / 2019, page 39 / 134 35 / 105 relative (e.g., a relative gain parameter) associated with the reference signal and the non-causally shifted target signal. For example, in response to the determination that the final temporal mismatch value is positive, the encoder may estimate a gain value to normalize or equalize the amplitude or power levels of the first audio signal relative to the second audio signal that is shifted by the non-causally shifted temporal mismatch value (e.g., an absolute value of the final temporal mismatch value). Alternatively, in response to the determination that the final temporal mismatch value is negative, the encoder may estimate a gain value to normalize or equalize the power or amplitude levels of the first non-causally shifted audio signal relative to the second audio signal.In some examples, the encoder may estimate a gain value to normalize or equalize the amplitude or power levels of the reference signal relative to the non-causal shifted target signal. In other examples, the encoder may estimate the gain value (e.g., a relative gain value) based on the reference signal relative to the target signal (e.g., the non-shifted target signal).
[0059] The encoder can generate at least one encoded signal (e.g., an average signal, a side signal, or both) based on the reference signal, the target signal, the non-causal time mismatch value, and the relative gain parameter. In other implementations, the encoder can generate at least one encoded signal (by Petition 870190112962, dated 05 / 11 / 2019, page 40 / 134 36 / 105 example, a mid-channel, a side-channel, or both) based on the reference channel and the target channel adjusted for temporal mismatch. The side-channel signal can correspond to a difference between first samples of the first frame of the first audio signal and selected samples of a selected frame of the second audio signal. The encoder can select the selected frame based on the final temporal mismatch value. Fewer bits can be used to encode the side-channel signal due to the reduced difference between the first samples and the selected samples compared to other samples of the second audio signal that correspond to a frame of the second audio signal that is received by the device at the same time as the first frame.A device transmitter can transmit at least one encoded signal, the non-causal temporal mismatch value, the relative gain parameter, the signal or reference channel indicator, or a combination thereof.
[0060] The encoder can generate at least one encoded signal (e.g., a mid-signal, a side signal, or both) based on the reference signal, the target signal, the non-causal time mismatch value, the relative gain parameter, low-band parameters of a particular frame of the first audio signal, high-band parameters of the particular frame, or a combination thereof. The particular frame may precede the first frame. Certain low-band parameters, high-band parameters, or a combination thereof, from one or more previous frames may be used to encode. Petition 870190112962, dated 05 / 11 / 2019, page 41 / 134 37 / 105 an average signal, a side signal, or both, from the first frame. Encoding the average signal, the side signal, or both, based on low-band parameters, high-band parameters, or a combination thereof, can improve estimates of the non-causal temporal mismatch value and inter-channel relative gain parameter. Low-band parameters, high-band parameters, or a combination thereof, may include a pitch parameter, a voice parameter, an encoder type parameter, a low-band power parameter, a high-band power parameter, a slope parameter, a pitch gain parameter, a FCB gain parameter, a coding mode parameter, a voice activity parameter, a noise estimate parameter, a signal-to-noise ratio parameter, a formant parameter, a speech / music decision parameter, the non-causal offset, the inter-channel gain parameter, or a combination thereof.A device transmitter can transmit at least one encoded signal, the non-causal temporal mismatch value, the relative gain parameter, the reference channel (or signal) indicator, or a combination thereof. In the present invention, terms such as determine, calculate, shift, adjust, etc., may be used to describe how one or more operations are performed. It should be noted that these terms should not be interpreted as limiting and that other techniques may be used to perform similar operations.
[0061] According to some implementations, the final temporal mismatch value (e.g., a Petition 870190112962, dated 05 / 11 / 2019, page 42 / 134 38 / 105 offset value) is a non-quantized value indicating the true offset between a target channel and a reference channel. Although all digital values are quantized due to the precision provided by the system that stores or uses the digital value, as used here, digital values are quantized if generated by a quantization operation to reduce a precision of the digital value (e.g., to reduce a range or bandwidth associated with the digital value) and are non-quantized otherwise. As a non-limiting example, the first audio signal could be the target channel and the second audio signal could be the reference channel. If the true offset between the target channel and the reference channel is thirty-seven samples, the target channel can be offset by thirty-seven samples in the encoder to generate an offset target channel, which is temporarily aligned with the reference channel.In other implementations, both channels can be shifted such that the relative offset between the channels is equal to the final offset value (37 samples in this example). This relative channel shift by the offset value achieves the effect of temporal alignment of the channels. A high-efficiency encoder can align the channels as much as possible to reduce coding entropy and thus increase coding efficiency, because coding entropy is sensitive to changes in offset between channels. The shifted target channel and the reference channel can be used to generate an average channel that is encoded and transmitted to one. Petition 870190112962, dated 05 / 11 / 2019, page 43 / 134 39 / 105 decoder as part of a bitstream. Additionally, the final time mismatch value can be quantized and transmitted to the decoder as part of the bitstream. For example, the final time mismatch value can be quantized using a floor of four, such that the quantized final time mismatch value is equal to nine (e.g., approximately 37 / 4).
[0062] The decoder can decode the middle channel to generate a decoded middle channel, and the decoder can generate a first channel and a second channel based on the decoded middle channel. For example, the decoder can upmix the decoded middle channel using stereo parameters included in the bitstream to generate the first channel and the second channel. The first and second channels can be temporarily aligned in the decoder; however, the decoder can shift one or more channels relative to each other based on the final quantized temporal mismatch value. For example, if the first channel matches the target channel (e.g., the first audio signal) in the encoder, the decoder can shift the first channel by thirty-six samples (e.g., 4*9) to generate a shifted first channel. Perceptually, the shifted first channel and the second channel are similar to the target channel and the reference channel, respectively.For example, if the offset of thirty-seven samples between the target channel and the reference channel in the encoder corresponds to an offset of 10 ms, then the offset of thirty-six... Petition 870190112962, dated 05 / 11 / 2019, p. 44 / 134 A shift of 40 / 105 samples between the first shifted channel and the second channel in the decoder is perceptually similar to, and may be perceptually indistinguishable from, a shift of thirty-seven samples.
[0063] Referring to Figure 1, a particular illustrative example of a system 100 is shown. The system 100 includes a first device 104 communicatively coupled, via a network 120, to a second device 106. The network 120 may include one or more wireless networks, one or more wired networks, or a combination thereof.
[0064] The first device 104 includes an encoder 114, a transmitter 110, and one or more input interfaces 112. A first input interface of the input interfaces 112 may be coupled to a first microphone 146. A second input interface of the input interface(s) 112 may be coupled to a second microphone 148. The first device 104 may also include a memory 153 configured to store analysis data, as described below. The second device 106 may include a decoder 118 and a memory 154. The second device 106 may be coupled to a first loudspeaker 142, a second loudspeaker 144, or both.
[0065] During operation, the first device 104 can receive a first audio signal 130 through the first input interface of the first microphone 146 and can receive a second audio signal 132 through the second input interface of the second microphone 148. The Petition 870190112962, dated 05 / 11 / 2019, page 45 / 134 41 / 105 The first audio signal 130 can correspond to either a right-channel signal or a left-channel signal. The second audio signal 132 can correspond to either a right-channel signal or a left-channel signal. As described here, the first audio signal 130 can correspond to a reference channel, and the second audio signal 132 can correspond to a target channel. However, it should be understood that in other implementations, the first audio signal 130 may correspond to the target channel, and the second audio signal 132 may correspond to the reference channel. In other implementations, there may be no reference and target channel assignment. In these cases, channel alignment in the encoder and channel misalignment in the decoder may be performed on one or both channels, such that the relative offset between the channels is based on an offset value.
[0066] The first microphone 146 and the second microphone 148 can receive audio from a sound source 152 (e.g., a user, a loudspeaker, ambient noise, a musical instrument, etc.). In a particular aspect, the first microphone 146, the second microphone 148, or both, can receive audio from multiple sound sources. The multiple sound sources may include a dominant (or most dominant) sound source (e.g., sound source 152) and one or more secondary sound sources. The one or more secondary sound sources may correspond to traffic, background music, another speaker, street noise, etc. The sound source 152 (e.g., the source Petition 870190112962, dated 05 / 11 / 2019, page 46 / 134 42 / 105 dominant sound) may be closer to the first microphone 146 than to the second microphone 148. Consequently, an audio signal from the sound source 152 may be received at the input interface(s) 112 via the first microphone 146 earlier than via the second microphone 148. This natural delay in acquiring a multichannel signal through multiple microphones may introduce a time shift between the first audio signal 130 and the second audio signal 132.
[0067] The first device 104 can store the first audio signal 130, the second audio signal 132, or both, in memory 153. The encoder 114 can determine a first offset value 180 (e.g., a non-causal offset value) indicative of the offset (e.g., a non-causal offset) of the first audio signal 130 relative to the second audio signal 132 for a first frame 190. The first offset value 180 can be a value (e.g., a non-quantized value) representing an offset between the reference channel (e.g., the first audio signal 130) and the target channel (e.g., the second audio signal 132) for the first frame 190. The first offset value 180 can be stored in memory 153 as analysis data.Encoder 114 can also determine a second offset value 184 indicating the offset of the first audio signal 130 relative to the second audio signal 132 for a second frame 192. The second frame 192 can follow (for example, be later in time than) the first frame 190. The second value of... Petition 870190112962, dated 05 / 11 / 2019, page 47 / 134 43 / 105 offset 184 can be a value (e.g., a non-quantized value) representing an offset between the reference channel (e.g., the first audio signal 130) and the target channel (e.g., the second audio signal 132) for the second frame 192. The second offset value 184 can also be stored in memory 153 as analysis data.
[0068] Thus, the offset values 180, 184 (e.g., the mismatch values) may be indicative of a quantity of temporal mismatch (e.g., time delay) between the first audio signal 130 and the second audio signal 132 for the first and second frames 190, 192, respectively. As referred to here, time delay may correspond to temporal lag. Temporal mismatch may indicate a time delay between the reception, via the first microphone 146, of the first audio signal 130 and the reception, via the second microphone 148, of the second audio signal 132. For example, a first value (e.g., a positive value) of the offset value 180, 184 may indicate that the second audio signal 132 is delayed relative to the first audio signal 130. In this example, the first audio signal 130 may correspond to a leading signal and the second audio signal 132 may correspond to a lagging signal.A second value (for example, a negative value) of the offset values 180, 184 may indicate that the first audio signal 130 is delayed relative to the second audio signal 132. In this example, the first audio signal 130 may... Petition 870190112962, dated 05 / 11 / 2019, page 48 / 134 44 / 105 corresponds to a delayed signal, and the second audio signal 132 may correspond to an advanced signal. A third value (e.g., 0) of the offset values 180, 184 may indicate no delay between the first audio signal 130 and the second audio signal 132.
[0069] Encoder 114 can quantize the first offset value 180 to generate a first quantized offset value 181. To illustrate, if the first offset value 180 (e.g., the true offset value) is equal to thirty-seven samples, encoder 114 can quantize the first offset value 180 based on a floor to generate the first quantized offset value 181. As a non-limiting example, if the floor is equal to four, the first quantized offset value 181 can be equal to nine (e.g., approximately 37 / 4). As described below, the first offset value 180 can be used to generate a first portion of a medium channel 191, and the first quantized offset value 181 can be encoded into a bit stream 160 and transmitted to the second device 106.As used herein, a portion of a signal or channel includes one or more frames of the signal or channel, one or more subframes of the signal or channel, one or more samples, bits, blocks, words, or other segments of the signal or channel, or any combination thereof. Similarly, the 114 encoder can quantize the second offset value 184 to generate a second quantized offset value 185. To illustrate, if the second offset value 184 is equal to thirty-six. Petition 870190112962, dated 05 / 11 / 2019, page 49 / 134 45 / 105 samples, encoder 114 can quantize the second offset value 184 based on the floor to generate the second quantized offset value 185. As a non-limiting example, the second quantized offset value 185 can also be equal to nine (e.g., 36 / 4). As described below, the second offset value 184 can be used to generate a second portion of the middle channel 193, and the second quantized offset value 185 can be encoded into the bit stream 160 and transmitted to the second device 106.
[0070] Encoder 114 can also generate a reference signal indicator based on offset values 180, 184. For example, encoder 114 can, in response to the determination that the first offset value 180 indicates a first value (e.g., a positive value), generate the reference signal indicator to have a first value (e.g., 0) indicating that the first audio signal 130 is a reference signal and that the second audio signal 132 corresponds to a target signal.
[0071] Encoder 114 can temporally align the first audio signal 130 and the second audio signal 132 based on offset values 180, 184. For example, for the first frame 190, encoder 114 can temporally shift the second audio signal 132 by the first offset value 180 to generate a second shifted audio signal that is temporally aligned with the first audio signal 130. Although the second audio signal 132 is described as Petition 870190112962, dated 05 / 11 / 2019, page 50 / 134 46 / 105 subjected to a time shift in the time domain, it should be understood that the second audio signal 132 can be subjected to a phase shift in the frequency domain to generate the second shifted audio signal 132. For example, the first shift value 180 can correspond to a frequency domain shift value. For the second frame 192, the encoder 114 can temporally shift the second audio signal 132 by the second shift value 184 to generate a second shifted audio signal that is temporally aligned with the first audio signal 130. Although the second audio signal 132 is described as subjected to a time shift in the time domain, it should be understood that the second audio signal 132 can be subjected to a phase shift in the frequency domain to generate the second shifted audio signal 132.For example, the second shift value of 184 could correspond to a frequency domain shift value.
[0072] Encoder 114 can generate one or more additional stereo parameters (e.g., other stereo parameters besides the offset values 180, 184) for each frame based on the reference channel samples and target channel samples. As a non-limiting example, encoder 114 can generate a first stereo parameter 182 for the first frame 190 and a second stereo parameter 186 for the second frame 192. Non-limiting examples of stereo parameters 182, 186 may include other offset values, parameters of Petition 870190112962, dated 05 / 11 / 2019, page 51 / 134 47 / 105 interchannel phase difference, interchannel level difference parameters, interchannel time difference parameters, interchannel correlation parameters, spectral slope parameters, interchannel gain parameters, interchannel voice parameters, or interchannel pitch parameters.
[0073] To illustrate, if stereo parameters 182, 186 correspond to gain parameters, for each frame, encoder 114 can generate a gain parameter (e.g., a CODEC gain parameter) based on samples of the reference signal (e.g., the first audio signal 130) and based on samples of the target signal (e.g., the second audio signal 132). For example, for the first frame 190, encoder 114 can select samples from the second audio signal 132 based on the first offset value 180 (e.g., the non-causal offset value). As referred to here, selecting samples from an audio signal based on an offset value can correspond to generating a modified audio signal (time-shifted or frequency-shifted) by adjusting (e.g., shifting) the audio signal based on the offset value and selecting samples from the modified audio signal.For example, encoder 114 can generate a second time-shifted audio signal by offsetting the second audio signal 132 based on the first offset value 180, and can select samples from the second time-shifted audio signal. Encoder 114 can, in response to the determination that the first audio signal 130 is the reference signal, ... Petition 870190112962, dated 05 / 11 / 2019, page 52 / 134 48 / 105 Determine the gain parameter of the selected samples based on the first samples of the first frame 190 of the first audio signal 130. As an example, the gain parameter could be based on one of the following equations: Equation 1 a3[f' _ Kg=olffc / CnJI “° ΣποοίΤαΓ^ίη)!' _ iff”1Ref tn} Target)3a Equation lb Equation Ic Equation Id Equation le If equation
[0074] where go corresponds to the relative gain parameter for downmix processing, Ref(n) corresponds to samples of the reference signal, Νχ corresponds to the first 180 offset value of the first 190 frame, and Targ(n + Νχ) corresponds to samples of the target signal. The gain parameter (gD) can be modified, for example, based on one of the laIf equations, to incorporate long-term smoothing / hysteresis logic to avoid large gain jumps between frames.
[0075] Encoder 114 can quantize stereo parameters 182, 186 to generate quantized stereo parameters 183, 187 which are encoded in bitstream 160 and transmitted to the second device 106. By Petition 870190112962, dated 05 / 11 / 2019, page 53 / 134 49 / 105 For example, encoder 114 can quantize the first stereo parameter 182 to generate a first quantized stereo parameter 183, and encoder 114 can quantize the second stereo parameter 186 to generate a second quantized stereo parameter 187. The quantized stereo parameters 183, 187 may have a lower resolution (i.e., less precision) than the stereo parameters 182, 186, respectively.
[0076] For each frame 190, 192, the encoder 114 can generate one or more encoded signals based on the offset values 180, 184, the other stereo parameters 182, 186, and the audio signals 130, 132. For example, for the first frame 190, the encoder 114 can generate a first portion of a mid-channel 191 based on the first offset value 180 (e.g., the non-quantized offset value), the first stereo parameter 182, and the audio signals 130, 132. Additionally, for the second frame 192, the encoder 114 can generate a second portion of the mid-channel 193 based on the second offset value 184 (e.g., the non-quantized offset value), the second stereo parameter 186, and the audio signals 130, 132.According to some implementations, encoder 114 can generate side channels (not shown) for each frame 190, 192 based on offset values 180, 184, other stereo parameters 182, 186, and audio signals 130, 132.
[0077] For example, encoder 114 can generate the middle channel portions 191, 193 based on one of the following equations: Petition 870190112962, dated 05 / 11 / 2019, p. 54 / 134 50 / 105 Μ = Ref(η) + ,^ΤατψΟι + iVt), Λί = Ref(η) + Targin + / VL), Equation 2a Equation 2b M = Ref(n — iV2) + Targ(n + iVL— N2^emAueN2 can have any arbitrary value, Equation 2c
[0078] where M corresponds to the mid-channel, gD corresponds to the relative gain parameter (e.g., stereo parameters 182, 186) for downmix processing, Ref(n) corresponds to samples of the reference signal, Νχ corresponds to the offset values 180, 184, and Targ(n + Νχ) corresponds to samples of the target signal.
[0079] Encoder 114 can generate side channels based on one of the following equations: S = Ref(n) - gDTarg(.n + Sei Equation 5'= jjy / íefÇTt) — 7'íít^(ti + AJl), Equation 3b S = Ref(n - N2) — gDTarg(n + Ni~ JV3)rem where N2 can have any arbitrary value, Equation 3c
[0080] where S corresponds to the side channel signal, gD corresponds to the relative gain parameter (e.g., stereo parameters 182, 186) for downmix processing, Ref(n) corresponds to samples of the reference signal, Νχ corresponds to the offset values 180, 184, and Targ(n + Νχ) corresponds to samples of the target signal.
[0081] Transmitter 110 can transmit bit stream 160, through network 120, to the second Petition 870190112962, dated 05 / 11 / 2019, p. 55 / 134 51 / 105 device 106. The first frame 190 and the second frame 192 can be encoded in bitstream 160. For example, the first portion of the middle channel 191, the first quantized offset value 181, and the first quantized stereo parameter 183 can be encoded in bitstream 160. Additionally, the second portion of the middle channel 193, the second quantized offset value 185, and the second quantized stereo parameter 187 can be encoded in bitstream 160. Side channel information can also be encoded in bitstream 160. Although not shown, additional information can also be encoded in bitstream 160 for each frame 190, 192. As a non-limiting example, a reference channel indicator can be encoded in bitstream 160 for each frame 190, 192.
[0082] Due to poor transmission conditions, some data encoded in bitstream 160 may be lost in transmission. Packet loss may occur due to poor transmission conditions, frame erasure may occur due to poor radio conditions, packets may arrive late due to high instability, etc. According to the non-limiting illustrative example, the second device 106 may receive the first frame 190 of bitstream 160 and the second portion of the middle channel 193 of the second frame 192. In this way, the second quantized offset value 185 and the second quantized stereo parameter 187 may be lost in transmission due to poor transmission conditions.
[0083] The second device 106 can therefore, Petition 870190112962, dated 05 / 11 / 2019, page 56 / 134 52 / 105 receive at least a portion of the bit stream 160 as transmitted by the first device 102. The second device 106 can store the received portion of the bit stream 160 in memory 154 (for example, in a buffer). For example, the first frame 190 can be stored in memory 154 and the second portion of the middle channel 193 of the second frame 192 can also be stored in memory 154.
[0084] Decoder 118 can decode the first frame 190 to generate a first output signal 126 that corresponds to the first audio signal 130 and to generate a second output signal 128 that corresponds to the second audio signal 132. For example, decoder 118 can decode the first portion of the mid-channel 191 to generate a first portion of a decoded mid-channel 170. Decoder 118 can also perform a transformation operation on the first portion of the decoded mid-channel 170 to generate a first portion of a frequency domain (FD) decoded mid-channel 171. Decoder 118 can upmix the first portion of the frequency domain decoded mid-channel 171 to generate a first frequency domain channel (not shown) associated with the first output signal 126 and a second frequency domain channel (not shown) associated with the second output signal 128.During upmixing, decoder 118 can apply the first quantized stereo parameter 183 to the first portion of the decoded frequency domain mid-channel 171.
[0085] It should be noted that, in other Petition 870190112962, dated 05 / 11 / 2019, page 57 / 134 In 53 / 105 implementations, the decoder 118 may not perform the transformation operation, but perform the upmix based on the middle channel, some stereo parameters (e.g., the downmix gain) and, additionally, if available, also based on a time-domain decoded side channel to generate the first time-domain channel (not shown) associated with the first output channel 126 and a second time-domain channel (not shown) associated with the second output channel 128.
[0086] If the first quantized shift value 181 corresponds to a frequency domain shift value, the decoder 118 can shift the second frequency domain channel by the first quantized shift value 181 to generate a second shifted frequency domain channel (not shown). The decoder 118 can perform an inverse transform operation on the first frequency domain channel to generate the first output signal 126. The decoder 118 can also perform an inverse transform operation on the second shifted frequency domain channel to generate the second output signal 128.
[0087] If the first quantized shift value 181 corresponds to a time-domain shift value, the decoder 118 can perform an inverse transform operation on the first frequency-domain channel to generate the first output signal 126. The decoder 118 can also perform an inverse transform operation on the second frequency-domain channel. Petition 870190112962, dated 05 / 11 / 2019, page 58 / 134 54 / 105 frequency to generate a second time-domain channel. Decoder 118 can shift the second time-domain channel by the first quantized shift value 181 to generate the second output signal 128. In this way, decoder 118 can use the first quantized shift value 181 to emulate a perceptible difference between the first output signal 126 and the second output signal 128. The first speaker 142 can output the first output signal 126, and the second speaker 144 can output the second output signal 128. In some cases, the inverse transformation operation can be omitted in implementations where the upmix was performed in the time domain to directly generate the first time-domain channel and the second time-domain channel, as described above.It should also be noted that the presence of a time-domain shift value in the 118 decoder may simply be a matter of indicating that the decoder is configured to perform time-domain shifting, and in some implementations, although a time-domain shift may be available in the 118 decoder (indicating that the decoder performs the shifting operation in the time domain), the encoder from which the bitstream was received may have performed a frequency-domain shift operation or a time-domain shift operation to align the channels.
[0088] If decoder 118 determines that the second frame 192 is unavailable for decoding operations (for example, determines that the second value Petition 870190112962, dated 05 / 11 / 2019, page 59 / 134 Since the 55 / 105 quantized shift 185 and the second quantized stereo parameter 187 are unavailable), the decoder 118 can generate the output signals 126, 128 for the second frame 192 based on the stereo parameters associated with the first frame 190. For example, the decoder 118 can estimate or interpolate the second quantized shift value 185 based on the first quantized shift value 181. Additionally, the decoder 118 can estimate or interpolate the second quantized stereo parameter 187 based on the first quantized stereo parameter 183.
[0089] After estimating the second quantized offset value 185 and the second quantized stereo parameter 187, the decoder 118 can generate the output signals 126, 128 for the second frame 192 in a similar way to how output signals 126, 128 are generated for the first frame 190. For example, the decoder 118 can decode the second portion of the mid-channel 193 to generate a second portion of the decoded mid-channel 172. The decoder 118 can also perform a transformation operation on the second portion of the decoded mid-channel 172 to generate a second frequency-domain decoded mid-channel 173. Based on the estimated quantized offset value and the estimated quantized stereo parameter 187, the decoder 118 can upmix the second frequency-domain decoded mid-channel 173, perform an inverse transformation on the upmixed signals, and shift the resulting signal to generate the Output signals 126, 128. An example of operations of Petition 870190112962, dated 05 / 11 / 2019, pp. 60 / 134 56 / 105 decoding is described in greater detail with regard to Figure 2.
[0090] System 100 can align the channels as much as possible in encoder 114 to reduce coding entropy and thus increase coding efficiency, because coding entropy is sensitive to shift changes between channels. For example, encoder 114 can use non-quantized shift values to precisely align the channels because non-quantized shift values have a relatively high resolution. In decoder 118, quantized stereo parameters can be used to emulate a perceptible difference between the output signals 126, 128 using a reduced number of bits compared to using non-quantized shift values, and missing stereo parameters (due to poor transmission) can be interpellated or estimated using stereo parameters from one or more previous frames.According to some implementations, offset values 180, 184 (e.g., non-quantized offset values) can be used to offset target channels in the frequency domain, and quantized offset values 181, 185 can be used to offset target channels in the time domain. For example, offset values used for time-domain stereo encoding may have a lower resolution than offset values used for frequency-domain stereo encoding.
[0091] Referring to Figure 2, a diagram Petition 870190112962, dated 05 / 11 / 2019, page 61 / 134 Figure 57 / 105 illustrating a particular implementation of the 118 decoder is shown. The 118 decoder includes a medium channel decoder 202, a transformer unit 204, an upmixer 206, an inverse transformer unit 210, an inverse transformer unit 212, and a shifter 214.
[0092] The bit stream 160 of Figure 1 can be provided to the decoder 118. For example, the first portion of the mid-channel 191 of the first frame 190 and the second portion of the mid-channel 193 of the second frame 192 can be provided to the mid-channel decoder 202. Additionally, stereo parameters 201 can be provided to the upmixer 206 and the shifter 214. The stereo parameters 201 can include the first quantized shift value 181 associated with the first frame 190 and the first quantized stereo parameter 183 associated with the first frame 190. As described above with respect to Figure 1, the second quantized shift value 185 associated with the second frame 192 and the second quantized stereo parameter 187 associated with the second frame 192 may not be received by the decoder 118 due to poor transmission conditions.
[0093] To decode the first frame 190, the mid-channel decoder 202 can decode the first portion of the mid-channel 191 to generate the first portion of the decoded mid-channel 170 (e.g., a time-domain mid-channel). According to some implementations, two asymmetric windows can be applied to the first portion of the decoded mid-channel 170 to generate a windowed portion of a mid-channel of Petition 870190112962, dated 05 / 11 / 2019, page 62 / 134 58 / 105 time domain. The first portion of the decoded mid-channel 170 is provided to the transformation unit 204. The transformation unit 204 can be configured to perform a transformation operation on the first portion of the decoded mid-channel 170 to generate the first portion of the frequency domain decoded mid-channel 171. The first portion of the frequency domain decoded mid-channel 171 is provided to the upmixer 206. According to some implementations, the windowing and transformation operation can be completely ignored and the first portion of the decoded mid-channel 170 (e.g., a time domain mid-channel) can be directly provided to the upmixer 206.
[0094] The upmixer 206 can upmix the first portion of the decoded frequency domain mid-channel 171 to generate a frequency domain channel portion 250 and a frequency domain channel portion 254. The upmixer 206 can apply the first quantized stereo parameter 183 to the first portion of the decoded frequency domain mid-channel 171 during upmix operations to generate the frequency domain channel portions 250, 254. According to an implementation where the first quantized offset value 181 includes a frequency domain offset (e.g., the first quantized offset value 181 corresponds to a first quantized frequency domain offset value 281), the upmixer 206 can perform a frequency domain offset (e.g., a phase shift) based on Petition 870190112962, dated 05 / 11 / 2019, pages 63 / 134 59 / 105 in the first quantized frequency domain shift value 281 to generate the frequency domain channel portion 254. The frequency domain channel portion 250 is provided to the inverse transform unit 210, and the frequency domain channel portion 254 is provided to the inverse transform unit 212. According to some implementations, the upmixer 206 can be configured to operate on time domain channels where stereo parameters (e.g., based on target gain values) can be applied in the time domain.
[0095] The inverse transform unit 210 can perform an inverse transform operation on the frequency domain channel portion 250 to generate a time domain channel portion 260. The time domain channel portion 260 is provided to the shifter 214. The inverse transform unit 212 can perform an inverse transform operation on the frequency domain channel portion 254 to generate a time domain channel portion 264. The time domain channel portion 264 is also provided to the shifter 214. In implementations where the upmix operation is performed in the time domain, the inverse transform operations after the upmix operation can be ignored.
[0096] According to the implementation in which the first quantized shift value 181 corresponds to a first quantized frequency domain shift value 281, the shifter 214 can bypass the shift operations and pass portions of the time domain channels 260, 264 as portions of the output signals 126, Petition 870190112962, dated 05 / 11 / 2019, pages 64 / 134 60 / 105 128, respectively. According to an implementation in which the first quantized offset value 181 includes a time-domain offset (for example, the first quantized offset value 181 corresponds to a first quantized time-domain offset value 291), the shifter 214 can shift the time-domain channel portion 264 by the first quantized time-domain offset value 291 to generate the second output signal portion 128.
[0097] In this way, decoder 118 can use quantized offset values having reduced precision (compared to the non-quantized offset values used in encoder 114) to generate the portions of output signals 126, 128 for the first frame 190. Using quantized offset values to shift output signal 128 relative to output signal 126 can restore the user's perception of the offset in encoder 114.
[0098] To decode the second frame 192, the mid-channel decoder 202 can decode the second portion of the mid-channel 193 to generate the second portion of the decoded mid-channel 172 (e.g., a time-domain mid-channel). According to some implementations, two asymmetric windows can be applied to the second portion of the decoded mid-channel 172 to generate a windowed portion of the time-domain mid-channel. The second portion of the decoded mid-channel 172 is provided to the transformation unit 204. The transformation unit 204 can be configured to perform an operation of Petition 870190112962, dated 05 / 11 / 2019, page 65 / 134 61 / 105 transformation on the second portion of the decoded mid-channel 172 to generate the second portion of the frequency-domain decoded mid-channel 173. The second portion of the frequency-domain decoded mid-channel 173 is provided to the upmixer 206. According to some implementations, the windowing and transformation operation can be completely ignored and the second portion of the decoded mid-channel 172 (e.g., a time-domain mid-channel) can be directly provided to the upmixer 206.
[0099] As described above with respect to Figure 1, the second quantized offset value 185 and the second quantized stereo parameter 187 may not be received by the decoder 118 due to poor transmission conditions. As a result, stereo parameters for the second frame 192 may not be accessible to the upmixer 206 and the shifter 214. The upmixer 206 includes a stereo parameter interpolator 208 that is configured to interpolate (or estimate) the second quantized offset value 185 based on the first quantized frequency domain offset value 281. For example, the stereo parameter interpolator 208 can generate a second interpolated frequency domain offset value 285 based on the first quantized frequency domain offset value 281.The stereo parameter interpolator 208 can also be configured to interpolate (or estimate) the second quantized stereo parameter 187 based on the first quantized stereo parameter 183. For example, the stereo parameter interpolator 208 can generate a second interpolated stereo parameter 287 based on... Petition 870190112962, dated 05 / 11 / 2019, page 66 / 134 62 / 105 first quantized stereo parameter 183.
[0100] The upmixer 206 can upmix the second portion of the decoded frequency domain mid-channel 173 to generate a portion of a frequency domain channel 252 and a portion of a frequency domain channel 256. The upmixer 206 can apply the second interpolated stereo parameter 287 to the second portion of the decoded frequency domain mid-channel 173 during upmix operations to generate the portions of the frequency domain channels 252, 256.According to an implementation where the first quantized shift value 181 includes a frequency domain shift (e.g., the first quantized shift value 181 corresponds to a first quantized frequency domain shift value 281), the upmixer 206 can perform a frequency domain shift (e.g., a phase shift) based on the second interpolated frequency domain shift value 285 to generate the frequency domain channel portion 256. The frequency domain channel portion 252 is provided to the inverse transform unit 210, and the frequency domain channel portion 256 is provided to the inverse transform unit 212.
[0101] The inverse transformer unit 210 can perform an inverse transformer operation on the frequency domain channel portion 252 to generate a time domain channel portion 262. The time domain channel portion 262 is provided to the shifter 214. The inverse transformer unit 212 can perform a Petition 870190112962, dated 05 / 11 / 2019, page 67 / 134 63 / 105 inverse transformation operation on the frequency domain channel portion 256 to generate a time domain channel portion 266. The time domain channel portion 266 is also provided to the shifter 214. In implementations where the upmixer 206 operates on the time domain channels, the output of the upmixer 206 can be provided to the shifter 214, and the inverse transformation units 210, 212 can be ignored or omitted.
[0102] The shifter 214 includes a shift value interpolator 216 that is configured to interpolate (or estimate) the second quantized shift value 185 based on the first quantized time-domain shift value 291. For example, the shift value interpolator 216 can generate a second interpolated time-domain shift value 295 based on the first quantized time-domain shift value 291. According to the implementation where the first quantized shift value 181 corresponds to the first quantized frequency-domain shift value 281, the shifter 214 can bypass the shift operations and pass the time-domain channel portions 262, 266 as the output signals 126, 128, respectively.According to the implementation where the first quantized shift value 181 corresponds to the first quantized time-domain shift value 291, the shifter 214 can shift the time-domain channel portion 266 by the second interpolated time-domain shift value 295 to generate the second output signal 128. Petition 870190112962, dated 05 / 11 / 2019, pp. 68 / 134 64 / 105
[0103] In this way, decoder 118 can approximate stereo parameters (e.g., offset values) based on stereo parameters or variation in stereo parameters of previous frames. For example, decoder 118 can extrapolate stereo parameters for frames that are lost during transmission (e.g., the second frame 192) from stereo parameters of one or more previous frames.
[0104] Referring to Figure 3, a diagram 300 for predicting stereo parameters of a missing frame in a decoder is shown. According to diagram 300, the first frame 190 can be successfully transmitted from encoder 114 to decoder 118, and the second frame 192 can be unsuccessfully transmitted from encoder 114 to decoder 118. For example, the second frame 192 can be lost in transmission due to poor transmission conditions.
[0105] Decoder 118 can generate the first portion of the decoded middle channel 170 from the first frame 190. For example, decoder 118 can decode the first portion of the middle channel 191 to generate the first portion of the decoded middle channel 170. Using the techniques described in relation to Figure 2, decoder 118 can also generate a first portion of a left channel 302 and a first portion of a right channel 304 based on the first portion of the decoded middle channel 170. The first portion of the left channel 302 can correspond to the first output signal 126, and the first portion of the right channel 304 can correspond to the second output signal 128. By Petition 870190112962, dated 05 / 11 / 2019, page 69 / 134 For example, the 65 / 105 decoder 118 can use the first quantized stereo parameter 183 and the first quantized offset value 181 to generate channels 302 and 304.
[0106] Decoder 118 can interpolate (or estimate) the second interpolated frequency domain offset value 285 (or the second interpolated time domain offset value 295) based on the first quantized offset value 181. According to other implementations, the second interpolated offset values 285, 295 can be estimated (i.e., interpolated or extrapolated) based on quantized offset values associated with two or more previous frames (e.g., the first frame 190 and at least one frame preceding the first frame or one frame following the second frame 192, one or more other frames in the bitstream 160, or any combination thereof). Decoder 118 can also interpolate (or estimate) the second interpolated stereo parameter 287 based on the first quantized stereo parameter 183.According to other implementations, the second interpolated stereo parameter 287 can be estimated based on quantized stereo parameters associated with two or more other frames (for example, the first frame 190 and at least one frame preceding or following the first frame).
[0107] Additionally, decoder 118 can interpolate (or estimate) a second portion of the decoded middle channel 306 based on the first portion of the decoded middle channel 170 (or middle channels associated with two or more previous frames). Using the techniques described with Petition 870190112962, dated 05 / 11 / 2019, pp. 70 / 134 66 / 105 in relation to Figure 2, the decoder 118 can also generate a second portion of the left channel 308 and a second portion of the right channel 310 based on the estimated second portion of the decoded middle channel 306. The second portion of the left channel 308 can correspond to the first output signal 126, and the second portion of the right channel 310 can correspond to the second output signal 128. For example, the decoder 118 can use the second interpolated stereo parameter 287 and the second interpolated frequency domain quantized shift value 285 to generate the right and left channels.
[0108] Referring to Figure 4A, a method 400 for decoding a signal is shown. The method 400 can be performed by the second device 106 in Figure 1, the decoder 118 in Figures 1 and 2, or both.
[0109] The 400 method involves receiving, in a decoder, a bit stream including a middle channel and a quantized value representing an offset between a first channel (e.g., a reference channel) associated with an encoder and a second channel (e.g., a target channel) associated with the encoder, in 402. The quantized value is based on a value of the offset. The value is associated with the encoder and has a higher precision than the quantized value.
[0110] The 400 method also includes decoding the middle channel to generate a decoded middle channel, in 404. The 400 method further includes generating a first channel (a first generated channel) based on the decoded middle channel, in 406, and generating a second channel (a second Petition 870190112962, dated 05 / 11 / 2019, page 71 / 134 67 / 105 generated channel) based on the decoded average channel and the quantized value, in 408. The first generated channel corresponds to the first channel associated with the encoder (e.g., the reference channel) and the second generated channel corresponds to the second channel associated with the encoder (e.g., the target channel). In some implementations, both the first and second channels may be based on the quantized offset value. In some implementations, the decoder may not explicitly identify target and reference channels before the offset operation.
[0111] Thus, method 400 of Figure 4A can enable encoder side channel alignment to reduce coding entropy and thus increase coding efficiency, because coding entropy is sensitive to shift changes between channels. For example, encoder 114 can use non-quantized shift values to precisely align channels because non-quantized shift values have a relatively high resolution. Quantized shift values can be transmitted to decoder 118 to reduce the utilization of data transmission resources. In decoder 118, the quantized shift parameters can be used to emulate a perceptible difference between output signals 126, 128.
[0112] Referring to Figure 4B, a 450 method for decoding a signal is shown. In some implementations, the 450 method in Figure 4B is a more complex version. Petition 870190112962, dated 05 / 11 / 2019, page 72 / 134 68 / 105 detailed of method 400 for decoding the audio signal of Figure 4A. Method 450 can be performed by the second device 106 of Figure 1, by the decoder 118 of Figures 1 and 2, or both.
[0113] The 450 method involves receiving, in a decoder, a bitstream from an encoder, in 452. The bitstream includes a middle channel and a quantized value representing an offset between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value may be based on a value (e.g., a non-quantized value) of the offset that has a higher precision than the quantized value. For example, referring to Figure 1, decoder 118 may receive bitstream 160 from encoder 114. Bitstream 160 may include the first portion of the middle channel 191 and the first quantized offset value 181 representing the offset between the first audio signal 130 (e.g., the reference channel) and the second audio signal 132 (e.g., the target channel). The first quantized shift value 181 may be based on the first shift value 180 (i.e., a non-quantized value).
[0114] The first offset value 180 may have greater precision than the first quantized offset value 181. For example, the first quantized offset value 181 may correspond to a low-resolution version of the first offset value 180. The first offset value may be used by the encoder 114 to temporarily match the target channel. Petition 870190112962, dated 05 / 11 / 2019, pp. 73 / 134 69 / 105 (for example, the second audio signal 132) and the reference channel (for example, the first audio signal 130).
[0115] Method 450 also includes decoding the middle channel to generate a decoded middle channel, in 454. For example, referring to Figure 2, the middle channel decoder 202 can decode the first portion of the middle channel 191 to generate the first portion of the decoded middle channel 170. Method 400 can also include performing a transformation operation on the decoded middle channel to generate a decoded frequency domain middle channel, in 456. For example, referring to Figure 2, the transformation unit 204 can perform a transformation operation on the first portion of the decoded middle channel 170 to generate the first portion of the decoded frequency domain middle channel 171.
[0116] The 450 method can also include upmixing the decoded frequency domain mid-channel to generate a first frequency domain channel portion and a second frequency domain channel, in 458. For example, referring to Figure 2, the upmixer 206 can upmix the first portion of the decoded frequency domain mid-channel 171 to generate frequency domain channel portion 250 and frequency domain channel portion 254. The 450 method can also include generating a first channel based on the first frequency domain channel portion, in 460. The first channel can correspond to the reference channel. For example, the inverse transform unit 210 can perform an inverse transform operation on the portion of Petition 870190112962, dated 05 / 11 / 2019, pp. 74 / 134 70 / 105 frequency domain channel 250 to generate the time domain channel portion 260, and the shifter 214 can pass the time domain channel portion 260 as a portion of the first output signal 126. The first output signal 126 can correspond to the reference channel (e.g., the first audio signal 130).
[0117] Method 450 may also include generating a second channel based on the second frequency domain channel, in 462. The second channel may correspond to the target channel. According to one implementation, the second frequency domain channel may be shifted in a frequency domain by the quantized value if the quantized value corresponds to a frequency domain shift. For example, referring to Figure 2, the upmixer 206 may shift the portion of the frequency domain channel 254 by the first quantized frequency domain shift value 281 to a second shifted frequency domain channel (not shown). The inverse transform unit 212 may perform an inverse transform on the shifted second frequency domain channel to generate a portion of the second output signal 128. The second output signal 128 may correspond to the target channel (e.g., the second audio signal 132).
[0118] According to another implementation, a time-domain version of the second frequency-domain channel can be shifted by the quantized value if the quantized value corresponds to a time-domain shift. For example, the inverse transformation unit Petition 870190112962, dated 05 / 11 / 2019, pp. 75 / 134 71 / 105 212 can perform an inverse transformation operation on the frequency domain channel portion 254 to generate the time domain channel portion 264. The shifter 214 can shift the time domain channel portion 264 by the first quantized time domain shift value 291 to generate a portion of the second output signal 128. The second output signal 128 can correspond to the target channel (e.g., the second audio signal 132).
[0119] Thus, method 450 of Figure 4B can allow the alignment of encoder side channels to reduce coding entropy and thus increase coding efficiency, because coding entropy is sensitive to shift changes between channels. For example, encoder 114 can use non-quantized shift values to precisely align the channels because non-quantized shift values have a relatively high resolution. Quantized shift values can be transmitted to decoder 118 to reduce the utilization of data transmission resources. In decoder 118, the quantized shift parameters can be used to emulate a perceptible difference between the output signals 126, 128.
[0120] Referring to Figure 5A, another method 500 of decoding a signal is shown. Method 500 can be performed by the second device 106 in Figure 1, the decoder 118 in Figures 1 and 2, or both.
[0121] Method 500 involves receiving at least a portion of a bit stream, in 502. The bit stream includes Petition 870190112962, dated 05 / 11 / 2019, page 76 / 134 72 / 105 a first frame and a second frame. The first frame includes a first portion of a mid-channel and a first value of a stereo parameter, and the second frame includes a second portion of the mid-channel and a second value of the stereo parameter.
[0122] Method 500 also includes decoding the first portion of the mid-channel to generate a first portion of a decoded mid-channel, in 504. Method 500 further includes generating a first portion of a left channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter, in 506, and generating a first portion of a right channel based at least on the first portion of the decoded mid-channel and the first value of the stereo parameter, in 508. The method also includes, in response to the second frame being unavailable for decoding operations, generating a second portion of the left channel and a second portion of the right channel based at least on the first value of the stereo parameter, in 510. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0123] According to one implementation, method 500 includes generating an interpolated value of the stereo parameter based on the first stereo parameter value and the second stereo parameter value in response to the second frame being available for decoding operations. According to another implementation, method 500 includes generating, in response to the second frame being unavailable for decoding operations, at least the second portion of the Petition 870190112962, dated 05 / 11 / 2019, page 77 / 134 73 / 105 left channel and the second portion of the right channel based at least on the first value of the stereo parameter, the first portion of the left channel and the first portion of the right channel.
[0124] According to one implementation, method 500 includes generating, in response to the second frame being unavailable for decoding operations, at least the second portion of the mid-channel and a second portion of a side-channel based on at least the first stereo parameter value, the first portion of the mid-channel, the first portion of the left-channel, or the first portion of the right-channel. Method 500 also includes generating, in response to the second frame being unavailable for decoding operations, the second portion of the left-channel and the second portion of the right-channel based on the second portion of the mid-channel, the second portion of the side-channel, and a third stereo parameter value. The third stereo parameter value is at least based on the first stereo parameter value, an interpolated stereo parameter value, and an encoding mode.
[0125] In this way, method 500 can allow decoder 118 to approximate stereo parameters (e.g., offset values) based on stereo parameters or variation in stereo parameters of previous frames. For example, decoder 118 can extrapolate stereo parameters for frames that are lost during transmission (e.g., the second frame 192) from stereo parameters of one or more previous frames.
[0126] Referring to Figure 5B, another method 550 Petition 870190112962, dated 05 / 11 / 2019, pp. 78 / 134 74 / 105 of a signal decoding is shown. In some implementations, the 550 method of Figure 5B is a more detailed version of the 500 method of audio signal decoding in Figure 5A. The 550 method can be performed by the second device 106 in Figure 1, by the decoder 118 in Figures 1 and 2, or both.
[0127] The 550 method involves receiving, in a decoder, at least a portion of a bitstream from an encoder, in 552. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a mid-channel and a first value of a stereo parameter, and the second frame includes a second portion of the mid-channel and a second value of the stereo parameter. For example, referring to Figure 1, the second device 106 can receive a portion of bitstream 160 from encoder 114. The bitstream includes the first frame 190 and the second frame 192. The first frame 190 includes the first portion of the mid-channel 191, the first quantized offset value 181, and the first quantized stereo parameter 183. The second frame 192 includes the second portion of the mid-channel 193, the second quantized offset value 185, and the second quantized stereo parameter 187.
[0128] The 550 method also includes decoding the first portion of the medium channel to generate a first portion of a decoded medium channel, in 554. For example, referring to Figure 2, the medium channel decoder 202 can decode the first portion of the medium channel 191 to generate the first portion of the decoded medium channel 170. The Petition 870190112962, dated 05 / 11 / 2019, pp. 79 / 134 75 / 105 method 550 may also include performing a transformation operation on the first portion of the decoded medium channel to generate a first portion of a decoded frequency domain medium channel, in 556. For example, referring to Figure 2, the transformation unit 204 may perform a transformation operation on the first portion of the decoded medium channel 170 to generate the first portion of the decoded frequency domain medium channel 171.
[0129] The 550 method can also include upmixing the first portion of the decoded frequency domain mid-channel to generate a first portion of a left frequency domain channel and a first portion of a right frequency domain channel, in 558. For example, referring to Figure 1, the upmixer 206 can upmix the first portion of the decoded frequency domain mid-channel 171 to generate frequency domain channel 250 and frequency domain channel 254. As described here, frequency domain channel 250 can be a left channel, and frequency domain channel 254 can be a right channel. However, in other implementations, frequency domain channel 250 can be a right channel, and frequency domain channel 254 can be a left channel.
[0130] The 550 method can also include generating a first portion of a left channel based at least on the first portion of the left frequency domain channel, the first value of the stereo parameter, at 560. For example, the upmixer 206 can use the first stereo parameter. Petition 870190112962, dated 05 / 11 / 2019, pages 80 / 134 76 / 105 quantized 183 to generate frequency domain channel 250. The inverse transform unit 210 can perform an inverse transform operation on frequency domain channel 250 to generate time domain channel 260, and the shifter 214 can pass time domain channel 260 as the first output signal 126 (e.g., the first portion of the left channel according to the 550 method).
[0131] The 550 method can also include generating a first portion of a right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter, in 562. For example, the upmixer 206 can use the first quantized stereo parameter 183 to generate the frequency domain channel 254. The inverse transform unit 212 can perform an inverse transform operation on the frequency domain channel 254 to generate the time domain channel 264, and the shifter 214 can pass (or selectively shift) the time domain channel 264 as the second output signal 128 (e.g., the first portion of the right channel according to the 550 method).
[0132] Method 550 also includes determining that the second frame is unavailable for decoding operations, in 564. For example, decoder 118 can determine that one or more portions of the second frame 192 are unavailable for decoding operations. To illustrate, the second quantized offset value 185 and the second quantized stereo parameter 187 can be lost in transmission (from the first device 104 to the Petition 870190112962, dated 05 / 11 / 2019, page 81 / 134 77 / 105 second device 106) based on poor transmission conditions. The 550 method also includes generating, based at least on the first value of the stereo parameter, a second portion of the left channel and a second portion of the right channel in response to the determination that the second frame is unavailable, in 566. The second portion of the left channel and the second portion of the right channel may correspond to a decoded version of the second frame.
[0133] For example, the stereo parameter interpolator 208 can interpolate (or estimate) the second quantized offset value 185 based on the first quantized frequency domain offset value 281. To illustrate, the stereo parameter interpolator 208 can generate the second interpolated frequency domain offset value 285 based on the first quantized frequency domain offset value 281. The stereo parameter interpolator 208 can also Interpolate (or estimate) the second quantized stereo parameter 187 based on the first quantized stereo parameter 183. For example, stereo parameter interpolator 208 can generate a second interpolated stereo parameter 287 based on the first quantized stereo parameter 183. 0134] 0 upmixer 206 can upmix the second The decoded mid-channel of frequency domain 173 is used to generate frequency domain channel 252 and frequency domain channel 256. The upmixer 206 can apply the second interpolated stereo parameter 287 to the second decoded mid-channel of frequency domain 173 during the... Petition 870190112962, dated 05 / 11 / 2019, page 82 / 134 78 / 105 upmix operations to generate frequency domain channels 252, 256. According to the implementation where the first quantized offset value 181 includes a frequency domain offset (e.g., the first quantized offset value 181 corresponds to a first quantized frequency domain offset value 281), the upmixer 206 can perform a frequency domain offset (e.g., a phase shift) based on the second interpolated frequency domain offset value 285 to generate frequency domain channel 256.
[0135] The inverse transform unit 210 can perform an inverse transform operation on the frequency domain channel 252 to generate the time domain channel 262, and the inverse transform unit 212 can perform an inverse transform operation on the frequency domain channel 256 to generate a time domain channel 266. The shift value interpolator 216 can interpolate (or estimate) the second quantized shift value 185 based on the first quantized time domain shift value 291. For example, the shift value interpolator 216 can generate the second interpolated time domain shift value 295 based on the first quantized time domain shift value 291.According to the implementation where the first quantized shift value 181 corresponds to the first quantized frequency domain shift value 281, the shifter 214 can bypass the shift operations and. Petition 870190112962, dated 05 / 11 / 2019, pp. 83 / 134 79 / 105 pass the time domain channels 262, 266 as the output signals 126, 128, respectively. According to the implementation where the first quantized shift value 181 corresponds to the first quantized time domain shift value 291, the shifter 214 can shift the time domain channel 266 by the second interpolated time domain shift value 295 to generate the second output signal 128.
[0136] In this way, the 550 method can allow the decoder 118 to interpolate (or estimate) stereo parameters for frames that are lost during transmission (for example, the second frame 192) based on stereo parameters for one or more previous frames.
[0137] Referring to Figure 6, a block diagram of a particular illustrative example of a device (e.g., a wireless communication device) is presented and generally designated 600. In various implementations, device 600 may have fewer or more components than illustrated in Figure 6. In one illustrative implementation, device 600 may correspond to the first device 104 of Figure 1, the second device 106 of Figure 1, or a combination thereof. In one illustrative implementation, device 600 may perform one or more operations described with reference to systems and methods in Figures 1-3, 4A, 4B, 5A, and 5B.
[0138] In a particular implementation, device 600 includes a processor 606 (for example, a central processing unit (CPU)). Device 600 Petition 870190112962, dated 05 / 11 / 2019, pages 84 / 134 80 / 105 may include one or more additional processors 610 (for example, one or more digital signal processors (DSPs)). The processors 610 may include a media encoder / decoder (CODEC) (for example, speech and music) 608, and an echo canceller 612. The media CODEC 608 may include the decoder 118, the encoder 114, or a combination thereof.
[0139] Device 600 may include a memory 153 and a CODEC 634. Although media CODEC 608 is illustrated as a component of processors 610 (e.g., dedicated circuits and / or executable programming code), in other implementations, one or more components of media CODEC 608, such as decoder 118, encoder 114, or a combination thereof, may be included in processor 606, in CODEC 634, in another processing component, or a combination thereof.
[0140] Device 600 may include transmitter 110 coupled to an antenna 642. Device 600 may include a display 628 coupled to a display controller 626. One or more loudspeakers 648 may be coupled to CODEC 634. One or more microphones 646 may be coupled, via input interface(s) 112, to CODEC 634. In a particular implementation, the loudspeakers 648 may include the first loudspeaker 142, the second loudspeaker 144 of Figure 1, or a combination thereof. In a particular implementation, the microphones 646 may include the first microphone 146, the second microphone 148 of Figure 1, or a combination thereof. CODEC 634 Petition 870190112962, dated 05 / 11 / 2019, page 85 / 134 81 / 105 may include an analog-to-digital converter (DAC) 602 and a digital-to-analog converter (ADC) 604.
[0141] Memory 153 may contain instructions 660 executable by processor 606, processors 610, CODEC 634, another processing unit of device 600, or a combination thereof, to perform one or more operations described with reference to Figures 1-3, 4A, 4B, 5A, 5B. Instructions 660 may be executable to cause the processor (e.g., processor 606, processors 606, CODEC 634, decoder 118, another processing unit of device 600, or a combination thereof) to perform method 400 of Figure 4A, method 450 of Figure 4B, method 500 of Figure 5A, method 550 of Figure 5B, or a combination thereof.
[0142] One or more components of the 600 device may be implemented via dedicated hardware (e.g., circuits), by a processor executing instructions to perform one or more tasks, or a combination thereof. As an example, the 153 memory or one or more components of the 606 processor, the 610 processors and / or the 634 CODEC may be a memory device, such as random access memory (RAM), magnetoresistive random access memory (MRAM), spin-torque transfer MRAM (STT-MRAM), flash memory, read-only memory (ROM), programmable read-only memory (FROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, a removable disk or a memory only of Petition 870190112962, dated 05 / 11 / 2019, page 86 / 134 82 / 105 compact disc (CD-ROM) reading. The memory device may include instructions (e.g., instructions 660) which, when executed by a computer (e.g., a processor in CODEC 634, processor 606 and / or processors 610), may cause the computer to perform one or more operations described with reference to Figures 1-3, 4A, 4B, 5A, 5B. As an example, memory 153 or one or more components of processor 606, processors 610 and / or CODEC 634 may be a non-transient, computer-readable medium that includes instructions (e.g., instructions 660) which, when executed by a computer (e.g., a processor in CODEC 634, processor 606 and / or processors 610), cause the computer to perform one or more operations described with reference to Figures 1-3, 4A, 4B, 5A, 5B.
[0143] In a particular implementation, the device 600 may be included in a system-in-package or system-on-a-chip device (for example, a mobile station modem (MSM)) 622. In a particular implementation, the processor 606, the processors 610, the display controller 626, the memory 153, the CODEC 634, and the transmitter 110 are included in a system-in-package or system-on-a-chip device 622. In a particular implementation, an input device 630, such as a touch screen and / or keyboard, and a power supply 644 are coupled to the system-on-a-chip device 622. Furthermore, in a particular implementation, as illustrated in Figure 6, the display 628, the input device 630, the speakers 648, the microphones 646, the antenna 642, and the Petition 870190112962, dated 05 / 11 / 2019, page 87 / 134 83 / 105 power supply 644 are external to the system-on-a-chip device 622. However, each of the display 628, the input device 630, the speakers 648, the microphones 646, the antenna 642 and the power supply 644 can be coupled to a component of the system-on-a-chip device 622, such as an interface or a controller.
[0144] Device 600 may include a cordless phone, a mobile communication device, a cell phone, a smartphone, a mobile phone, a laptop, a desktop computer, a computer, a tablet, a set-top box, a personal digital assistant (PDA), a display device, a television, a game console, a music player, a radio, a video player, an entertainment unit, a communication device, a fixed location data unit, a personal media player, a digital video player, a digital video disc (DVD), a tuner, a camera, a navigation device, a set-top box system, an encoder system, or any combination thereof.
[0145] In a particular implementation, one or more components of the systems and devices disclosed herein may be integrated into a decoding device or system (for example, an electronic device, a CODEC or a processor therein), into an encoding device or system, or both. In other implementations, one or more components of the systems and devices disclosed herein may be integrated into a cordless phone, a tablet, a desktop computer, a laptop, a set-top box, a Petition 870190112962, dated 05 / 11 / 2019, pages 88 / 134 84 / 105 music player, a video player, an entertainment unit, a television, a game console, a navigation device, a communication device, a personal digital assistant (PDA), a fixed location data unit, a personal media player or other type of device.
[0146] In conjunction with the techniques described herein, a first apparatus includes means for receiving a bit stream. The bit stream includes a medium channel and a quantized value representing an offset between a reference channel associated with an encoder and a target channel associated with the encoder. The quantized value is based on a value of the offset. The value is associated with the encoder and has a higher precision than the quantized value. For example, the means for receiving the bit stream may include the second device 106 of Figure 1, a receiver (not shown) for the second device 106, the decoder 118 of Figure 1, 2 or 6, the antenna 642 of Figure 6, one or more other circuits, devices, components, modules, or a combination thereof.
[0147] The first apparatus may also include means for decoding the medium channel to generate a decoded medium channel. For example, the means for decoding the medium channel may include the decoder 118 of Figures 1, 2 or 6, the medium channel decoder 202 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a Petition 870190112962, dated 05 / 11 / 2019, pp. 89 / 134 85 / 105 combination of the same.
[0148] The first device may also include means for generating a first channel based on the decoded middle channel. The first channel corresponds to the reference channel. For example, the means for generating the first channel may include the decoder 118 of Figures 1, 2 or 6, the inverse transformer unit 210 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0149] The first device may also include means for generating a second channel based on the decoded middle channel and the quantized value. The second channel corresponds to the target channel. The means for generating the second channel may include the decoder 118 of Figures 1, 2 or 6, the inverse transformation unit 212 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0150] In conjunction with the techniques described herein, a second apparatus includes means for receiving a bit stream from an encoder. The bit stream may include an average channel and a quantized value representing an offset between a reference channel associated with the encoder and a target channel associated with the encoder. Petition 870190112962, dated 05 / 11 / 2019, pp. 90 / 134 The 86 / 105 quantized value may be based on an offset value that has greater precision than the quantized value. For example, the means for receiving the bit stream may include the second device 106 of Figure 1, a receiver (not shown) of the second device 106, the decoder 118 of Figure 1, 2 or 6, the antenna 642 of Figure 6, one or more other circuits, devices, components, modules, or a combination thereof.
[0151] The second device may also include means for decoding the medium channel to generate a decoded medium channel. For example, the means for decoding the medium channel may include the decoder 118 of Figures 1, 2 or 6, the medium channel decoder 202 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0152] The second apparatus may also include means for performing a transformation operation on the decoded medium channel to generate a decoded frequency domain medium channel. For example, the means for performing the transformation operation may include the decoder 118 of Figures 1, 2 or 6, the transformation unit 204 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof. Petition 870190112962, dated 05 / 11 / 2019, pp. 91 / 134 87 / 105
[0153] The second apparatus may also include means for upmixing the decoded frequency domain middle channel to generate a first frequency domain channel and a second frequency domain channel. For example, the means for upmixing may include the decoder 118 of Figures 1, 2 or 6, the upmixer 206 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0154] The second device may also include means for generating a first channel based on the first frequency domain channel. The first channel may correspond to the reference channel. For example, the means for generating the first channel may include the decoder 118 of Figures 1, 2 or 6, the inverse transformer unit 210 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0155] The second device may also include means for generating a second channel based on the second frequency domain channel. The second channel may correspond to the target channel. If the quantized value corresponds to a frequency domain shift, the second frequency domain channel may be shifted by a domain of Petition 870190112962, dated 05 / 11 / 2019, pp. 92 / 134 88 / 105 frequency by the quantized value. If the quantized value corresponds to a time-domain shift, a time-domain version of the second frequency-domain channel can be shifted by the quantized value. The means for generating the second channel may include the decoder 118 of Figures 1, 2, or 6, the inverse transform unit 212 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0156] In conjunction with the techniques described herein, a third apparatus includes means for receiving at least a portion of a bit stream. The bit stream includes a first frame and a second frame. The first frame includes a first portion of a medium channel and a first value of a stereo parameter, and the second frame includes a second portion of the medium channel and a second value of the stereo parameter. The means for reception may include the second device 106 of Figure 1, a receiver (not shown) for the second device 106, the decoder 118 of Figure 1, 2 or 6, the antenna 642 of Figure 6, one or more other circuits, devices, components, modules, or a combination thereof.
[0157] The third device may also include means for decoding the first portion of the medium channel to generate a first portion of a decoded medium channel. For example, the means for decoding Petition 870190112962, dated 05 / 11 / 2019, pages 93 / 134 89 / 105 may include decoder 118 of Figures 1, 2 or 6, medium channel decoder 202 of Figure 2, processor 606 of Figure 6, processors 610 of Figure 6, CODEC 634 of Figure 6, instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0158] The third device may also include means for generating a first portion of a left channel based at least on the first portion of the decoded middle channel and the first value of the stereo parameter. For example, the means for generating the first portion of the left channel may include the decoder 118 of Figures 1, 2 or 6, the inverse transformer unit 210 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0159] The third device may also include means for generating a first portion of a right channel based at least on the first portion of the decoded middle channel and the first value of the stereo parameter. For example, the means for generating the first portion of the right channel may include the decoder 118 of Figures 1, 2 or 6, the inverse transformer unit 212 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a Petition 870190112962, dated 05 / 11 / 2019, pages 94 / 134 90 / 105 processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0160] The third device may also include means for generating, in response to the second frame being unavailable for decoding operations, a second portion of the left channel and a second portion of the right channel based at least on the first value of the stereo parameter. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame. The means for generating the second portion of the left channel and the second portion of the right channel may include the decoder 118 of Figures 1, 2 or 6, the stereo shift value interpolator 216 of Figure 2, the stereo parameter interpolator 208 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0161] In conjunction with the techniques described herein, a fourth device includes means for receiving at least a portion of a bitstream from an encoder. The bitstream may include a first frame and a second frame. The first frame may include a first portion of a medium channel and a first value of a stereo parameter, and the second frame may include a second portion of the medium channel and a second value of the stereo parameter. The means for reception may include the second device 106 of Figure 1, a receiver (not shown) of the second Petition 870190112962, dated 05 / 11 / 2019, pages 95 / 134 91 / 105 device 106, the decoder 118 of Figure 1, 2 or 6, the antenna 642 of Figure 6, one or more other circuits, devices, components, modules, or a combination thereof.
[0162] The fourth apparatus may also include means for decoding the first portion of the medium channel to generate a first portion of a decoded medium channel. For example, the means for decoding the first portion of the medium channel may include the decoder 118 of Figures 1, 2 or 6, the medium channel decoder 202 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0163] The fourth apparatus may also include means for performing a transformation operation on the first portion of the decoded medium channel to generate a first portion of a decoded frequency domain medium channel. For example, the means for performing the transformation operation may include the decoder 118 of Figures 1, 2 or 6, the transformation unit 204 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0164] The fourth device may also include means for upmixing the first portion of the mid-channel of Petition 870190112962, dated 05 / 11 / 2019, pages 96 / 134 92 / 105 frequency domain decoded to generate a first portion of a left frequency domain channel and a first portion of a right frequency domain channel. For example, the means to perform upmixing may include the decoder 118 of Figures 1, 2 or 6, the upmixer 206 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0165] The fourth device may also include means for generating a first portion of a left channel based at least on the first portion of the left frequency domain channel and the first value of the stereo parameter. For example, the means for generating the first portion of the left channel may include the decoder 118 of Figures 1, 2 or 6, the inverse transformer unit 210 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0166] The fourth device may also include means for generating a first portion of a right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter. For example, the means for generating the first portion of the right channel may include the 118 decoder of Petition 870190112962, dated 05 / 11 / 2019, pp. 97 / 134 93 / 105 Figures 1, 2 or 6, the inverse transformation unit 212 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0167] The fourth device may also include means for generating, based at least on the first stereo parameter value, a second left-channel portion and a second right-channel portion in response to a determination that the second frame is unavailable. The second left-channel portion and the second right-channel portion may correspond to a decoded version of the second frame. The means for generating the second left-channel portion and the second right-channel portion may include the decoder 118 of Figures 1, 2, or 6, the stereo offset value interpolator 216 of Figure 2, the stereo parameter interpolator 208 of Figure 2, the shifter 214 of Figure 2, the processor 606 of Figure 6, the processors 610 of Figure 6, the CODEC 634 of Figure 6, the instructions 660 of Figure 6, executable by a processor, one or more other circuits, devices, components, modules, or a combination thereof.
[0168] It should be noted that several functions performed by one or more components of the systems and devices disclosed here are described as being performed by specific components or modules. This division of components and modules is for illustrative purposes only. Petition 870190112962, dated 05 / 11 / 2019, pages 98 / 134 94 / 105 In an alternative implementation, a function performed by a specific component or module can be divided among multiple components or modules. Furthermore, in another alternative implementation, two or more components or modules can be integrated into a single component or module. Each component or module can be implemented using hardware (e.g., a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a DSP, a controller, etc.), software (e.g., instructions executable by a processor), or any combination thereof.
[0169] Referring to Figure 7, a block diagram of a particular illustrative example of a 700 base station is presented. In various implementations, the 700 base station may have more or fewer components than illustrated in Figure 7. In one illustrative example, the 700 base station may include the second device 106 of Figure 1. In another illustrative example, the 700 base station may operate according to one or more of the methods or systems described with reference to Figures 1-3, 4A, 4B, 5A, 5B, and 6.
[0170] Base station 700 may be part of a wireless communication system. The wireless communication system may include multiple base stations and multiple wireless devices. The wireless communication system may be a Long Term Evolution (LTE) system, a Code Division Multiple Access (CDMA) system, a Global System for Mobile Communications (GSM) system, a wireless local area network (WLAN) system, or some other system. Petition 870190112962, dated 05 / 11 / 2019, pages 99 / 134 95 / 105 another wireless system. A CDMA system can implement Wideband CDMA (WCDMA), CDMA IX, Optimized Evolution Data (EVDO), Time Division Synchronous CDMA (TDSCDMA), or some other version of CDMA.
[0171] Wireless devices may also be referred to as user equipment (UE), a mobile station, a terminal, an access terminal, a subscriber unit, a station, etc. Wireless devices may include a mobile phone, a smartphone, a tablet, a wireless modem, a personal digital assistant (PDA), a handheld device, a laptop, a smartbook, a netbook, a tablet, a cordless phone, a wireless local circuit (WLL) station, a Bluetooth device, etc. Wireless devices may include or correspond to device 600 in Figure 6.
[0172] Several functions can be performed by one or more components of the 700 base station (and / or other components not shown), such as sending and receiving messages and data (e.g., audio data). In one particular example, the 700 base station includes a 706 processor (e.g., a CPU). The 700 base station may include a 710 transcoder. The 710 transcoder may include a 708 audio CODEC. For example, the 710 transcoder may include one or more components (e.g., circuits) configured to perform 708 audio CODEC operations. As another example, the 710 transcoder may be configured to execute one or more computer-readable instructions to perform 708 audio CODEC operations. Although the Petition 870190112962, dated 05 / 11 / 2019, pages 100 / 134 96 / 105 Audio codec 708 may be illustrated as a component of transcoder 710; in other examples, one or more components of audio codec 708 may be included in processor 706, another processing component, or a combination thereof. For example, a decoder 738 (e.g., a Vocoder decoder) may be included in a reception data processor 764. As another example, an encoder 736 (e.g., a Vocoder encoder) may be included in a transmission data processor 782. Encoder 736 may include encoder 114 of Figure 1. Decoder 738 may include decoder 118 of Figure 1.
[0173] The 710 transcoder can be used to transcode messages and data between two or more networks. The 710 transcoder can be configured to convert audio message and data from a first format (e.g., a digital format) to a second format. For example, the 738 decoder can decode signals encoded in a first format, and the 736 encoder can encode signals encoded in a second format. Additionally or alternatively, the 710 transcoder can be configured to perform data rate adaptation. For example, the 710 transcoder can perform downconversion at a data rate or upconversion at a data rate without changing the format of the audio data. For example, the 710 transcoder can perform downconversion of 64 kbit / s signals to 16 kbit / s signals. Petition 870190112962, dated 05 / 11 / 2019, pages 101 / 134 97 / 105
[0174] Base station 700 may include a memory 732. The memory 732, as a computer-readable storage device, may contain instructions. The instructions may include one or more instructions that are executable by the processor 706, the transcoder 710, or a combination thereof, to perform one or more operations described with reference to the methods and systems of Figures 1-3, 4A, 4B, 5A, 5B, 6.
[0175] Base station 700 may include multiple transmitters and receivers (e.g., transceivers), such as a first transceiver 752 and a second transceiver 754, coupled to an antenna array. The antenna array may include a first antenna 742 and a second antenna 744. The antenna array may be configured to communicate wirelessly with one or more wireless devices, such as device 600 in Figure 6. For example, the second antenna 744 may receive a data stream 714 (e.g., a bit stream) from a wireless device. The data stream 714 may include messages, data (e.g., encoded speech data), or a combination thereof.
[0176] Base station 700 may include a 760 network connection, such as a backhaul connection. The 760 network connection may be configured to communicate with a core network or one or more base stations on the wireless communication network. For example, base station 700 may receive a second data stream (e.g., messages or audio data) from a core network via the 760 network connection. Base station 700 may process the Petition 870190112962, dated 05 / 11 / 2019, pages 102 / 134 98 / 105 second data stream to generate text messages or audio data and provide the messages or audio data to one or more wireless devices via one or more antennas of the antenna array or to another base station via the 7 60 network connection. In a particular implementation, the 7 60 network connection may be a Wide Area Network (WAN) connection, as an illustrative and non-limiting example. In some implementations, the core network may include or correspond to a Public Switched Telephone Network (PSTN), a packet-frame network, or both.
[0177] Base station 700 may include a media gateway 770 that is connected to network link 760 and processor 706. The media gateway 770 may be used to convert media streams from different telecommunications technologies. For example, the media gateway 770 may convert between different transmission protocols, different encoding schemes, or both. For illustration, the media gateway 770 may convert from PCM signals to RTP (Real-Time Transport Protocol) signals, as an illustrative and non-limiting example. The media gateway 770 may convert data between packet-switched networks (e.g., a Voice over Internet Protocol (VoIP) network, an IP Multimedia Subsystem (IMS), a fourth-generation (4G) wireless network such as LTE, WiMax, and UMB, etc.).), circuit-switched networks (e.g., a PSTN), and hybrid networks (e.g., a second-generation (2G) wireless network, such as GSM, GPRS, and EDGE, a third-generation (3G) wireless network, such as WCDMA, EV-DO, and HSPA, etc.). Petition 870190112962, dated 05 / 11 / 2019, pp. 103 / 134 99 / 105
[0178] Additionally, the 770 media gateway may include a transcoder, such as the 710 transcoder, and may be configured to transcode data when codecs are incompatible. For example, the 770 media gateway may transcode between an Adaptive Multirate (AMR) codec and a G.711 codec, as an illustrative and non-limiting example. The 770 media gateway may include a router and a plurality of physical interfaces. In some implementations, the 770 media gateway may also include a controller (not shown). In a specific implementation, the media gateway controller may be external to the 770 media gateway, external to the 700 base station, or both. The media gateway controller may control and coordinate operations of multiple media gateways.The 770 media gateway can receive control signals from the media gateway controller and can act as a bridge between different transmission technologies, adding service to end-user resources and connections.
[0179] Base station 700 may include a demodulator 762 which is coupled to transceivers 752, 754, receiver data processor 764 and processor 706, and receiver data processor 764 may be used by processor 706. Demodulator 762 may be configured to demodulate modulated signals received from transceivers 752, 754 and to provide demodulated data to receiver data processor 764. Receiver data processor 764 may be used to extract a message or audio data from the demodulated data and send it to Petition 870190112962, dated 05 / 11 / 2019, pages 104 / 134 100 / 105 message or audio data to processor 706.
[0180] Base station 700 may include a 782 transmission data processor and a 784 multi-input multiple-output (MIMO) transmission processor. The 782 transmission data processor may be connected to the 706 processor and the 784 transmission MIMO processor. The 784 transmission MIMO processor may be coupled to transceivers 752, 754, and the 706 processor. In some implementations, the 784 transmission MIMO processor may be connected to the 770 media gateway. The 782 transmission data processor may be configured to receive messages or audio data from the 706 processor and encode the messages or audio data based on an encoding scheme, such as CDMA or OFDM (orthogonal frequency division multiplexing), as an illustrative and non-limiting example. The 782 transmission data processor may provide the encoded data to the 784 transmission MIMO processor.
[0181] Encoded data can be multiplexed with other data, such as pilot data, using CDMA or OFDM techniques to generate multiplexed data. The multiplexed data can be modulated (e.g., symbol-mapped) by the 782 transmission data processor, based on a specific modulation scheme (e.g., binary phase-shift keying (BPSK), quadrature phase-shift keying (QSPK), M-ary phase-shift keying (M-PSK), M-ary quadrature amplitude modulation (M-QAM), etc.) to generate Petition 870190112962, dated 05 / 11 / 2019, pages 105 / 134 101 / 105 modulation symbols. In a specific implementation, codec data and other encoded data can be modulated using different modulation schemes. The data rate, encoding, and modulation for each data stream can be determined by instructions executed by the 706 processor.
[0182] The 784 transmission MIMO processor can be configured to receive modulation symbols from the 782 transmission data processor and can also process the modulation symbols and perform spatial filtering (beamforming) on the data. For example, the 784 transmission MIMO processor can apply spatial filtering weights to the modulation symbols.
[0183] During operation, the second antenna 744 of base station 700 can receive a data stream 714. The second transceiver 754 can receive the data stream 714 from the second antenna 744 and can provide the data stream 714 to the demodulator 762. The demodulator 762 can demodulate modulated signals from the data stream 714 and provide demodulated data to the receiver data processor 764. The receiver data processor 764 can extract audio data from the demodulated data and provide the extracted audio data to the processor 706.
[0184] Processor 706 can provide audio data to transcoder 710 for transcoding. Transcoder 738 of transcoder 710 can decode audio data from a first format into decoded audio data, and encoder 736 can encode the decoded audio data into a second format. In some Petition 870190112962, dated 05 / 11 / 2019, pages 106 / 134 In implementations 102 / 105, the 736 encoder can encode audio data using a higher data rate (e.g., upconversion) or a lower data rate (e.g., downconversion) than that received from the wireless device. In other implementations, the audio data may not be transcoded. Although transcoding (e.g., decoding and encoding) is illustrated as being performed by a 710 transcoder, transcoding operations (e.g., decoding and encoding) can be performed by multiple base station components 700. For example, decoding can be performed by the receiver data processor 764 and encoding can be performed by the transmission data processor 782. In other implementations, the 706 processor can provide the audio data to the media gateway 770 for conversion to another transmission protocol, encoding scheme, or both.The 770 media gateway can provide the converted data to another base station or core network via the 760 network connection.
[0185] Encoded audio data generated in encoder 736 can be provided to transmission data processor 782 or to network connection 760 via processor 706. Transcoded audio data from transcoder 710 can be provided to transmission data processor 782 for encoding according to a modulation scheme, such as OFDM, to generate modulation symbols. Transmission data processor 782 can provide modulation symbols to the MIMO processor of Petition 870190112962, dated 05 / 11 / 2019, pp. 107 / 134 103 / 105 transmission 784 for additional processing and spatial filtering. The transmission MIMO processor 784 can apply spatial filtering weights and can provide modulation symbols to one or more antennas in the antenna array, such as the first antenna 742 through the first transceiver 752. In this way, the base station 700 can provide a transcoded data stream 716, which corresponds to the data stream 714 received from the wireless device, to another wireless device. The transcoded data stream 716 may have a different encoding format, data rate, or both, than the data stream 714. In other implementations, the transcoded data stream 716 may be provided to the network link 760 for transmission to another base station or a core network.
[0186] Those skilled in the art would further appreciate that the various illustrative logic blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processing device such as a hardware processor, or combinations of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as executable hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but these decisions Petition 870190112962, dated 05 / 11 / 2019, pages 108 / 134 104 / 105 of implementation should not be interpreted as causes for a departure from the scope of the present invention.
[0187] The steps of a method or algorithm described in connection with the implementations disclosed herein may be incorporated directly into hardware, in a software module executed by a processor, or in a combination of both. A software module may reside in a memory device, such as random access memory (RAM), magnetoresistive random access memory (MRAM), spin-torque transfer MRAM (STT-MRAM), flash memory, read-only memory (ROM), programmable read-only memory (FROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, a removable disk, or a compact disc read-only memory (CD-ROM). An exemplary memory device is coupled to the processor such that the processor can read information from and write information to the memory device.Alternatively, the memory device may be an integral part of the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and storage medium may reside as discrete components in a computing device or a user terminal.
[0188] The above description of the implementations Petition 870190112962, dated 05 / 11 / 2019, pages 109 / 134 The disclosed 105 / 105 implementations are provided to enable one skilled in the art to produce or utilize the disclosed implementations. Various modifications to these implementations will be readily apparent to those skilled in the art, and the principles defined herein can be applied to other implementations without departing from the scope of the invention. Thus, the present invention intends to be limited to the implementations shown in this document, but should have the broadest possible scope consistent with the principles and novel features, as defined by the following claims.
Claims
CLAIMS 1. Apparatus (700) comprising: a receiver (752, 754) configured to receive at least one portion of a bitstream (160), the bitstream comprising a first frame (190) and a second frame (192), the first frame including a first portion of a medium channel (191) and a first value of a stereo parameter (183), the second frame including a second portion of the medium channel (193) and a second value of the stereo parameter (187), wherein the at least one portion of a bitstream further comprises a quantized value (181) representing an offset between a reference channel associated with an encoder and a target channel associated with the encoder, wherein the quantized value is based on an offset value, the offset value associated with the encoder and having a greater precision than the quantized value and wherein the encoder is configured to generate at least one received portion of the bitstream (160);and a decoder (118) configured to: decode the first portion of the middle channel to generate a first portion of a decoded middle channel (170); generate a first portion of a left channel (302) based at least on the first portion of the decoded middle channel (170) and the first value of the stereo parameter (183); generate a first portion of a right channel (304) based at least on the first portion of the decoded middle channel and the first value of the stereo parameter (183);and the device characterized by the fact that: in response to the second frame being unavailable for decoding operations, it generates a second portion of the left channel (308) and a second portion of the right channel (310) based at least on the first value of the stereo parameter (183) and based on the quantized value, wherein the second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.; 2. Apparatus, according to claim 1, characterized in that the decoder is further configured to, in response to the second frame being available for decoding operations, generate an interpolated value of the stereo parameter based on the first value of the stereo parameter and the second value of the stereo parameter (287).
3. Device according to claim 1, characterized in that the decoder is further configured to, in response to the second frame being unavailable for decoding operations, generate at least the second portion of the middle channel and a second portion of a side channel based at least on the first value of the stereo parameter, the first portion of the middle channel, the first portion of the left channel or the first portion of the right channel.
4. Apparatus, according to claim 3, characterized in that the decoder is further configured to, in response to the second frame being unavailable for decoding operations, generate the second portion of the left channel and the second portion of the right channel based on the second portion of the middle channel, the second portion of the side channel and a third value of the stereo parameter, and preferably in that the third value of the stereo parameter is at least based on the first value of the stereo parameter, an interpolated value of the stereo parameter and an encoding mode.
5. Device, according to claim 1, characterized in that the decoder is additionally configured to, in response to Petition 870250038999, dated 05 / 13 / 2025, page 15 / 22 3 / 8, the second frame being unavailable for decoding operations, generate at least the second portion of the left channel and the second portion of the right channel based at least on the first value of the stereo parameter, in the first portion of the left channel and in the first portion of the right channel.
6. Apparatus, according to claim 1, characterized in that the decoder is further configured to: perform a transformation operation on the first portion of the decoded mid-channel to generate a first portion of a decoded frequency-domain mid-channel; upmix the first portion of the decoded frequency-domain mid-channel based on the first value of the stereo parameter to generate a first portion of a left frequency-domain channel and a first portion of a right frequency-domain channel; perform a first time-domain operation on the first portion of the left frequency-domain channel to generate the first portion of the left channel; and perform a second time-domain operation on the first portion of the right frequency-domain channel to generate the first portion of the right channel.
7. Apparatus, according to claim 6, characterized in that, in response to the second frame being unavailable for decoding operations, the decoder is configured to: generate a second portion of the decoded medium channel based on the first portion of the decoded medium channel; perform a second transformation operation on the second portion of the decoded medium channel to generate a second portion of the decoded frequency domain medium channel; Petition 870250038999, dated 05 / 13 / 2025, p.16 / 22 4 / 8 perform an upmix of the second portion of the decoded frequency domain medium channel to generate a second portion of the left frequency domain channel and a second portion of the right frequency domain channel; perform a third time domain operation on the second portion of the left frequency domain channel to generate the second portion of the left channel; and perform a fourth time domain operation on the second portion of the right frequency domain channel to generate the second portion of the right channel.
8. Apparatus, according to claim 7, characterized in that the decoder is additionally configured to estimate the second value of the stereo parameter based on the first value of the stereo parameter, wherein the second estimated value of the stereo parameter is used to upmix the second portion of the decoded frequency domain mid-channel, or wherein the decoder is additionally configured to interpolate the second value of the stereo parameter based on the first value of the stereo parameter.where the second interpolated value of the stereo parameter is used to upmix the second portion of the decoded frequency domain mid-channel, or where the decoder is configured to perform an interpolation operation on the first portion of the decoded mid-channel to generate the second portion of the decoded mid-channel, or where the decoder is configured to perform an estimation operation on the first portion of the decoded mid-channel to generate the second portion of the decoded mid-channel.
9. Apparatus, according to claim 1, characterized by the fact that the stereo parameter comprises an interchannel phase difference parameter or an interchannel level difference parameter or an interchannel time difference parameter or an interchannel correlation parameter or a spectral slope parameter or an interchannel gain parameter or an interchannel voice parameter or an interchannel pitch parameter and wherein the receiver and decoder are integrated into a mobile device or wherein the receiver and decoder are integrated into a base station.
10. Method comprising: receiving, in a decoder (118), at least one portion of a bitstream, the bitstream comprising a first frame (190) and a second frame (192), the first frame including a first portion of a middle channel (191) and a first value of a stereo parameter (183), the second frame including a second portion of the middle channel (193) and a second value of the stereo parameter (187), the at least one portion of a bitstream further comprising a quantized value (181) representing an offset between a reference channel associated with an encoder and a target channel associated with the encoder, wherein the quantized value is based on an offset value, the offset value associated with the encoder and having a greater precision than the quantized value and wherein the encoder is configured to generate the at least one received portion of the bitstream (160);decode the first portion of the middle channel to generate a first portion of a decoded middle channel (170); generate a first portion of a left channel (302) based at least on the first portion of the decoded middle channel (170) and the first value of the stereo parameter (183); generate a first portion of a right channel (304) based at least on the first portion of the decoded middle channel and the first value of the stereo parameter (183); and the method characterized by the fact that: in response to the second frame being unavailable for decoding operations, generate a second portion of the left channel (308) and a second portion of the right channel (310) based at least on the first value of the stereo parameter (183) and based on the quantized value (181), wherein the second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.; 11. Method according to claim 10, characterized in that it further comprises: performing a transformation operation on the first portion of the decoded mid-channel to generate a first portion of a decoded frequency-domain mid-channel; upmixing the first portion of the decoded frequency-domain mid-channel based on the first value of the stereo parameter to generate a first portion of a left frequency-domain channel and a first portion of a right frequency-domain channel; performing a first time-domain operation on the first portion of the left frequency-domain channel to generate the first portion of the left channel; and performing a second time-domain operation on the first portion of the right frequency-domain channel to generate the first portion of the right channel.
12. Method, according to claim 11, characterized in that it further comprises, in response to the second frame being unavailable for decoding operations: Petition 870250038999, dated 05 / 13 / 2025, p.19 / 22 7 / 8 generate a second portion of the decoded mid-channel based on the first portion of the decoded mid-channel; perform a second transformation operation on the second portion of the decoded mid-channel to generate a second portion of the decoded frequency-domain mid-channel; upmix the second portion of the decoded frequency-domain mid-channel to generate a second portion of the left frequency-domain channel and a second portion of the right frequency-domain channel; perform a third time-domain operation on the second portion of the left frequency-domain channel to generate the second portion of the left channel; and perform a fourth time-domain operation on the second portion of the right frequency-domain channel to generate the second portion of the right channel.
13. Method, according to claim 12, characterized in that it further comprises estimating the second value of the stereo parameter based on the first value of the stereo parameter, wherein the second estimated value of the stereo parameter is used to upmix the second portion of the decoded frequency domain mid-channel; or further comprises interpolating the second value of the stereo parameter based on the first value of the stereo parameter, wherein the second interpolated value of the stereo parameter is used to upmix the second portion of the decoded frequency domain mid-channel; or further comprises performing an interpolation operation on the first portion of the decoded mid-channel to generate the second portion of the decoded mid-channel; or further comprises performing an estimation operation in Petition 870250038999, dated 05 / 13 / 2025, p.20 / 22 8 / 8 first portion of the decoded middle channel to generate the second portion of the decoded middle channel.
14. Method according to claim 10, characterized in that the decoder is integrated into a mobile device or in that the decoder is integrated into a base station.
15. Memory characterized in that it comprises instructions stored therein which, when executed by a processor within a decoder, cause the processor to perform the method as defined in any one of claims 10 to 14.