Stereo parameters for stereo decoding
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-04-27
- Publication Date
- 2026-08-14
AI Technical Summary
然而,相较于传输低精确度数据,传输指示音频信号的对准的高精确度数据会使用增加的传输资源
Smart Images

Figure CN116665682B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on April 27, 2018, with application number "201880030918.7" and invention title "Stereo Parameters for Stereo Decoding".
[0002] Priority Claim
[0003] This application claims priority to co-owned U.S. Provisional Patent Application No. 62 / 505,041, filed May 11, 2017, entitled “STEREOPARAMETERS FOR STEREO DECODING,” and U.S. Non-Provisional Patent Application No. 15 / 962,834, filed April 25, 2018, entitled “STEREO PARAMETERS FOR STEREO DECODING,” the entire contents of each of which are expressly incorporated herein by reference. Technical Field
[0004] This invention generally relates to decoding audio signals. Background Technology
[0005] Technological advancements have led to smaller and more powerful computing devices. For example, a variety of portable personal computing devices exist today, including cordless phones such as mobile phones and smartphones, tablets, and laptops, which are small, lightweight, and easily carried by users. These devices can transmit voice and data packets via wireless networks. Furthermore, many of these devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Moreover, these devices can process executable instructions that can access the internet, including software applications such as web browser applications. Therefore, these devices can contain significant computing power.
[0006] A computing device may include or be coupled to multiple microphones to receive audio signals. Typically, the sound source is closer to a first microphone than to a second microphone. Therefore, due to the respective distances between the first and second microphones and the sound source, a second audio signal received from the second microphone may be delayed relative to a first audio signal received from the first microphone. In other embodiments, the first audio signal may be delayed relative to the second audio signal. In stereo coding, audio signals from the microphones may be encoded to produce a center channel signal and one or more side channel signals. The center channel signal may correspond to the sum of the first and second audio signals. The side channel signals may correspond to the difference between the first and second audio signals. Due to the delay in receiving the second audio signal relative to the first audio signal, the first audio signal may not be aligned with the second audio signal. The delay may be indicated by an encoded shift value (e.g., stereo parameters) transmitted to the decoder. Precise alignment of the first and second audio signals enables efficient encoding for transmission to the decoder. However, transmitting high-precision data indicating the alignment of the audio signals uses increased transmission resources compared to transmitting low-precision data. It can also encode other stereo parameters that indicate the characteristics between the first audio signal and the second audio signal and transmit them to the decoder.
[0007] The decoder can reconstruct the first and second audio signals based at least on the center channel signal and stereo parameters, which are received at the decoder via a bitstream comprising a series of frames. The accuracy at the decoder during audio signal reconstruction can be based on the accuracy of the encoder. For example, encoded high-precision shift values can be received at the decoder, enabling the decoder to reproduce delays in the reconstructed versions of the first and second audio signals with high accuracy. If the shift values are unavailable at the decoder, for example when frames of data transmitted via the bitstream are corrupted due to noisy transmission conditions, then the shift values can be requested and retransmitted to the decoder to accurately reproduce the delays between the audio signals. For example, the decoder's accuracy in reproducing delays can exceed the limits of human audible perception to detect changes in delay. Summary of the Invention
[0008] According to one embodiment of the present invention, a device includes a receiver configured to receive at least a portion of a bitstream. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a center channel and a first value of a stereo parameter, and the second frame includes a second portion of the center channel and a second value of the stereo parameter. The device further includes a decoder configured to decode the first portion of the center channel to produce a decoded first portion of the center channel. The decoder is further configured to produce a first portion of a left channel based at least on the first portion of the decoded center channel and the first value of the stereo parameter, and to produce a first portion of a right channel based at least on the first portion of the decoded center channel and the first value of the stereo parameter. The decoder is further configured to produce the second portion of the left channel and the second portion of the right channel based at least on the first value of the stereo parameter in response to the second frame being unavailable for decoding. The second portion of the left channel and the second portion of the right channel correspond to decoded versions of the second frame.
[0009] According to another embodiment, a method for decoding a signal includes receiving at least a portion of a bitstream. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a center channel and a first value of a stereo parameter, and the second frame includes a second portion of the center channel and a second value of the stereo parameter. The method further includes decoding the first portion of the center channel to generate a decoded first portion of the center channel. The method further includes generating a first portion of a left channel based at least on the first portion of the decoded center channel and the first value of the stereo parameter, and generating a first portion of a right channel based at least on the first portion of the decoded center channel and the first value of the stereo parameter. The method further includes generating the second portion of the left channel and the second portion of the right channel based at least on the first value of the stereo parameter in response to the second frame being unavailable for decoding. The second portion of the left channel and the second portion of the right channel correspond to decoded versions of the second frame.
[0010] According to another embodiment, a non-transitory computer-readable medium includes instructions that, when executed by a processor within a decoder, cause the processor to perform an operation comprising receiving at least a portion of a bitstream. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a center channel and a first value of a stereo parameter, and the second frame includes a second portion of the center channel and a second value of the stereo parameter. The operation further includes decoding the first portion of the center channel to generate a decoded first portion of the center channel. The operation further includes generating a first portion of a left channel based at least on the first portion of the decoded center channel and the first value of the stereo parameter, and generating a first portion of a right channel based at least on the first portion of the decoded center channel and the first value of the stereo parameter. The operation also includes generating the second portion of the left channel and the second portion of the right channel based at least on the first value of the stereo parameter in response to the second frame being unavailable for decoding. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0011] According to another embodiment, an apparatus includes means for receiving at least a portion of a bitstream. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a center channel and a first value of a stereo parameter, and the second frame includes a second portion of the center channel and a second value of the stereo parameter. The apparatus further includes means for decoding the first portion of the center channel to generate a decoded first portion of the center channel. The apparatus further includes means for generating a first portion of a left channel based at least on the first portion of the decoded center channel and the first value of the stereo parameter, and means for generating a first portion of a right channel based at least on the first portion of the decoded center channel and the first value of the stereo parameter. The apparatus further includes means for generating the second portion of the left channel and the second portion of the right channel based at least on the first value of the stereo parameter in response to the second frame being unavailable for decoding. The second portion of the left channel and the second portion of the right channel correspond to decoded versions of the second frame.
[0012] According to another embodiment, a device includes a receiver configured to receive at least a portion of a bitstream from an encoder. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a center channel and a first value of a stereo parameter. The second frame includes a second portion of the center channel and a second value of the stereo parameter. The device further includes a decoder configured to decode the first portion of the center channel to produce a decoded first portion of the center channel. The decoder is further configured to perform a transform operation on the first portion of the decoded center channel to produce a decoded first portion of the frequency domain center channel. The decoder is further configured to overmix the first portion of the decoded frequency domain center channel to produce a first portion of a left frequency domain channel and a first portion of a right frequency domain channel. The decoder is further configured to produce the first portion of the left channel based at least on the first portion of the left frequency domain channel and the first value of the stereo parameter. The decoder is further configured to produce the first portion of the right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter. The decoder is also configured to determine that the second frame is not available for decoding. The decoder is further configured to generate a second portion of the left channel and a second portion of the right channel in response to determining that the second frame is unavailable, based at least on the first value of the stereo parameters. The second portion of the left channel and the second portion of the right channel correspond to the decoded version of the second frame.
[0013] According to another embodiment, a method for decoding a signal includes receiving at least a portion of a bitstream from an encoder at a decoder. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a center channel and a first value of a stereo parameter. The second frame includes a second portion of the center channel and a second value of the stereo parameter. The method further includes decoding the first portion of the center channel to generate a first portion of a decoded center channel. The method further includes performing a transform operation on the first portion of the decoded center channel to generate a first portion of a decoded frequency domain center channel. The method further includes upmixing the first portion of the decoded frequency domain center channel to generate a first portion of a left frequency domain channel and a first portion of a right frequency domain channel. The method further includes generating the first portion of the left channel based at least on the first portion of the left frequency domain channel and the first value of the stereo parameter. The method further includes generating the first portion of the right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter. The method further includes determining that the second frame is not available for decoding. The method further includes generating a second portion of the left channel and a second portion of the right channel in response to determining that the second frame is unavailable, based at least on the first value of the stereo parameters. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0014] According to another embodiment, a non-transitory computer-readable medium includes instructions that, when executed by a processor within a decoder, cause the processor to perform an operation comprising receiving at least a portion of a bitstream from an encoder. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a center channel and a first value for a stereo parameter. The second frame includes a second portion of the center channel and a second value for the stereo parameter. The operation further includes decoding the first portion of the center channel to generate a first portion of a decoded center channel. The operation further includes performing a transform operation on the first portion of the decoded center channel to generate a first portion of a decoded frequency domain center channel. The operation further includes upmixing the first portion of the decoded frequency domain center channel to generate a first portion of a left frequency domain channel and a first portion of a right frequency domain channel. The operation further includes generating a first portion of the left channel based at least on the first portion of the left frequency domain channel and the first value of the stereo parameter. The operation further includes generating a first portion of the right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter. The operation further includes determining that the second frame is not available for decoding. The operation further includes generating a second portion of the left channel and a second portion of the right channel in response to determining that the second frame is unavailable, based at least on the first value of the stereo parameters. The second portion of the left channel and the second portion of the right channel correspond to the decoded version of the second frame.
[0015] According to another embodiment, an apparatus includes means for receiving at least a portion of a bitstream from an encoder. The bitstream includes a first frame and a second frame. The first frame includes a first portion of a center channel and a first value of a stereo parameter. The second frame includes a second portion of the center channel and a second value of the stereo parameter. The apparatus further includes means for decoding the first portion of the center channel to generate a first portion of a decoded center channel. The apparatus further includes means for performing a transform operation on the first portion of the decoded center channel to generate a first portion of a decoded frequency domain center channel. The apparatus further includes means for upmixing the first portion of the decoded frequency domain center channel to generate a first portion of a left frequency domain channel and a first portion of a right frequency domain channel. The apparatus further includes means for generating a first portion of a left channel based at least on the first portion of the left frequency domain channel and the first value of the stereo parameter. The apparatus further includes means for generating a first portion of a right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameter. The apparatus further includes means for determining that the second frame is not available for decoding. The device further includes means for generating a second portion of the left channel and a second portion of the right channel, at least based on the first value of the stereo parameters, in response to a determination that the second frame is unavailable. The second portion of the left channel and the second portion of the right channel correspond to a decoded version of the second frame.
[0016] According to another embodiment, a device includes a receiver and a decoder. The receiver is configured to receive a bitstream comprising an encoded intermediate channel and a quantized value, the quantized value representing a shift between a reference channel associated with an encoder and a target channel associated with the encoder. The quantized value is based on the shift value. The shift value is associated with the encoder and has greater accuracy than the quantized value. The decoder is configured to decode the encoded intermediate channel to generate a decoded intermediate channel, and to generate a first channel based on the decoded intermediate channel. The decoder is further configured to generate a second channel based on the decoded intermediate channel and the quantized value. The first channel corresponds to the reference channel and the second channel corresponds to the target channel.
[0017] According to another embodiment, a method for decoding a signal includes receiving a bitstream at a decoder comprising an intermediate channel and a quantized value, the quantized value representing a shift between a reference channel associated with an encoder and a target channel associated with the encoder. The quantized value is a value based on the shift. The value is associated with the encoder and has greater accuracy than the quantized value. The method further includes decoding the intermediate channel to generate a decoded intermediate channel. The method further includes generating a first channel based on the decoded intermediate channel and generating a second channel based on the decoded intermediate channel and the quantized value. The first channel corresponds to the reference channel and the second channel corresponds to the target channel.
[0018] According to another embodiment, a non-transitory computer-readable medium includes instructions that, when executed by a processor within a decoder, cause the processor to perform an operation comprising receiving at the decoder a bitstream comprising an intermediate channel and a quantized value, the quantized value representing a shift between a reference channel associated with an encoder and a target channel associated with the encoder. The quantized value is a value based on the shift. The value is associated with the encoder and has greater accuracy than the quantized value. The operation further includes decoding the intermediate channel to generate a decoded intermediate channel. The operation further includes generating a first channel based on the decoded intermediate channel and generating a second channel based on the decoded intermediate channel and the quantized value. The first channel corresponds to the reference channel and the second channel corresponds to the target channel.
[0019] According to another embodiment, an apparatus includes means for receiving at a decoder a bitstream comprising an intermediate channel and a quantized value, the quantized value representing a shift between a reference channel associated with an encoder and a target channel associated with the encoder. The quantized value is a value based on the shift. The value is associated with the encoder and has greater accuracy than the quantized value. The apparatus further includes means for decoding the intermediate channel to generate a decoded intermediate channel. The apparatus further includes means for generating a first channel based on the decoded intermediate channel, and means for generating a second channel based on the decoded intermediate channel and the quantized value. The first channel corresponds to the reference channel and the second channel corresponds to the target channel.
[0020] According to another embodiment, a device includes a receiver configured to receive a bit stream from an encoder. The bit stream includes an intermediate channel and a quantized value representing a shift between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value is a value based on the shift, which has greater accuracy than the quantized value. The device further includes a decoder configured to decode the intermediate channel to generate a decoded intermediate channel. The decoder is further configured to perform a transform operation on the decoded intermediate channel to generate a decoded frequency domain intermediate channel. The decoder is further configured to overmix the decoded frequency domain intermediate channel to generate a first frequency domain channel and a second frequency domain channel. The decoder is also configured to generate a first channel based on the first frequency domain channel. The first channel corresponds to the reference channel. The decoder is further configured to generate a second channel based on the second frequency domain channel. The second channel corresponds to the target channel. If the quantized value corresponds to a frequency domain shift, then the second frequency domain channel is shifted by the quantized value in the frequency domain, and if the quantized value corresponds to a time domain shift, then the time domain version of the second frequency domain channel is shifted by the quantized value.
[0021] According to another embodiment, a method includes receiving a bitstream from an encoder at a decoder. The bitstream includes an intermediate channel and a quantized value representing a shift between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value is a value based on the shift, which has greater accuracy than the quantized value. The method further includes decoding the intermediate channel to generate a decoded intermediate channel. The method further includes performing a transform operation on the decoded intermediate channel to generate a decoded frequency domain intermediate channel. The method further includes upmixing the decoded frequency domain intermediate channel to generate a first frequency domain channel and a second frequency domain channel. The method further includes generating a first channel based on the first frequency domain channel. The first channel corresponds to the reference channel. The method further includes generating a second channel based on the second frequency domain channel. The second channel corresponds to the target channel. If the quantized value corresponds to a frequency domain shift, then the second frequency domain channel is shifted by the quantized value in the frequency domain, and if the quantized value corresponds to a time domain shift, then the time domain version of the second frequency domain channel is shifted by the quantized value.
[0022] According to another embodiment, a non-transitory computer-readable medium includes instructions for decoding a signal. When executed by a processor within a decoder, the instructions cause the processor to perform an operation including receiving a bit stream from an encoder. The bit stream includes an intermediate channel and a quantized value representing a shift between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value is a value based on the shift, which has greater precision than the quantized value. The operation further includes decoding the intermediate channel to generate a decoded intermediate channel. The operation further includes performing a transform operation on the decoded intermediate channel to generate a decoded frequency domain intermediate channel. The operation further includes upmixing the decoded frequency domain intermediate channel to generate a first frequency domain channel and a second frequency domain channel. The operation further includes generating a first channel based on the first frequency domain channel. The first channel corresponds to the reference channel. The operation further includes generating a second channel based on the second frequency domain channel. The second channel corresponds to the target channel. If the quantized value corresponds to a frequency domain shift, then the second frequency domain channel is shifted by the quantized value in the frequency domain, and if the quantized value corresponds to a time domain shift, then the time domain version of the second frequency domain channel is shifted by the quantized value.
[0023] According to another embodiment, an apparatus includes means for receiving a bit stream from an encoder. The bit stream includes an intermediate channel and a quantized value, the quantized value representing a shift between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value is a value based on the shift, which has greater accuracy than the quantized value. The apparatus further includes means for decoding the intermediate channel to generate a decoded intermediate channel. The apparatus further includes means for performing a transform operation on the decoded intermediate channel to generate a decoded frequency domain intermediate channel. The apparatus further includes means for upmixing the decoded frequency domain intermediate channel to generate a first frequency domain channel and a second frequency domain channel. The apparatus further includes means for generating a first channel based on the first frequency domain channel. The first channel corresponds to the reference channel. The apparatus further includes means for generating a second channel based on the second frequency domain channel. The second channel corresponds to the target channel. If the quantized value corresponds to a frequency domain shift, then the second frequency domain channel is shifted by the quantized value in the frequency domain, and if the quantized value corresponds to a time domain shift, then the time domain version of the second frequency domain channel is shifted by the quantized value.
[0024] After reviewing the entire application, other embodiments, advantages and features of the invention will become apparent, the entire application comprising the following sections: description of drawings, detailed description and claims. Attached Figure Description
[0025] Figure 1 A block diagram of a specific illustrative instance of a system containing a decoder, which is operable to estimate stereo parameters of missing frames and decode audio signals using quantized stereo parameters;
[0026] Figure 2 For illustration Figure 1 A diagram of the decoder;
[0027] Figure 3 A diagram illustrating an example of how to predict the stereo parameters of missing frames at the decoder;
[0028] Figure 4A These are non-limiting illustrative examples of methods for decoding audio signals;
[0029] Figure 4B for Figure 4A A non-limiting illustrative example of a more detailed version of a method for decoding audio signals;
[0030] Figure 5A This is another non-limiting illustrative example of a method for decoding audio signals;
[0031] Figure 5B for Figure 5A A non-limiting illustrative example of a more detailed version of a method for decoding audio signals;
[0032] Figure 6 The block diagram is a specific illustrative example of a device containing a decoder, which estimates the stereo parameters of missing frames and uses the quantized stereo parameters to decode the audio signal; and
[0033] Figure 7 The diagram shows a base station that can be operated to estimate the stereo parameters of missing frames and use the quantized stereo parameters to decode the audio signal. Detailed Implementation
[0034] Specific aspects of the invention are described below with reference to the accompanying drawings. In this description, common features are indicated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular embodiments and are not intended to limit the embodiments. For example, unless the context clearly indicates otherwise, the singular forms “a / an” and “the” are intended to also include the plural forms. It can be further understood that the terms “comprises” and “comprising” are used interchangeably with “includes” or “including”. Additionally, it should be understood that the terms “wherein” and “where” are used interchangeably. As used herein, ordinal terms (e.g., “first,” “second,” “third,” etc.) used to modify elements such as structures, components, operations, etc., do not themselves indicate any priority or order of said element relative to another element, but merely distinguish said element from another element having the same name (if the ordinal term were not used). As used herein, the term “set” refers to one or more of a particular element, and the term “multiple” refers to multiple (e.g., two or more) of a particular element.
[0035] In this invention, terms such as “determine,” “calculate,” “shift,” “adjust,” etc., are used to describe how one or more operations are performed. It should be noted that such terms should not be considered limiting, and other techniques can be used to perform similar operations. Additionally, as mentioned herein, “generate,” “calculate,” “use,” “select,” “access,” and “determine” are used interchangeably. For example, “generate,” “calculate,” or “determine” a parameter (or signal) can refer to actively generating, calculating, or determining the parameter (or signal), or it can refer to using, selecting, or accessing the parameter (or signal) that has already been generated, for example, by another component or device.
[0036] This invention discloses systems and apparatus operable to encode multiple audio signals. The apparatus may include an encoder configured to encode multiple audio signals. Multiple recording devices—such as multiple microphones—can be used to capture multiple audio signals simultaneously in time. In some instances, multiple audio signals (or multichannel audio) can be synthesized (e.g., artificially) by multiplexing several audio channels recorded at the same time or at different times. As illustrative examples, simultaneous recording or multiplexing of audio channels can produce 2-channel configurations (i.e., stereo: left and right), 5.1-channel configurations (left, right, center, left surround, right surround, and low frequency emphasis (LFE) channels), 7.1-channel configurations, 7.1+4-channel configurations, 22.2-channel configurations, or N-channel configurations.
[0037] An audio capture device in a telephone conference room (or telepresence room) may include multiple microphones for acquiring spatial audio. The spatial audio may include speech and encoded and transmitted background audio. Depending on how the multiple microphones are arranged and the location of a given source (e.g., a speaker) relative to the microphones and the room dimensions, speech / audio from the source (e.g., the speaker) may arrive at the microphones at different times. For example, the proximity of the sound source (e.g., the speaker) to a first microphone associated with the device may be greater than the proximity to a second microphone associated with the device. Therefore, sound from the sound source may arrive at the first microphone earlier than it arrives at the second microphone. The device may receive a first audio signal via the first microphone and a second audio signal via the second microphone.
[0038] Mid-side (MS) decoding and parametric stereo (PS) decoding are stereo decoding techniques that offer improved performance compared to dual-mono decoding. In dual-mono decoding, the left (L) channel (or signal) and right (R) channel (or signal) are decoded independently without utilizing inter-channel correlation. MS decoding reduces redundancy between related L / R channel pairs by transforming the left and right channels into sum and difference channels (e.g., side channels) before decoding. The sum and difference signals are decoded either waveform-based or based on a model from MS decoding. The number of bits required for the sum signal is relatively greater than that for the side signals. PS decoding reduces redundancy in each subband by transforming the L / R signals into a sum signal and a set of side parameters. Side parameters can indicate inter-channel intensity difference (IID), inter-channel phase difference (IPD), inter-channel time difference (ITD), side or residual prediction gain, etc. The sum signal is waveform-decoded and transmitted along with the side parameters. In a hybrid system, side channels may be waveform-decoded in the lower frequency band (e.g., less than 2 kHz) and PS-decoded in the upper frequency band (e.g., greater than or equal to 2 kHz), where inter-channel phase preservation is less perceptually critical. In some implementations, PS decoding may also be used in the lower frequency band prior to waveform decoding to reduce inter-channel redundancy.
[0039] MS decoding and PS decoding can be performed in the frequency domain, sub-band domain, or time domain. In some instances, the left and right channels may be uncorrelated. For example, the left and right channels may contain uncorrelated synthesized signals. When the left and right channels are uncorrelated, the decoding efficiency of MS decoding, PS decoding, or both can approach that of dual-mono decoding.
[0040] Depending on the recording configuration, there may be time shifts between the left and right channels, as well as other spatial effects such as echo and room reverberation. Without compensation for the time shifts and phase mismatches between channels, the sum and difference channels can contain considerable energy, thus reducing the decoding gain associated with MS or PS techniques. This reduction in decoding gain can be based on the amount of time (or phase) shift. The considerable energy of the sum and difference signals can limit the use of MS decoding in certain frames where the channels are time-shifted but highly correlated. In stereo decoding, the middle channel (e.g., the sum channel) and side channels (e.g., the difference channel) can be generated based on the following formula:
[0041] M = (L + R) / 2, S = (LR) / 2, Formula 1
[0042] Where M corresponds to the center channel, S corresponds to the side channel, L corresponds to the left channel, and R corresponds to the right channel.
[0043] In some situations, the center channel and side channels can be generated based on the following formula:
[0044] M = c(L + R), S = c(LR), Formula 2
[0045] Where c corresponds to the frequency-dependent composite value. Generating the center and side channels based on Formula 1 or Formula 2 can be called "downmixing". The reverse process of generating the left and right channels from the center and side channels based on Formula 1 or Formula 2 can be called "upmixing".
[0046] In some cases, the middle channel can be based on other formulas, for example:
[0047] M=(L+g D R) / 2, or formula 3
[0048] M = g1L + g2R (Formula 4)
[0049] Where g1 + g2 = 1.0, and where g D Here is the gain parameter. In other instances, downmixing can be performed in the frequency band, where mid(b) = c1L(b) + c2R(b), where c1 and c2 are complex numbers, and side(b) = c3L(b) - c4R(b), where c3 and c4 are complex numbers.
[0050] A specific approach for selecting between MS decoding and dual-mono decoding for a particular frame may include generating an intermediate signal and a side signal, calculating the energy of the intermediate and side signals, and determining whether to perform MS decoding based on the energy. For example, MS decoding may be performed in response to determining that the energy ratio of the side signal to the intermediate signal is less than a threshold. For illustration, if the right channel is shifted for at least a first time (e.g., about 0.001 seconds or 48 samples at 48 kHz), then for a spoken speech frame, the first energy of the intermediate signal (corresponding to the sum of the left and right signals) may be equivalent to the second energy of the side signal (corresponding to the difference between the left and right signals). When the first energy is equivalent to the second energy, a higher number of bits can be used to encode the side channel, thereby reducing the decoding performance of MS decoding compared to dual-mono decoding. When the first energy is equivalent to the second energy (e.g., when the ratio of the first energy to the second energy is greater than or equal to a threshold), dual-mono decoding can therefore be used. In alternative approaches, a decision between MS decoding and dual-mono decoding can be made for a specific frame based on a comparison of thresholds and normalized cross-correlation values for the left and right channels.
[0051] In some instances, the encoder may determine a mismatch value indicating the amount of time misalignment between a first audio signal and a second audio signal. As used herein, the terms "time shift value," "shift value," and "mismatch value" are used interchangeably. For example, the encoder may determine a time shift value indicating a shift (e.g., time mismatch) of the first audio signal relative to the second audio signal. The time mismatch value may correspond to the amount of time delay between the reception of the first audio signal at a first microphone and the reception of the second audio signal at a second microphone. Furthermore, the encoder may determine the time mismatch value on a frame-by-frame basis—for example, based on each 20-millisecond (ms) speech / audio frame. For example, the time mismatch value may correspond to the amount of time delay of a second frame of the second audio signal relative to a first frame of the first audio signal. Alternatively, the time mismatch value may correspond to the amount of time delay of a first frame of the first audio signal relative to a second frame of the second audio signal.
[0052] When the sound source is closer to the first microphone than to the second microphone, the frames of the second audio signal may be delayed relative to the frames of the first audio signal. In this case, the first audio signal may be referred to as the "reference audio signal" or "reference channel," and the delayed second audio signal may be referred to as the "target audio signal" or "target channel." Alternatively, when the sound source is closer to the second microphone than to the first microphone, the frames of the first audio signal may be delayed relative to the frames of the second audio signal. In this case, the second audio signal may be referred to as the reference audio signal or reference channel, and the delayed first audio signal may be referred to as the target audio signal or target channel.
[0053] Depending on the location of the sound source (e.g., a speaker) in the conference room or telepresence room, or how the location of the sound source (e.g., a speaker) changes relative to the microphone, the reference channel and target channel can change from frame to frame; similarly, the time delay value can also change from frame to frame. However, in some implementations, the time mismatch value can always be positive to indicate the amount of delay of the "target" channel relative to the "reference" channel. Furthermore, the time mismatch value can correspond to a "non-causal shift" value, which delays the target channel by "pulling back" in time to such that the target channel is aligned (e.g., maximized alignment) with the "reference" channel. A downmixing algorithm for determining the center and side channels can be performed on the reference channel and the non-causal shifted target channel.
[0054] The encoder can determine the time mismatch value based on a reference audio channel and multiple time mismatch values applied to the target audio channel. For example, the first frame X of the reference audio channel can be received at a first time (m1). The first specific frame Y of the target audio channel can be received at a second time (n1) corresponding to the first time mismatch value, for example, shift1 = n1 - m1. Additionally, the second frame of the reference audio channel can be received at a third time (m2). The second specific frame of the target audio channel can be received at a fourth time (n2) corresponding to the second time mismatch value, for example, shift2 = n2 - m2.
[0055] The device can perform a framing or buffering algorithm at a first sampling rate (e.g., a 32kHz sampling rate (i.e., 640 samples per frame)) to generate frames (e.g., 20ms samples). The encoder can estimate a time mismatch value (e.g., shift1) equal to zero samples in response to determining that the first frame of the first audio signal and the second frame of the second audio signal arrive at the device at the same time. The left channel (e.g., corresponding to the first audio signal) and the right channel (e.g., corresponding to the second audio signal) can be time-aligned. In some cases, even when aligned, the left and right channels may differ in energy due to various reasons (e.g., microphone calibration).
[0056] In some instances, the left and right channels may be temporally misaligned for various reasons (e.g., the speaker's sound source may be closer to one microphone than to the other, and the distance between the two microphones may be greater than a threshold (e.g., 1 to 20 cm)). The position of the sound source relative to the microphones can introduce different delays in the left and right channels. Additionally, there may be gain, energy, or level differences between the left and right channels.
[0057] In some instances, where more than two channels exist, a reference channel is initially selected based on the channel's level or energy, and subsequently refined based on the time mismatch values between different channel pairs—e.g., t1(ref, ch2), t2(ref, ch3), t3(ref, ch4), ...—where ch1 is initially the reference channel and t1(.), t2(.), etc., are functions used to estimate the mismatch values. If all time mismatch values are positive, then ch1 is considered the reference channel. If any of the mismatch values is negative, then the reference channel is reconfigured to the channel associated with the mismatch value that caused the negative value, and the above process continues until the optimal selection of the reference channel is achieved (e.g., based on maximizing the decorrelation of the maximum number of side channels). Hysteresis can be used to overcome any abrupt changes in the reference channel selection.
[0058] In some instances, when multiple speakers alternate (e.g., without overlap), the arrival time of audio signals from multiple sound sources (e.g., speakers) at the microphone can vary. In such cases, the encoder can dynamically adjust the time mismatch value based on the speaker to identify the reference channel. In other instances, multiple speakers may speak simultaneously, which can cause varying time mismatch values depending on which speaker is loudest, closest to the microphone, etc. In such cases, the identification of the reference and target channels can be based on the varying time shift value in the current frame and the estimated time mismatch value in the previous frame, and on the energy or time evolution of the first and second audio signals.
[0059] In some instances, when the first and second audio signals potentially exhibit little (e.g., no) correlation, the two signals may be synthesized or artificially generated. It should be understood that the examples described herein are illustrative and may provide guidance in determining the relationship between the first and second audio signals in similar or different situations.
[0060] The encoder can generate comparison values (e.g., difference or cross-correlation values) based on comparisons between a first frame of a first audio signal and multiple frames of a second audio signal. Each of the multiple frames can correspond to a specific time mismatch value. The encoder can generate a first estimated time mismatch value based on the comparison values. For example, the first estimated time mismatch value can correspond to a comparison value indicating a higher time similarity (or lower difference) between the first frame of the first audio signal and a corresponding first frame of the second audio signal.
[0061] An encoder can determine a final time mismatch value by refining a series of estimated time mismatch values in multiple stages. For example, the encoder can first estimate a "provisional" time mismatch value based on comparison values generated from stereo preprocessed and resampled versions of a first audio signal and a second audio signal. The encoder can generate interpolated comparison values associated with the time mismatch value immediately following the estimated "provisional" time mismatch value. The encoder can determine a second estimated "interpolated" time mismatch value based on the interpolated comparison values. For example, the second estimated "interpolated" time mismatch value can correspond to a specific interpolated comparison value that indicates higher temporal similarity (or lower difference) compared to the remaining interpolated comparison values and the first estimated "provisional" time mismatch value. If the second estimated "interpolated" time mismatch value of the current frame (e.g., the first frame of the first audio signal) differs from the final time mismatch value of the previous frame (e.g., a frame of the first audio signal preceding the first frame), then the "interpolated" time mismatch value of the current frame is further "corrected" to improve the time similarity between the first audio signal and the shifted second audio signal. Specifically, the third estimated "corrected" time mismatch value corresponds to a more accurate measure of time similarity by examining the second estimated "interpolated" time mismatch value of the current frame and the final estimated time mismatch value of the previous frame. The third estimated "corrected" time mismatch value is further adjusted to estimate the final time mismatch value by limiting any spurious changes in the time mismatch value between frames, and is further controlled to prevent a switch from a negative time mismatch value to a positive time mismatch value (or vice versa) in two successive (or consecutive) frames as described herein.
[0062] In some instances, the encoder may prevent switching between positive and negative time mismatch values in consecutive frames or in adjacent frames, or vice versa. For example, the encoder may set the final time mismatch value to a specific value (e.g., 0) indicating no time shift, based on an estimated "interpolated" or "corrected" time mismatch value for the first frame and a corresponding estimated "interpolated" or "corrected" or final time mismatch value in a specific frame preceding the first frame. For illustration, the encoder may set the final time mismatch value of the current frame to indicate no time shift, i.e., shift1 = 0, in response to determining that an estimated "provisional" or "interpolated" or "corrected" time mismatch value for the current frame (e.g., the first frame) is positive and another estimated "provisional" or "interpolated" or "corrected" or "final" estimated time mismatch value for the previous frame (e.g., a frame preceding the first frame) is negative. Alternatively, the encoder may also set the final time mismatch value of the current frame to indicate no time shift, i.e., shift1 = 0, in response to determining that an estimated "provisional", "interpolated", or "corrected" time mismatch value of the current frame (e.g., the first frame) is negative and another estimated "provisional", "interpolated", "corrected", or "final" estimated time mismatch value of the previous frame (e.g., a frame preceding the first frame) is positive.
[0063] The encoder can select a frame of the first or second audio signal as a "reference" or "target" based on a time mismatch value. For example, in response to determining that the final time mismatch value is positive, the encoder can generate a reference channel or signal indicator with a first value (e.g., 0) indicating that the first audio signal is a "reference" signal and the second audio signal is a "target" signal. Alternatively, in response to determining that the final time mismatch value is negative, the encoder can generate a reference channel or signal indicator with a second value (e.g., 1) indicating that the second audio signal is a "reference" signal and the first audio signal is a "target" signal.
[0064] The encoder can estimate a relative gain (e.g., a relative gain parameter) associated with a reference signal and a non-causally shifted target signal. For example, in response to determining that the final time mismatch is positive, the encoder can estimate a gain value used to normalize or equalize the amplitude or power level of a first audio signal relative to a second audio signal, said gain value being offset to the non-causally shifted time mismatch value (e.g., the absolute value of the final time mismatch). Alternatively, in response to determining that the final time mismatch is negative, the encoder can estimate a gain value used to normalize or equalize the amplitude or power level of the non-causally shifted first audio signal relative to the second audio signal. In some instances, the encoder can estimate a gain value used to normalize or equalize the amplitude or power level of a "reference" signal relative to a non-causally shifted "target" signal. In other instances, the encoder can estimate a gain value (e.g., a relative gain value) based on the reference signal relative to the target signal (e.g., the unshifted target signal).
[0065] The encoder can generate at least one encoded signal (e.g., an intermediate signal, a side signal, or both) based on a reference signal, a target signal, a non-causal time mismatch value, and a relative gain parameter. In other embodiments, the encoder can generate at least one encoded signal (e.g., an intermediate channel, a side channel, or both) based on a reference channel and a time-mismatch-adjusted target channel. The side signal may correspond to the difference between a first sample of a first frame of a first audio signal and a selected sample of a selected frame of a second audio signal. The encoder can select the selected frame based on the final time mismatch value. Since the difference between the first sample and the selected sample is reduced compared to other samples of the second audio signal corresponding to a frame of the second audio signal received by the device simultaneously with the first frame, fewer bits can be used to encode the side channel signal. The transmitter of the device can transmit at least one encoded signal, a non-causal time mismatch value, a relative gain parameter, a reference channel or signal indicator, or a combination thereof.
[0066] The encoder can generate at least one encoded signal (e.g., an intermediate signal, a side signal, or both) based on a reference signal, a target signal, a non-causal time mismatch value, a relative gain parameter, low-frequency band parameters of a specific frame of a first audio signal, high-frequency band parameters of a specific frame, or a combination thereof. The specific frame may precede the first frame. Certain low-frequency band parameters, high-frequency band parameters, or a combination thereof from one or more previous frames can be used to encode the intermediate signal, side signal, or both of the first frame. Encoding the intermediate signal, side signal, or both based on low-frequency band parameters, high-frequency band parameters, or a combination thereof can improve the estimation of the non-causal time mismatch value and the inter-channel relative gain parameter. Low-frequency band parameters, high-frequency band parameters, or a combination thereof may include pitch parameters, phonation parameters, decoder type parameters, low-frequency band energy parameters, high-frequency band energy parameters, tilt parameters, pitch gain parameters, FCB gain parameters, decoding mode parameters, speech activity parameters, noise estimation parameters, signal-to-noise ratio parameters, formant parameters, speech / music decision parameters, non-causal shift, inter-channel gain parameters, or a combination thereof. The transmitter of the device can transmit at least one encoded signal, a non-causal time mismatch value, a relative gain parameter, a reference channel (or signal) indicator, or a combination thereof. In this invention, terms such as "determine," "calculate," "shift," "adjust," etc., are used to describe how one or more operations are performed. It should be noted that such terms should not be considered limiting, and other techniques can be used to perform similar operations.
[0067] According to some implementations, the final time mismatch value (e.g., shift value) is an "unquantized" value indicating a "true" shift between the target channel and the reference channel. While all digital values are "quantized" due to the precision provided by the system storing or using the digital values, as used herein, digital values are "quantized" if they are produced by quantization operations used to reduce the precision of the digital values (e.g., to reduce the range or bandwidth associated with the digital values), and otherwise "unquantized." As a non-limiting example, the first audio signal may be the target channel, and the second audio signal may be the reference channel. If the true shift between the target channel and the reference channel is thirty-seven samples, then the target channel may be shifted by thirty-seven samples at the encoder to produce a shifted target channel that is time-aligned with the reference channel. In other implementations, both channels may be shifted such that the relative shift between the channels is equal to the final shift value (37 samples in this example). This relative shift of the channels to the shift value achieves the effect of time-aligning the channels. High-efficiency encoders align channels as closely as possible to reduce decoding entropy and thus increase decoding efficiency, as decoding entropy is sensitive to shift changes between channels. The shifted target and reference channels are used to generate the encoded intermediate channels, which are transmitted as part of the bitstream to the decoder. Additionally, the final time mismatch value can be quantized and transmitted as part of the bitstream to the decoder. For example, a "floor" of four can be used to quantize the final time mismatch value, such that the quantized final time mismatch value equals nine (e.g., approximately 37 / 4).
[0068] The decoder can decode the intermediate channel to produce a decoded intermediate channel, and the decoder can generate the first and second channels based on the decoded intermediate channel. For example, the decoder can use stereo parameters contained in the bitstream to upmix the decoded intermediate channel to produce the first and second channels. The first and second channels can be time-aligned at the decoder; however, the decoder can shift one or more of the channels relative to each other based on a quantized final time mismatch value. For example, if the first channel corresponds to the target channel (e.g., the first audio signal) at the encoder, then the decoder can shift the first channel by thirty-six samples (e.g., 4*9) to produce a shifted first channel. Perceptually, the shifted first and second channels are similar to the target channel and the reference channel, respectively. For example, if a thirty-seven-sample shift between the target channel and the reference channel at the encoder corresponds to a 10ms shift, then a thirty-six-sample shift between the shifted first and second channels at the decoder is perceptually similar to and can be perceptually indistinguishable from the thirty-seven-sample shift.
[0069] See Figure 1This illustrates a specific example of system 100. System 100 includes a first device 104 communicatively coupled to a second device 106 via a network 120. The network 120 may include one or more wireless networks, one or more wired networks, or a combination thereof.
[0070] The first device 104 includes an encoder 114, a transmitter 110, and one or more input interfaces 112. A first input interface of the input interfaces 112 may be coupled to a first microphone 146. A second input interface of the input interfaces 112 may be coupled to a second microphone 148. The first device 104 may also include a memory 153 configured to store analyzed data, as described below. The second device 106 may include a decoder 118 and a memory 154. The second device 106 may be coupled to a first speaker 142, a second speaker 144, or both.
[0071] During operation, the first device 104 may receive a first audio signal 130 from a first microphone 146 via a first input interface, and a second audio signal 132 from a second microphone 148 via a second input interface. The first audio signal 130 may correspond to either a right channel signal or a left channel signal. The second audio signal 132 may correspond to the other of the right channel signal or the left channel signal. As described herein, the first audio signal 130 may correspond to a reference channel, and the second audio signal 132 may correspond to a target channel. However, it should be understood that in other embodiments, the first audio signal 130 may correspond to a target channel, and the second audio signal 132 may correspond to a reference channel. In other embodiments, there may be no assignment of a reference channel and a target channel. In such cases, channel alignment at the encoder and channel dealignment at the decoder may be performed on either or both of the channels, such that the relative shift between the channels is based on a shift value.
[0072] First microphone 146 and second microphone 148 can receive audio from sound source 152 (e.g., user, speaker, ambient noise, musical instrument, etc.). In certain aspects, first microphone 146, second microphone 148, or both can receive audio from multiple sound sources. The multiple sound sources may include a primary (or most primary) sound source (e.g., sound source 152) and one or more secondary sound sources. The one or more secondary sound sources may correspond to traffic, background music, another speaker, street noise, etc. Sound source 152 (e.g., the primary sound source) may be closer to first microphone 146 than to second microphone 148. Therefore, the time at which an audio signal is received from sound source 152 via first microphone 146 at input interface 112 may be earlier than the time at which an audio signal is received from sound source 152 via second microphone 148. This natural delay in the multi-channel signal obtained via multiple microphones can introduce a time shift between the first audio signal 130 and the second audio signal 132.
[0073] The first device 104 may store a first audio signal 130, a second audio signal 132, or both in memory 153. The encoder 114 may determine a first shift value 180 (e.g., a non-causal shift value) indicating a shift (e.g., a non-causal shift) of the first audio signal 130 relative to the second audio signal 132 for the first frame 190. The first shift value 180 may be a value (e.g., an unquantized value) representing a shift between a reference channel (e.g., the first audio signal 130) and a target channel (e.g., the second audio signal 132) for the first frame 190. The first shift value 180 may be stored in memory 153 as analysis data. The encoder 114 may also determine a second shift value 184 indicating a shift of the first audio signal 130 relative to the second audio signal 132 for the second frame 192. The second frame 192 may be after the first frame 190 (e.g., later in time than the first frame 190). The second shift value 184 may be a value (e.g., an unquantized value) representing the shift between the reference channel (e.g., the first audio signal 130) and the target channel (e.g., the second audio signal 132) for the second frame 192. The second shift value 184 may also be stored in the memory 153 as analysis data.
[0074] Therefore, shift values 180 and 184 (e.g., mismatch values) can indicate the amount of time mismatch (e.g., time delay) between the first audio signal 130 and the second audio signal 132 for the first frame 190 and the second frame 192, respectively. As mentioned herein, "time delay" can correspond to "temporal delay." Time mismatch can indicate the time delay between the reception of the first audio signal 130 via the first microphone 146 and the reception of the second audio signal 132 via the second microphone 148. For example, a first value (e.g., a positive value) of shift values 180 and 184 can indicate that the second audio signal 132 is delayed relative to the first audio signal 130. In this example, the first audio signal 130 can correspond to a preamble signal and the second audio signal 132 can correspond to a hysteresis signal. A second value (e.g., a negative value) of shift values 180 and 184 can indicate that the first audio signal 130 is delayed relative to the second audio signal 132. In this example, the first audio signal 130 may correspond to a hysteresis signal and the second audio signal 132 may correspond to a preamble signal. A third value (e.g., 0) of the shift values 180 and 184 may indicate that there is no delay between the first audio signal 130 and the second audio signal 132.
[0075] Encoder 114 can quantize a first shift value 180 to produce a first quantized shift value 181. For illustration, if the first shift value 180 (e.g., a true shift value) equals thirty-seven samples, then encoder 114 can quantize the first shift value 180 based on a threshold to produce the first quantized shift value 181. As a non-limiting example, if the threshold equals four, then the first quantized shift value 181 can equal nine (e.g., approximately 37 / 4). As described below, the first shift value 180 can be used to produce a first portion 191 of the middle channel, and the first quantized shift value 181 can be encoded into bitstream 160 and transmitted to the second device 106. As used herein, a “portion” of a signal or channel includes: one or more frames of a signal or channel; one or more subframes of a signal or channel; one or more samples, bits, blocks, words, or other segments of a signal or channel; or any combination thereof. Similarly, encoder 114 can quantize the second shift value 184 to produce a second quantized shift value 185. For illustrative purposes, if the second shift value 184 equals thirty-six samples, then encoder 114 can quantize the second shift value 184 based on a minimum threshold to produce the second quantized shift value 185. As a non-limiting example, the second quantized shift value 185 can also equal nine (e.g., 36 / 4). As described below, the second shift value 184 can be used to produce a second portion 193 of the middle channel, and the second quantized shift value 185 can be encoded into bitstream 160 and transmitted to the second device 106.
[0076] The encoder 114 may also generate a reference signal indicator based on shift values 180, 184. For example, the encoder 114 may generate a reference signal indicator having a first value (e.g., 0) indicating that the first audio signal 130 is a “reference” signal and the second audio signal 132 corresponds to a “target” signal in response to determining that the first shift value 180 indicates a first value (e.g., a positive value).
[0077] Encoder 114 can time-align the first audio signal 130 and the second audio signal 132 based on shift values 180 and 184. For example, for a first frame 190, encoder 114 can time-shift the second audio signal 132 by a first shift value 180 to produce a time-shifted second audio signal aligned with the first audio signal 130. Although the second audio signal 132 is described as undergoing a time shift in the time domain, it should be understood that the second audio signal 132 can undergo a phase shift in the frequency domain to produce the time-shifted second audio signal 132. For example, the first shift value 180 can correspond to a frequency domain shift value. For a second frame 192, encoder 114 can time-shift the second audio signal 132 by a second shift value 184 to produce a time-shifted second audio signal aligned with the first audio signal 130. Although the second audio signal 132 is described as undergoing a time shift in the time domain, it should be understood that the second audio signal 132 may undergo a phase shift in the frequency domain to produce the shifted second audio signal 132. For example, the second shift value 184 may correspond to a frequency domain shift value.
[0078] Encoder 114 may generate one or more additional stereo parameters (e.g., stereo parameters other than shift values 180 and 184) for each frame based on samples from the reference channel and samples from the target channel. As a non-limiting example, encoder 114 may generate a first stereo parameter 182 for the first frame 190 and a second stereo parameter 186 for the second frame 192. Non-limiting examples of stereo parameters 182 and 186 may include other shift values, inter-channel phase difference parameters, inter-channel level difference parameters, inter-channel time difference parameters, inter-channel correlation parameters, spectral tilt parameters, inter-channel gain parameters, inter-channel phonation parameters, or inter-channel pitch parameters.
[0079] For illustrative purposes, if stereo parameters 182 and 186 correspond to gain parameters, then for each frame, encoder 114 can generate gain parameters (e.g., decoder-decoder gain parameters) based on samples of a reference signal (e.g., the first audio signal 130) and samples of a target signal (e.g., the second audio signal 132). For example, for the first frame 190, encoder 114 can select samples of the second audio signal 132 based on a first shift value 180 (e.g., a non-causal shift value). As mentioned herein, selecting samples of the audio signal based on a shift value can correspond to generating a modified (e.g., time-shifted or frequency-shifted) audio signal by adjusting (e.g., shifting) the audio signal based on the shift value and selecting samples of the modified audio signal. For example, encoder 114 can generate a time-shifted second audio signal by shifting the second audio signal 132 based on the first shift value 180, and can select samples of the time-shifted second audio signal. Encoder 114 may determine a gain parameter of a selected sample based on a first sample of a first frame 190 of the first audio signal 130 in response to determining that the first audio signal 130 is a reference signal. As an example, the gain parameter may be based on one of the following equations:
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086] Where g D Corresponding to the relative gain parameter used for downmixing, Ref(n) corresponds to a sample of the "reference" signal, N1 corresponds to the first shift value 180 of the first frame 190, and Targ(n+N1) corresponds to a sample of the "target" signal. The gain parameter (g) can be modified, for example, based on one of equations 1a to 1f. D This is to incorporate long-term smoothing / hysteresis logic to avoid large gain jumps between frames.
[0087] Encoder 114 can quantize stereo parameters 182 and 186 to generate quantized stereo parameters 183 and 187 that are encoded into bit stream 160 and transmitted to second device 106. For example, encoder 114 can quantize the first stereo parameter 182 to generate a first quantized stereo parameter 183, and encoder 114 can quantize the second stereo parameter 186 to generate a second quantized stereo parameter 187. Quantized stereo parameters 183 and 187 may have lower resolution (e.g., less precision) compared to stereo parameters 182 and 186, respectively.
[0088] For each frame 190, 192, encoder 114 can generate one or more encoded signals based on shift values 180, 184, other stereo parameters 182, 186, and audio signals 130, 132. For example, for the first frame 190, encoder 114 can generate a first portion 191 of the center channel based on a first shift value 180 (e.g., an unquantized shift value), a first stereo parameter 182, and audio signals 130, 132. Additionally, for the second frame 192, encoder 114 can generate a second portion 193 of the center channel based on a second shift value 184 (e.g., an unquantized shift value), a second stereo parameter 186, and audio signals 130, 132. According to some embodiments, encoder 114 can generate side channels (not shown) for each frame 190, 192 based on shift values 180, 184, other stereo parameters 182, 186, and audio signals 130, 132.
[0089] For example, encoder 114 may generate portions 191 and 193 of the middle channel based on one of the following equations:
[0090] M = Ref(n) + g D Targ(n+N1), Equation 2aM=Ref(n)+Targ(n+N1), Equation 2bM=Ref(n-N2)+Targ(n+N1-N2), where N2 can take any arbitrary value, Equation 2c
[0091] Where M corresponds to the middle channel, g D Corresponding to the relative gain parameters used for downmixing (e.g., stereo parameters 182, 186), Ref(n) corresponds to a sample of the “reference” signal, N1 corresponds to shift values 180, 184, and Targ(n+N1) corresponds to a sample of the “target” signal.
[0092] Encoder 114 can generate a side channel based on one of the following equations:
[0093] S = Ref(n) - g D Targ(n+N1), Equation 3a
[0094] S = g D Ref(n) - Targ(n+N1), Equation 3b
[0095] S = Ref(n-N2)-g D Targ(n+N1-N2), where N2 can take any arbitrary value.
[0096] Equation 3c
[0097] Where S corresponds to the side channel signal, g D Corresponding to the relative gain parameters used for downmixing (e.g., stereo parameters 182, 186), Ref(n) corresponds to a sample of the “reference” signal, N1 corresponds to shift values 180, 184, and Targ(n+N1) corresponds to a sample of the “target” signal.
[0098] Transmitter 110 can transmit bitstream 160 to second device 106 via network 120. First frame 190 and second frame 192 can be encoded into bitstream 160. For example, the first portion 191 of the center channel, the first quantized shift value 181, and the first quantized stereo parameter 183 can be encoded into bitstream 160. Additionally, the second portion 193 of the center channel, the second quantized shift value 185, and the second quantized stereo parameter 187 can be encoded into bitstream 160. Side channel information can also be encoded into bitstream 160. Although not shown, additional information can also be encoded into bitstream 160 for each frame 190, 192. As a non-limiting example, a reference channel indicator can be encoded into bitstream 160 for each frame 190, 192.
[0099] Due to poor transmission conditions, some data encoded into bitstream 160 may be lost during transmission. Packet loss may occur due to poor transmission conditions, frame erasure may occur due to poor radio conditions, packets may arrive late due to high jitter, and so on. According to a non-limiting illustrative example, the second device 106 may receive the second portion 193 of the center channel of the first frame 190 and the second frame 192 of bitstream 160. Therefore, the second quantized shift value 185 and the second quantized stereo parameter 187 may be lost during transmission due to poor transmission conditions.
[0100] The second device 106 can therefore receive at least a portion of the bit stream 160 transmitted by the first device 102. The second device 106 can store the received portion of the bit stream 160 in memory 154 (e.g., in a buffer). For example, the first frame 190 can be stored in memory 154, and the second portion 193 of the middle channel of the second frame 192 can also be stored in memory 154.
[0101] Decoder 118 can decode the first frame 190 to generate a first output signal 126 corresponding to the first audio signal 130, and a second output signal 128 corresponding to the second audio signal 132. For example, decoder 118 can decode a first portion 191 of the center channel to generate a first portion 170 of the decoded center channel. Decoder 118 can also perform a transformation operation on the first portion 170 of the decoded center channel to generate a first portion 171 of the frequency-domain (FD) decoded center channel. Decoder 118 can upmix the first portion 171 of the FD decoded center channel to generate a first frequency-domain channel (not shown) associated with the first output signal 126 and a second frequency-domain channel (not shown) associated with the second output signal 128. During upmixing, decoder 118 can apply a first quantized stereo parameter 183 to the first portion 171 of the FD decoded center channel.
[0102] It should be noted that in other implementations, the decoder 118 may not perform a transformation operation, but instead perform an upmixing based on the middle channel, some stereo parameters (such as downmixing gain), and additionally based on the decoded side channel in the time domain when available, to produce a first time-domain channel (not shown) associated with the first output channel 126 and a second time-domain channel (not shown) associated with the second output channel 128.
[0103] If the first quantized shift value 181 corresponds to a frequency domain shift value, then the decoder 118 can shift the second frequency domain channel by the first quantized shift value 181 to generate a second shifted frequency domain channel (not shown). The decoder 118 can perform an inverse transform operation on the first frequency domain channel to generate a first output signal 126. The decoder 118 can also perform an inverse transform operation on the second shifted frequency domain channel to generate a second output signal 128.
[0104] If the first quantized shift value 181 corresponds to a time-domain shift value, then the decoder 118 can perform an inverse transform operation on the first frequency domain channel to generate a first output signal 126. The decoder 118 can also perform an inverse transform operation on the second frequency domain channel to generate a second time-domain channel. The decoder 118 can shift the second time-domain channel by the first quantized shift value 181 to generate a second output signal 128. Therefore, the decoder 118 can use the first quantized shift value 181 to simulate a perceptible difference between the first output signal 126 and the second output signal 128. The first speaker 142 can output the first output signal 126, and the second speaker 144 can output the second output signal 128. In some implementations where the inverse transform operation can be omitted to directly generate the first and second time-domain channels, as described above, the inverse transform operation can be performed in the time domain. It should also be noted that the presence of a time-domain shift value at decoder 118 may simply indicate that the decoder is configured to perform a time-domain shift, and in some implementations, although a time-domain shift may be available at decoder 118 (indicating that the decoder performs a shift operation in the time domain), the encoder supplied with the received bit stream may have performed a frequency-domain shift operation or a time-domain shift operation for channel alignment.
[0105] If decoder 118 determines that the second frame 192 is unavailable for decoding (e.g., determining that the second quantized shift value 185 and the second quantized stereo parameter 187 are unavailable), then decoder 118 may generate output signals 126, 128 for the second frame 192 based on the stereo parameters associated with the first frame 190. For example, decoder 118 may estimate or interpolate the second quantized shift value 185 based on the first quantized shift value 181. Additionally, decoder 118 may estimate or interpolate the second quantized stereo parameter 187 based on the first quantized stereo parameter 183.
[0106] After estimating the second quantized shift value 185 and the second quantized stereo parameter 187, the decoder 118 can generate output signals 126 and 128 for the second frame 192 in a similar manner to generating output signals 126 and 128 for the first frame 190. For example, the decoder 118 can decode the second portion 193 of the center channel to generate the second portion 172 of the decoded center channel. The decoder 118 can also perform a transform operation on the second portion 172 of the decoded center channel to generate the second frequency-domain decoded center channel 173. Based on the estimated quantized shift value and the estimated quantized stereo parameter 187, the decoder 118 can upmix the second frequency-domain decoded center channel 173, perform an inverse transform on the upmixed signal, and shift the resulting signal to generate output signals 126 and 128. Regarding Figure 2 A more detailed description of an example of the decoding operation.
[0107] System 100 can align the channels as closely as possible at encoder 114 to reduce decoding entropy and thus increase decoding efficiency, because decoding entropy is sensitive to shift changes between channels. For example, encoder 114 can use unquantized shift values to accurately align the channels because unquantized shift values have relatively high resolution. At decoder 118, compared to using unquantized shift values, quantized stereo parameters can be used to simulate the perceptible difference between output signals 126 and 128 using a reduced number of bits, and stereo parameters from one or more previous frames can be used to interpolate or estimate missing stereo parameters (due to poor transmission). According to some embodiments, shift values 180 and 184 (e.g., unquantized shift values) can be used to shift the target channel in the frequency domain, and quantized shift values 181 and 185 can be used to shift the target channel in the time domain. For example, shift values used for time-domain stereo coding can have lower resolution than shift values used for frequency-domain stereo coding.
[0108] See Figure 2 The illustration shows a specific implementation of the decoder 118. The decoder 118 includes a center channel decoder 202, a conversion unit 204, an upmixer 206, an inverse conversion unit 210, an inverse conversion unit 212, and a shifter 214.
[0109] Can Figure 1 The bitstream 160 is provided to the decoder 118. For example, a first portion 191 of the center channel of the first frame 190 and a second portion 193 of the center channel of the second frame 192 can be provided to the center channel decoder 202. Additionally, stereo parameters 201 can be provided to the upmixer 206 and the shifter 214. Stereo parameters 201 may include a first quantized shift value 181 associated with the first frame 190 and a first quantized stereo parameter 183 associated with the first frame 190. As mentioned above regarding... Figure 1 As described, due to poor transmission conditions, decoder 118 may not receive the second quantized shift value 185 associated with the second frame 192 and the second quantized stereo parameter 187 associated with the second frame 192.
[0110] To decode the first frame 190, the center channel decoder 202 can decode a first portion 191 of the center channel to produce a first portion 170 of the decoded center channel (e.g., a time-domain center channel). According to some embodiments, two asymmetric windows can be applied to the first portion 170 of the decoded center channel to produce a windowed portion of the time-domain center channel. The first portion 170 of the decoded center channel is provided to the transformation unit 204. The transformation unit 204 can be configured to perform a transformation operation on the first portion 170 of the decoded center channel to produce a first portion 171 of the frequency-domain decoded center channel. The first portion 171 of the frequency-domain decoded center channel is provided to the upmixer 206. According to some embodiments, the windowing and transformation operations can be completely skipped, and the first portion 170 of the decoded center channel (e.g., a time-domain center channel) can be directly provided to the upmixer 206.
[0111] Upmixer 206 can upmix a first portion 171 of the frequency-domain decoded intermediate channel to generate portions of frequency-domain channel 250 and frequency-domain channel 254. Upmixer 206 can apply a first quantized stereo parameter 183 during upmixing to the first portion 171 of the frequency-domain decoded intermediate channel to generate portions of frequency-domain channels 250 and 254. According to an embodiment where the first quantized shift value 181 includes a frequency-domain shift (e.g., the first quantized shift value 181 corresponds to a first quantized frequency-domain shift value 281), upmixer 206 can perform a frequency-domain shift (e.g., phase shift) based on the first quantized frequency-domain shift value 281 to generate a portion of frequency-domain channel 254. A portion of frequency-domain channel 250 is provided to inverse transform unit 210, and a portion of frequency-domain channel 254 is provided to inverse transform unit 212. According to some implementations, the upmixer 206 can be configured to operate the time-domain channels where stereo parameters (e.g., based on a target gain value) can be applied in the time domain.
[0112] Inverse transform unit 210 can perform an inverse transform operation on a portion of frequency domain channel 250 to generate a portion of time domain channel 260. The portion of time domain channel 260 is provided to shifter 214. Inverse transform unit 212 can perform an inverse transform operation on a portion of frequency domain channel 254 to generate a portion of time domain channel 264. The portion of time domain channel 264 is also provided to shifter 214. In embodiments where upmixing is performed in the time domain, the inverse transform operation following the upmixing operation can be skipped.
[0113] According to the embodiment where the first quantized shift value 181 corresponds to the first quantized frequency domain shift value 281, the shifter 214 can bypass the shift operation and pass portions of the time domain channels 260 and 264 as portions of the output signals 126 and 128, respectively. According to the embodiment where the first quantized shift value 181 includes a time domain shift (e.g., the first quantized shift value 181 corresponds to the first quantized time domain shift value 291), the shifter 214 can shift a portion of the time domain channel 264 by the first quantized time domain shift value 291 to generate a portion of the second output signal 128.
[0114] Therefore, decoder 118 can use a quantized shift value with reduced accuracy (compared to the unquantized shift value used at encoder 114) to generate portions of output signals 126, 128 for the first frame 190. Using the quantized shift value to shift output signal 128 relative to output signal 126 restores the user's perception of the shift at encoder 114.
[0115] To decode the second frame 192, the center channel decoder 202 can decode the second portion 193 of the center channel to produce a second portion 172 of the decoded center channel (e.g., a time-domain center channel). According to some embodiments, two asymmetric windows can be applied to the second portion 172 of the decoded center channel to produce a windowed portion of the time-domain center channel. The second portion 172 of the decoded center channel is provided to the transformation unit 204. The transformation unit 204 can be configured to perform a transformation operation on the second portion 172 of the decoded center channel to produce a second portion 173 of the frequency-domain decoded center channel. The second portion 173 of the frequency-domain decoded center channel is provided to the upmixer 206. According to some embodiments, the windowing and transformation operations can be completely skipped, and the second portion 172 of the decoded center channel (e.g., a time-domain center channel) can be directly provided to the upmixer 206.
[0116] As mentioned above Figure 1 As described, due to poor transmission conditions, decoder 118 may not receive the second quantized shift value 185 and the second quantized stereo parameter 187. As a result, the stereo parameters for the second frame 192 may not be accessible by upmixer 206 and shifter 214. Upmixer 206 includes a stereo parameter interpolator 208 configured to interpolate (or estimate) the second quantized shift value 185 based on the first quantized frequency domain shift value 281. For example, stereo parameter interpolator 208 may generate the second interpolated frequency domain shift value 285 based on the first quantized frequency domain shift value 281. Stereo parameter interpolator 208 may also be configured to interpolate (or estimate) the second quantized stereo parameter 187 based on the first quantized stereo parameter 183. For example, stereo parameter interpolator 208 may generate the second interpolated stereo parameter 287 based on the first quantized stereo parameter 183.
[0117] Upmixer 206 can upmix the second portion 173 of the frequency-domain decoded intermediate channel to produce portions of frequency-domain channel 252 and frequency-domain channel 256. Upmixer 206 can apply a second interpolated stereo parameter 287 during upmixing to the second portion 173 of the frequency-domain decoded intermediate channel to produce portions of frequency-domain channels 252 and 256. According to an embodiment where the first quantized shift value 181 includes a frequency-domain shift (e.g., the first quantized shift value 181 corresponds to a first quantized frequency-domain shift value 281), upmixer 206 can perform a frequency-domain shift (e.g., phase shift) based on the second interpolated frequency-domain shift value 285 to produce a portion of frequency-domain channel 256. A portion of frequency-domain channel 252 is provided to inverse transform unit 210, and a portion of frequency-domain channel 256 is provided to inverse transform unit 212.
[0118] Inverse transform unit 210 performs an inverse transform operation on a portion of frequency domain channel 252 to generate a portion of time domain channel 262. The portion of time domain channel 262 is provided to shifter 214. Inverse transform unit 212 performs an inverse transform operation on a portion of frequency domain channel 256 to generate a portion of time domain channel 266. The portion of time domain channel 266 is also provided to shifter 214. In an embodiment where upmixer 206 operates on the time domain channel, the output of upmixer 206 can be provided to shifter 214, and inverse transform units 210 and 212 can be skipped or omitted.
[0119] Shifter 214 includes a shift value interpolator 216 configured to interpolate (or estimate) a second quantized shift value 185 based on a first quantized time-domain shift value 291. For example, shift value interpolator 216 can generate a second interpolated time-domain shift value 295 based on the first quantized time-domain shift value 291. According to an embodiment where the first quantized shift value 181 corresponds to a first quantized frequency-domain shift value 281, shifter 214 can bypass the shift operation and pass portions of time-domain channels 262 and 266 as portions of output signals 126 and 128, respectively. According to an embodiment where the first quantized shift value 181 corresponds to the first quantized time-domain shift value 291, shifter 214 can shift a portion of time-domain channel 266 by the second interpolated time-domain shift value 295 to generate a second output signal 128.
[0120] Therefore, decoder 118 can approximate stereo parameters (e.g., shift values) based on stereo parameters or variations of stereo parameters from previous frames. For example, decoder 118 can extrapolate stereo parameters for frames lost during transmission (e.g., second frame 192) from stereo parameters of one or more previous frames.
[0121] See Figure 3The diagram 300 illustrates the stereo parameters used to predict missing frames at the decoder. According to diagram 300, the first frame 190 may be successfully transmitted from encoder 114 to decoder 118, while the second frame 192 may not be successfully transmitted. For example, the second frame 192 may be lost during transmission due to poor transmission conditions.
[0122] Decoder 118 can generate a first portion 170 of the decoded middle channel from the first frame 190. For example, decoder 118 can decode a first portion 191 of the middle channel to generate the first portion 170 of the decoded middle channel. (The last sentence appears to be incomplete and possibly refers to a technical detail about using a middle channel.) Figure 2 In the case of the described technology, decoder 118 can also generate a first portion 302 of the left channel and a first portion 304 of the right channel based on the first portion 170 of the decoded middle channel. The first portion 302 of the left channel may correspond to the first output signal 126, and the first portion 304 of the right channel may correspond to the second output signal 128. For example, decoder 118 may use a first quantized stereo parameter 183 and a first quantized shift value 181 to generate channels 302 and 304.
[0123] Decoder 118 may interpolate (or estimate) a second interpolated frequency domain shift value 285 (or a second interpolated time domain shift value 295) based on a first quantized shift value 181. According to other embodiments, the second interpolated shift values 285 and 295 may be estimated (e.g., interpolated or extrapolated) based on quantized shift values associated with two or more previous frames (e.g., the first frame 190 and at least one frame before the first frame or after the second frame 192, one or more other frames in bitstream 160, or any combination thereof). Decoder 118 may also interpolate (or estimate) a second interpolated stereo parameter 287 based on a first quantized stereo parameter 183. According to other embodiments, the second interpolated stereo parameter 287 may be estimated based on quantized stereo parameters associated with two or more other frames (e.g., the first frame 190 and at least one frame before or after the first frame).
[0124] Additionally, decoder 118 can interpolate (or estimate) the second portion 306 of the decoded intermediate channel based on the first portion 170 of the decoded intermediate channel (or the intermediate channel associated with two or more previous frames). When using... Figure 2In the case of the described technique, decoder 118 can also generate a second portion 308 of the left channel and a second portion 310 of the right channel based on the estimated second portion 306 of the decoded middle channel. The second portion 308 of the left channel may correspond to the first output signal 126, and the second portion 310 of the right channel may correspond to the second output signal 128. For example, decoder 118 may use a second interpolated stereo parameter 287 and a second interpolated frequency domain quantization shift value 285 to generate the left and right channels.
[0125] See Figure 4A Method 400 for decoding signals is demonstrated. Method 400 can be derived from... Figure 1 The second device 106 Figure 1 and 2 Decoder 118 or both are executed.
[0126] Method 400 includes: at 402, receiving at a decoder a bitstream comprising an intermediate channel and quantized values, the quantized values representing a shift between a first channel (e.g., a reference channel) associated with the encoder and a second channel (e.g., a target channel) associated with the encoder. The quantized values are shift-based values. These values are associated with the encoder and have greater accuracy than the quantized values.
[0127] Method 400 further includes: at 404, decoding the intermediate channel to generate a decoded intermediate channel. Method 400 further includes: at 406, generating a first channel (first generated channel) based on the decoded intermediate channel; and at 408, generating a second channel (second generated channel) based on the decoded intermediate channel and a quantized value. The first generated channel corresponds to a first channel associated with the encoder (e.g., a reference channel), and the second generated channel corresponds to a second channel associated with the encoder (e.g., a target channel). In some embodiments, both the first and second channels may be based on shifted quantized values. In some embodiments, the decoder may not explicitly identify the reference and target channels prior to the shift operation.
[0128] therefore, Figure 4A Method 400 enables alignment of the encoder side channels to reduce decoding entropy and thus increase decoding efficiency, because decoding entropy is sensitive to shift changes between channels. For example, encoder 114 can use unquantized shift values to accurately align the channels because unquantized shift values have relatively high resolution. Quantized shift values can be transmitted to decoder 118 to reduce data transmission resource usage. At decoder 118, quantized shift parameters can be used to simulate a perceptible difference between output signals 126 and 128.
[0129] See Figure 4B Method 450 for decoding signals is demonstrated. In some implementations, Figure 4B Method 450 is Figure 4A A more detailed version of method 400 for decoding audio signals. Method 450 can be derived from... Figure 1 The second device 106 Figure 1 and 2 Decoder 118 or both are executed.
[0130] Method 450 includes: at 452, receiving a bitstream from the encoder at the decoder. The bitstream includes a middle channel and quantized values, the quantized values representing the shift between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized values may be based on the shifted values (e.g., unquantized values), which have greater accuracy than the quantized values. For example, see... Figure 1 Decoder 118 may receive bitstream 160 from encoder 114. Bitstream 160 may include a first portion 191 of a middle channel and a first quantized shift value 181, the first quantized shift value 181 representing a shift between a first audio signal 130 (e.g., a reference channel) and a second audio signal 132 (e.g., a target channel). The first quantized shift value 181 may be based on a first shift value 180 (e.g., an unquantized value).
[0131] The first shift value 180 can have greater accuracy than the first quantized shift value 181. For example, the first quantized shift value 181 can correspond to a lower resolution version of the first shift value 180. The first shift value can be used by the encoder 114 to time-match the target channel (e.g., the second audio signal 132) with the reference channel (e.g., the first audio signal 130).
[0132] Method 450 also includes: at 454, decoding the intermediate channel to produce the decoded intermediate channel. For example, see... Figure 2 The middle channel decoder 202 can decode the first portion 191 of the middle channel to produce the first portion 170 of the decoded middle channel. Method 400 further includes: at 456, performing a transform operation on the decoded middle channel to produce a decoded frequency domain middle channel. For example, see... Figure 2 The transformation unit 204 can perform a transformation operation on the first part 170 of the decoded intermediate channel to generate the first part 171 of the frequency domain decoded intermediate channel.
[0133] Method 450 may further include: at 458, upmixing the decoded frequency domain intermediate channel to generate a first portion of the frequency domain channel and a second frequency domain channel. For example, see... Figure 2Upmixer 206 can upmix a first portion 171 of the frequency-domain decoded intermediate channel to generate a portion of frequency-domain channel 250 and a portion of frequency-domain channel 254. Method 450 may further include: at 460, generating a first channel based on the first portion of the frequency-domain channel. The first channel may correspond to a reference channel. For example, inverse transform unit 210 may perform an inverse transform operation on said portion of frequency-domain channel 250 to generate a portion of time-domain channel 260, and shifter 214 may pass said portion of time-domain channel 260 as a portion of a first output signal 126. The first output signal 126 may correspond to a reference channel (e.g., a first audio signal 130).
[0134] Method 450 may further include: at 462, generating a second channel based on the second frequency domain channel. The second channel may correspond to the target channel. According to one embodiment, if the quantized value corresponds to a frequency domain shift, then the second frequency domain channel may be shifted in the frequency domain to reach the quantized value. For example, see... Figure 2 The upmixer 206 can shift a portion of the frequency domain channel 254 to the first quantized frequency domain shift value 281 to a second shifted frequency domain channel (not shown). The inverse transform unit 212 can perform an inverse transform on the second shifted frequency domain channel to generate a portion of the second output signal 128. The second output signal 128 may correspond to a target channel (e.g., the second audio signal 132).
[0135] According to another embodiment, if the quantized value corresponds to a time-domain shift, then the time-domain version of the second frequency-domain channel can be shifted to the quantized value. For example, inverse transform unit 212 can perform an inverse transform operation on the portion of frequency-domain channel 254 to produce a portion of time-domain channel 264. Shifter 214 can shift a portion of time-domain channel 264 to the first quantized time-domain shift value 291 to produce a portion of the second output signal 128. The second output signal 128 may correspond to a target channel (e.g., the second audio signal 132).
[0136] therefore, Figure 4B Method 450 may facilitate alignment of the encoder side channels to reduce decoding entropy and thus increase decoding efficiency, because decoding entropy is sensitive to shift changes between channels. For example, encoder 114 can use unquantized shift values to accurately align the channels because unquantized shift values have relatively high resolution. Quantized shift values can be transmitted to decoder 118 to reduce data transmission resource usage. At decoder 118, quantized shift parameters can be used to simulate a perceptible difference between output signals 126 and 128.
[0137] See Figure 5A This demonstrates another method 500 for decoding signals. Method 500 can be derived from... Figure 1 The second device 106 Figure 1and 2 Decoder 118 or both are executed.
[0138] Method 500 includes: at 502, receiving at least a portion of a bit stream. The bit stream includes a first frame and a second frame. The first frame includes a first portion of the center channel and a first value of the stereo parameter, and the second frame includes a second portion of the center channel and a second value of the stereo parameter.
[0139] Method 500 further includes: at 504, decoding a first portion of the center channel to generate a decoded first portion of the center channel. Method 500 further includes: at 506, generating a first portion of the left channel based at least on the first portion of the decoded center channel and a first value of the stereo parameters; and at 508, generating a first portion of the right channel based at least on the first portion of the decoded center channel and the first value of the stereo parameters. The method also includes: at 510, in response to a second frame being unavailable for decoding, generating a second portion of the left channel and a second portion of the right channel based at least on the first value of the stereo parameters. The second portion of the left channel and the second portion of the right channel correspond to the decoded version of the second frame.
[0140] According to one embodiment, method 500 includes: generating interpolated values of stereo parameters based on a first value and a second value of stereo parameters in response to a second frame being available for decoding operation. According to another embodiment, method 500 includes: generating at least a second portion of the left channel and a second portion of the right channel based at least on the first value of the stereo parameters, a first portion of the left channel, and a first portion of the right channel in response to a second frame being unavailable for decoding operation.
[0141] According to one embodiment, method 500 includes: generating at least a second portion of the center channel and a second portion of the side channel based on at least a first value of stereo parameters, a first portion of the center channel, a first portion of the left channel, or a first portion of the right channel, in response to a second frame being unavailable for decoding operation. Method 500 further includes: generating a second portion of the left channel and a second portion of the right channel based on the second portion of the center channel, the second portion of the side channel, and a third value of the stereo parameters, in response to a second frame being unavailable for decoding operation. The third value of the stereo parameters is based at least on the first value of the stereo parameters, an interpolated value of the stereo parameters, and the decoding mode.
[0142] Therefore, method 500 enables decoder 118 to approximate stereo parameters (e.g., shift values) based on stereo parameters or variations of stereo parameters from previous frames. For example, decoder 118 can extrapolate stereo parameters for frames lost during transmission (e.g., second frame 192) from stereo parameters of one or more previous frames.
[0143] See Figure 5BThis demonstrates another method for decoding signals, 550. In some implementations, Figure 5B Method 550 is Figure 5A A more detailed version of method 500 for decoding audio signals. Method 550 can be derived from... Figure 1 The second device 106 Figure 1 and 2 Decoder 118 or both are executed.
[0144] Method 550 includes: at 552, receiving at least a portion of a bitstream from an encoder at a decoder. The bitstream includes a first frame and a second frame. The first frame includes a first portion of the center channel and a first value of the stereo parameters, and the second frame includes a second portion of the center channel and a second value of the stereo parameters. For example, see... Figure 1 The second device 106 can receive a portion of the bit stream 160 from the encoder 114. The bit stream includes a first frame 190 and a second frame 192. The first frame 190 includes a first portion 191 of the center channel, a first quantized shift value 181, and a first quantized stereo parameter 183. The second frame 192 includes a second portion 193 of the center channel, a second quantized shift value 185, and a second quantized stereo parameter 187.
[0145] Method 550 further includes: at 554, decoding the first portion of the intermediate channel to produce the first portion of the decoded intermediate channel. For example, see... Figure 2 The middle channel decoder 202 can decode the first portion 191 of the middle channel to produce the first portion 170 of the decoded middle channel. Method 550 may further include: at 556, performing a transform operation on the first portion of the decoded middle channel to produce the first portion of the decoded frequency domain middle channel. For example, see... Figure 2 The transformation unit 204 can perform a transformation operation on the first part 170 of the decoded intermediate channel to generate the first part 171 of the frequency domain decoded intermediate channel.
[0146] Method 550 may further include: at 558, upmixing the first portion of the decoded frequency domain middle channel to generate the first portion of the left frequency domain channel and the first portion of the right frequency domain channel. For example, see... Figure 1 The upmixer 206 can upmix the first portion 171 of the frequency-domain decoded intermediate channel to produce frequency-domain channels 250 and 254. As described herein, frequency-domain channel 250 can be the left channel and frequency-domain channel 254 can be the right channel. However, in other embodiments, frequency-domain channel 250 can be the right channel and frequency-domain channel 254 can be the left channel.
[0147] Method 550 may further include: at 560, generating a first portion of the left channel based at least on a first portion of the left frequency domain channel and a first value of the stereo parameter. For example, upmixer 206 may use a first quantized stereo parameter 183 to generate frequency domain channel 250. Inverse transform unit 210 may perform an inverse transform operation on frequency domain channel 250 to generate time domain channel 260, and shifter 214 may pass time domain channel 260 as a first output signal 126 (e.g., the first portion of the left channel according to method 550).
[0148] Method 550 may further include: at 562, generating a first portion of the right channel based at least on a first portion of the right frequency domain channel and a first value of the stereo parameter. For example, upmixer 206 may use a first quantized stereo parameter 183 to generate frequency domain channel 254. Inverse transform unit 212 may perform an inverse transform operation on frequency domain channel 254 to generate time domain channel 264, and shifter 214 may pass (or selectively shift) time domain channel 264 as a second output signal 128 (e.g., the first portion of the right channel according to method 550).
[0149] Method 550 further includes: at 564, determining that the second frame is unavailable for decoding. For example, decoder 118 may determine that one or more portions of the second frame 192 are unavailable for decoding. For illustrative purposes, the second quantized shift value 185 and the second quantized stereo parameter 187 may be lost during transmission (from the first device 104 to the second device 106) due to poor transmission conditions. Method 550 further includes: at 566, in response to determining that the second frame is unavailable, generating a second portion of the left channel and a second portion of the right channel based at least on a first value of the stereo parameter. The second portion of the left channel and the second portion of the right channel may correspond to a decoded version of the second frame.
[0150] For example, stereo parameter interpolator 208 can interpolate (or estimate) a second quantized shift value 185 based on a first quantized frequency domain shift value 281. For illustration, stereo parameter interpolator 208 can generate a second interpolated frequency domain shift value 285 based on the first quantized frequency domain shift value 281. Stereo parameter interpolator 208 can also interpolate (or estimate) a second quantized stereo parameter 187 based on a first quantized stereo parameter 183. For example, stereo parameter interpolator 208 can generate a second interpolated stereo parameter 287 based on the first quantized stereo parameter 183.
[0151] Upmixer 206 can upmix the second frequency-domain decoded middle channel 173 to generate frequency-domain channels 252 and 256. Upmixer 206 can apply a second interpolated stereo parameter 287 to the second frequency-domain decoded middle channel 173 during upmixing to generate frequency-domain channels 252 and 256. According to an embodiment where the first quantized shift value 181 includes a frequency-domain shift (e.g., the first quantized shift value 181 corresponds to a first quantized frequency-domain shift value 281), upmixer 206 can perform a frequency-domain shift (e.g., phase shift) based on the second interpolated frequency-domain shift value 285 to generate frequency-domain channel 256.
[0152] Inverse transform unit 210 can perform an inverse transform operation on frequency domain channel 252 to generate time domain channel 262, and inverse transform unit 212 can perform an inverse transform operation on frequency domain channel 256 to generate time domain channel 266. Shift value interpolator 216 can interpolate (or estimate) a second quantized shift value 185 based on a first quantized time domain shift value 291. For example, shift value interpolator 216 can generate a second interpolated time domain shift value 295 based on the first quantized time domain shift value 291. According to the embodiment where the first quantized shift value 181 corresponds to the first quantized frequency domain shift value 281, shifter 214 can bypass the shift operation and pass time domain channels 262 and 266 as output signals 126 and 128, respectively. According to the embodiment where the first quantized shift value 181 corresponds to the first quantized time-domain shift value 291, the shifter 214 can shift the time-domain channel 266 by the second interpolated time-domain shift value 295 to generate the second output signal 128.
[0153] Therefore, method 550 enables decoder 118 to interpolate (or estimate) stereo parameters for frames lost during transmission (e.g., second frame 192) based on stereo parameters for one or more previous frames.
[0154] See Figure 6 This is a block diagram depicting a specific illustrative example of a device (e.g., a wireless communication device), and the device is generally designated as 600. In various embodiments, device 600 may have a greater... Figure 6 The illustrated components are fewer or more components. In the illustrative embodiment, device 600 may correspond to... Figure 1 First device 104 Figure 1 The second device 106 or a combination thereof. In an illustrative embodiment, device 600 may perform the reference... Figures 1 to 3 The system and methods described in 4A, 4B, 5A and 5B are one or more operations.
[0155] In a particular embodiment, device 600 includes a processor 606 (e.g., a central processing unit (CPU)). Device 600 may include one or more additional processors 610 (e.g., one or more digital signal processors (DSPs)). Processor 610 may include a media (e.g., speech and music) coder-decoder (CODEC) 608 and an echo canceller 612. Media CODEC 608 may include decoder 118, encoder 114, or a combination thereof.
[0156] Device 600 may include memory 153 and CODEC 634. Although media CODEC 608 is illustrated as a component of processor 610 (e.g., dedicated circuitry and / or executable programmable code), in other embodiments, one or more components of media CODEC 608, such as decoder 118, encoder 114, or a combination thereof, may be included in processor 606, CODEC 634, another processing component, or a combination thereof.
[0157] Device 600 may include a transmitter 110 coupled to antenna 642. Device 600 may include a display 628 coupled to display controller 626. One or more speakers 648 may be coupled to CODEC 634. One or more microphones 646 may be coupled to CODEC 634 via input interface 112. In a particular embodiment, speaker 648 may include Figure 1 The first loudspeaker 142 Figure 1 The second speaker 144 or a combination thereof. In a particular embodiment, the microphone 646 may include... Figure 1 The first microphone 146 Figure 1 The second microphone 148 or a combination thereof. The CODEC 634 may include a digital-to-analog converter (DAC) 602 and an analog-to-digital converter (ADC) 604.
[0158] Memory 153 may contain a processing unit, or a combination thereof, that can be executed by processor 606, processor 610, CODEC 634, device 600, or to perform reference operations. Figures 1 to 3 Instruction 660 describes one or more operations as described in 4A, 4B, 5A, and 5B. Instruction 660 is executable to cause a processor (e.g., processor 606, CODEC 634, decoder 118, another processing unit of device 600, or a combination thereof) to perform... Figure 4A Method 400 Figure 4B Method 450 Figure 5A Method 500 Figure 5B Method 550 or a combination thereof.
[0159] One or more components of device 600 may be implemented via dedicated hardware (e.g., a circuit system), by a processor that executes instructions to perform one or more tasks, or a combination thereof. As an example, one or more components of memory 153 or processor 606, processor 610, and / or CODEC 634 may be memory devices, such as random access memory (RAM), magnetoresistive random access memory (MRAM), spin-torque transfer MRAM (STT-MRAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, or compact disc read-only memory (CD-ROM). The memory device may include instructions (e.g., instruction 660) that, when executed by a computer (e.g., the processor, processor 606, and / or processor 610 in CODEC 634), cause the computer to perform reference... Figures 1 to 3 The operations described in 4A, 4B, 5A, and 5B. As an example, memory 153 or processor 606, processor 610, and / or one or more components of CODEC 634 may be non-transitory computer-readable media containing instructions (e.g., instruction 660) that, when executed by a computer (e.g., the processor, processor 606, and / or processor 610 in CODEC 634), cause the computer to perform reference... Figures 1 to 3 One or more operations described by 4A, 4B, 5A, and 5B.
[0160] In a particular embodiment, device 600 may be included in a system-in-package or system-on-a-chip device (e.g., a mobile station modem (MSM)) 622. In a particular embodiment, processor 606, processor 610, display controller 626, memory 153, CODEC 634, and transmitter 110 are included in the system-in-package or system-on-a-chip device 622. In a particular embodiment, input devices 630, such as a touchscreen and / or keypad, and power supply 644 are coupled to the system-on-a-chip device 622. Furthermore, in a particular embodiment, such as... Figure 6 As illustrated, the display 628, input device 630, speaker 648, microphone 646, antenna 642, and power supply 644 are external to the system-on-a-chip device 622. However, each of the display 628, input device 630, speaker 648, microphone 646, antenna 642, and power supply 644 may be coupled to components of the system-on-a-chip device 622, such as an interface or controller.
[0161] Device 600 may include a wireless telephone, mobile communication device, mobile phone, smartphone, cellular phone, laptop computer, desktop computer, computer, tablet computer, set-top box, personal digital assistant (PDA), display device, television, game console, music player, radio, video player, entertainment unit, communication device, fixed location data unit, personal media player, digital video player, digital video disc (DVD) player, tuner, camera, navigation device, decoder system, encoder system, or any combination thereof.
[0162] In certain embodiments, one or more components of the systems and apparatuses disclosed herein may be integrated into a decoding system or device (e.g., an electronic device, a CODEC, or a processor therein), integrated into an encoding system or device, or both. In other embodiments, one or more components of the systems and apparatuses disclosed herein may be integrated into a wireless telephone, tablet computer, desktop computer, laptop computer, set-top box, music player, video player, entertainment unit, television, game console, navigation device, communication device, personal digital assistant (PDA), fixed location data unit, personal media player, or another type of device.
[0163] In conjunction with the techniques described herein, the first device includes means for receiving a bit stream. The bit stream includes a middle channel and a quantized value, the quantized value representing the shift between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized value is a shift-based value. This value is associated with the encoder and has greater accuracy than the quantized value. For example, the means for receiving the bit stream may include: Figure 1 The second device 106; the receiver of the second device 106 (not shown); Figure 1 , 2 Or the decoder 118 of 6; Figure 6 Antenna 642; one or more other circuits, devices, components, modules; or combinations thereof.
[0164] The first device may also include means for decoding the middle channel to produce a decoded middle channel. For example, the means for decoding the middle channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 The middle channel decoder 202; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0165] The first device may further include means for generating a first channel based on the decoded intermediate channel. The first channel corresponds to a reference channel. For example, the means for generating the first channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Inverse transformation unit 210; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0166] The first device may further include means for generating a second channel based on the decoded intermediate channel and the quantized value. The second channel corresponds to the target channel. The means for generating the second channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Inverse transformation unit 212; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0167] In conjunction with the techniques described herein, the second device includes means for receiving a bit stream from an encoder. The bit stream may include a middle channel and quantized values, the quantized values representing a shift between a reference channel associated with the encoder and a target channel associated with the encoder. The quantized values may be based on shift values that have greater accuracy than the quantized values. For example, the means for receiving the bit stream may include: Figure 1 The second device 106; the receiver of the second device 106 (not shown); Figure 1 , 2 Or the decoder 118 of 6; Figure 6 Antenna 642; one or more other circuits, devices, components, modules; or combinations thereof.
[0168] The second device may also include means for decoding the middle channel to produce a decoded middle channel. For example, the means for decoding the middle channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 The middle channel decoder 202; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0169] The second device may also include means for performing a transformation operation on the decoded intermediate channel to generate a decoded frequency domain intermediate channel. For example, the means for performing the transformation operation may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Transformation unit 204; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0170] The second device may further include means for upmixing the decoded frequency domain middle channel to generate a first frequency domain channel and a second frequency domain channel. For example, the means for upmixing may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 The supermixer 206; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0171] The second device may further include means for generating a first channel based on a first frequency domain channel. The first channel may correspond to a reference channel. For example, the means for generating the first channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Inverse transformation unit 210; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0172] The second device may further include means for generating a second channel based on the second frequency domain channel. The second channel may correspond to a target channel. If the quantized value corresponds to a frequency domain shift, then the second frequency domain channel may be shifted by the quantized value in the frequency domain. If the quantized value corresponds to a time domain shift, then the time domain version of the second frequency domain channel may be shifted by the quantized value. The means for generating the second channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Inverse transformation unit 212; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0173] In conjunction with the techniques described herein, the third device includes means for receiving at least a portion of a bitstream. The bitstream includes a first frame and a second frame. The first frame includes a first portion of the center channel and a first value of the stereo parameters, and the second frame includes a second portion of the center channel and a second value of the stereo parameters. The means for receiving may include: Figure 1 The second device 106; the receiver of the second device 106 (not shown); Figure 1 , 2 Or the decoder 118 of 6; Figure 6Antenna 642; one or more other circuits, devices, components, modules; or combinations thereof.
[0174] The third device may also include means for decoding the first portion of the middle channel to produce the decoded first portion of the middle channel. For example, the means for decoding may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 The middle channel decoder 202; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0175] The third device may further include means for generating a first portion of the left channel based at least on a first portion of the decoded middle channel and a first value of stereo parameters. For example, the means for generating the first portion of the left channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Inverse transformation unit 210; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0176] The third device may further include means for generating a first portion of the right channel based at least on a first portion of the decoded middle channel and a first value of stereo parameters. For example, the means for generating the first portion of the right channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Inverse transformation unit 212; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0177] The third device may further include means for generating a second portion of the left channel and a second portion of the right channel, at least based on a first value of stereo parameters, in response to the second frame being unavailable for decoding. The second portion of the left channel and the second portion of the right channel correspond to the decoded version of the second frame. The means for generating the second portion of the left channel and the second portion of the right channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Stereo shift value interpolator 216; Figure 2 Stereo parameter interpolator 208; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0178] In conjunction with the techniques described herein, the fourth device includes means for receiving at least a portion of a bit stream from an encoder. The bit stream may include a first frame and a second frame. The first frame may include a first portion of the center channel and a first value of a stereo parameter, and the second frame may include a second portion of the center channel and a second value of a stereo parameter. The means for receiving may include: Figure 1 The second device 106; the receiver of the second device 106 (not shown); Figure 1 , 2 Or the decoder 118 of 6; Figure 6 Antenna 642; one or more other circuits, devices, components, modules; or combinations thereof.
[0179] The fourth device may also include means for decoding the first portion of the middle channel to produce the decoded first portion of the middle channel. For example, the means for decoding the first portion of the middle channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 The middle channel decoder 202; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0180] The fourth device may further include means for performing a transformation operation on the first portion of the decoded intermediate channel to generate the first portion of the decoded frequency domain intermediate channel. For example, the means for performing the transformation operation may include: Figure 1 , 2Or the decoder 118 of 6; Figure 2 Transformation unit 204; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0181] The fourth device may further include means for upmixing the first portion of the decoded frequency domain center channel to generate a first portion of the left frequency domain channel and a first portion of the right frequency domain channel. For example, the means for upmixing may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 The supermixer 206; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0182] The fourth device may further include means for generating a first portion of the left channel based at least on a first portion of the left frequency domain channel and a first value of stereo parameters. For example, the means for generating the first portion of the left channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Inverse transformation unit 210; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0183] The fourth device may further include means for generating the first portion of the right channel based at least on the first portion of the right frequency domain channel and the first value of the stereo parameters. For example, the means for generating the first portion of the right channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Inverse transformation unit 212; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0184] The fourth device may further include means for generating a second portion of the left channel and a second portion of the right channel, at least based on a first value of stereo parameters, in response to a determination that the second frame is unavailable. The second portion of the left channel and the second portion of the right channel may correspond to a decoded version of the second frame. The means for generating the second portion of the left channel and the second portion of the right channel may include: Figure 1 , 2 Or the decoder 118 of 6; Figure 2 Stereo shift value interpolator 216; Figure 2 Stereo parameter interpolator 208; Figure 2 Shifter 214; Figure 6 The processor 606; Figure 6 The processor is 610; Figure 6 CODEC 634; Figure 6 Instruction 660, which can be executed by a processor; one or more other circuits, devices, components, modules; or combinations thereof.
[0185] It should be noted that the various functions performed by one or more components of the systems and apparatuses disclosed herein are described as being performed by certain components or modules. This division of components and modules is for illustrative purposes only. In alternative embodiments, functions performed by a particular component or module may be divided among multiple components or modules. Furthermore, in alternative embodiments, two or more components or modules may be integrated into a single component or module. Each component or module may be implemented using hardware (e.g., field-programmable gate array (FPGA) devices, application-specific integrated circuits (ASICs), DSPs, controllers, etc.), software (e.g., instructions executable by a processor), or any combination thereof.
[0186] See Figure 7 This is a block diagram depicting a specific illustrative example of a base station 700. In various embodiments, the base station 700 may have more than Figure 7 The illustrated components may include more or fewer components. In an illustrative example, base station 700 may contain... Figure 1 The second device 106. In an illustrative example, the base station 700 may be based on a reference. Figures 1 to 3 Operate by one or more of the methods or systems described in 4A, 4B, 5A, 5B and 6.
[0187] Base station 700 may be part of a wireless communication system. The wireless communication system may include multiple base stations and multiple wireless devices. The wireless communication system may be a Long Term Evolution (LTE) system, a Code Division Multiple Access (CDMA) system, a Global System for Mobile Communications (GSM) system, a Wireless Local Area Network (WLAN) system, or another wireless system. The CDMA system may implement Wideband CDMA (WCDMA), CDMA 1X, Evolution-Data Optimized (EVDO), Time Division Synchronous CDMA (TD-SCDMA), or another version of CDMA.
[0188] Wireless devices can also be referred to as user equipment (UE), mobile stations, terminals, access terminals, subscriber units, stations, etc. Wireless devices can include cellular phones, smartphones, tablet computers, wireless modems, personal digital assistants (PDAs), handheld devices, laptop computers, smart notebook computers, mini-notebook computers, tablet computers, cordless phones, wireless local loop (WLL) stations, Bluetooth devices, etc. Wireless devices can include or correspond to... Figure 6 Device 600.
[0189] One or more components of base station 700 may perform (and / or in other components not shown) various functions, such as sending and receiving messages and data (e.g., audio data). In a particular instance, base station 700 includes processor 706 (e.g., CPU). Base station 700 may include transcoder 710. Transcoder 710 may include audio CODEC 708. For example, transcoder 710 may include one or more components (e.g., circuitry) configured to perform operations of audio CODEC 708. As another example, transcoder 710 may be configured to execute one or more computer-readable instructions to perform operations of audio CODEC 708. Although audio CODEC 708 is illustrated as a component of transcoder 710, in other instances, one or more components of audio CODEC 708 may be included in processor 706, another processing component, or a combination thereof. For example, decoder 738 (e.g., vocoder decoder) may be included in receiver data processor 764. As another example, encoder 736 (e.g., vocoder encoder) may be included in data transmission processor 782. Encoder 736 may include Figure 1 The encoder 114. The decoder 738 may contain... Figure 1 Decoder 118.
[0190] Transcoder 710 can be used to transcode messages and data between two or more networks. Transcoder 710 can be configured to convert message and audio data from a first format (e.g., digital format) to a second format. For illustration, decoder 738 can decode an encoded signal having the first format, and encoder 736 can encode the decoded signal into an encoded signal having the second format. Alternatively, transcoder 710 can be configured to perform data rate adaptation. For example, transcoder 710 can down-convert or up-convert the data rate without changing the format of the audio data. For illustration, transcoder 710 can down-convert a 64 kbit / s signal to a 16 kbit / s signal.
[0191] Base station 700 may include memory 732. For example, memory 732 of a computer-readable storage device may contain instructions. The instructions may include instructions executable by processor 706, transcoder 710, or a combination thereof to perform a reference. Figures 1 to 3 Methods and systems described in 4A, 4B, 5A, 5B, and 6, and instructions for one or more operations.
[0192] Base station 700 may include multiple transmitters and receivers (e.g., transceivers) coupled to an antenna array, such as a first transceiver 752 and a second transceiver 754. The antenna array may include a first antenna 742 and a second antenna 744. The antenna array may be configured to wirelessly connect to one or more wireless devices—e.g., Figure 6The device 600 is for communication. For example, the second antenna 744 can receive a data stream 714 (e.g., a bit stream) from the wireless device. The data stream 714 may contain messages, data (e.g., encoded speech data), or a combination thereof.
[0193] Base station 700 may include network connection 760, such as a backhaul connection. Network connection 760 may be configured to communicate with one or more base stations of a core network or wireless communication network. For example, base station 700 may receive a second data stream (e.g., message or audio data) from the core network via network connection 760. Base station 700 may process the second data stream to generate message or audio data and provide the message or audio data to one or more wireless devices via one or more antennas of an antenna array, or provide the message or audio data to another base station via network connection 760. In certain embodiments, as an illustrative and non-limiting example, network connection 760 may be a wide area network (WAN) connection. In some embodiments, the core network may include or correspond to a public switched telephone network (PSTN), a packet backbone, or both.
[0194] Base station 700 may include media gateway 770 coupled to network connection 760 and processor 706. Media gateway 770 may be configured to convert media streams between different telecommunications technologies. For example, media gateway 770 may convert between different transport protocols, different decoding schemes, or both. For illustrative purposes, as an illustrative and non-limiting example, media gateway 770 may convert from PCM signals to Real-Time Transport Protocol (RTP) signals. Media gateway 770 may convert data between: packet-switched networks (e.g., Voice Over Internet Protocol (VoIP) networks, IP Multimedia Subsystem (IMS), fourth-generation (4G) wireless networks such as LTE, WiMax, and UMB, etc.); circuit-switched networks (e.g., PSTN); and hybrid networks (e.g., second-generation (2G) wireless networks such as GSM, GPRS, and EDGE; third-generation (3G) wireless networks such as WCDMA, EV-DO, and HSPA, etc.).
[0195] Additionally, media gateway 770 may include a transcoder, such as transcoder 710, and may be configured to transcode data when decoder-decoder incompatibility occurs. For example, as an illustrative and non-limiting example, media gateway 770 may perform transcoding between an Adaptive Multi-Rate (AMR) decoder-decoder and a G.711 decoder-decoder. Media gateway 770 may include a router and multiple physical interfaces. In some embodiments, media gateway 770 may also include a controller (not shown). In certain embodiments, the media gateway controller may be external to media gateway 770, external to base station 700, or both. The media gateway controller can control and coordinate the operation of multiple media gateways. Media gateway 770 may receive control signals from the media gateway controller and can be used to bridge different transmission technologies and add services to end-user capabilities and connectivity.
[0196] Base station 700 may include demodulator 762 coupled to transceivers 752, 754, receiver data processor 764, and processor 706, and receiver data processor 764 may be coupled to processor 706. Demodulator 762 may be configured to demodulate modulated signals received from transceivers 752, 754, and provide demodulated data to receiver data processor 764. Receiver data processor 764 may be configured to extract message or audio data from the demodulated data and transmit the message or audio data to processor 706.
[0197] Base station 700 may include a transmission data processor 782 and a transmission multiple-input multiple-output (MIMO) processor 784. Transmission data processor 782 may be coupled to processor 706 and transmission MIMO processor 784. Transmission MIMO processor 784 may be coupled to transceivers 752, 754 and processor 706. In some embodiments, transmission MIMO processor 784 may be coupled to media gateway 770. As an illustrative and non-limiting example, transmission data processor 782 may be configured to receive message or audio data from processor 706 and decode the message or audio data based on a decoding scheme such as CDMA or orthogonal frequency-division multiplexing (OFDM). Transmission data processor 782 may provide the decoded data to transmission MIMO processor 784.
[0198] CDMA or OFDM technologies can be used to multiplex decoded data with other data, such as pilot data, to produce multiplexed data. The multiplexed data can then be modulated (i.e., symbol mapped) by the data transmission processor 782 based on a specific modulation scheme (e.g., binary phase-shift keying (BPSK), quadrature phase-shift keying (QSPK), M-ary phase-shift keying (M-PSK), M-ary quadrature amplitude modulation (M-QAM), etc.) to produce modulated symbols. In certain embodiments, different modulation schemes can be used to modulate the decoded data and other data. The data rate, decoding, and modulation for each data stream can be determined by instructions executed by the processor 706.
[0199] The transport MIMO processor 784 can be configured to receive modulation symbols from the transport data processor 782, and can further process the modulation symbols and perform beamforming on the data. For example, the transport MIMO processor 784 can apply beamforming weights to the modulation symbols.
[0200] During operation, the second antenna 744 of base station 700 can receive data stream 714. The second transceiver 754 can receive data stream 714 from the second antenna 744 and provide data stream 714 to demodulator 762. Demodulator 762 can demodulate the modulated signal of data stream 714 and provide the demodulated data to receiver data processor 764. Receiver data processor 764 can extract audio data from the demodulated data and provide the extracted audio data to processor 706.
[0201] Processor 706 may provide audio data to transcoder 710 for transcoding. Transcoder 710's decoder 738 may decode the audio data from a first format into decoded audio data, and encoder 736 may encode the decoded audio data into a second format. In some embodiments, encoder 736 may use a higher data rate (e.g., up-conversion) or a lower data rate (e.g., down-conversion) to encode the audio data relative to the data rate received from the wireless device. In other embodiments, audio data may not be transcoded. Although transcoding (e.g., decoding and encoding) is illustrated as being performed by transcoder 710, transcoding operations (e.g., decoding and encoding) may be performed by multiple components of base station 700. For example, decoding may be performed by receiver data processor 764, and encoding may be performed by transport data processor 782. In other embodiments, processor 706 may provide audio data to media gateway 770 for conversion to another transport protocol, decoding scheme, or both. Media gateway 770 may provide the converted data to another base station or core network via network connection 760.
[0202] Encoded audio data generated at encoder 736 can be provided to transmission data processor 782 or network connection 760 via processor 706. Transcoded audio data from transcoder 710 can be provided to transmission data processor 782 for decoding according to, for example, an OFDM modulation scheme to generate modulation symbols. Transmission data processor 782 can provide the modulation symbols to transmission MIMO processor 784 for further processing and beamforming. Transmission MIMO processor 784 can apply beamforming weights and can provide the modulation symbols to one or more antennas of an antenna array, such as first antenna 742, via first transceiver 752. Therefore, base station 700 can provide transcoded data stream 716, corresponding to data stream 714 received from a wireless device, to another wireless device. Transcoded data stream 716 may have a different encoding format, data rate, or both than data stream 714. In other embodiments, transcoded data stream 716 can be provided to network connection 760 for transmission to another base station or core network.
[0203] Those skilled in the art will further understand that the various illustrative logic blocks, configurations, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processing device such as a hardware processor, or a combination of both. The various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in varying ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention.
[0204] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be embodied directly in hardware, in software modules executed by a processor, or in a combination of both. The software modules can reside in memory devices such as: random access memory (RAM), magnetoresistive random access memory (MRAM), spin torque transfer MRAM (STT-MRAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, or compact optical disc read-only memory (CD-ROM). An exemplary memory device is coupled to a processor, allowing the processor to read information from and write information to the memory device. Alternatively, the memory device can be integrated with the processor. The processor and storage media can reside in an application-specific integrated circuit (ASIC). The ASIC can reside in a computing device or user terminal. Alternatively, the processor and storage media can reside as discrete components in a computing device or user terminal.
[0205] A prior description of the disclosed embodiments is provided to enable those skilled in the art to make or use the disclosed embodiments. Various modifications to these embodiments will be readily apparent to those skilled in the art without departing from the scope of the invention, and the principles defined herein can be applied to other embodiments. Therefore, the invention is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope that may be consistent with the principles and novel features as defined by the appended claims.
Claims
1. A device for decoding audio signals, comprising: A receiver configured to receive at least a portion of a bitstream, the bitstream including a first frame and a second frame, the first frame including a first portion of a center channel and a first quantized stereo parameter, the second frame including a second portion of the center channel and a second quantized stereo parameter, wherein the first quantized stereo parameter has a lower resolution than the first stereo parameter, and the second quantized stereo parameter has a lower resolution than the second stereo parameter. as well as The decoder is configured as follows: Decode the first portion of the middle channel to produce the first portion of the decoded middle channel; The first portion of the left channel is generated based at least on the first portion of the decoded middle channel and the first quantized stereo parameters; The first part of the right channel is generated based at least on the first part of the decoded middle channel and the first quantized stereo parameters; as well as In response to the second frame being unavailable for decoding: The second quantized stereo parameter is estimated based on stereo parameters from one or more previous frames; The second portion of the middle channel and the second portion of the side channel are generated based at least on the stereo parameters of the one or more previous frames; as well as The second part of the left channel and the second part of the right channel are generated based at least on the third stereo parameters, the second part of the middle channel and the second part of the side channel, wherein the third stereo parameters are based at least on the first quantized stereo parameters, the estimated second quantized stereo parameters and the decoding mode, and the second part of the left channel and the second part of the right channel correspond to the decoded version of the second frame.
2. The device according to claim 1, wherein, The stereo parameters of the one or more previous frames include the first quantized stereo parameters.
3. The device of claim 2, wherein the decoder is configured to estimate the second quantized stereo parameters by interpolating the first quantized stereo parameters.
4. The device of claim 2, wherein the decoder is configured to estimate the second quantized stereo parameters by extrapolating the first quantized stereo parameters.
5. The device of claim 1, wherein the decoder is further configured to: A transformation operation is performed on the first portion of the decoded intermediate channel to generate the first portion of the decoded frequency domain intermediate channel; Based on the first quantized stereo parameters, the first portion of the decoded frequency domain middle channel is upmixed to generate the first portion of the left frequency domain channel and the first portion of the right frequency domain channel. Perform a first time-domain operation on the first portion of the left frequency domain channel to generate the first portion of the left channel; as well as A second time-domain operation is performed on the first portion of the right frequency domain channel to generate the first portion of the right channel.
6. The device according to claim 5, wherein, In response to the second frame being unavailable for the decoding operation, the decoder is configured to: Perform a second transformation operation on the second portion of the middle channel to generate the second portion of the decoded frequency domain middle channel; The second portion of the decoded frequency domain middle channel is upmixed to generate the second portion of the left frequency domain channel and the second portion of the right frequency domain channel; A third time-domain operation is performed on the second portion of the left frequency domain channel to generate the second portion of the left channel; as well as A fourth time-domain operation is performed on the second portion of the right frequency domain channel to produce the second portion of the right channel.
7. The device according to claim 6, wherein, The estimated second quantized stereo parameters are used to upmix the second portion of the decoded frequency domain middle channel.
8. The device of claim 6, wherein the decoder is configured to perform an interpolation operation on the first portion of the decoded intermediate channel to produce the second portion of the decoded intermediate channel.
9. The device of claim 1, wherein the first quantized stereo parameter is a quantized value representing a shift between a reference channel associated with the encoder and a target channel associated with the encoder, the quantized value being based on the shift value, the shift value being associated with the encoder and having greater accuracy than the quantized value.
10. The device according to claim 1, wherein, The first stereo parameter and the second stereo parameter include the inter-channel phase difference parameter.
11. The device according to claim 1, wherein, The first stereo parameter and the second stereo parameter include the inter-channel level difference parameter.
12. The device according to claim 1, wherein, The first stereo parameter and the second stereo parameter include the inter-channel time difference parameter.
13. The device according to claim 1, wherein, The first stereo parameter and the second stereo parameter include inter-channel correlation parameters.
14. The device according to claim 1, wherein, The first stereo parameter and the second stereo parameter include the spectral tilt parameter.
15. The device according to claim 1, wherein, The first stereo parameter and the second stereo parameter include inter-channel gain parameters.
16. The device according to claim 1, wherein, The first stereo parameter and the second stereo parameter include inter-channel sound generation parameters.
17. The device of claim 1, wherein the first quantized stereo parameter and the second quantized stereo parameter include inter-channel pitch parameters.
18. The device according to claim 1, wherein, The receiver and the decoder are integrated into the mobile device.
19. The device according to claim 1, wherein, The receiver and the decoder are integrated into the base station.
20. A method for decoding an audio signal, comprising: At least a portion of a bitstream is received at the decoder, the bitstream including a first frame and a second frame, the first frame including a first portion of the center channel and a first quantized stereo parameter, the second frame including a second portion of the center channel and a second quantized stereo parameter, wherein the first quantized stereo parameter has a lower resolution than the first stereo parameter, and the second quantized stereo parameter has a lower resolution than the second stereo parameter. Decode the first portion of the middle channel to produce the first portion of the decoded middle channel; The first portion of the left channel is generated based at least on the first portion of the decoded middle channel and the first quantized stereo parameters; The first part of the right channel is generated based at least on the first part of the decoded middle channel and the first quantized stereo parameters; as well as In response to the second frame being unavailable for decoding: The second quantized stereo parameter is estimated based on stereo parameters from one or more previous frames; The second portion of the middle channel and the second portion of the side channel are generated based at least on the stereo parameters of the one or more previous frames; as well as The second part of the left channel and the second part of the right channel are generated based at least on the third stereo parameters, the second part of the middle channel and the second part of the side channel, wherein the third stereo parameters are based at least on the first quantized stereo parameters, the estimated second quantized stereo parameters and the decoding mode, and the second part of the left channel and the second part of the right channel correspond to the decoded version of the second frame.
21. The method according to claim 20, wherein, The stereo parameters of the one or more previous frames include the first quantized stereo parameters.
22. The method of claim 21, wherein estimating the second quantized stereo parameter includes interpolating the first quantized stereo parameter.
23. The method of claim 21, wherein estimating the second quantized stereo parameter includes extrapolating the first quantized stereo parameter.
24. The method of claim 20, further comprising: A transformation operation is performed on the first portion of the decoded intermediate channel to generate the first portion of the decoded frequency domain intermediate channel; Based on the first quantized stereo parameters, the first portion of the decoded frequency domain middle channel is upmixed to generate the first portion of the left frequency domain channel and the first portion of the right frequency domain channel. Perform a first time-domain operation on the first portion of the left frequency domain channel to generate the first portion of the left channel; as well as A second time-domain operation is performed on the first portion of the right frequency domain channel to generate the first portion of the right channel.
25. The method of claim 24, further comprising, in response to the second frame being unavailable for the decoding operation: Perform a second transformation operation on the second portion of the middle channel to generate the second portion of the decoded frequency domain middle channel; The second portion of the decoded frequency domain middle channel is upmixed to generate the second portion of the left frequency domain channel and the second portion of the right frequency domain channel; A third time-domain operation is performed on the second portion of the left frequency domain channel to generate the second portion of the left channel; as well as A fourth time-domain operation is performed on the second portion of the right frequency domain channel to produce the second portion of the right channel.
26. The method of claim 22, further comprising: An interpolation operation is performed on the first portion of the decoded intermediate channel to produce the second portion of the decoded intermediate channel.
27. The method of claim 20, wherein the first quantized stereo parameter is a quantized value representing a shift between a reference channel associated with the encoder and a target channel associated with the encoder, the quantized value being based on the shift value, the shift value being associated with the encoder and having greater accuracy than the quantized value.
28. The method of claim 20, wherein the decoder is integrated into a mobile device.
29. The method of claim 20, wherein the decoder is integrated into the base station.
30. A device for decoding audio signals, comprising: A component for receiving at least a portion of a bitstream, the bitstream including a first frame and a second frame, the first frame including a first portion of a center channel and a first quantized stereo parameter, the second frame including a second portion of the center channel and a second quantized stereo parameter, wherein the first quantized stereo parameter has a lower resolution than the first stereo parameter, and the second quantized stereo parameter has a lower resolution than the second stereo parameter. Components for decoding the first portion of the middle channel to produce the first portion of the decoded middle channel; Components for generating a first portion of the left channel based at least on the first portion of the decoded middle channel and the first quantized stereo parameters; Components for generating a first portion of the right channel based at least on the first portion of the decoded middle channel and the first quantized stereo parameters; as well as In response to the second frame being unavailable for decoding: A component for estimating the second quantized stereo parameter based on stereo parameters from one or more previous frames; Components for generating a second portion of the middle channel and a second portion of the side channel based at least on the stereo parameters of the one or more previous frames; as well as A component for generating the second portion of the left channel and the second portion of the right channel based at least on a third stereo parameter, the second portion of the middle channel and the second portion of the side channel, wherein the third stereo parameter is based at least on a first quantized stereo parameter, an estimated second quantized stereo parameter and a decoding mode, and the second portion of the left channel and the second portion of the right channel correspond to the decoded version of the second frame.
31. A computer-readable medium having program code recorded thereon, wherein, The program code can be executed by one or more processors to perform the method according to any one of claims 20-29.
Citation Information
Patent Citations
Stereo parameters for stereo decoding
CN110622242A
Stereo parameters for stereo decoding
CN116665682A