Modification of inter-channel phase difference parameters
By modifying the phase difference parameter value between channels, the signal delay misalignment problem caused by microphone distance in stereo coding was solved, thereby improving coding efficiency and decoding gain.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2017-12-11
- Publication Date
- 2026-05-29
AI Technical Summary
In stereo coding, the different distances between the microphone and the sound source cause misalignment in the delay between the first and second audio signals, increasing the phase difference between the frequency domain versions of the audio signals and affecting coding efficiency.
The receiver decodes the intermediate channel and performs transformation operations to modify the phase difference parameter value between the channels. The upmixer is used to perform upmixing operations to generate the frequency domain left channel and the frequency domain right channel. The inverse transformation unit then converts them into time domain channels to achieve channel alignment.
It improves the decoding gain of stereo coding, reduces time shift and phase mismatch between channels, and enhances coding efficiency.
Smart Images

Figure CN116033328B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 201780080408.6, filed on December 11, 2017, entitled "Modification of Inter-channel Phase Difference Parameter".
[0002] Cross-references to related applications
[0003] This application claims the benefit of jointly owned U.S. Provisional Patent Application No. 62 / 448,297, filed January 19, 2017, entitled “Multiple Signal Coding and Inter-Channel Parameter Modification,” and U.S. Non-Provisional Patent Application No. 15 / 836,618, filed December 8, 2017, entitled “Inter-Channel Phase Difference Parameter Modification,” the contents of each of which are expressly incorporated herein by reference in their entirety. Technical Field
[0004] This invention generally relates to the encoding of multiple audio signals. Background Technology
[0005] Technological advancements have led to smaller and more powerful computing devices. For example, various portable personal computing devices exist today, including cordless phones (such as mobile phones and smartphones), tablets, and laptops. These portable personal computing devices are small, lightweight, and easily carried by the user. These devices can transmit voice and data packets via wireless networks. Furthermore, many of these devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Moreover, such devices can process executable instructions, including software applications, such as web browser applications used to access the Internet. Therefore, these devices can contain significant computing power.
[0006] A computing device may include or be coupled to multiple microphones to receive audio signals. Generally, the sound source is closer to the first microphone than to a second microphone with multiple microphones. Therefore, due to the corresponding distance between the microphone and the sound source, the second audio signal received from the second microphone may be delayed relative to the first audio signal received from the first microphone. In other embodiments, the first audio signal may be delayed relative to the second audio signal. In stereo coding, the audio signal from the microphone may be encoded to generate an intermediate channel signal and one or more side channel signals. The intermediate channel signal may correspond to the sum of the first audio signal and the second audio signal. The side channel signals may correspond to the difference between the first audio signal and the second audio signal. Due to the delay in receiving the second audio signal relative to the first audio signal, the first audio signal may not be aligned with the second audio signal. This misalignment of the first audio signal relative to the second audio signal may increase the difference between the two audio signals. As the difference increases, the phase difference between the frequency domain versions of the audio signals may become less correlated. Summary of the Invention
[0007] In a particular embodiment, an apparatus includes a receiver configured to receive an encoded bitstream comprising an encoded intermediate channel and stereo parameters. The stereo parameters include an inter-channel phase difference (IPD) parameter value and a mismatch value, the mismatch value indicating an amount of time misalignment between an encoder-side reference channel and an encoder-side target channel. The apparatus further includes an intermediate channel decoder configured to decode the encoded intermediate channel to generate a decoded intermediate channel. The apparatus further includes a transform unit configured to perform a transform operation on the decoded intermediate channel to generate a frequency-domain decoded intermediate channel. The apparatus further includes a stereo parameter adjustment unit configured to modify at least a portion of the IPD parameter value based on the mismatch value to generate modified IPD parameter values. The apparatus further includes an upmixer configured to perform an upmixing operation on the frequency-domain decoded intermediate channel to generate a frequency-domain left channel and a frequency-domain right channel. The modified IPD parameter value is applied to the frequency-domain decoded intermediate channel during the upmixing operation. The apparatus further includes a first inverse transform unit configured to perform a first inverse transform operation on the frequency domain left channel to generate a time domain left channel. The apparatus further includes a second inverse transform unit configured to perform a second inverse transform operation on the frequency domain right channel to generate a time domain right channel.
[0008] In another specific embodiment, a method for decoding an audio channel includes receiving, at a decoder, an encoded bitstream comprising an encoded intermediate channel and stereo parameters. The stereo parameters include an inter-channel phase difference (IPD) parameter value and a mismatch value, the mismatch value indicating the amount of time misalignment between an encoder-side reference channel and an encoder-side target channel. The method further includes decoding the encoded intermediate channel to generate a decoded intermediate channel and performing a transform operation on the decoded intermediate channel to generate a frequency-domain decoded intermediate channel. The method further includes modifying at least a portion of the IPD parameter value based on the mismatch value to generate a modified IPD parameter value. The method further includes performing a upmixing operation on the frequency-domain decoded intermediate channel to generate a frequency-domain left channel and a frequency-domain right channel. The modified IPD parameter value is applied to the frequency-domain decoded intermediate channel during the upmixing operation. The method further includes performing a first inverse transform operation on the frequency-domain left channel to generate a time-domain left channel and performing a second inverse transform operation on the frequency-domain right channel to generate a time-domain right channel.
[0009] In another specific embodiment, a non-transitory computer-readable medium includes instructions that, when executed by a processor within a decoder, cause the processor to perform operations including: decoding an encoded intermediate channel to generate a decoded intermediate channel. The encoded intermediate channel is contained in an encoded bitstream received by the decoder. The encoded bitstream further includes stereo parameters including an inter-channel phase difference (IPD) parameter value and a mismatch value, the mismatch value indicating an amount of time misalignment between an encoder-side reference channel and an encoder-side target channel. The operation further includes performing a transform operation on the decoded intermediate channel to generate a frequency-domain decoded intermediate channel. The operation further includes modifying at least a portion of the IPD parameter value based on the mismatch value to generate modified IPD parameter values. The operation further includes performing a up-mixing operation on the frequency-domain decoded intermediate channel to generate a frequency-domain left channel and a frequency-domain right channel. The modified IPD parameter value is applied to the frequency-domain decoded intermediate channel during the up-mixing operation. The operation further includes performing a first inverse transform operation on the frequency domain left channel to generate a time domain left channel and performing a second inverse transform operation on the frequency domain right channel to generate a time domain right channel.
[0010] In another specific embodiment, an apparatus includes means for receiving an encoded bitstream comprising an encoded intermediate channel and stereo parameters. The stereo parameters include an inter-channel phase difference (IPD) parameter value and a mismatch value, the mismatch value indicating an amount of time misalignment between an encoder-side reference channel and an encoder-side target channel. The apparatus further includes means for decoding the encoded intermediate channel to generate a decoded intermediate channel and means for performing a transform operation on the decoded intermediate channel to generate a frequency-domain decoded intermediate channel. The apparatus further includes means for modifying at least a portion of the IPD parameter value based on the mismatch value to generate a modified IPD parameter value. The apparatus further includes means for performing a upmixing operation on the frequency-domain decoded intermediate channel to generate a frequency-domain left channel and a frequency-domain right channel. The modified IPD parameter value is applied to the frequency-domain decoded intermediate channel during the upmixing operation. The apparatus further includes means for performing a first inverse transform operation on the frequency-domain left channel to generate a time-domain left channel and means for performing a second inverse transform operation on the frequency-domain right channel to generate a time-domain right channel.
[0011] After reviewing the entire application, other embodiments, advantages and features of the invention will become apparent, the entire application comprising the following sections: description of drawings, detailed description and claims. Attached Figure Description
[0012] Figure 1 It is a block diagram of a specific illustrative example of a system including an encoder operable to modify the inter-channel phase difference (IPD) parameter and a decoder operable to modify the IPD parameter;
[0013] Figure 2 This is an explanation Figure 1 A diagram of an instance of an encoder;
[0014] Figure 3 This is an explanation Figure 1 A diagram of instances of the decoder;
[0015] Figure 4 It is a specific instance of a method for determining IPD information;
[0016] Figure 5 It is a specific instance of a method for decoding bitstreams;
[0017] Figure 6 This is a block diagram of a specific illustrative example of a device including an encoder operable to modify IPD parameters and a decoder operable to modify IPD parameters; and
[0018] Figure 7It is a block diagram of a specific illustrative example of a base station that includes an encoder operable to modify IPD parameters and a decoder operable to modify IPD parameters. Detailed Implementation
[0019] Specific aspects of the invention are described below with reference to the drawings. In this specification, common components are indicated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular embodiments and are not intended to limit the embodiments. For example, unless the context clearly indicates otherwise, the singular forms “a” and “described” are intended to also include the plural forms. It will be further understood that the terms “comprising” and “including” are used interchangeably with “including” or “containing”. Additionally, it should be understood that the term “wherein” is used interchangeably with “where”. As used herein, ordinal terms used to modify elements such as structures, components, operations, etc. (e.g., “first,” “second,” “third,” etc.) do not themselves indicate any priority or order of an element with respect to another element, but merely distinguish an element from another element having the same name (unless ordinal terms are used). As used herein, the term “group” refers to one or more of a particular element, and the term “a plurality” refers to multiple of a particular element (e.g., two or more).
[0020] In this invention, terms such as “determine,” “calculate,” “shift,” and “adjust” are used to describe how one or more operations are performed. It should be noted that these terms should not be construed as limiting and other techniques can be used to perform similar operations. Additionally, as mentioned herein, “generate,” “calculate,” “use,” “select,” “access,” and “determine” are used interchangeably. For example, “generate,” “calculate,” or “determine” a parameter (or signal) can refer to actively generating, calculating, or determining a parameter (or signal), or it can refer to using, selecting, or accessing a parameter (or signal) that has been generated, for example, by another component or device.
[0021] Systems and apparatuses operable to encode multiple audio signals are disclosed. The apparatus may include an encoder configured to encode multiple audio signals. Multiple recording devices (e.g., multiple microphones) can be used to simultaneously capture multiple audio signals in a timely manner. In some instances, multiple audio signals (or multi-channel audio) can be synthetically (e.g., artificially) generated by multiplexing several audio channels recorded simultaneously or not simultaneously. As illustrative examples, parallel recording or multiplexing of audio channels can produce 2-channel configurations (i.e., stereo: left and right), 5.1-channel configurations (left, right, center, left surround, right surround, and low frequency emphasis (LFE) channels), 7.1-channel configurations, 7.1+4-channel configurations, 22.2-channel configurations, or N-channel configurations.
[0022] An audio capture device in a telephone conference room (or telepresence room) may include multiple microphones for acquiring spatial audio. Spatial audio may include speech and encoded, transmitted background audio. Depending on how the microphones are configured and the location of a given source (e.g., a speaker) relative to the microphones and the room size, speech / audio from said source (e.g., a speaker) may arrive at the multiple microphones at different times. For example, a sound source (e.g., a speaker) may be closer to a first microphone associated with the device than a second microphone associated with the device. Therefore, sound from the sound source may arrive at the first microphone earlier than at the second microphone. The device may receive a first audio signal via the first microphone and a second audio signal via the second microphone.
[0023] Mid-side (MS) decoding and parametric stereo (PS) decoding are improved stereo decoding techniques that offer superior performance compared to dual single-channel decoding. In dual single-channel decoding, the left (L) channel (or signal) and right (R) channel (or signal) are decoded independently without utilizing inter-channel correlation. Before decoding, MS decoding reduces redundancy between correlated L / R channel pairs by transforming the left and right channels into sum and difference channels (e.g., side channels). The sum and difference signals are decoded either by waveform decoding or based on a model in MS decoding. The sum signal requires relatively more bits than the side signals. PS decoding reduces redundancy in each subband by transforming the L / R signals into a sum signal and a set of side parameters. Side parameters can indicate inter-channel intensity difference (IID), inter-channel phase difference (IPD), inter-channel time difference (ITD), side or residual prediction gain, etc. The sum signal is the decoded waveform and is transmitted along with the side parameters. In hybrid systems, side channels can be waveform decoded in lower frequency bands (e.g., less than 2 kHz) and PS-decoded in higher frequency bands (e.g., greater than or equal to 2 kHz), where inter-channel phase preservation is less perceptually critical. In some implementations, PS-decoding can also be performed in the lower frequency band prior to waveform decoding to reduce inter-channel redundancy.
[0024] MS and PS decoding can be performed in the frequency domain or subband domain. In some instances, the left and right channels may be uncorrelated. For example, the left and right channels may contain uncorrelated composite signals. When the left and right channels are uncorrelated, the decoding efficiency of MS, PS, or both can approach that of dual single-channel decoding.
[0025] Depending on the recording configuration, there may be a time shift and other spatial effects (such as echo and reverberation) between the left and right channels. If the time shift and phase mismatch between the channels are not compensated, the sum channel and the difference channel may contain comparable energy that reduces the decoding gain associated with the MS or PS techniques. The reduction in decoding gain may be based on the amount of the time (or phase) shift. The comparable energy of the sum signal and the difference signal may limit the use of MS decoding in certain frames where the channels are time-shifted but highly correlated. In stereo decoding, the middle channel (such as the sum channel) and the side channel (such as the difference channel) may be generated based on the following formulas:
[0026] M = (L + R) / 2, S = (L - R) / 2, Equation 1
[0027] where M corresponds to the middle channel, S corresponds to the side channel, L corresponds to the left channel, and R corresponds to the right channel.
[0028] In some situations, the middle channel and the side channel may be generated based on the following formulas:
[0029] M = c (L + R), S = c (L - R), Equation 2
[0030] where c corresponds to a frequency-dependent composite value. Generating the middle channel and the side channel based on Equation 1 or Equation 2 may be referred to as "downmixing". The reverse process of generating the left channel and the right channel from the middle channel and the side channel based on Equation 1 or Equation 2 may be referred to as "upmixing".
[0031] In some situations, the middle channel may be based on other formulas, such as:
[0032] M = (L + g D R) / 2 or Equation 3
[0033] M = g1L + g2R Equation 4
[0034] where g1 + g2 = 1.0, and where g D is a gain parameter. In other instances, downmixing may be performed in a frequency band where mid(b) = c1L(b) + c2R(b), where c1 and c2 are complex numbers, where side(b) = c3L(b) - c4R(b), and where c3 and c4 are complex numbers.
[0035] A specific approach for selecting between MS decoding and dual-single-channel decoding for a particular frame may include: generating an intermediate signal and a side signal, calculating the energy of the intermediate signal and the side signal, and determining whether to perform MS decoding based on the energy. For example, MS decoding may be performed in response to a determination that the energy ratio of the side signal to the intermediate signal is less than a threshold. To illustrate, if the right channel is shifted by a first time (e.g., approximately 0.001 seconds or 48 samples at 48 kHz), then the first energy of the intermediate signal (corresponding to the sum of the left and right signals) may be equivalent to the second energy of the side signal of the spoken speech frame (corresponding to the difference between the left and right signals). When the first energy is equivalent to the second energy, a higher number of bits can be used to encode the side channel, thereby reducing the decoding performance of MS decoding relative to dual-single-channel decoding. When the first energy is equivalent to the second energy (e.g., when the ratio of the first energy to the second energy is greater than or equal to a threshold), dual-single-channel decoding can therefore be used. In alternative approaches, the choice between MS decoding and dual-single-channel decoding for a specific frame can be determined by comparing the thresholds of the left and right channels and the normalized cross-correlation values.
[0036] In some instances, the encoder may determine a mismatch value indicating an amount of time misalignment between a first audio signal and a second audio signal. As used herein, the terms "time shift value," "shift value," and "mismatch value" are used interchangeably. For example, the encoder may determine a time shift value indicating a shift (e.g., time mismatch) in the first audio signal relative to the second audio signal. The time mismatch value may correspond to the amount of time delay between the reception of the first audio signal at a first microphone and the reception of the second audio signal at a second microphone. Furthermore, the encoder may determine the time mismatch value on a frame-by-frame basis, for example, based on every 20 milliseconds (ms) of speech / audio frames. For example, the time mismatch value may correspond to the amount of time delay of a second frame of the second audio signal relative to a first frame of the first audio signal. Alternatively, the time mismatch value may correspond to the amount of time delay of a first frame of the first audio signal relative to a second frame of the second audio signal.
[0037] When the sound source is closer to the first microphone than to the second microphone, the frames of the second audio signal may be delayed relative to the frames of the first audio signal. In this case, the first audio signal may be referred to as the "reference audio signal" or "reference channel," and the delayed second audio signal may be referred to as the "target audio signal" or "target channel." Alternatively, when the sound source is closer to the second microphone than to the first microphone, the frames of the first audio signal may be delayed relative to the frames of the second audio signal. In this case, the second audio signal may be referred to as the reference audio signal or reference channel, and the delayed first audio signal may be referred to as the target audio signal or target channel.
[0038] Depending on the location of the sound source (e.g., a speaker) in the conference room or telepresence room and how the sound source's (e.g., the speaker's) position changes relative to the microphone, the reference channel and target channel can change from one frame to another; similarly, the time delay value can also change from one frame to another. However, in some implementations, the time mismatch value can always be positive to indicate the amount of delay of the "target" channel relative to the "reference" channel. Furthermore, the time mismatch value can correspond to a "non-causal shift" value, through which the delayed target channel is "pulled back" in time, aligning the target channel with the "reference" channel (e.g., maximizing alignment). A demixing algorithm to determine intermediate and side channels can be performed on the reference channel and the non-causal shifted target channel.
[0039] The encoder can determine the time mismatch value based on a reference audio channel and multiple time mismatch values applied to the target audio channel. For example, the first frame X of the reference audio channel can be received at a first time (m1). The first specific frame Y of the target audio channel can be received at a second time (n1) corresponding to the first time mismatch value (e.g., shift1 = n1 - m1). Additionally, the second frame of the reference audio channel can be received at a third time (m2). The second specific frame of the target audio channel can be received at a fourth time (n2) corresponding to the second time mismatch value (e.g., shift2 = n2 - m2).
[0040] The device can perform a framing or buffering algorithm at a first sampling rate (e.g., a 32 kHz sampling rate (i.e., 640 samples per frame)) to generate frames (e.g., 20 ms samples). In response to the device's determination that the first frame of the first audio signal and the second frame of the second audio signal arrive simultaneously, the encoder can estimate a time mismatch value (e.g., shift1) equal to 0 samples. The left channel (e.g., corresponding to the first audio signal) and the right channel (e.g., corresponding to the second audio signal) can be aligned in time. In some situations, even when aligned, the left and right channels may differ in energy due to various reasons (e.g., microphone calibration).
[0041] In some instances, the left and right channels may be temporally misaligned for various reasons (e.g., compared to one microphone, the speaker's sound source may be closer to one microphone, and the distance between the two microphones may be greater than a threshold (e.g., 1 to 20 cm)). The position of the sound source relative to the microphones can introduce different delays in the left and right channels. Additionally, there may be gain differences, energy differences, or level differences between the left and right channels.
[0042] In some instances, where more than two channels exist, the reference channel is initially selected based on the channel's level or energy, and subsequently refined based on the time mismatch values between different channel pairs (e.g., t1(ref,ch2), t2(ref,ch3), t3(ref,ch4), ..., t3(ref,chN)), where ch1 is the initial reference channel and t1(.), t2(.), etc., are functions of the estimated mismatch values. If all time mismatch values are positive, then ch1 is considered the reference channel. If any of the mismatch values is negative, then the reference channel is reconfigured to the channel associated with the mismatch value that produced the negative value, and the above process continues until the optimal selection of the reference channel is achieved (i.e., based on maximizing the decorrelation of the maximum number of side channels). Hysteresis can be used to overcome any abrupt changes in the selection of the reference channel.
[0043] In some instances, when multiple speakers speak alternately (e.g., without overlap), the arrival times of audio signals from multiple sound sources (e.g., speakers) at the microphone can vary. In this situation, the encoder can dynamically adjust the time mismatch value based on the speaker to identify the reference channel. In other instances, multiple speakers may speak simultaneously, depending on which speaker is loudest, closest to the microphone, etc., which can produce varying time mismatch values. In this situation, the identification of the reference and target channels can be based on the varying time shift value in the current frame and the estimated time mismatch value in previous frames, as well as the energy or time evolution of the first and second audio signals.
[0044] In some instances, when two signals may exhibit little (e.g., no) correlation, a first audio signal and a second audio signal may be synthesized or artificially generated. It should be understood that the examples described herein are illustrative and can be instructive in determining the relationship between a first audio signal and a second audio signal in similar or different contexts.
[0045] The encoder can generate comparison values (e.g., difference or cross-correlation values) based on comparisons between a first frame of a first audio signal and multiple frames of a second audio signal. Each of the multiple frames can correspond to a specific time mismatch value. The encoder can generate a first estimated time mismatch value based on the comparison values. For example, the first estimated time mismatch value can correspond to a comparison value indicating a higher time similarity (or lower difference) between the first frame of the first audio signal and a corresponding first frame of the second audio signal.
[0046] An encoder can determine a final time mismatch value by optimizing a series of estimated time mismatch values in multiple stages. For example, the encoder can first estimate a "provisional" time mismatch value based on comparison values generated from stereo preprocessed and resampled versions of a first audio signal and a second audio signal. The encoder can generate interpolated comparison values associated with time mismatch values close to the estimated "provisional" time mismatch value. The encoder can determine a second estimated "interpolated" time mismatch value based on the interpolated comparison values. For example, the second estimated "interpolated" time mismatch value can correspond to a specific interpolated comparison value indicating a higher time similarity (or lower difference) than the remaining interpolated comparison values and the first estimated "provisional" time mismatch value. If the second estimated "interpolated" time mismatch value of the current frame (e.g., the first frame of the first audio signal) differs from the final time mismatch value of the previous frame (e.g., a frame of the first audio signal preceding the first frame), then the "interpolated" time mismatch value of the current frame is further "corrected" to improve the temporal similarity between the first audio signal and the shifted second audio signal. Specifically, the third estimated "corrected" time mismatch value corresponds to a more accurate measure of temporal similarity by examining the second estimated "interpolated" time mismatch value of the current frame and the final estimated time mismatch value of the previous frame. The third estimated "corrected" time mismatch value is further adjusted to estimate the final time mismatch value by limiting any spurious changes in the time mismatch value between frames, and is further controlled to prevent a switch from a negative time mismatch value to a positive time mismatch value (or vice versa) in two consecutive (or connected) frames as described herein.
[0047] In some instances, the encoder may control the switching between positive and negative time mismatch values, or vice versa, in consecutive or adjacent frames. For example, the encoder may set the final time mismatch value to a specific value (e.g., 0), which indicates no time shift based on the estimated "interpolated" or "corrected" time mismatch value of the first frame and the corresponding estimated "interpolated" or "corrected" or final time mismatch value in a specific frame preceding the first frame. To illustrate, in response to a determination that one of the estimated "provisional" or "interpolated" or "corrected" time mismatch values of the current frame is positive and the other of the estimated "provisional" or "interpolated" or "corrected" or "final" estimated time mismatch values of the previous frame (e.g., a frame preceding the first frame) is negative, the encoder may set the final time mismatch value of the current frame (e.g., the first frame) to indicate no time shift, i.e., shift1 = 0. Alternatively, in response to a determination that one of the estimated "provisional," "interpolated," or "corrected" time mismatch values for the current frame is negative and the other of the estimated "provisional," "interpolated," "corrected," or "final" time mismatch values for the previous frame (e.g., the frame preceding the first frame) is positive, the encoder may also set the final time mismatch value for the current frame (e.g., the first frame) to indicate no time shift, i.e., shift1 = 0.
[0048] The encoder can select a frame of either the first or second audio signal as a "reference" or "target" based on a time mismatch value. For example, in response to determining that the final time mismatch value is positive, the encoder can generate a reference channel or signal indicator with a first value (e.g., 0), indicating that the first audio signal is a "reference" signal and the second audio signal is a "target" signal. Alternatively, in response to determining that the final time mismatch value is negative, the encoder can generate a reference channel or signal indicator with a second value (e.g., 1), indicating that the second audio signal is a "reference" signal and the first audio signal is a "target" signal.
[0049] The encoder can estimate a relative gain (e.g., a relative gain parameter) associated with a reference signal and a non-causal shifted target signal. For example, in response to determining that the final time mismatch value is positive, the encoder can estimate a gain value to normalize or equalize the amplitude or power level of the first audio signal offset relative to the second audio signal by the non-causal time mismatch value (e.g., the absolute value of the final time mismatch value). Alternatively, in response to determining that the final time mismatch value is negative, the encoder can estimate a gain value to normalize or equalize the amplitude or power level of the non-causal shifted first audio signal relative to the second audio signal. In some instances, the encoder can estimate a gain value to normalize or equalize the amplitude or power level of the "reference" signal relative to the non-causal shifted "target" signal. In other instances, the encoder can estimate a gain value (e.g., a relative gain value) based on a reference signal relative to the target signal (e.g., the unshifted target signal).
[0050] The encoder can generate at least one encoded signal (e.g., an intermediate signal, a side signal, or both) based on a reference signal, a target signal, a non-causal time mismatch value, and a relative gain parameter. In other embodiments, the encoder can generate at least one encoded signal (e.g., an intermediate channel, a side channel, or both) based on a reference channel and a time mismatch-adjusted target channel. The side signal can correspond to the difference between a first sample of a first frame of a first audio signal and a selected sample of a selected frame of a second audio signal. The encoder can select the selected frame based on the final time mismatch value. Due to the reduced difference between the first sample and the selected sample, fewer bits are available to encode the side channel signal compared to other samples of the second audio signal corresponding to a frame received by the device simultaneously with the first frame. The transmitter of the device can transmit at least one encoded signal, a non-causal time mismatch value, a relative gain parameter, a reference channel, or a signal indicator, or a combination thereof.
[0051] The encoder can generate at least one encoded signal (e.g., an intermediate signal, a side signal, or both) based on a reference signal, a target signal, a non-causal time mismatch value, a relative gain parameter, low-frequency band parameters of a specific frame of a first audio signal, high-frequency band parameters of a specific frame, or a combination thereof. The specific frame may precede the first frame. Certain low-frequency band parameters, high-frequency band parameters, or combinations thereof from one or more preceding frames can be used to encode the intermediate signal, side signal, or both of the first frame. Encoding the intermediate signal, side signal, or both based on low-frequency band parameters, high-frequency band parameters, or combinations thereof can improve the estimation of the non-causal time mismatch value and the inter-channel relative gain parameter. Low-frequency band parameters, high-frequency band parameters, or combinations thereof may include pitch parameters, speech parameters, decoder type parameters, low-frequency band energy parameters, high-frequency band energy parameters, tilt parameters, pitch gain parameters, FCB gain parameters, decoding mode parameters, speech activity parameters, noise assessment parameters, signal-to-noise ratio parameters, formant parameters, speech / music decision parameters, non-causal shift, inter-channel gain parameters, or combinations thereof. The transmitter of the device can transmit at least one encoded signal, a non-causal time mismatch value, a relative gain parameter, a reference channel (or signal) indicator, or a combination thereof. In this invention, terms such as "determine," "calculate," "shift," and "adjust" are used to describe how one or more operations are performed. It should be noted that these terms should not be construed as limiting and that other techniques can be used to perform similar operations.
[0052] See Figure 1 This document discloses specific illustrative examples of the system and generally designates the system as 100. System 100 includes a first device 104 communicatively coupled to a second device 106 via a network 120. The network 120 may include one or more wireless networks, one or more wired networks, or a combination thereof.
[0053] The first device 104 includes an encoder 114, a transmitter 110, and one or more input interfaces 112. A first input interface of the input interfaces 112 is coupled to a first microphone 146, and a second input interface of the input interfaces 112 is coupled to a second microphone 148. (Regarding...) Figure 2 A non-limiting example of the architecture of encoder 114 is described. The second device 106 includes receiver 115 and decoder 118. Regarding... Figure 3 A non-limiting example of the architecture of decoder 118 is described. A second device 106 is coupled to a first loudspeaker 142 and to a second loudspeaker 144.
[0054] During operation, the first device 104 receives a reference channel 130 (e.g., a first audio signal) from a first microphone 146 via a first input interface and a target channel 132 (e.g., a second audio signal) from a second microphone 148 via a second input interface. The reference channel 130 corresponds to either a left or right channel, and the target channel 132 corresponds to the other. A sound source 152 (e.g., a user, speaker, ambient noise, musical instrument, etc.) may be closer to the first microphone 146 than to the second microphone 148. Therefore, the audio signal from the sound source 152 may be received at the input interface 112 via the first microphone 146 at an earlier time than via the second microphone 148. This inherent delay in multi-channel signal acquisition via multiple microphones can introduce a time misalignment between the reference channel 130 and the target channel 132. Therefore, the target channel 132 may be adjusted (e.g., time-shifted) to be substantially aligned with the reference channel 130.
[0055] Encoder 114 is configured to determine a mismatch value 116 (e.g., a non-causal shift value) indicating the amount of time misalignment between reference channel 130 and target channel 132. According to one embodiment, mismatch value 116 indicates the amount of time misalignment in the time domain. According to another embodiment, mismatch value 116 indicates the amount of time misalignment in the frequency domain. Encoder 114 is configured to adjust target channel 132 according to mismatch value 116 to produce adjusted target channel 134. Because target channel 132 is adjusted according to mismatch value 116, adjusted target channel 134 is substantially aligned with reference channel 130.
[0056] Encoder 114 is configured to estimate stereo parameters 162 based on the frequency domain versions of the adjusted target channel 134 and reference channel 130. According to one embodiment, a mismatch value 116 is included in the stereo parameters 162. Stereo parameters 162 also include an inter-channel phase difference (IPD) parameter value 164 and an inter-channel time difference (ITD) parameter value 166. According to one embodiment, the mismatch value 116 is similar to (e.g., the same value) the ITD parameter value 166. The IPD parameter value 164 can indicate the phase difference between channels 130 and 134 on a per-band basis.
[0057] According to one implementation, encoder 114 modifies IPD parameter value 164 based on time mismatch value 116 to produce modified IPD parameter value 165. For example, encoder 114 may modify IPD parameter value 164 to produce modified IPD parameter value 165 in response to determination that the absolute value of mismatch value 116 meets a threshold. The determination of whether IPD parameter value 164 should be modified may be based on short-term and long-term IPD values.
[0058] According to one embodiment, encoder 114 sets one or more of the IPD parameter values 164 to zero to produce a modified IPD parameter value 165. According to another embodiment, encoder 114 smooths one or more of the IPD parameter values 164 in time to produce a modified IPD parameter value 165.
[0059] For illustration, encoder 114 may determine IPD information based on mismatch value 116. The IPD information may indicate how IPD parameter value 164 should be modified, and IPD parameter value 164 may indicate the phase difference between the frequency domain version of reference channel 130 and the frequency domain version of adjusted target channel 134 in different frequency bands (b). According to one embodiment, modifying IPD parameter value 164 includes setting one or more of the IPD parameter values 164 to zero (or other gain value). According to another embodiment, modifying IPD parameter value 164 may include temporally smoothing one or more of the IPD parameter values 164. According to one embodiment, IPD parameter values using residual decoding (e.g., IPD parameters with lower frequency bands (b)) are modified, while IPD parameter values with higher frequency bands remain unchanged.
[0060] Encoder 114 can determine whether the mismatch value 116 meets a first mismatch threshold (e.g., an upper mismatch threshold). If encoder 114 determines that the mismatch value 116 meets (e.g., is greater than) the first mismatch threshold, then encoder 114 is configured to modify the IPD parameter value 164 for each frequency band (b) associated with the frequency domain version of the adjusted target channel 134. Therefore, if the time misalignment between channels 130, 132 is large (e.g., greater than the first mismatch threshold), then shifting the target channel 132 to improve the time alignment of the target channel 130 with the reference channel 132 can result in a large change in the IPD parameter value between one frame and the next after the shift. For example, the time shift of the target channel 132 can cause the target channel 132 to shift much larger than the time distance that can be indicated by the IPD parameter value 164. For illustration, the IPD parameter value 164 can indicate a value from -pi to pi. However, the time shift can be greater than said range. Therefore, if the mismatch value 116 is greater than the first mismatch threshold, the encoder 114 can determine that the IPD parameter value 164 does not have a specific correlation. As a result, the IPD parameter value 164 can be set to zero (or smoothed over time over several frames).
[0061] Encoder 114 can also determine whether the mismatch value 116 meets a second mismatch threshold (e.g., a lower mismatch threshold). If encoder 114 determines that the mismatch value 116 fails to meet (e.g., is less than) the second mismatch threshold, then encoder 114 is configured to bypass modification of the IPD parameter value 164. Therefore, if the time misalignment between channels 130 and 132 is small (e.g., less than the second mismatch threshold), then shifting the target channel 132 to improve the time alignment between the target channel 130 and the reference channel 132 can result in a small change in the IPD parameter value 164 between one frame and the next after the shift. As a result, the change indicated by the IPD parameter value 164 can be more significant, and the IPD parameter value 164 for each frequency band (b) can remain unchanged.
[0062] Encoder 114 may modify the IPD parameter value 164 of a subset of frequency bands (b) associated with the frequency domain version of target channel 132 in response to a first determination that mismatch value 116 fails to meet a first mismatch threshold and in response to a determination that mismatch value 116 meets a second mismatch threshold. According to one embodiment, the IPD parameter value 164 may be modified (e.g., set to zero or time-smoothed) for frequency bands (b) associated with residual decoding in response to mismatch value 116 failing to meet the first mismatch threshold and meeting the second mismatch threshold. According to another embodiment, the IPD parameter value 164 used for selecting frequency bands (b) may be modified in response to mismatch value 116 failing to meet the first mismatch threshold and meeting the second mismatch threshold.
[0063] Encoder 114 is configured to perform upmixing operations on the adjusted target channel 134 (or a frequency domain version of the adjusted target channel 134) and reference channel 130 (or a frequency domain version of the reference channel 130) using IPD parameter value 164, modified IPD parameter value 165, etc. For example, encoder 114 can generate intermediate channel 262 and side channel 264 at least in part based on the upmixing operation. Regarding... Figure 2 The generation of intermediate channel 262 and side channel 264 is described in more detail. Encoder 114 is further configured to encode intermediate channel 262 to generate encoded intermediate channel 340, and encoder is configured to encode side channel 264 to generate encoded side channel 342.
[0064] Bitstream 248 (e.g., an encoded bitstream) includes an encoded intermediate channel 340, an encoded side channel 342, and stereo parameters 162. According to one embodiment, modified IPD parameter values 165 are not included in bitstream 248, and decoder 118 adjusts IPD parameter values 164 to produce modified IPD parameter values (e.g., regarding...). Figure 3(As described). According to another embodiment, the modified IPD parameter value 165 is included in the bit stream 248. The transmitter 110 is configured to transmit the bit stream 248 to the second device 106 via the network 120.
[0065] Receiver 115 is configured to receive bit stream 248. (See also: Regarding...) Figure 3 As described, decoder 118 is configured to perform decoding operations on bitstream 248 to generate left channel 126 and right channel 128. One or more speakers are configured to output left channel 126 and right channel 128. For example, second device 106 may output left channel 126 via first loudspeaker 142, and second device 106 may output right channel 128 via second loudspeaker 144. In an alternative embodiment, left channel 126 and right channel 128 may be transmitted as a stereo signal pair to a single output loudspeaker.
[0066] System 100 can modify IPD parameters based on mismatch value 116 to reduce artifacts during the decoding phase. For example, to reduce the introduction of artifacts that can be caused by decoding IPD parameter values that do not contain relevant information, encoder 114 can generate IPD information (e.g., one or more flags, IPD parameter values with predefined patterns, IPD parameter values set to zero in the low-frequency band) indicating whether encoder 114 should modify (e.g., time-smooth) IPD parameters, indicating which IPD parameters should be modified, etc.
[0067] See Figure 2 The diagram illustrates a specific embodiment of encoder 114A. Encoder 114A may correspond to... Figure 1 The encoder 114A includes a transformation unit 202, a stereo parameter estimator 206, a downmixer, a stereo parameter adjustment unit 111, an inverse transformation unit 213, an intermediate channel encoder 216, a side channel encoder 210, a side channel modifier 230, an inverse transformation unit 232, and a multiplexer 252.
[0068] Reference channel 130 and adjusted target channel 134 are provided to conversion unit 202. Adjusted target channel 134 is generated by shifting target channel 132 (e.g., non-causally shifting) with a mismatch value 116. Encoder 114A can determine whether a time shift operation should be performed on target channel 132 based on the mismatch value 116, and can determine a decoding mode to generate adjusted target channel 134. In some embodiments, if the mismatch value 116 is not used to shift target channel 132 in time, then adjusted target channel 134 may be time-shifted identically to target channel 132.
[0069] Transform unit 202 is configured to perform a first transform operation on reference channel 130 to generate frequency-domain reference channel 258, and transform unit 202 is configured to perform a second transform operation on adjusted target channel 134 to generate frequency-domain adjusted target channel 256. The transform operation may include Discrete Fourier Transform (DFT) operation, Fast Fourier Transform (FFT) operation, etc. According to some embodiments, Quadrature Mirror Filterbank (QMF) operation (using filter banks, such as complex low-delay filter banks) can be used to split the input signal (e.g., reference channel 130 and adjusted target channel 134) into multiple sub-bands. Encoder 114A may be configured to determine, based on a first time-shift operation, whether a second time-shift (e.g., non-causal) operation should be performed on the frequency-domain adjusted target channel 256 in the transform domain to generate a modified version of the frequency-domain adjusted target channel 256.
[0070] Frequency domain reference channel 258 and frequency domain adjusted target channel 256 are provided to stereo parameter estimator 206. Stereo parameter estimator 206 is configured to extract (e.g., generate) stereo parameters 162 based on frequency domain reference channel 258 and frequency domain adjusted target channel 256. For example, IID(b) may depend on the energy E of the left channel in frequency band (b). L (b) and the energy E of the right channel in frequency band (b). R (b). For example, IID(b) can be expressed as 20 × log 10 (E L (b) / E R (b)). The IPD estimated and transmitted at the encoder provides an estimate of the phase difference between the left and right channels in band (b) in the frequency domain. Stereo parameter 162 may include additional (or alternative) parameters, such as ICC, ITD, etc. Stereo parameter 162 can be transmitted to Figure 1 The second device 106 can also be provided to the downmixer 207. The downmixer 207 includes an intermediate channel generator 212 and a side channel generator 208. In some embodiments, stereo parameters 162 are provided to the side channel encoder 210.
[0071] Stereo parameter 162 is also provided to stereo parameter adjustment unit 111. Stereo parameter adjustment unit 111 is configured to modify IPD parameter value 164 (e.g., stereo parameter 162) based on mismatch value 116 to produce modified IPD parameter value 165. Alternatively or additionally, stereo parameter adjustment unit 111 is configured to determine the residual gain (e.g., residual gain value) to be applied to the residual channel (e.g., side channel 264). In some embodiments, stereo parameter adjustment unit 111 may also determine the value of an IPD flag (not shown). The value of the IPD flag indicates whether the IPD parameter values for one or more frequency bands should be ignored or set to zero. For example, when the IPD flag is confirmed, the IPD parameter values for one or more frequency bands may be ignored or set to zero. The stereo parameter adjustment unit 111 can provide IPD information (e.g., modified IPD parameter value 165, IPD parameter value 164, IPD flag or a combination thereof) to the downmixer 207 (e.g., side channel generator 208) and the side channel modifier 230.
[0072] Frequency domain reference channel 258 and frequency-adjusted target channel 256 are provided to demixer 207. According to some embodiments, stereo parameters 162 are provided to intermediate channel generator 212. The intermediate channel generator 212 of demixer 207 is configured to generate a frequency domain intermediate channel M based on frequency domain reference channel 258 and frequency-adjusted target channel 256. fr (b) 266. According to some implementation schemes, a frequency domain channel 266 is also generated based on stereo parameters 162.
[0073] Frequency domain intermediate channel M fr (b) 266 is provided from intermediate channel generator 212 to inverse transform unit 213 (e.g., DFT synthesizer) and side channel modifier 230. Inverse transform unit 213 is configured to perform an inverse transform operation on frequency domain intermediate channel 266 to generate intermediate channel 262 (e.g., time domain intermediate channel). The inverse transform operation may include inverse discrete fourier transform (IDFT) operation, inverse discrete cosine transform (IDCT) operation, etc. According to one embodiment, inverse transform unit 213 synthesizes frequency domain intermediate channel 266 to generate intermediate channel 262. Intermediate channel 262 is provided to intermediate channel encoder 216. Intermediate channel encoder 216 is configured to encode intermediate channel 262 to generate encoded intermediate channel 340. Encoded intermediate channel 340 is provided to multiplexer 252.
[0074] The side channel generator 208 of the downmixer 207 is configured to generate a frequency domain side channel S based on the frequency domain reference channel 258, the frequency domain adjusted target channel 256, the stereo parameter 162, and the modified IPD parameter value 165. fr (b) 270. In each frequency band (e.g., interval) of the frequency domain side channel 270, the gain parameter (g) may be different and may be based on the inter-channel level difference (e.g., based on the stereo parameter 162). For example, the frequency domain side channel 270 may be expressed as (L fr (b) - c(b)×R fr (b)) / (1+c(b)), where c(b) can be ILD(b) or depends on ILD(b) (e.g., c(b) = 10^(ILD(b) / 20)). A frequency-domain side channel 270 is provided to the side channel modifier 230. The side channel modifier 230 modifies the IPD parameter value 165. The side channel modifier 230 is configured to generate a modified side channel 268 (e.g., a frequency-domain modified side channel) based on the frequency-domain side channel 270, the frequency-domain intermediate channel 266, and the modified IPD parameter value 165.
[0075] Inverse transform unit 232 is configured to perform an inverse transform operation on the modified side channel 268 to generate a side channel 264 (e.g., a time-domain side channel). The inverse transform operation may include IDFT operations, IDCT operations, etc. According to one embodiment, inverse transform unit 232 synthesizes the modified side channel 268 to generate a side channel 264. The side channel 264 is provided to the side channel encoder 210. In response to a residual decoding enable signal 254, the side channel encoder 210 is activated and configured to encode the side channel 264 to generate an encoded side channel 342. If the residual decoding enable signal 254 indicates that residual coding is disabled, then the side channel encoder 210 may generate the encoded side channel 342 for one or more frequency bands.
[0076] The encoded intermediate channel 340, the encoded side channel 342, and the stereo parameter 162 are provided to the multiplexer 252. The multiplexer 252 is configured to generate a bitstream 248 based on the encoded intermediate channel 340, the encoded side channel 342, and the stereo parameter 162.
[0077] Encoder 114A can modify IPD parameters based on mismatch value 116 to reduce artifacts during the decoding phase. For example, to reduce the introduction of artifacts that can be caused by decoding IPD parameter values that do not contain relevant information, encoder 114A can generate IPD information (e.g., one or more flags, IPD parameter values with predefined patterns, IPD parameter values set to zero in the low-frequency band) indicating whether encoder 114A should modify (e.g., time-smooth) IPD parameters, indicating which IPD parameters should be modified, etc.
[0078] See Figure 3 The diagram illustrates a specific embodiment of decoder 118A. Decoder 118A may correspond to... Figure 1 The decoder 118A includes an intermediate channel decoder 302, a side channel decoder 304, a conversion unit 306, a conversion unit 308, an upmixer 310, a stereo parameter adjustment unit 312, an inverse conversion unit 318, an inverse conversion unit 320, and an inter-channel alignment unit 322.
[0079] Bitstream 248 is provided to decoder 118A, and decoder 118A is configured to decode portions of bitstream 248 to generate left channel 126 and right channel 128. Bitstream 248 includes an encoded intermediate channel 340, an encoded side channel 342, and stereo parameters 162. According to one embodiment, a demultiplexer (not shown) can extract the encoded intermediate channel 340, the encoded side channel 342, and stereo parameters 162 from bitstream 248. The encoded intermediate channel 340 is provided to intermediate channel decoder 302, the encoded side channel 342 is provided to side channel decoder 304, and stereo parameters 162 are provided to stereo parameter adjustment unit 312. Stereo parameters 162 include at least IPD parameter value 164, ITD parameter value 166, and mismatch value 116.
[0080] Intermediate channel decoder 302 is configured to decode encoded intermediate channel 340 to produce decoded intermediate channel 344 (e.g., time-domain intermediate channel m). CODED (t)). The decoded intermediate channel 344 is provided to the transformation unit 306. The transformation unit 306 is configured to perform a transformation operation on the decoded intermediate channel 344 to generate a frequency-domain decoded intermediate channel 348. The transformation operation may include a discrete cosine transform (DCT) operation, a discrete fourier transform (DFT) operation, a fast fourier transform (FFT) operation, etc. The frequency-domain decoded intermediate channel 348 is provided to the upmixer 310.
[0081] Side-channel decoder 304 is configured to decode the encoded side-channel 342 to produce a decoded side-channel 346. The decoded side-channel 346 is provided to transform unit 308. Transform unit 308 is configured to perform a second transform operation on the decoded side-channel 346 to produce a frequency-domain decoded side-channel 350. The second transform operation may include DCT, DFT, FFT, etc. The frequency-domain decoded side-channel 350 is also provided to upmixer 310. Although the decoding operation of the encoded side-channel 342 has been described, in one embodiment, decoder 118A may receive an IPD flag indicating whether decoder 118A should process or ignore residual signal information in one or more frequency bands. Therefore, when the IPD flag indicates that residual information in one or more frequency bands should be ignored, the decoding operation of the encoded side-channel 342 can be bypassed (for one or more frequency bands).
[0082] Stereo parameters 162, encoded into bitstream 248, are provided to stereo parameter adjustment unit 312. Stereo parameter adjustment unit 312 includes comparison unit 314 and modification unit 316. Comparison unit 314 is configured to compare the absolute value of mismatch value 116 with a threshold. Modification unit 316 is configured to modify at least a portion of IPD parameter value 164 in response to a determination that the absolute value of mismatch value 116 satisfies (e.g., is greater than) the threshold, to generate modified IPD parameter value 352. For illustration, the determination of whether IPD parameter value 352 should be modified can be expressed using the following pseudocode:
[0083] for( b=0; b <nbands; b++ )
[0084] {
[0085] if( b<= maxband&&res_coding_Active == FALSE )
[0086] {
[0087] g = gLB; / * a fixed threshold * /
[0088] }
[0089] else
[0090] {
[0091] g = pSideGain[b]; / * a per-band side gain value * /
[0092] }
[0093] if( b <ipd_band_max )
[0094] {
[0095] c = (1+g) / (1-g);
[0096] if( b <res_pred_band_min
[0097] &&res_coding_Active == TRUE
[0098] &&|(ITD mismatch value)|>80.0 )
[0099] {
[0100] / * modify the IPD parameters * /
[0101] alpha = 0;
[0102] beta = (atan2(sin(alpha), (cos(alpha) + 2*c)));
[0103] }
[0104] else
[0105] {
[0106] / * Don't modify the IPD parameters * /
[0107] alpha = pIpd[b];
[0108] beta = (atan2(sin(alpha), (cos(alpha) + 2*c)));
[0109] }
[0110] }
[0111] As a non-limiting example, the modification unit 316 can generate a modified IPD parameter value 352 by setting one or more of the IPD parameter values 164 to zero. As another non-limiting example, the modification unit 316 can generate a modified IPD parameter value 352 by temporally smoothing one or more of the IPD parameter values 164. The modified IPD parameter value 352 is provided to the upmixer 310. According to one embodiment, the stereo parameter adjustment unit 312 is configured to modify the IPD parameter value 164 based on the availability of the coded side channel 342. According to another embodiment, the stereo parameter adjustment unit 312 is configured to modify the IPD parameter value 164 based on the bit rate associated with the bit stream 248.
[0112] According to another embodiment, the stereo parameter adjustment unit 312 is configured to modify the IPD parameter value 164 based on the sound parameters, packet loss determination associated with the previous frame, speech / music classification, or another parameter. As a non-limiting example, in response to a determination that the previous frame was lost in transmission, the stereo parameter adjustment unit 312 may modify the IPD parameter value 164 to produce a modified IPD parameter value 352.
[0113] Upmixer 310 is configured to perform upmixing on the frequency-domain decoded intermediate channel 348 to generate a frequency-domain left channel 354 and a frequency-domain right channel 356. Modified IPD parameter values 352 and other stereo parameters 162 (e.g., ILD, residual prediction gain, etc.) are applied to the frequency-domain decoded intermediate channel 348 during the upmixing operation. According to some embodiments, upmixer 310 performs upmixing on the frequency-domain decoded intermediate channel 348 and the frequency-domain decoded side channel 350 to generate frequency-domain channels 354 and 356. In this scenario, modified IPD parameter values 352 are applied to the frequency-domain decoded intermediate channel 348 and the frequency-domain decoded side channel 350 during the upmixing operation. The frequency-domain left channel 354 is provided to inverse transform unit 318, and the frequency-domain right channel 356 is provided to inverse transform unit 320.
[0114] Inverse transform unit 318 is configured to perform a first inverse transform operation on the frequency domain left channel 354 to generate a time domain left channel 358. For example, the first inverse transform operation may include an inverse discrete cosine transform (IDCT) operation, an inverse discrete fourfold transform (IDFT) operation, an inverse fast fourfold transform (IFFT) operation, etc. According to one embodiment, inverse transform unit 318 is configured to perform a synthesis windowing operation on the frequency domain left channel 354 to generate a time domain left channel 358. The time domain left channel 358 is provided to inter-channel alignment unit 322. Inverse transform unit 320 is configured to perform a second inverse transform operation on the frequency domain right channel 356 to generate a time domain right channel 360. For example, the second inverse transform operation may include an IDCT operation, an IDFT operation, an IFFT operation, etc. According to one embodiment, inverse transform unit 320 is configured to perform a synthesis windowing operation on the frequency domain right channel 356 to generate a time domain right channel 368. The time-domain right channel 360 is also provided to the inter-channel alignment unit 322.
[0115] The ITD parameter value 166 of stereo parameter 162 is provided to the inter-channel alignment unit 322. According to... Figure 3In the illustrated example, the stereo parameter adjustment unit 312 provides the ITD parameter value 166 to the inter-channel alignment unit 322. In other embodiments, the ITD parameter value 166 is provided directly to the inter-channel alignment unit 322. According to one embodiment, the inter-channel alignment unit 322 is configured to adjust the time-domain right channel 360 based on the ITD parameter value 166 to generate a right channel 128 and to transmit the time-domain left channel 358 as the left channel 126. According to one embodiment, the inter-channel alignment unit 322 is configured to adjust the time-domain left channel 358 based on the ITD parameter value 166 to generate a left channel 126 and to transmit the time-domain right channel 360 as the right channel 128.
[0116] Decoder 118A can generate channels 126 and 128 with reduced artifacts compared to channels that do not have the modified IPD parameter value 352. For example, in order to reduce the introduction of artifacts that can be caused by decoding IPD parameter values that do not contain relevant information (e.g., IPD parameter value 164), decoder 118A can modify IPD parameter value 164 to temporally smooth out irrelevant IPD parameter value 164 that may otherwise cause artifacts.
[0117] refer to Figure 4 This demonstrates method 400 for determining IPD information. Method 400 can be derived from... Figure 1 First device 104 Figure 2 The encoder 114A or a combination thereof is used for execution.
[0118] Method 400 includes performing a first transform operation on the reference channel at the encoder at 402 to generate a frequency domain reference channel. For example, see... Figure 2 The transformation unit 202 performs a first transformation operation on the reference channel 130 to generate a frequency domain reference channel 258.
[0119] Method 400 also includes performing a second transform operation on the adjusted version of the target channel at 404 to generate a frequency-domain adjusted target channel. For example, see... Figure 2 The transformation unit 202 performs a second transformation operation on the adjusted target channel 134 (e.g., an adjusted version of the target channel 132 based on the mismatch value 116) to generate the frequency-domain adjusted target channel 256.
[0120] Method 400 also includes determining at 406 a mismatch value indicating the amount of time misalignment between the reference channel and the target channel. For example, the reference... Figure 1 Encoder 114 determines a mismatch value 116 indicating the amount of time misalignment between reference channel 130 and target channel 132.
[0121] Method 400 also includes determining IPD information based on the mismatch value at 408. The IPD information indicates that at least a portion of the IPD parameters should be modified, and the IPD parameters indicate the phase difference between the frequency-domain reference channel and the frequency-domain adjusted target channel in different frequency bands. For example, see... Figure 2 The stereo parameter adjustment unit 111 determines at least a portion of the IPD parameter value 164 to be modified based on the mismatch value 116.
[0122] According to one embodiment, method 400 includes setting one or more of the IPD parameter values 164 to zero to modify the IPD parameter value 164. According to one embodiment, method 400 includes temporally smoothing one or more of the IPD parameter values 164 to modify the IPD parameter value 164. According to one embodiment, method 400 includes determining that the mismatch value 116 meets a first mismatch threshold. Method 400 may further include modifying the IPD parameter value 164 of each frequency band associated with the frequency-domain adjusted target channel 256 in response to determining that the mismatch value 116 meets the first mismatch threshold. According to one embodiment, method 400 includes determining that the mismatch value 116 fails to meet a second mismatch threshold. Method 400 may further include bypassing the modification of the IPD parameter value 164 in response to the determination that the mismatch value 116 fails to meet the second mismatch threshold.
[0123] According to one implementation, method 400 includes determining that mismatch value 116 fails to meet a first mismatch value and determining that mismatch value 116 meets a second mismatch value. Method 400 may further include modifying IPD parameter values 164 of a subset of the frequency band associated with the frequency-domain adjusted target channel 256 in response to determining that mismatch value 116 fails to meet the first mismatch threshold and in response to determining that mismatch value 116 meets the second mismatch threshold.
[0124] Method 400 also includes transmitting a bit stream based on IPD information at 410. For example, see... Figure 1 Transmitter 110 can transmit the bit stream to the second device 106.
[0125] Figure 4 Method 400 can modify the IPD parameter value based on the mismatch value 116 to reduce artifacts during the decoding phase. For example, to reduce the introduction of artifacts that can be caused by decoding IPD parameter values that do not contain relevant information, method 400 can enable the generation of IPD information (e.g., one or more flags, IPD parameter values with predefined patterns, IPD parameter values set to zero in the low-frequency band) that indicate whether the encoder 114A should modify (e.g., smooth in time) the IPD parameters, and which IPD parameters should be modified.
[0126] refer to Figure 5 This demonstrates method 500 for decoding bitstreams. Method 400 can be derived from... Figure 1 The second device 106 Figure 3 The decoder 300 or a combination thereof is executed.
[0127] Method 500 includes receiving, at 502, an encoded bitstream containing encoded intermediate channels and stereo parameters at the decoder. The stereo parameters include IPD parameter values and mismatch values indicating the amount of time misalignment between the encoder-side reference channel and the encoder-side target channel. For example, see... Figure 1 Receiver 115 receives bit stream 248 containing encoded intermediate channel 340, encoded side channel 342 and stereo parameters 162.
[0128] Method 500 also includes decoding the coded intermediate channel at 504 to produce a decoded intermediate channel. For example, see... Figure 3 Intermediate channel decoder 302 decodes the encoded intermediate channel 340 to produce a decoded intermediate channel 344. Method 500 also includes performing a transform operation on the decoded intermediate channel at 506 to produce a frequency-domain decoded intermediate channel. For example, see... Figure 3 The transformation unit 306 performs a transformation operation on the decoded intermediate channel 344 to generate the frequency domain decoded intermediate channel 348.
[0129] Method 500 further includes, at 508, modifying at least a portion of the IPD parameter value based on the mismatch value to produce a modified IPD parameter value. For example, see... Figure 3 The comparison unit 314 compares the absolute value of the mismatch value 116 with a threshold. The modification unit 316 modifies at least a portion of the IPD parameter value 164 to produce a modified IPD parameter value 352 in response to the determination that the absolute value of the mismatch value 116 satisfies (e.g., is greater than) the threshold.
[0130] Method 500 also includes performing a upmixing operation on the frequency-domain decoded intermediate channel at 510 to generate a frequency-domain left channel and a frequency-domain right channel. Modified IPD parameters are applied to the frequency-domain decoded intermediate channel during the upmixing operation. For example, see... Figure 3 During the upmixing process, the upmixer 310 applies the modified IPD parameter value to the frequency-domain decoded intermediate channel 348 to generate the frequency-domain left channel 354 and the frequency-domain right channel 356.
[0131] Method 500 includes performing a first inverse transform operation on the frequency domain left channel at position 512 to generate the time domain left channel. For example, see... Figure 3 Inverse transform unit 318 performs a first inverse transform operation on the frequency domain left channel 354 to generate the time domain left channel 358. Method 500 also includes performing a second inverse transform operation on the frequency domain right channel at 514 to generate the time domain right channel. For example, see... Figure 3The inverse transform unit 520 performs a second inverse transform operation on the frequency domain right channel 356 to generate the time domain right channel 360.
[0132] Method 500 also includes outputting at least one of the left or right channels at point 516. The left channel is associated with the time-domain left channel, and the right channel is associated with the time-domain right channel. For example, see Figure 1 The first loudspeaker 142 outputs a left channel 126 associated with the time-domain left channel 358, and the second loudspeaker 144 outputs a right channel 128 associated with the time-domain right channel 360.
[0133] Figure 5 Method 500 enables the generation of channels 126 and 128 with reduced artifacts compared to channels that do not have the modified IPD parameter value 352. For example, to reduce the introduction of artifacts that can be caused by decoding IPD parameter values that do not contain relevant information (e.g., IPD parameter value 164), decoder 118A can modify IPD parameter value 164 to temporally smooth out irrelevant IPD parameter value 164 that may further cause artifacts.
[0134] See Figure 6 This is a block diagram depicting a specific illustrative example of a device (e.g., a wireless communication device), and the device is generally designated as 600. In various embodiments, with... Figure 6 Compared to the description herein, device 600 may have fewer or more components. In the illustrative embodiment, device 600 may correspond to Figure 1 First device 104 Figure 1 The second device 106 or a combination thereof. In an illustrative embodiment, device 600 may perform the reference... Figures 1 to 5 The system and methods described one or more operations.
[0135] In a particular embodiment, device 600 includes a processor 606 (e.g., a central processing unit, CPU). Device 600 includes one or more additional processors 610 (e.g., one or more digital signal processors, DSPs). Processor 610 may include a media (e.g., voice and music) codec (CODEC) 608 and an echo canceller 612. Media CODEC 608 includes a decoder 118A and an encoder 114A. Encoder 114A includes a stereo parameter adjustment unit 111, and decoder 118A includes a stereo parameter adjustment unit 312.
[0136] Device 600 includes memory 153 and CODEC 634. Although media CODEC 608 is described as a component of processor 610 (e.g., dedicated circuitry and / or executable program code), in other embodiments, one or more components of media CODEC 608 (e.g., decoder 118A, encoder 114A, or a combination thereof) may be included in processor 606, CODEC 634, another processing component, or a combination thereof.
[0137] Device 600 includes a transmitter 110 and a receiver 115. Transmitter 110 and receiver 115 are coupled to antenna 642. Device 600 includes a display 628 coupled to display controller 626. One or more speakers 648 are coupled to CODEC 634. One or more microphones 646 can be coupled to CODEC 634 via input interface 112. In a particular embodiment, speaker 648 includes... Figure 1 A first loudspeaker 142, a second loudspeaker 144, or a combination thereof. In a particular embodiment, microphone 646 includes... Figure 1 The first microphone 146, the second microphone 148, or a combination thereof. The CODEC 634 includes a digital-to-analog converter (DAC) 602 and an analog-to-digital converter (ADC) 604.
[0138] Memory 153 contains instructions 660, which can be executed by processor 606, processor 610, CODEC 634, encoder 114A, decoder 118A, another processing unit of device 600, or a combination thereof, to perform reference operations. Figures 1 to 5 The one or more operations described.
[0139] One or more components of device 600 may be implemented via dedicated hardware (e.g., circuitry) by a processor that executes instructions for performing one or more tasks or combinations thereof. As an example, one or more components of memory 153 or processor 606, processor 610, and / or CODEC 634 may be memory devices, such as random access memory (RAM), magnetoresistive random access memory (MRAM), spin-torque transfer MRAM (STT-MRAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, removable disk, or compact disc read-only memory (CD-ROM). The memory device may contain instructions (e.g., instruction 660) that, when executed by a computer (e.g., the processor, processor 606, encoder 114A, decoder 118A, and / or processor 610 in CODEC 634), cause the computer to perform reference... Figures 1 to 5 The described one or more operations. As an example, memory 153 or processor 606, processor 610, encoder 114A, decoder 118A and / or one or more components of CODEC 634 may be non-transitory computer-readable media containing instructions (e.g., instruction 660) that, when executed by a computer (e.g., the processor, processor 606 and / or processor 610 in CODEC 634), cause the computer to perform reference... Figures 1 to 5 The one or more operations described.
[0140] In certain embodiments, device 600 may be included in a system-in-package (SoC) or system-on-a-chip (SoC) device (e.g., a mobile station modem (MSM)) 622. In certain embodiments, processor 606, processor 610, display controller 626, memory 153, CODEC 634, transmitter 110, and receiver 115 are included in the SoC or SoC device 622. In certain embodiments, input devices 630, such as a touchscreen and / or keypad, and power supply 644 are coupled to the SoC device 622. Furthermore, in certain embodiments, such as... Figure 6 As described, the display 628, input device 630, speaker 648, microphone 646, antenna 642, and power supply 644 are external to the system-on-a-chip device 622. However, each of the display 628, input device 630, speaker 648, microphone 646, antenna 642, and power supply 644 may be coupled to components of the system-on-a-chip device 622, such as an interface or controller.
[0141] Device 600 may include: wireless telephone, mobile communication device, mobile phone, smartphone, cellular telephone, notebook computer, desktop computer, computer, tablet computer, set-top box, personal digital assistant (PDA), display device, television, game console, music player, radio, video player, entertainment unit, communication device, fixed location data unit, personal media player, digital video player, digital video disc (DVD) player, tuner, camera, navigation device, decoder system, encoder system, or any combination thereof.
[0142] In certain embodiments, one or more components of the systems and apparatuses disclosed herein may be integrated into a decoding system or device (e.g., an electronic device, a CODEC, or a processor therein), an encoding system or device, or both. In other embodiments, one or more components of the systems and apparatuses disclosed herein may be integrated into a wireless telephone, tablet computer, desktop computer, laptop computer, set-top box, music player, video player, entertainment unit, television, game console, navigation device, communication device, personal digital assistant (PDA), fixed location data unit, personal media player, or another type of device.
[0143] In conjunction with the techniques disclosed above, an apparatus includes means for receiving a coded bitstream comprising a coded intermediate channel and stereo parameters. The stereo parameters include IPD parameter values and mismatch values indicating the amount of misalignment between an encoder-side reference channel and an encoder-side target channel. For example, the means for receiving may include... Figure 1 and 6 Receiver 115 Figure 6 Antenna 642, other processors, circuits, hardware components or combinations thereof.
[0144] The device further includes means for decoding the encoded intermediate channel to produce a decoded intermediate channel. For example, the means for decoding may include... Figure 1 Decoder 118 Figure 1 and 3 Intermediate channel decoder 302, Figure 1 and 6 Decoder 118A, Figure 6 Processor 610, Figure 6 The processor 606 can be made by Figure 6 The processor components execute instructions 660, other processors, circuits, hardware components, or combinations thereof.
[0145] The device further includes means for performing a transform operation on the decoded intermediate channel to generate a frequency-domain decoded intermediate channel. For example, the means for performing the transform operation may include... Figure 1 Decoder 118 Figure 1 and 3 Transformation unit 306 Figure 1 and 6 Decoder 118A, Figure 6 Processor 610, Figure 6 The processor 606 can be made by Figure 6 The processor components execute instructions 660, other processors, circuits, hardware components, or combinations thereof.
[0146] The device further includes means for modifying at least a portion of the IPD parameter values based on the mismatch value to produce modified IPD parameter values. For example, the means for modification may include... Figure 1 Decoder 118 Figure 1 , 3 and the stereo parameter adjustment unit 312 of 6, Figure 1 and 6 Decoder 118A, Figure 6 Processor 610, Figure 6 The processor 606 can be made by Figure 6 The processor components execute instructions 660, other processors, circuits, hardware components, or combinations thereof.
[0147] The device further includes means for performing a upmixing operation on the frequency-domain decoded intermediate channel to generate a frequency-domain left channel and a frequency-domain right channel. Modified IPD parameter values are applied to the frequency-domain decoded intermediate channel during the upmixing operation. For example, the means for performing the upmixing operation may include... Figure 1 Decoder 118 Figure 1 and 3 310 upmixer Figure 1 and 6 Decoder 118A, Figure 6 Processor 610, Figure 6 The processor 606 can be made by Figure 6 The processor components execute instructions 660, other processors, circuits, hardware components, or combinations thereof.
[0148] The device further includes means for performing a first inverse transform operation on the frequency domain left channel to generate a time domain left channel. For example, the means for performing the first inverse transform operation may include... Figure 1 Decoder 118 Figure 1 and 3 Inverse transformation unit 318 Figure 1 and 6 Decoder 118A, Figure 6 Processor 610, Figure 6 The processor 606 can be made by Figure 6 The processor components execute instructions 660, other processors, circuits, hardware components, or combinations thereof.
[0149] The device further includes means for performing a second inverse transform operation on the frequency-domain right channel to generate a time-domain right channel. For example, the means for performing the second inverse transform operation may include... Figure 1 Decoder 118 Figure 1 and 3 Inverse transform unit 320, Figure 1 and 6 Decoder 118A, Figure 6 Processor 610, Figure 6 The processor 606 can be made by Figure 6 The processor components execute instructions 660, other processors, circuits, hardware components, or combinations thereof.
[0150] The device further includes means for outputting at least one of a left channel or a right channel, the left channel being associated with a time-domain left channel and the right channel being associated with a time-domain right channel. For example, the means for outputting may include... Figure 1 The first loudspeaker 142 Figure 1 The second loudspeaker 144 Figure 6 The speaker 648, other processors, circuits, hardware components or combinations thereof.
[0151] refer to Figure 7 A block diagram depicting a specific illustrative example of a base station 700 is provided. In various embodiments, base station 700 can be compared to... Figure 7The description indicates whether it has more or fewer components. In the illustrative example, base station 700 may be based on... Figure 4 Method 400 Figure 5 Method 500 or both can be used.
[0152] Base station 700 can be part of a wireless communication system. The wireless communication system can include multiple base stations and multiple wireless devices. The wireless communication system can be a Long Term Evolution (LTE) system, a fourth-generation (4G) LTE system, a fifth-generation (5G) system, a Code Division Multiple Access (CDMA) system, a Global System for Mobile Communications (GSM) system, a wireless local area network (WLAN) system, or any other wireless system. The CDMA system can implement Wideband CDMA (WCDMA), CDMA 1X, Evolution-Data Optimized (EVDO), Time Division Synchronous CDMA (TD-SCDMA), or some other version of CDMA.
[0153] Wireless devices can also be referred to as user equipment (UE), mobile station, terminal, access terminal, user unit, workstation, etc. These wireless devices may include: cellular phones, smartphones, tablet computers, wireless modems, personal digital assistants (PDAs), handheld devices, laptops, smartbooks, mini-notebooks, tablet computers, wired phones, wireless local loop (WLL) stations, Bluetooth devices, etc. Wireless devices may include or correspond to... Figure 6 Device 600.
[0154] Various functions can be performed by one or more components of base station 700 (and / or in other components not shown), such as sending and receiving messages and data (e.g., audio data). In a particular instance, base station 700 includes processor 706 (e.g., CPU). Base station 700 may include transcoder 710. Transcoder 710 may include audio CODEC 708 (e.g., voice and music CODEC). For example, transcoder 710 may include one or more components (e.g., circuitry) configured to perform operations on audio CODEC 708. As another example, transcoder 710 is configured to execute one or more computer-readable instructions to perform operations on audio CODEC 708. Although audio CODEC 708 is described as a component of transcoder 710, in other instances, one or more components of audio CODEC 708 may be included in processor 706, another processing component, or a combination thereof. For example, decoder 118 (e.g., vocoder decoder) may be included in receiver data processor 764. As another example, encoder 114 (e.g., vocoder encoder) may be included in data transmission processor 782.
[0155] Transcoder 710 serves to transcode messages and data between two or more networks. Transcoder 710 is configured to convert message and audio data from a first format (e.g., digital format) to a second format. For illustration, decoder 118 can decode an encoded signal having the first format, and encoder 114 can encode the decoded signal into an encoded signal having the second format. Alternatively or additionally, transcoder 710 is configured to perform data rate adaptation. For example, transcoder 710 can down-convert or up-convert the data rate without changing the format of the audio data. For illustration, transcoder 710 can down-convert a 64 kbit / s signal to a 16 kbit / s signal. Audio CODEC 708 may include encoder 114 and decoder 118. Decoder 118 may include stereo parameter adjuster 618.
[0156] Base station 700 includes memory 732. Memory 732 (an example of a computer-readable storage device) may contain instructions. The instructions may include instructions that can be executed by processor 706, transcoder 710, or a combination thereof. Figure 4 Method 400 Figure 5 Method 500 or one or more instructions of both. Base station 700 may include multiple transmitters and receivers (e.g., transceivers) coupled to an array of antennas, such as a first transceiver 752 and a second transceiver 754. The antenna array may include a first antenna 742 and a second antenna 744. The antenna array is configured to communicate wirelessly with one or more wireless devices, such as... Figure 6The device 600. For example, the second antenna 744 can receive a data stream 714 (e.g., a bit stream) from the wireless device. The data stream 714 may contain messages, data (e.g., encoded voice data), or a combination thereof.
[0157] Base station 700 may include network connection 760, such as a backhaul connection. Network connection 760 is configured to communicate with one or more base stations of a core network or wireless communication network. For example, base station 700 may receive a second data stream (e.g., message or audio data) from the core network via network connection 760. Base station 700 may process the second data stream to generate message or audio data and provide the message or audio data to one or more wireless devices via one or more antennas of an antenna array, or provide it to another base station via network connection 760. In certain embodiments, as an illustrative and non-limiting example, network connection 760 may be a wide area network (WAN) connection. In some embodiments, the core network may include or correspond to a Public Switched Telephone Network (PSTN), a backbone network, or both.
[0158] Base station 700 may include media gateway 770 coupled to network connection 760 and processor 706. Media gateway 770 is configured to switch between media streaming transmissions of different telecommunications technologies. For example, media gateway 770 may switch between different transport protocols, different decoding schemes, or both. For illustration, as an illustrative and non-limiting example, media gateway 770 may switch from PCM signals to Real-Time Transport Protocol (RTP) signals. Media gateway 770 may switch data between packet-switched networks (e.g., Voice Over Internet Protocol (VoIP) networks, IP Multimedia Subsystem (IMS), fourth-generation (4G) wireless networks such as LTE, WiMax, and UMB, fifth-generation (5G) wireless networks, etc.), circuit-switched networks (e.g., PSTN), and hybrid networks (e.g., second-generation (2G) wireless networks such as GSM, GPRS, and EDGE, third-generation (3G) wireless networks such as WCDMA, EV-DO, and HSPA, etc.).
[0159] Additionally, media gateway 770 may include a transcoder, such as transcoder 710, configured to transcode data when codecs are incompatible. For example, as an illustrative and non-limiting example, media gateway 770 may perform transcoding between an Adaptive Multi-Rate (AMR) codec and a G.711 codec. Media gateway 770 may include a router and multiple physical interfaces. In some embodiments, media gateway 770 may also include a controller (not shown). In certain embodiments, the media gateway controller may be external to media gateway 770, external to base station 700, or external to both. The media gateway controller can control and coordinate the operation of multiple media gateways. Media gateway 770 may receive control signals from the media gateway controller and may act as a bridge between different transmission technologies, and may add services to end-user capabilities and connectivity.
[0160] Base station 700 may include a demodulator 762 coupled to transceiver 752, transceiver 754, receiver data processor 764, and processor 706, and receiver data processor 764 may be coupled to processor 706. Demodulator 762 is configured to demodulate modulated signals received from transceivers 752 and 754, and is configured to provide demodulated data to receiver data processor 764. Receiver data processor 764 is configured to extract message or audio data from the demodulated data and transmit the message or audio data to processor 706.
[0161] Base station 700 may include a transmission data processor 782 and a transmission multiple-input multiple-output (MIMO) processor 784. Transmission data processor 782 may be coupled to processor 706 and transmission MIMO processor 784. Transmission MIMO processor 784 may be coupled to transceiver 752, transceiver 754, and processor 706. In some embodiments, transmission MIMO processor 784 may be coupled to media gateway 770. As an illustrative and non-limiting example, transmission data processor 782 is configured to receive message or audio data from processor 706 and decode the message or audio data based on a decoding scheme such as CDMA or orthogonal frequency-division multiplexing (OFDM). Transmission data processor 782 may provide decoded data to transmission MIMO processor 784.
[0162] CDMA or OFDM technologies can be used to multiplex decoded data with other data, such as pilot data, to generate multiplexed data. The multiplexed data can then be modulated (i.e., symbol-mapped) by the data processor 782 based on a specific modulation scheme (e.g., Binary Phase-Shift Keying (BPSK), Quadrature Phase-Shift Keying (QSPK), M-ary Phase-Shift Keying (M-PSK), M-ary Quadrature Amplitude Modulation (M-QAM), etc.) to generate modulated symbols. In certain embodiments, different modulation schemes can be used to modulate decoded data and other data. The data rate, decoding, and modulation for each data stream can be determined by instructions executed by the processor 706.
[0163] The transport MIMO processor 784 is configured to receive modulation symbols from the transport data processor 782 and can further process the modulation symbols and perform beamforming on the data. For example, the transport MIMO processor 784 can apply beamforming weights to the modulation symbols.
[0164] During operation, the second antenna 744 of base station 700 can receive data stream 714. The second transceiver 754 can receive data stream 714 from the second antenna 744 and can provide data stream 714 to demodulator 762. Demodulator 762 can demodulate the modulated signal of data stream 714 and provide the demodulated data to receiver data processor 764. Receiver data processor 764 can extract audio data from the demodulated data and provide the extracted audio data to processor 706.
[0165] Processor 706 may provide audio data to transcoder 710 for transcoding. Transcoder 710's decoder 118 may decode the audio data from a first format into decoded audio data, and encoder 114 may encode the decoded audio data into a second format. In some embodiments, encoder 114 may encode the audio data using a higher data rate (e.g., up-conversion) or a lower data rate (e.g., down-conversion) than the data rate received from the wireless device. In other embodiments, the audio data may be untranscoded. Although transcoding (e.g., decoding and encoding) is described as being performed by transcoder 710, transcoding operations (e.g., decoding and encoding) may be performed by multiple components of base station 700. For example, decoding may be performed by receiver data processor 764, and encoding may be performed by transport data processor 782. In other embodiments, processor 706 may provide audio data to media gateway 770 for conversion into another transport protocol, decoding scheme, or both. Media gateway 770 may provide the converted data to another base station or core network via network connection 760.
[0166] Encoded audio data, such as transcoded data, generated at encoder 114 can be provided to transmission data processor 782 or network connection 760 via processor 706. Transcoded audio data from transcoder 710 can be provided to transmission data processor 782 for decoding according to a modulation scheme such as OFDM to generate modulation symbols. Transmission data processor 782 can provide the modulation symbols to transmission MIMO processor 784 for further processing and beamforming. Transmission MIMO processor 784 can apply beamforming weights and can provide the modulation symbols to one or more antennas of an antenna array, such as first antenna 742, via first transceiver 752. Therefore, base station 700 can provide transcoded data stream 716, corresponding to data stream 714 received from a wireless device, to another wireless device. Transcoded data stream 716 may have a different encoding format, data rate, or both than data stream 714. In other embodiments, transcoded data stream 716 can be provided to network connection 760 for transmission to another base station or core network.
[0167] It should be noted that the various functions performed by one or more components of the systems and apparatuses disclosed herein are described as being performed by certain components or modules. This division of components and modules is for illustrative purposes only. In alternative embodiments, functions performed by a particular component or module may be divided among multiple components or modules. Furthermore, in alternative embodiments, two or more components or modules may be integrated into a single component or module. Each component or module may be implemented using hardware (e.g., field-programmable gate array (FPGA) devices, application-specific integrated circuits (ASICs), DSPs, controllers, etc.), software (e.g., instructions executable by a processor), or any combination thereof.
[0168] Those skilled in the art will further understand that the various illustrative logic blocks, configurations, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processing device such as a hardware processor, or a combination of both. The foregoing has generally described various illustrative components, blocks, configurations, modules, circuits, and steps in terms of functionality. Whether this functionality is implemented as hardware or executable software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in varying ways for each specific application, but these implementation decisions should not be interpreted as departing from the scope of the invention.
[0169] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be directly embodied in hardware, in a software module executed by a processor, or a combination of both. The software module can reside in a memory device, such as random access memory (RAM), magnetoresistive random access memory (MRAM), spin torque transfer MRAM (STT-MRAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, or optical disc read-only memory (CD-ROM). An exemplary memory device is coupled to a processor so that the processor can read information from and write information to the memory device. Alternatively, the memory device can be integrated with the processor. The processor and storage medium can reside in an application-specific integrated circuit (ASIC). The ASIC can reside in a computing device or user terminal. In an alternative example, the processor and storage medium can reside as discrete components in a computing device or user terminal.
[0170] The prior description of the disclosed embodiments is provided to enable those skilled in the art to make or use the disclosed embodiments. Various modifications to these embodiments will readily apparent to those skilled in the art, and the principles defined herein can be applied to other embodiments without departing from the scope of the invention. Therefore, the invention is not intended to be limited to the embodiments shown herein, but should be accorded the widest scope consistent with the principles and novel features as defined in the following claims.
Claims
1. An apparatus for decoding an audio channel, comprising: A receiver configured to receive a encoded bitstream comprising at least encoded intermediate channels and stereo parameters, the stereo parameters including inter-channel phase difference (IPD) parameter values and mismatch values, the mismatch values indicating the amount of time misalignment between the encoder-side reference channel and the encoder-side target channel; A stereo parameter adjustment unit configured to modify at least a portion of the IPD parameter values based on the mismatch value to produce modified IPD parameter values; and An upmixer configured to perform upmixing on a frequency-domain decoded intermediate channel to generate a frequency-domain left channel and a frequency-domain right channel, wherein modified IPD parameter values are applied to the frequency-domain decoded intermediate channel during the upmixing operation, and the frequency-domain decoded intermediate channel corresponds to the decoded version of the encoded intermediate channel.
2. The apparatus of claim 1, wherein the stereo parameter adjustment unit is further configured to modify at least a portion of the IPD parameter value based on the availability of the coded side channel contained in the coded bitstream and the mismatch value to generate the modified IPD parameter value.
3. The apparatus of claim 1 or 2, wherein the stereo parameter adjustment unit is further configured to set one or more of the IPD parameter values to zero.
4. The apparatus according to claim 1, further comprising: An intermediate channel decoder, configured to decode the encoded intermediate channel to produce a decoded intermediate channel; and A transformation unit configured to perform a transformation operation on the decoded intermediate channel to generate the frequency-domain decoded intermediate channel.
5. The apparatus according to claim 1, further comprising: A first inverse transform unit is configured to perform a first inverse transform operation on the frequency domain left channel to generate a time domain left channel; and The second inverse transform unit is configured to perform a second inverse transform operation on the frequency domain right channel to generate a time domain right channel.
6. The apparatus according to claim 5, further comprising: One or more speakers configured to output at least one of a left channel or a right channel, the left channel being associated with the time-domain left channel and the right channel being associated with the time-domain right channel.
7. The apparatus of claim 6, wherein the stereo parameters include an inter-channel time difference (ITD) parameter value as the mismatch value, and further comprises: Inter-channel alignment unit, configured to: The time-domain right channel is adjusted based on the ITD parameter value to generate the right channel; or The time-domain left channel is adjusted based on the ITD parameter value to generate the left channel.
8. The apparatus of claim 7, wherein the inter-channel alignment unit is included in the upmixer.
9. The apparatus of claim 1, wherein the stereo parameter adjustment unit is configured to: Compare the absolute value of the mismatch value with the threshold; and At least a portion of the IPD parameter value is modified in response to the determination that the absolute value of the mismatch value satisfies the threshold.
10. The apparatus according to claim 2, further comprising: A side channel decoder configured to decode the encoded side channel to generate a decoded side channel, the encoded side channel being contained in the encoded bitstream; and A second transformation unit is configured to perform a second transformation operation on the decoded side channel to generate a frequency-domain decoded side channel.
11. The apparatus of claim 1, wherein the stereo parameter adjustment unit is further configured to modify the IPD parameter value based on the bit rate associated with the encoded bit stream.
12. The apparatus of claim 1, wherein the stereo parameter adjustment unit is further configured to modify the IPD parameter value based on vocal parameters, packet loss determination associated with a previous frame, speech / music classification, or another parameter.
13. The apparatus of claim 1, wherein the stereo parameter adjustment unit is further configured to smooth one or more of the IPD parameter values over time.
14. The apparatus of claim 1, wherein the mismatch value indicates one of the following: an amount of time misalignment in the frequency domain, or an amount of time misalignment in the time domain.
15. The apparatus of claim 1, wherein the stereo parameter adjustment unit is integrated into a mobile device or base station.
16. A method for decoding an audio channel, the method comprising: At the decoder, a encoded bitstream is received that includes at least encoded intermediate channels and stereo parameters, wherein the stereo parameters include inter-channel phase difference (IPD) parameter values and mismatch values, wherein the mismatch values indicate the amount of time misalignment between the encoder-side reference channel and the encoder-side target channel. Modify at least a portion of the IPD parameter values based on the mismatch value to generate modified IPD parameter values; and An upmixing operation is performed on the frequency-domain decoded intermediate channel to generate a frequency-domain left channel and a frequency-domain right channel. The modified IPD parameter values are applied to the frequency-domain decoded intermediate channel during the upmixing operation, and the frequency-domain decoded intermediate channel corresponds to the decoded version of the encoded intermediate channel.
17. The method of claim 16, wherein modifying at least a portion of the IPD parameter value based on the mismatch value to generate the modified IPD parameter value further comprises: The modified IPD parameter values are generated by modifying at least a portion of the IPD parameter values based on the availability of the coded side channels contained in the coded bitstream and the mismatch value.
18. The method of claim 16 or 17 further comprises setting one or more of the IPD parameter values to zero.
19. An apparatus for decoding an audio channel, comprising: A unit for receiving a encoded bitstream that includes at least an encoded intermediate channel and stereo parameters, wherein the stereo parameters include an inter-channel phase difference (IPD) parameter value and a mismatch value, the mismatch value indicating the amount of time misalignment between the encoder-side reference channel and the encoder-side target channel; Units for modifying at least a portion of the IPD parameter values based on the mismatch value to generate modified IPD parameter values; and A unit for performing upmixing on a frequency-domain decoded intermediate channel to generate a frequency-domain left channel and a frequency-domain right channel, wherein the modified IPD parameter value is applied to the frequency-domain decoded intermediate channel during the upmixing operation, and the frequency-domain decoded intermediate channel corresponds to the decoded version of the encoded intermediate channel.
20. The device of claim 19, wherein the unit of modifying at least a portion of the IPD parameter value based on the mismatch value to generate the modified IPD parameter value further comprises: The unit that generates the modified IPD parameter value is modified by modifying at least a portion of the IPD parameter value based on the availability of the coded side channel contained in the coded bitstream and the mismatch value.
21. The device according to claim 19 or 20, further comprising: A unit for setting one or more of the IPD parameter values to zero.
22. A non-transitory computer-readable medium comprising instructions that, when executed by a processor within a decoder, cause the processor to perform the operation of the method according to any one of claims 16-18.